顯示具有 統計 標籤的文章。 顯示所有文章
顯示具有 統計 標籤的文章。 顯示所有文章

2010年2月19日 星期五

[統計] PK formula (Pollaczek–Khinchine formula)

cite: http://en.wikipedia.org/wiki/Pollaczek%E2%80%93Khinchine_formula

PK formula for M/G/1 queue
可以看到只要moment fitted
我們就可以預期得到一樣的 long term waiting time
不過問題是
我們怎麼知道真實世界的moment是怎樣呢?

Pollaczek–Khinchine formula

From Wikipedia, the free encyclopedia

Jump to: navigation, search

The Pollaczek-Khinchine formula is used in queuing theory to determine the mean time spent waiting in the queue to be serviced (the queuing delay) and the mean end-to-end time through the system. The formula is applicable in a single server situation with arrivals distributed according to a Poisson distribution and a general service time distribution. [Known as a M/G/1 system in Kendall's notation.] The formula was developed by Felix Pollaczek and Aleksandr Khinchin.

[edit] Formula

The formula states that the mean queuing delay is given by:
F_q=\frac{1}{\lambda_s}\times \frac{\rho}{1-\rho}\times\frac{1+C_s^2}{2}
The average time in the system, F, is given by:
F=F_q+\frac{1}{\lambda_s}
In the above equations, the variables are defined as:
λs=rate of service
λa=rate of arrival
\rho=\frac{\lambda_a}{\lambda_s}, which is called "traffic intensity," ranges between 0 and 1, and is the mean fraction of time that the server is busy. [If the arrival rate λa is greater than or equal to the service rate λs, the queuing delay becomes infinite.]
Cs is the coefficient of variation of the service time (the ratio of its standard deviation to its mean). This equals σsλs, where σs is the standard deviation of the service time, as the mean service time is 1/λs. Cs = 0 when service times are constant, and Cs = 1 when service times follow an Exponential distribution.

[edit] Examples

If ρ equals 0.5, that is the server is busy 50% of the time, and Cs = 1, then
F_q=\frac{1}{\lambda_s}\times \frac{\rho}{1-\rho}\times\frac{1+C_s^2}{2}
F_q=\frac{1}{\lambda_s}\times \frac{0.5}{0.5}\times\frac{1+1}{2}
F_q=\frac{1}{\lambda_s}
That is, when the server is busy only half the time, the mean queuing time equals the mean service time. That may help explain the long wait at the post office!

As ρ increases, the mean queuing time increases rapidly. If the server is 90% utilized, then the mean queuing delay is nine times the mean service time.

Note that, if the service time is always the same (Cs = 0), then the mean queuing delay is half what it would be if the service time were exponentially distributed (Cs = 1).

[統計] Beta distribution

Beta distribution是最常用的一個分配
我們看到的一般是
General beta (α,β,min, max)
前面兩個是shape parameters 後面兩個是 range parameters

其他相關的可以參考wiki連結
http://en.wikipedia.org/wiki/Beta_distribution

[轉錄] Histogram 製作方法

實用參考連結。

Creating a Histogram in Excel

Sales Forecast Example - Part III

In Part II of this Monte Carlo Simulation example, we completed the actual simulation. (If you haven't already, Download the example spreadsheet). We ended up with a column of 5000 possible values (observations) for our single response variable, profit. The last step is to analyze the results. We will start off by creating a histogram in Excel, a graphical method for visualizing the results.

We can glean a lot of information from this histogram:

  • It looks like profit will be positive, most of the time.
  • The uncertainty is quite large, varying between -1000 to 3400.
  • The distribution does not look like a perfect Normal distribution.
  • There doesn't appear to be outliers, truncation, multiple modes, etc.

The histogram tells a good story, but in many cases, we want to estimate the probability of being below or above some value, or between a set of specification limits.

Creating a Histogram in Excel

Method 1: Using the Histogram Tool in the Analysis Tool-Pak.

This is probably the easiest method, but you have to re-run the tool each to you do a new simulation. AND, you still need to create an array of bins (which will be discussed below).

Method 2: Using the FREQUENCY function in Excel.

This is the method used in the spreadsheet for the sales forecast example. One of the reasons I like this method is that you can make the histogram dynamic, meaning that every time you re-run the MC simulation, the chart will automatically update. This is how you do it:

Step 1: Create an array of bins

The figure below shows how to easily create a dynamic array of bins. This is a basic technique for creating an array of N evenly spaced numbers.

To create the dynamic array, enter the following formulas:
B6 = $B$2
B7 = B6+($B$3-$B$2)/5
Then, copy cell B7 down to B11

Array of Bins in Excel
Figure 2: A dynamic array of 5 bins.

After you create the array of bins, you can go ahead and use the Histogram tool, or you can proceed with the next step.

Step 2: Use Excel's FREQUENCY formula

The next figure is a screen shot from the example Monte Carlo simulation. I'm not going to explain the FREQUENCY function in detail since you can look it up in the Excel's help file. But, one thing to remember is that it is an array function, and after you enter the formula, you will need to press Ctrl+Shift+Enter. Note that the simulation results (Profit) are in column G and there are 5000 data points ( Points: J5=COUNT(G:G) ).

The Formula for the Count column:
FREQUENCY(data_array,bins_array)

a) Select cells J8:J48
b) Enter the array formula: {=FREQUENCY(G:G,I8:I48)}
c) Press Ctrl+Shift+Enter

Layout for Creating a Scaled Histogram
Figure 3: Layout in Excel for Creating a Dynamic Scaled Histogram.

Creating a Scaled Histogram

If you want to compare your histogram with a probability distribution, you will need to scale the histogram so that the area under the curve is equal to 1 (one of the properties of probability distributions). Histograms normally include the count of the data points that fall into each bin on the y-axis, but after scaling, the y-axis will be the frequency (a not-so-easy-to-interpret number that in all practicality you can just not worry about). The frequency doesn't represent probability!

To scale the histogram, use the following method:
Scaled = (Count/Points) / (BinSize)

a) K8 = (J8/$J$5)/($I$9-$I$8)
b) Copy cell K8 down to K48
c) Press F9 to force a recalculation (may take a while)

Step 3: Create the Histogram Chart

Bar Chart, Line Chart, or Area Chart:

To create the histogram, just create a bar chart using the Bins column for the Labels and the Count or Scaled column as the Values. Tip: To reduce the spacing between the bars, right-click on the bars and select "Format Data Series...". Then go to the Options tab and reduce the Gap. Figure 1 above was created this way.

A More Flexible Histogram Chart

One of the problems with using bar charts and area charts is that the numbers on the x-axis is actually just labels. This can make it very difficult to overlay data that uses a different number of points or to show the proper scale when bins are not all the same size. However, you CAN use a scatter plot to create a histogram. After creating a line using the Bins column for the X Values and Count or Scaled column for the Y Values, add Y Error Bars to the line that extend down to the x-axis (by setting the Percentage to 100%). You can right-click on these error bars to change the line widths, color, etc.

Histogram Via Error Bars

2008年5月27日 星期二

[電影] 決勝21點-主持人遊戲

決戰21點不愧是商業片
把真人真事改編成這樣應該算成功了!
看完電影以後,發現裡面的問題造成許多討論
原題:
『有三個門給玩家選擇一個門,玩家知道
三個門後面一個是大獎,其他兩個沒有獎,
主持人知道哪個門後面有大獎,
並且在玩家選完後他會開一個沒有獎的門,
然後詢問玩家是否要換自己的選擇?』

很多人一直覺得主持人打開門後,機率變成兩個門各1/2
這並沒有考慮到主持人開門的動作以及主持人擁有的資訊
該動作將部分的資訊傳達了出來~~
假設一開始有ABC三個門,玩家選擇了A門,主持人開了C門
此時身為玩家的你,應該要堅持選擇A還是更換選擇?

解法:
應該要換門,因為在B門的機率2/3比在A門1/3大

簡單來說...一般人需要知道
一開始沒有辦法對哪個門後面有車子有資訊所以在每個門機率都是1/3
但是根據遊戲規則:
1. 今天主持人知道哪個門後面有大獎
2. 玩家選的門一定不能開
3. 大獎一定不能開

所以我們可以計算主持人開C門的機率為:
pr(大獎在A門情況下主持人開C門)+ (1)
pr(大獎在B門情況下主持人開C門)+ (2)
pr(大獎在C門請況下主持人開C門) (3)

=1/3(獎在A門)*1/2(B和C都可以選) (1)
+1/3(獎在B門)*1 (根據規則B不能選)(2)
+1/3(獎在C門)*0 (根據規則C不能選)(3)
=1/6+1/3
=1/2

而機率的定義是= (發生的相對頻率或次數)/(所有發生的頻率或次數)
現在可能開C門的情況之下可能發生的是(2)和(3)兩種情況

所以懂條件機率的可以輕易算出在B門是2/3 A門是1/3
不懂的人可以根據上面機率的定義瞭解
pr(獎在A門的機率並且主持人開了C門)= 1/6 ÷ 1/2 = 1/3
pr(獎在B門的機率並且主持人開了C門)= 1/3 ÷ 1/2 = 2/3

希望大家可以一次看懂..


2008年5月22日 星期四

Sampling Distribution

今天被人問到『那你覺得sampling distribution為什麼可以求出來?』
一時之間還真不知道該怎樣解釋..很擔心自己學的和他學的名詞會有落差
用google查一下..

原文來自於Wikipeida Sampling distribution
中文是我自己翻譯的,看的懂就好..我沒有太花心思去翻譯

In statistics, a sampling distribution is the probability distribution, under repeated sampling of the population, of a given statistic (a numerical quantity calculated from the data values in a sample).
統計上樣本分配一種機率分配,根據從母體產生的試驗(樣本)中產生之統計量的分配

The formula for the sampling distribution depends on the distribution of the population, the statistic being considered, and the sample size used. A more precise formulation would speak of the distribution of the statistic as that for all possible samples of a given size, not just "under repeated sampling".
樣本分配的公式主要受到以下幾點因素影響:母體分配、使用的統計量、以及使用的樣本數
可以更精準的說是『固定樣本數下所有可能樣本得到的統計量分配』而非單純重複的抽樣而已

For example, consider a very large normal population (one that follows the so-called bell curve). Assume we repeatedly take samples of a given size from the population and calculate the sample mean (\bar x, the arithmetic mean of the data values) for each sample. Different samples will lead to different sample means. The distribution of these means is the "sampling distribution of the sample mean" (for the given sample size). This distribution will be normal since the population was normal. (According to the central limit theorem, if the population is not normal but "sufficiently well behaved", the sampling distribution of the sample mean will still be approximately normal provided the sample size is sufficiently large.)
例如:假設一個很大的常態分配。假設重複的得到固定size的樣本
並且計算樣本的平均值已計算統計量,因為重複試驗造成每次得到的樣本平均值會有所差異
這些樣本平均值的分配就是樣本分配。
該分配是常態分配,因為母體是常態,甚至於如果母體非常態分配,只要樣本量夠大,根據中央極限定理,也會得到常態分配的值。

後話:
自己一開始就把人家要做的東西推翻是一種很不好的習慣,要改進!就算母體是一個參數,樣本還是可以推導出分配,不過到底意義在哪裡..我覺得是可以好好思考的地方~~

2008年4月29日 星期二

R Language 介紹

R language 是一個統計的語言
雖然他介面比SAS與SPSS等軟體簡單
可是對於研究人員來說似乎是一個更為體貼的軟體
畢竟我們想要的分析
很多時候都需要大量的計算與模擬
如果使用一般的統計軟體
必須要統計軟體內建好相關的function
可是對於沒有function的問題怎麼辦?

其實SAS也可以寫一些相關的程式來減輕自己的負擔
但是相對於R Language來說
大量的後驗機率積分和驗證都可以利用程式寫好
大量的計算功夫也可以透過程式處理好
這樣方便的軟體真的是很方便

我最近看到Baysian Statistics裡面的paper
幾乎大部分都偏向使用R來處理...
之後會整理一些使用的小心得
相信使用過的人都會說讚~~

2008年4月22日 星期二

統計,改變了世界


雖然現在的心情總是擔心自己無法survived
畢竟許多事情將會是第一次面對
不過擔心之餘
有一種挑戰的熱血隨之而生
彷彿想要騎腳踏車超過汽車的那種心情
anyway...說了那麼多心情
其實今天是想要說這本不錯的書
-統計,改變了世界 The Lady Tasting Tea
薩爾斯伯格 David Salsburg

很久以前從讀書俱樂部選的書
最近很愛看它-不管是吃飯、捷運上、或休息
每次看都可以給自己一些雀躍感和好奇感..

這本書從許多有趣的實驗設計開始
介紹了一連串統計發展的歷史和野史
以前上課的時候總是在學著相關內容
怎樣應用、怎樣檢定、怎樣分析
但是對於背後的發展不太瞭解

透過這本書...我可以慢慢地去瞭解相關的歷史
畢竟之後自己也是要走這個領域的
這本書對我來說
只是個開端,他引出了我許多的好奇心
以及征服感

這本書從Pearson, Carl的故事開始
一直發展到Fisher Ronald Aylmer
以及相關的打壓面對還有尼曼的產生
在我現在的時間點..我可能去找到本書中提到的問題去看
因為有層級貝氏的概念..其實自己也很高興能看到不同的觀點
畢竟從機率的假設開始..統計一直就存在著不同哲學觀念的看法
我知道自己應該不是能夠留名的大師
畢竟他們可是從小就發光發熱的
fisher在大學的時候就發表科學論文
我只希望能夠透過paper和期刊
透過努力和思考
去多瞭解這個世界的運轉道理罷了

真的很幸運自己能夠看到這本書
我對自己說了很多次『我要加油!』
這條路就像老師說的..不好走
但是希望我能夠走的下去
之後可能會陸續整理一些相關的想法吧