Example: air traffic controller

Prior Distributions - stats.org.uk

Prior DistributionsThere are three main ways of choosing a Prior . Subjective Objective and informative NoninformativeSubjectiveAs mentioned previously, the Prior may be de-termined subjectively. In this case the Prior ex-presses the experimenter s personal probabilitythat lies in (essentially) any given subset of .See Chapter 3 of Berger,Statistical DecisionTheory and Bayesian Analysisfor a discussionof methods for subjectively choosing a and informativeThe experimenter may have information or datathat can be used to help formulate a could take at least two data on the distribution of pa-rameter from experiments done Prior to theone being example of 1 is as follows. A companywants to estimate the proportion of all partsproduced on a particular day that are defec-tive. They will take only a sample of the day sproduction to estimate this production records we have 1, 2, .. , N,which are (approximate) proportions of defec-tive parts for a sequence may fit a distribution to the observations 1.

Objective and informative The experimenter may have information or data that can be used to help formulate a prior. This could take at least two forms:

Tags:

  Information, Distribution, Prior, Prior distributions

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Prior Distributions - stats.org.uk

1 Prior DistributionsThere are three main ways of choosing a Prior . Subjective Objective and informative NoninformativeSubjectiveAs mentioned previously, the Prior may be de-termined subjectively. In this case the Prior ex-presses the experimenter s personal probabilitythat lies in (essentially) any given subset of .See Chapter 3 of Berger,Statistical DecisionTheory and Bayesian Analysisfor a discussionof methods for subjectively choosing a and informativeThe experimenter may have information or datathat can be used to help formulate a could take at least two data on the distribution of pa-rameter from experiments done Prior to theone being example of 1 is as follows. A companywants to estimate the proportion of all partsproduced on a particular day that are defec-tive. They will take only a sample of the day sproduction to estimate this production records we have 1, 2, .. , N,which are (approximate) proportions of defec-tive parts for a sequence may fit a distribution to the observations 1.

2 , Nand use this as a Prior . The usualoptions are available for fitting the distribu-tion: parametric methods, histograms, kerneldensity estimates, than having observations on the param-eter itself, one may have previous data thatmerely contains information about . In thiscase we could use the previous posterior asthe Prior for the upcoming experiment. This issummarized in the following maxim: Today s posterior is tomorrow s Prior . 52 Let s look at the maxim a bit closer. Definethe following quantities: Y1: the random vector whose value,y1,has already been observed Y2: the random vector we are to observein the next experiment : the Prior previous to the first experiment(in which we obtainedy1) f1(y1| ): the distribution ofY1given f2(y2| ): the distribution ofY2given f(y1,y2| ): the joint distribution ofY1andY2given 53 Two possible approaches for inferring havingobserved bothy1andy2 the posterior, 1( |y1), from experi-ment 1 as the Prior leading into experi-ment 2 where we will observe a value ofY2.

3 (This is literally what the maxim says.) (y1,y2) as one big set of data withlikelihoodfand compute posterior ( |y1,y2) f(y1,y2| ) ( ).When do the two approaches coincide?54In approach 1, the posterior after experiment2 is 2( |y2) 1( |y1)f2(y2| ) f1(y1| )f2(y2| ) ( ).In general, approaches 1 and 2 will give thesame posterior only whenf(y1,y2| ) =f1(y1| )f2(y2| ) ,which requiresY1andY2to be , in order for the maxim to be true in thestrictest sense, the data not yet observed shouldbe independent of those already priorsA noninformative Prior is one that expressesignorance as to the value of . Other termsfor a noninformative Prior are reference Prior ,diffuse Prior and vague general, a noninformative Prior is one whichis dominated by the likelihood function. Inother words, such a Prior does not change much over the region inwhich the likelihood is appreciable, and does not assume large values outside Prior having the two properties above is saidto belocally uniform.

4 Box and Tiao,BayesianInference in Statistical Analysisgive an excel-lent account of locally uniform Jeffreys noninformative priorSuppose that = ( 1, .. , p). TheFisher in-formation matrix,I( ), is thep pmatrix with(i, j) element E[ 2logf(Y| ) i j].The Jeffreys noninformative Prior is ( ) det(I( ))1/2,where det(A) denotes the determinant of motivation for this Prior is a certain invari-ance argument. Consider a 1-1 transformationof the parameter: =h( ).Now, if our Prior for is , then the corre-sponding density of =h( ) isg( ) = (h 1( ))|J( )|,whereJis the Jacobian of the a method for finding a noninforma-tive Prior . Now, supposeMPis used to obtaina Prior for . Jeffreys argued that ifMPisused to find a Prior for =h( ), then itshould be true that ( ) =g( ) ,wheregis defined at the bottom of the previ-ous Jeffreys Prior satisfies this property. Let scheck this in the casep= 1. We have ( ) { E[ 2logf(Y| ) 2]}1/2andg( ) = (h 1( )) dh 1( )d.

5 58 The Fisher information when we use the pa-rameterization =h( ) is E[ 2logf(Y|h 1( )) 2],and so the Jeffreys Prior for is proportionalto the square root of the last expression, whichis E[ 2logf(Y|h 1( )) h 1( )2 h 1( )2 2]= ( h 1( ) )2E[ 2logf(Y|h 1( )) h 1( )2].But the last expression is proportional tog2( )(as defined on the previous page), and henceJeffreys invariance property 7 Jeffreys Prior for the binomial ex-perimentWe havelogf(y| ) = log(ny)+ylog +(n y)log(1 ), logf(y| ) =y (n y)(1 ),and 2logf(y| ) 2= y 2 (n y)(1 ) , E[ 2logf(Y| ) 2]=n 2+(n n )(1 )2=n (1 ).So, the Jeffreys noninformative Prior for isproportional to [ (1 )] 1/2, and hence mustbe a Beta(1/2,1/2) 8 Jeffreys Prior for normal randomsampleLetY= (Y1, .. , Yn), whereY1, .. , Ynare ( 1, 22). We havef(y| ) =(1 2 2)nexp 12 22n i=1(yi 1)2 ,andlogf(y| ) = nlog( 2 2) 12 22n i=1(yi 1) , logf(y| ) 1=1 22n i=1(yi 1), logf(y| ) 2= n 2+1 32n i=1(yi 1)2,61 2logf(y| ) 21= n 22, 2logf(y| ) 22=n 22 3 42n i=1(yi 1)2,and 2logf(y| ) 1 2= 2 32n i=1(yi 1).

6 The Fisher information matrix is thus[n/ 2200 2n/ 22].The determinant of this matrix is 2n2/ 42, andhence the Jeffreys noninformative Prior is suchthat ( 1, 2) 1 22I( , )( 1)I(0, )( 2).62 The Prior at the bottom of the previous pageis calledimpropersince it is not is an example of the unfortunate fact thatJeffreys noninformative Prior is sometimes that the form of Jeffreys Prior in this caseimplies that 1and 2are a priori independentwith 1( 1) = constant 1and 2( 2) =1 22I(0, )( 2).So, it is the Prior for the location parameter 1that is improper. The Prior for 2is


Related search queries