Example: quiz answers

Prior vs Likelihood vs Posterior Posterior Predictive ...

Prior vs Likelihood vs PosteriorPosterior Predictive DistributionPoisson DataStatistics 220 Spring 2005 Copyrightc 2005 by Mark E. IrwinChoosing the Likelihood ModelWhile much thought is put into thinking about priors in a Bayesian Analysis,the data ( Likelihood ) model can have a big that need to be made involve Independence vs Exchangable vs More Complex Dependence Tail size, Normal vstdf Probability of eventsChoosing the Likelihood Model1 Example: Probability of God s ExistanceTwo different analyses - both using the priorP[God] =P[No God] = Ratio Components:Di=P[Datai|God]P[Datai|No God]Evidence (Datai)Di- UnwinDi- ShermerRecognition of of moral of natural miracles (prayers)21 Extranatural miracles (resurrection) the Likelihood Model2P[God|Data]: Unwin:23 Shermer: even starting with the same Prior , the difference beliefs about what thedata says gives quite diff

Choosing the Likelihood Model While much thought is put into thinking about priors in a Bayesian Analysis, the data (likelihood) model can have a big efiect.

Tags:

  Posterior, Predictive, Likelihood, Likelihood vs posterior posterior predictive

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Prior vs Likelihood vs Posterior Posterior Predictive ...

1 Prior vs Likelihood vs PosteriorPosterior Predictive DistributionPoisson DataStatistics 220 Spring 2005 Copyrightc 2005 by Mark E. IrwinChoosing the Likelihood ModelWhile much thought is put into thinking about priors in a Bayesian Analysis,the data ( Likelihood ) model can have a big that need to be made involve Independence vs Exchangable vs More Complex Dependence Tail size, Normal vstdf Probability of eventsChoosing the Likelihood Model1 Example: Probability of God s ExistanceTwo different analyses - both using the priorP[God] =P[No God] = Ratio Components:Di=P[Datai|God]P[Datai|No God]Evidence (Datai)Di- UnwinDi- ShermerRecognition of of moral of natural miracles (prayers)21 Extranatural miracles (resurrection) the Likelihood Model2P[God|Data]: Unwin:23 Shermer: even starting with the same Prior , the difference beliefs about what thedata says gives quite different Posterior is based on an analysis published in an July 2004 Scientific Americanarticle.

2 (Available on the course web site on the Articles page.)Stephen D. Unwin is a risk management consultant who has done work inphysics on quantum gravity. He is author of the bookThe Probability Shermer is the publisher ofSkepticand a regular contributor toScientific Americanas author of the the Likelihood Model3 Note in the article, Bayes rule is presented asP[God|Data] =P[God]DP[God]D+P[No God]See if you can show that this is equivalent to the normal version of Bayes rule, under the assumption that the components of the data model the Likelihood Model4 Prior vs Likelihood vs PosteriorThe Posterior distribution can be seen as a compromise between the priorand the dataIn general, this can be seen based on the two well known relationshipsE[ ] =E[E[ |y]](1)Var( ) =E[Var( |y)] + Var(E[ |y])(2)The first equation says that our Prior mean is the average of all possibleposterior means (averaged over all possible data sets).

3 The second says that the Posterior variance is, on average, smaller than theprior variance. The size of the difference depends on the variability of theposterior vs Likelihood vs Posterior5 This can be exhibited more precisely using examples Binomial Model - Conjugate Prior Beta(a, b)y| Bin(n, ) |y Beta(a+y, b+n y) Prior mean:E[ ] =aa+b= MLE: =ynPrior vs Likelihood vs Posterior6 Then the Posterior mean satisfiesE[ |y] =a+ba+b+n +na+b+n = a weighted average of the Prior mean and the sample proportion (MLE) Prior variance:Var( ) = (1 )a+b+ 1 Posterior variance:Var( |y) = (1 )a+b+n+ 1So ifnis large enough, the Posterior variance will be smaller than theprior vs Likelihood vs Posterior7 Normal Model - Conjugate Prior , fixed variance N( 0, 20)yi| iid N( , 2);i= 1.

4 , n |y N( n, 2n)Then n=1 20 0+n 2 y1 20+n 2and1 2n=1 20+n 2So the Posterior mean is a weighted average of the Prior mean and asample mean of the vs Likelihood vs Posterior8 The Posterior mean can be thought of in two other ways n= 0+ ( y 0) 20 2n+ 20= y ( y 0) 2n 2n+ 20 The first case has nas the Prior mean adjusted towards the sampleaverage of the second case has the sample averageshrunktowards the Prior most problems, the Posterior mean can be thought of as a shrinkageestimator, where the estimate just based on the data is shrunk towardthe Prior mean. The form of the shrinkage may not be able to be writtenout in as quite a nice form for more general vs Likelihood vs Posterior9In this example the Posterior variance is never bigger than the priorvariance as1 2n=1 20+n 2 1 20and1 2n n 2 The first part of this is thought of asPosterior Precision = Prior Precision + Data PrecisionThe first inequality gives 2n 20 The second inequality gives 2n 2nPrior vs Likelihood vs Posterior10n= 5, y= 2, a= 4, b= p( |y)PosteriorLikelihoodPriorE[ ] =47 =25E[ |y] =612= ( ) =398= Var( |y) =152= ( ) = SD( |y)

5 = vs Likelihood vs Posterior11n= 20, y= 8, a= 4, b= p( |y)PosteriorLikelihoodPriorE[ ] =47 =25E[ |y] =1227= ( ) =398= Var( |y) = ( ) = SD( |y) = vs Likelihood vs Posterior12n= 100, y= 40, a= 4, b= p( |y)PosteriorLikelihoodPriorE[ ] =47 =25E[ |y] =44107= ( ) =398= Var( |y) = ( ) = SD( |y) = vs Likelihood vs Posterior13n= 1000, y= 400, a= 4, b= p( |y)PosteriorLikelihoodPriorE[ ] =47 =25E[ |y] =4041007= ( ) =398= Var( |y) = ( ) = SD( |y) = vs Likelihood vs Posterior14 PredictionAnother useful summary is the Posterior Predictive distribution of a futureobservation, yp( y|y) = p( y|y, )p( |y)d In many situations, ywill be conditionally independent ofygiven.

6 Thusthe distribution in this case reduces top( y|y) = p( y| )p( |y)d In many situations this can be difficult to calculate, though it is often easywith a conjugate example, with Binomial-Beta model, the Posterior distribution of thesuccess probability isBeta(a1, b1)(for somea1, b1). Then the distributionof the number of successes inmnew trials isp( y|y) = p( y| )p( |y)d = 10(m y) y(1 )m y (a1+b1) (a1) (b1) a1 1(1 )b1 1d =(m y) (a1+b1) (a1) (b1) (a1+ y) (b1+m y) (a1+b1+m)Which is an example of the Beta-Binomial mean this distribution isE[ y|y] =ma1a1+b1=m Prediction16 This can be gotten by applyingE[ y|y] =E[E[ y| ]|y] =E[m |y]The variance of this can be gotten byVar( y|y) = Var(E[ y| ]|y) +E[Var( y| )|y]= Var(m |y) +E[m (1 )|y]This is of the formm (1 ){1 + (m 1)}

7 2}Prediction17 One way of thinking about this is that there is two pieces of uncertainty inpredicting a new about the true success of an observation from its expected valueThis is more clear with the Normal-Normal model with fixed variance. Aswe saw earlier, the Posterior distribution is of the form |y N( n, 2n)Thenp( y|y) = 1 2 exp( 12 2( y )2)1 n 2 exp( 12 2n( n)2)d Prediction18A little bit of calculus will show that this reduces to a normal density withmeanE[ y|y] = n=E[E[ y| n]|y] =E[ |y]and varianceVar( y|y) = 2n+ 2= Var(E[ y| ]|y) +E[Var( y| )|y]= Var( n|y) +E[ 2|y]Prediction19An analogue to this is the variance for prediction in linear regression.

8 It isexactly of this formVar( y|x) = 2(1 +1n+(x x)2(n 1)s2x)Simulating the Posterior Predictive distributionThis is easy to do, assuming that you can simulate from the posteriordistribution of the parameter, which is usually do it involves two ifrom |y;i= 1, .. , yifrom y| i(= y| i, y);i= 1, .. , mThe pairs( i, yi)are draws from the joint distribution , y|y. Therefore the yiare draws from y| interest in the Posterior Predictive distribution? You might want to do predictions. For example, what will happen to astock in 6 months. Model checking: Is your model reasonable?There are a number of ways of doing this.

9 Future observations could becompared with the Posterior Predictive option might be something along the lines of cross validation. Fitthe model with part of the data and compare the remaining observationto the Posterior Predictive distribution calculated from the sample usedfor One Parameter ModelsPoissonExample: Prussian Cavalry Fatailities Due to Horse Kicks10 Prussian cavalry corp were monitored for 20 years (200 Corp-Years) andthe number of fatalities due to horse kicks were recordedx= # DeathsNumber of Corp-Years withxFatalities01091652223341 Letyi, i= 1, .. ,200be the number of deaths in thatyiiid Poisson( ).

10 (This has been shown to be a gooddescription for this data). Then the MLE for is = y=122200= can be seen fromp(y| ) =200 i=11yi! yie yie n = n ye n Instead lets take a Bayesian approach. For a Prior , lets use Gamma( , )p( ) = ( ) 1e Note that this is a conjugate Prior for .Prediction23 The Posterior density satisfiesp( |y) n ye n 1e = n y+ 1e (n+ ) which is proportional to aGamma( +n y, +n)densityThe mean and variance of aGamma( , )areE[ ] = Var( ) = 2So the Posterior mean and variance in this analysis areE[ |y] = +n y +nVar( |y) = +n y( +n)2 Similarly to before, the Posterior mean is a weighted average of the priormean and the MLE (weights andn).


Related search queries