Example: biology

Solutions to Selected Exercises - University of Florida

1 categorical DATA ANALYSIS, 3rd editionSolutions to Selected ExercisesAlan AgrestiVersion August 3, 2012, Alan Agresti 2012 This file contains Solutions and hints to Solutions for some of the Exercises inCategoricalData Analysis, third edition, by Alan Agresti (John Wiley, & Sons, 2012). Thesolutionsgiven are partly those that are also available at the ~aa/cda2 many of the odd-numbered Exercises in the second editionof thebook (some of which are now even-numbered). I intend to expand the document withadditional Solutions , when I have report errors in these Solutions to the author (Department of Statistics, Univer-sity of Florida , Gainesville, Florida 32611-8545, e-mail so theycan be corrected in future revisions of this site.)

Solutions to Selected Exercises Alan Agresti Version August 3, 2012, Alan Agresti 2012 This file contains solutions and hints to solutions for some of the exercises in Categorical DataAnalysis, third edition, by Alan Agresti (John Wiley, & Sons, 2012). The solutions given are partly those that are also available at the website www.stat.ufl.edu ...

Tags:

  Solutions, Categorical, Dataanalysis, Categorical dataanalysis

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Solutions to Selected Exercises - University of Florida

1 1 categorical DATA ANALYSIS, 3rd editionSolutions to Selected ExercisesAlan AgrestiVersion August 3, 2012, Alan Agresti 2012 This file contains Solutions and hints to Solutions for some of the Exercises inCategoricalData Analysis, third edition, by Alan Agresti (John Wiley, & Sons, 2012). Thesolutionsgiven are partly those that are also available at the ~aa/cda2 many of the odd-numbered Exercises in the second editionof thebook (some of which are now even-numbered). I intend to expand the document withadditional Solutions , when I have report errors in these Solutions to the author (Department of Statistics, Univer-sity of Florida , Gainesville, Florida 32611-8545, e-mail so theycan be corrected in future revisions of this site.)

2 The authorregrets that he cannot providestudents with more detailed Solutions or with Solutions of other Exercises not in this 11. a. nominal, b. ordinal, c. interval, d. nominal, e. ordinal, f. nominal,3. varies from batch to batch, so the counts come from a mixture of binomials ratherthan a singlebin(n, ). Var(Y) =E[Var(Y| )] + Var[E(Y| )]> E[Var(Y| )] =E[n (1 )].7. a. ( ) = 20, so it is not close to = Wald statisticz= ( .5) (0)/20 = . Wald CI is (0)/20 = , or ( , ). These are not ( .5) (.5)/20 = ,P < Score CI is ( , ).

3 D. Test statistic 2(20) log(20/10) = , df= 1. The CI is (exp( ),1) =( , ). = 2(.5)20=. The chi-squared goodness-of-fit test of the null hypothesis that the binomial propor-tions equal ( , ) has expected frequencies ( , ), andX2= basedondf= 1. TheP-value is , giving moderate evidence against the The sample mean is Fitted probabilities for the truncated distribution , , , , The estimated expected frequencies are , , , , and , and the PearsonX2= withdf= 3 ( withdf= 2 if we truncateat 3 and above).

4 The fit seems With the binomial test the smallest possibleP-value, fromy= 0 ory= 5, is22(1/2)5= 1/16. Since this exceeds , it is impossible to rejectH0, and thus P(TypeI error) = 0. With the large-sample score test,y= 0 andy= 5 are the only outcomesto giveP ( , withy= 5,z= ( ) ( )/5 = andP= ).Thus, for that test, P(Type I error) =P(Y= 0) +P(Y= 5) = 1 a. No outcome can giveP .05, and hence one never WhenT= 2, midP-value = and one rejectsH0. Thus, P(Type I error) = P(T= 2) = of the two tests are and ; P(Type I error) = P(T= 2) = withboth P(Type I error) = E[P(Type I error|T)] = (5/8)( ) = Randomized tests arenot sensible for practical Var( ) = (1 )/ndecreases as moves toward 0 or 1 from a.

5 Var(Y) =n (1 ), Var(Y) =PVar(Yi)+2Pi<jCov(Yi, Yj) =n (1 )+2 (1 ) n2!> n (1 ).c. Var(Y) =E[Var(Y| )] + Var[E(Y| )] =E[n (1 )] + Var(n ) =n nE( 2) +[n2E( 2) n2 2] =n +(n2 n)[E( 2) 2] n 2=n (1 )+(n2 n)V ar( )> n (1 ).18. This is the binomial probability ofysuccesses andk 1 failures iny+k 1 trialstimes the probability of a failure at the next Using results shown in Sec. , Cov(nj, nk)/qVar(nj)Var(nk) = n j k/qn j(1 j)n k(1 k).Whenc= 2, 1= 1 2and correlation simplifies to a. For binomial,m(t) =E(etY) =Py ny ( et)y(1 )n y= (1 + et)n, som (0) =n.

6 2 log[(prob. underH0)/(prob. underHa)], so (prob. underH0)/(prob. underHa) = exp( to/2).22. a. ( ) = exp( n ) Pyi, soL( ) = n + (Pyi) log( ) andL ( ) = n+(Pyi)/ = 0 yields = (Pyi) (i)zw= ( y 0)/q y/n, (ii)zs= ( y 0)/q 0/n, (iii) 2[ n 0+ (Pyi) log( 0) +n y (Pyi) log( y)].c. (i) y z /2q y/n, (ii) all 0such that|zs| z /2, (iii) all 0such that the LR statistic 21( ).23. Conditional onn=y1+y2,y1has abin(n, ) distribution with = 1/( 1+ 2),which is underH0. The large sample score test usesz= (y1/n ) ( ) ( , u) denotes a CI for ( , the score CI), then the CI for /(1 ) = 1/ 2is[ /(1 ), u/(1 u)].

7 ( ) = (1 )/n is a concave function of , so if is random,g(E ) Eg( ) byJensen s inequality. Now is the expected value of for a distribution putting proba-bilityn/(n+z2 /2) at and probabilityz2 /2/(n+z2 /2) at 1 a. The likelihood-ratio (LR) CI is the set of 0for testingH0: = 0suchthat LR statistic = 2 log[(1 0)n/(1 )n] z2 /2, with = Solving for 0,nlog(1 0) z2 /2/2, or (1 0) exp( z2 /2/2n), or 0 1 exp( z2 /2/2n). Us-ing exp(x) = 1 +x+..for smallx, the upper bound is roughly 1 (1 ) = 22/2n= 2 Solve for (0 )/q (1 )/n= z If we form theP-value using the right tail, then midP-value = j/2 + j+1+.

8 Thus,E(midP-value) =Pj j( j/2 + j+1+ ) = (Pj j)2/2 = 1 The right-tail midP-value equalsP(T > to) + (1/2)p(to) = 1 P(T to) +(1/2)p(to) = 1 Fmid(to).29. a. The kernel of the log likelihood isL( ) =n1log 2+n2log[2 (1 )]+n3log(1 ) L/ = 2n1/ +n2/ n2/(1 ) 2n3/(1 ) = 0 and solve for .b. Find the expectation usingE(n1) =n 2, etc. Then, the asymptotic variance is theinverse information = (1 )/2n, and thus the estimatedSE=q (1 ) The estimated expected counts are [n 2, 2n (1 ),n(1 )2]. Compare these to theobserved counts (n1,n2,n3) usingX2orG2, withdf= (3 1) 1 = 1, since 1 parameteris Since 2L/ 2= (2n11/ 2) n12/ 2 n12/(1 )2 n22/(1 )2,the information is its negative expected value, which is2n 2/ 2+n (1 )/ 2+n (1 )/(1 )2+n(1 )/(1 )2,which simplifies ton(1 + )/ (1 ).

9 The asymptotic standard error is the square rootof the inverse information, orq (1 )/n(1 + ).32. c. Let =n1/n, and (1 ) =n2/n, and denote the null probabilities in the twocategories by 0and (1 0). Then,X2= (n1 n 0)2/n 0+ (n2 n(1 2))2/n(1 0)=n[( 0)2(1 0) + ((1 ) (1 0))2 0]/ 0(1 0),which equals ( 0)2/[ 0(1 0)/n] = Let X be a random variable that equals j0/ jwith probability j. By Jensen sinequality, since the negative log function is convex,E( logX) log(EX). Hence,E( logX) =P jlog( j/pj0) log[P j( j0/ j)] = log(P j0) = log(1) = 2nE( logX) IfY1is 2withdf= 1and ifY2is independent 2withdf= 2, then themgfofY1+Y2is the product of themgfs, which ism(t) = (1 2t) ( 1+ 2)/2, which is themgfofa 2withdf= 1+ The Bayes estimator is (n1+ )/(n+ + ), in which >0, >0.

10 No properprior leads to the ML estimate,n1/n. The ML estimator is the limit of Bayes estimatorsas and both converge to This happens with the improper prior, proportional to [ 1(1 1)] 1, which we getfrom the beta density by taking the improper settings = = ( |C) = 1/4. It is unclear from the wording, but presumably if tested means tested positive thenP( C|+) = 2/3. Sensitivity =P(+|C) = 1 P( |C) = 3 =P( | C) = 1 P(+| C) can t be determined from information a. Relative (i) 1= 2, so 1/ 2= (ii) 1 = a.


Related search queries