Transcription of 5.5.3 Convergence in Distribution - 國立臺灣大學
1 Convergence in DistributionDefinition sequence of random variables,X1, X2, .., converges in Distribution to a random variableXiflimn FXn(x) =FX(x)at all pointsxwhereFX(x) is (Maximum of uniforms)IfX1, X2, ..are iid uniform(0,1) andX(n)= max1 i nXi, let us examine ifX(n)convergesin , we have for any >0,P(|Xn 1| ) =P(X(n) 1 )=P(Xi 1 , i= 1, .. , n) = (1 )n,which goes to 0. However, if we take =t/n, we then haveP(X(n) 1 t/n) = (1 t/n)n e t,which, upon rearranging, yieldsP(n(1 X(n)) t) 1 e t;that is, the random variablen(1 X(n)) converges in Distribution to an exponential(1) that although we talk of a sequence of random variables converging in Distribution , itis really the cdfs that converge, not the random variables.
2 In this very fundamental wayconvergence in Distribution is quite different from Convergence in probability or convergencealmost the sequence of random variables,X1, X2, .., converges in probability to a random variableX, the sequence also converges in Distribution sequence of random variables,X1, X2, .., converges in probability to a constant ifand only if the sequence also converges in Distribution to . That is, the statementP(|Xn |> ) 0 for every >0is equivalent toP(Xn x) 0 ifx < 1 ifx > .Theorem (Central limit theorem)LetX1, X2, ..be a sequence of iid random variables whose mgfs exist in a neighborhood of0 (that is,MXi(t) exists for|t|< h, for some positiveh).
3 LetEXi= and VarXi= 2>0.(Both and 2are finite since the mgf exists.) Define Xn= (1n) ni=1Xi. LetGn(x) denotethe cdf of n( Xn )/ . Then, for anyx, < x < ,limn Gn(x) = x 1 2 e y2/2dy;that is, n( Xn )/ has a limiting standard normal (Stronger form of the central limit theorem)LetX1, X2, ..be a sequence of iid random variables withEXi= and 0<VarXi= 2< . Define Xn= (1n) ni=1Xi. LetGn(x) denote the cdf of n( Xn )/ . Then, for anyx, < x < ,limn Gn(x) = x 1 2 e y2/2dy;that is, n( Xn )/ has a limiting standard normal proof is almost identical to that of Theorem , except that characteristic functionsare used instead of (Normal approximation to the negative binomial )SupposeX1.
4 , Xnare a random sample from a negative binomial (r, p) Distribution . RecallthatEX=r(1 p)p,VarX=r(1 p)p22and the central limit theorem tells us that n( X r(1 p)/p) r(1 p)/p2is approximatelyN(0,1). The approximate probability calculation are much easier than theexact calculations. For example, ifr= 10,p=12, andn= 30, an exact calculation would beP( X 11) =P(30 i=1Xi 330)=330 x=0(300 +x 1x)(12)300+x= Xis negative binomial (nr, p). The CLT gives us the approximationP( X 11) =P( 30( X 10) 20 30(11 10) 20) P(Z ) =. (Slutsky s theorem)IfXn Xin Distribution andYn a, a constant, in probability, then(a)YnXn aXin Distribution .
5 (b)Xn+Yn X+ain (Normal approximation with estimated variance)Suppose that n( Xn ) N(0,1),but the value is unknown. We knowSn in probability. By Exercise , /Sn 1in probability. Hence, Slutsky s theorem tells us n( Xn )Sn= Sn n( Xn ) N(0,1). The Delta MethodFirst, we look at one motivation example. Example (Estimating the odds)Suppose we observeX1, X2, .. , Xnindependent Bernoulli(p) random variables. The typical3parameter of interest isp, but another population isp1 p. As we would estimatepby p= iXi/n, we might consider using p1 pas an estimate ofp1 p. But what are the propertiesof this estimator?
6 How might we estimate the variance of p1 p?DefinitionIf a functiong(x) has derivatives of orderr, that is,g(r)(x) =drdxrg(x) exists, then for anyconstanta, the Taylor polynomial of orderraboutaisTr(x) =r i=0g(i)(a)i!(x a) (Taylor)Ifg(r)(a) =drdxrg(x)|x=aexists, thenlimx ag(x) Tr(x)(x a)r= we are interested in approximations, we are just going to ignore the remainder. Thereare, however, many explicit forms, one useful one beingg(x) Tr(x) = xag(r+1)(t)r!(x t) we consider the multivariate case of Taylor series. LetT1, .. , Tkbe random variableswith means 1, .. , k, and defineT= (T1, .. , Tk) and = ( 1.)
7 , k). Suppose thereis a differentiable functiong(T) (an estimator of some parameter) for which we want anapproximate estimate of variance. Defineg i( ) = tig(t)|t1= 1,..,tk= first-order Taylor series expansion ofgabout isg(t) =g( ) +k i=1g i( )(ti i) + our statistical approximation we forget about the remainder and writeg(t) g( ) +k i=1g i( )(ti i).4 Now, take expectation on both sides to getE g(T) g(theta) +k i=1g i(theta)E (Ti i) =g(theta).We can now approximate the variance ofg(T) byVar g(T) E ([g(T) g(theta)]2) E ((k i=1g i(theta)(Ti i)2)=k i=1[g i(theta)]2 Var Ti+ 2 i>jg i( )g j(theta)Cov (Ti, Tj).
8 This approximation is very useful because it gives us a variance formula for a general function,using only simple variance and (Continuation of Example )In our above notation, takeg(p) =p1 p, sog (p) =1(1 p)2andVar( p1 p) [g (p)]2 Var( p)[1(1 p)2]2p(1 p)n=pn(1 p)3,giving us an approximation for the variance of our (Approximate mean and variance)SupposeXis a random variable withE X= 6= 0. If we want to estimate a functiong( ),a first-order approximation would give usg(X) =g( ) +g ( )(X ).If we useg(X) as an estimator ofg( ), we can say that approximatelyE g(X) g( ),andVar g(X) [g ( )]2 Var (Delta method)LetYnbe a sequence of random variables that satisfies n(Yn ) N(0, 2) in a given functiongand a specific value of , suppose thatg ( ) exists and is not 0.
9 Then n[g(Yn) g( )] N(0, 2[g ( )2])in :The Taylor expansion ofg(Yn) aroundYn= isg(Yn) =g( ) +g ( )(Yn ) + remainder,where the remainder 0 asYn . SinceYn in probability it follows that theremainder 0 in probability. By applying Slutsky s theorem (a),g ( ) n(Yn ) g ( )X,whereX N(0, 2). Therefore n[g(Yn) g( )] g ( ) n(Yn ) N(0, 2[g ( )]2). ExampleSuppose now that we have the mean of a random sample X. For 6= 0, we have n(1 X 1 ) N(0,(1 )4 Var X1).in are two extensions of the basic Delta method that we need to deal with to completeour treatment. The first concerns the possibility thatg ( ) = 0.
10 (Second-order Delta Method)LetYnbe a sequence of random variables that satisfies n(Yn ) N(0, 2) in a given functiongand a specific value of , suppose thatg ( ) = 0 andg ( ) exists andis not 0. Thenn[g(Yn) g( )] 2g ( )2 216in we consider the extension of the basic Delta method to the multivariate , .. ,Xnbe a random sample withE(Xij) = iand Cov(Xik, Xjk) = ij. For a givenfunctiongwith continuous first partial derivatives and a specific value of = ( 1, .. , p)for which 2= ij g( ) i g( ) j>0, n[g( X1, .. , Xp) g( 1, .. , p)] N(0, 2)in