Example: confidence

LECTURE NOTES 7 1 Stochastic convergence - stat.cmu.edu

LECTURE NOTES 71 Stochastic convergenceStochastic convergence refers to convergence of sequences of random variables. Recall fromthe last NOTES that we are concerned with two types of convergence in probability and con-vergence in distribution. convergence in probability implies convergence in distribution butthe reverse is not, in general, N(0,n 1) forn= 1,2,.. ThenYnP 0. We also have thatYn ZwhereZisdegenerate at 0, that isP(Z= 0) = 1. Also, nYn N(0,1). In fact, a stronger statementis true: nYnd=N(0,1) for suppose thatYn N(n,1).

LECTURE NOTES 7 1 Stochastic convergence Stochastic convergence refers to convergence of sequences of random variables. Recall from the last notes that we are concerned with two types of convergence in probability and con-

Tags:

  Lecture, Notes, Lecture notes, Convergence

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of LECTURE NOTES 7 1 Stochastic convergence - stat.cmu.edu

1 LECTURE NOTES 71 Stochastic convergenceStochastic convergence refers to convergence of sequences of random variables. Recall fromthe last NOTES that we are concerned with two types of convergence in probability and con-vergence in distribution. convergence in probability implies convergence in distribution butthe reverse is not, in general, N(0,n 1) forn= 1,2,.. ThenYnP 0. We also have thatYn ZwhereZisdegenerate at 0, that isP(Z= 0) = 1. Also, nYn N(0,1). In fact, a stronger statementis true: nYnd=N(0,1) for suppose thatYn N(n,1).

2 ThenYndoes not converge to say that a sequenceYnconverges toYin quadratic mean if:E(Yn Y)2 0,asn . This is once again a convergence of values of a sequence of random variables. Infact, convergence in quadratic mean = convergence in probability since by Chebyshev sinequality we know that:P(|Yn Y| ) E(Yn Y)2 2 0,asn . Usually, we are concerned with convergence in quadratic mean to a means thatE(Yn c)2 0. This implies thatYnP say that a sequenceYnconverges toYin`1if:E|Yn Y| 0,asn . convergence in quadratic mean = convergence in`1.

3 To prove this we canjust use the Cauchy-Schwarz inequality:E|Yn Y| E(Yn Y)2 0,asn .We say thatYnconverges almost surely tocifP( limn Yn=c) = is stronger than convergence in The Central Limit Theorem (CLT)LetX1,..,Xnbe with mean and variance 2. LetXn=1n ni=1 XiandZn= n( n ) .Note thatE[Zn] = 0 and Var[Zn] = 1 Znconverges in distribution to a standard Gaussian. That is,Zn ZwhereZ N(0,1).We can use the CLT to approximate probability calcuations. For example:P(a X b) =P( n(a ) Zn n(b ) ) P( n(a ) Z n(b ) )= ( n(b ) ) ( n(a ) )where is the cdf of a standard 2 LetTn= n( n )swheres2=n 1 i(Xi Xn)2.

4 ThenTn N(0,1).(The theorem is also true if we defines2= (n 1) 1 i(Xi Xn)2. We ll see later why wemight want to do this.) It then follows thatP(a X b) ( n(b )s) ( n(a )s).If we define the distance between the CDF of the average, and the CDF of a Gaussianappropriately, we can ask how far the two CDFs are for afinite sample sizen. The answerisC/ nfor some constantC. These results are typically called Berry-Esseen bounds. Theyassure us that the convergence to normality can happen quite quickly in some we average a collection of independent random vectors then they will converge in distri-bution to a multivariate thatYnconverges in distribution to a Gaussian, one can ask about functions some regularity conditions these also converge to a Gaussian, and the delta methodtells us how to compute the mean and variance of the new Gaussian.

5 In detail, IfYn N( , 2) andris a smooth function, thenr(Yn) N(r( ),(r ( ))2 2).3 OPandoPIn statistics and machine learning, we make use first, thatan=o(1) means thatan 0 asn .an=o(bn) means thatan/bn=o(1).an=O(1) means thatanis eventually bounded, that is, for all largen,|an| Cfor someC > (bn) means thatan/bn=O(1).We writean bnif bothan/bnandbn/anare eventually bounded. In computer science thisis written asan= (bn) but we prefer usingan bnsince, in statistics, often denotessomething we move on to the probabilistic versions.

6 Say thatYn=oP(1) ifYnP 0. Say thatYn=oP(an) if,Yn/an=oP(1).Say thatYn=OP(1) if, for every >0, there is aC >0 such thatP(|Yn|> C) .Say thatYn=OP(an) ifYn/an=OP(1).Let s use Hoeffding s inequality to show that sample proportions areOP(1/ n) within thethe true mean. LetY1,..,Ynbe coin flips {0,1}. Letp=P(Yi= 1). Let pn=1nn i= will show that: pn p=oP(1) and pn p=OP(1/ n).We have thatP(| pn p|> ) 2e 2n 2 0and so pn p=oP(1). Also,P( n| pn p|> C) =P(| pn p|>C n) 2e 2C2< 3if we pickClarge enough.

7 Hence, n( pn p) =OP(1) and so pn p=OP(1 n).Make sure you can prove the following:OP(1)oP(1) =oP(1)OP(1)OP(1) =OP(1)oP(1) +OP(1) =OP(1)OP(an)oP(bn) =oP(anbn)OP(an)OP(bn) =OP(anbn)4


Related search queries