Transcription of THE PROBABLE ERROROF A MEAN Introduction
1 THE PROBABLE ERROR OF A MEANBy STUDENTI ntroductionAny experiment may he regarded as forming an individual of a population of experiments which might he performed under the same conditions. A seriesof experiments is a sample drawn from this any series of experiments is only of value in so far as it enables us toform a judgment as to the statistical constants of the population to which theexperiments belong. In a greater number of cases the question finally turns onthe value of a mean, either directly, or as the mean differencebetween the the number of experiments be very large, we may have precise informationas to the value of the mean, but if our sample be small, we have two sources ofuncertainty: (1) owing to the error of random sampling themean of our seriesof experiments deviates more or less widely from the mean of the population,and (2) the sample is not sufficiently large to determine what is the law ofdistribution of individuals.
2 It is usual, however, to assume a normal distribution,because, in a very large number of cases, this gives an approximation so closethat a small sample will give no real information as to the manner in whichthe population deviates from normality: since some law of distribution musthe assumed it is better to work with a curve whose area and ordinates aretabled, and whose properties are well known. This assumption is accordinglymade in the present paper, so that its conclusions are not strictly applicable topopulations known not to be normally distributed; yet it appears PROBABLE thatthe deviation from normality must be very extreme to load to serious error. Weare concerned here solely with the first of these two sources of usual method of determining the probability that the mean of the pop-ulation lies within a given distance of the mean of the sampleis to assume anormal distribution about the mean of the sample with a standard deviationequal tos/ n, wheresis the standard deviation of the sample, and to use thetables of the probability , as we decrease the number of experiments, the value of the standarddeviation found from the sample of experiments becomes itself subject to anincreasing error, until judgments reached in this way may become routine work there are two ways of dealing with this difficulty.
3 (1) an ex-periment may he repeated many times, until such a long seriesis obtained thatthe standard deviation is determined once and for all with sufficient value can then he used for subsequent shorter series of similar experiments.(2) Where experiments are done in duplicate in the natural course of the work,the mean square of the difference between corresponding pairs is equal to thestandard deviation of the population multiplied by 2. We call thus combine1together several series of experiments for the purpose of determining the stan-dard deviation. Owing however to secular change, the value obtained is nearlyalways too low, successive experiments being positively are other experiments, however, which cannot easily be repeated veryoften; in such cases it is sometimes necessary to judge of thecertainty of theresults from a very small sample, which itself affords the only indication of thevariability.
4 Some chemical, many biological, and most agricultural and large-scale experiments belong to this class, which has hitherto been almost outsidethe range of statistical , although it is well known that the method of using the normal curveis only trustworthy when the sample is large , no one has yettold us veryclearly where the limit between large and small samplesis to be aim of the present paper is to determine the point at whichwe may usethe tables of the probability integral in judging of the significance of the meanof a series of experiments, and to furnish alternative tables for use when thenumber of experiments is too paper is divided into the following nine sections:I. The equation is determined of the curve which represents the frequency dis-tribution of standard deviations of samples drawn from a normal There is shown to be no kind of correlation between the mean and thestandard deviation of such a The equation is determined of the curve representing the frequency distri-bution of a quantityz, which is obtained by dividing the distance between themean of a sample and the mean of the population by the standarddeviation ofthe The curve found in I is The curve found in III is The two curves are compared with some actual Tables of the curves found in III are given for samples ofdifferent and IX.
5 The tables are explained and some instances are given of their 1 Samples ofnindividuals are drawn out of a population distributed normally, tofind an equation which shall represent the frequency of the standard deviationsof these the standard deviation found from a (all thesebeing measured from the mean of the population), thens2=S(x21)n S(x1)n 2=S(x21)n S(x21)n2 2S(x1x2) for all samples and dividing by the number of sampleswe get themoan value ofs2, which we will write s2: s2=n 2n n 2n2= 2(n 1)n,where 2is the second moment coefficient in the original normal distribution ofx: sincex1,x2, etc. are not correlated and the distribution is normal, productsinvolving odd powers ofx1vanish on summing, so that2S(x1x2)n2is equal to Rrepresent theRth moment coefficient of the distribution ofs2aboutthe end of the range wheres2= 0,M 1= 2(n 1) S(x21)n S(x1)n 2= S(x21)n 2 2S(x21)n S(x1)n 2+ S(x1)n 4=S(x41)n2+2S(x21x22)n2 2S(X41)n3 4S(x21x22)n3+S(x41)n4+6S(x21x22)n4+ other terms involving odd powers ofx1, etc.
6 Whichwill vanish on (x41) hasnterms, butS(x21x22) has12n(n 1), hence summing for allsamples and dividing by the number of samples, we getM 2= 4n+ 22(n 1)n 2 4n2 2 22(n 1)n2+ 4n3+ 3 22(n 1)n3= 4n3{n2 2n+ 1}+ 22n3(n 1){n2 2n+ 3}.Now since the distribution ofxis normal, 4= 3 22, henceM 2= 22(n 1)n3{3n 3 +n2 2n+ 3}= 22(n 1)(n+ 1) a similar tedious way I findM 3= 32(n 1)(n+ 1)(n+ 3)n3andM 4= 42(n 1)(n+ 1)(n+ 3)(n+ 5) law of formation of these moment coefficients appears to bea simpleone, but I have not seen my way to a general nowMRbe theRth moment coefficient ofs2about its mean, we haveM2= 22(n 1)n3{(n+ 1) (n 1)}= 2 22(n 1) 32 (n 1)(n+ 1)(n+ 3)n3 3(n 1)n.(2(n 1)n2 (n 1)3n3 = 32(n 1)n3{n2+ 4n+ 3 6n+ 6 n2+ 2n 1}= 8 32(n 1)n3,M4= 42n4 (n 1)(n+ 1)(n+ 3)(n+ 5) 32(n 1)2 12(n 1)3 (n 1)4 = 42(n 1)n4{n3+ 9n2+ 23n+ 15 32n+ 32 12n2+ 24n 12 n3+ 3n2 3n+ 1}=12 42(n 1)(n+ 3) 1=M23M32=8n 1, 2=M4M22=3(n+ 3)n 1), 2 2 3 1 6 =1n 1{6(n+ 3) 24 6(n 1)}= a curve of Prof.
7 Pearson s Type III may he expected to fit thedistribution equation referred to an origin at the zero end of the curvewill bey=Cxpe x,where = 2M2M3=4 22(n 1)n38n2 22(n 1)=n2 2andp=4 1 1 =n 12 1 =n the equation becomesy=Cxn 32e nx2 2,which will give the distribution area of this curve isCR 0xn 32e nx2 2dx=I(say). The first momentcoefficient about the end of the range will therefore beCR 0xn 12e nx2 2dxI=hC 2 2nxn 12e nx2 2ix= x=0I+CR 0n 1n 2xn 32e nx2 first part vanishes at each limit and the second is equal ton 1n 2II=n 1n we see that the higher moment coefficients will he formed bymultiplyingsuccessively byn+1n 2,n+3n 2etc., just as appeared to he the law of formationofM 2,M 3,M 4, it is PROBABLE that the curve found represents the theoretical distri-bution ofs2; so that although we have no actual proof we shall assume it todoso in what distribution ofsmay he found from this, since the frequency ofsisequal to that ofs2and all that we must do is to compress the base line ify1= (s2) be the frequency curve ofs2andy2= (s) be the frequency curve ofs,theny1d(s2) =y2ds,y2ds= 2y1sds, y2= 2Cs(s2)n 32e ns22 the distribution reduces toy2= 2 Csn 2e ns22 2e s22 2will give the frequency distribution of standarddeviations of samples ofn, taken out of a population distributed normally withstandard deviation 2.
8 The constantAmay he found by equating the area ofthe curve as follows:Area =AZ 0xn 2e nx22 2dx. LetIprepresentZ 0xpe nx22 2dx. ThenIp= 2nZ 0xp 1ddx e nx22 2 dx= 2nh xp 1e nx22 2ix= x=0+ 2n(p 1)Z 0xp 2e x22 2dx= 2n(p 1)Ip 2,since the first part vanishes at both continuing this process we findIn 2= 2n n 22(n 3)(n 5).. 2= 2n n 22(n 3)(n 5).. even or 0e nx22 2dx=r 2n ,andI1isZ 0xe nx22sigma2dx= 2ne nx22 2 x= x=0= ifnbe even,A=Area(n 3)(n 5).. 2 2n n 12,while isnbe oddA=Area(n 3)(n 5).. 2n n the equation may be writteny=N(n 3)(n 5).. 2 n 2 n 12xn 2e nx22 2(neven)ory=N(n 3)(n 5).. n 2 n 12xn 2e nx22 2(nodd)whereNas usual represents the total IITo show that there is no correlation between (a) the distance of the meanof a sample from the mean of the population and (b) the standard deviation ofa sample with normal distribution.
9 (1) Clearly positive and negative positions of the mean of the sample areequally likely, and hence there cannot be correlation between the absolute valueof the distance of the mean from the mean of the population andthe standard6deviation, but (2) there might be correlation between the square of the distanceand the square of the standard deviation. Letu2= S(x1)n 2ands2=S(x21)n S(x1)n ifm 1,M 1be the mean values ofu2andsz, we have by the preceding partM 1= 2(n 1)nandm 1= (x21)n S(x1)n 2 S(x1)n 4= S(x21)n 2+ 2S(x1x2).S(x21)n3 S(x41)n4 6S(x21x22)n4 other terms of odd order which will vanish on for all values and dividing by the number of cases we getRu2s2 u2 s2+m1M1= 4n2+ 22(n 1)n2 4n3 3 22(n 1)n3,whereRu2s2is the correlation u2 s2+ 22(n 1)n2= 22(n 1)n3{3 +n 3}= 22(n 1)
10 U2 s2= 0, or there is no correlation IIITo find the equation representing the frequency distribution of the means ofsamples ofndrawn from a normal population, the mean being expressed interms of the standard deviation of the havey=C n 1sn 2e nx22 2as the equation representing the distributionofs, the standard deviation of a sample ofn, when the samples are drawn froma normal population with standard the means of these samples ofnare distributed according to the equa-tion1y=p(n)Np(2 ) e nx22 2,and we have shown that there is no correlation betweenx, the distance of themean of the sample, ands, the standard deviation of the ,Theory of Errors of Observations, Part II, let us supposexmeasured in terms ofs, let us find the distributionofz= we havey1= (x) andy2= (z) as the equations representing thefrequency ofxand ofzrespectively, theny1dx=y2dz=y3dxs, y2= (n)sp(2 ) e ns2z22 2is the equation representing the distribution ofzfor samples ofnwith the chance thatslies betweensands+dsisRs+dssC n 1sn 2e ns22 2dsR 0C n 1sn 2e ns22 2dswhich represents theNin the above the distribu