Transcription of Statistical uncertainty and error propagation - Aalto
1 Statistical uncertainty and error propagationMartin VermeerMarch 27, 2014 IntroductionThis lecture is about some of the basics of uncertainty : the use of variances and covari-ances to express Statistical uncertainty ( error bars , error ellipses); we are concentratingon the uncertainties of geographic co-ordinates. This makes it necessary to talk aboutgeodeticdatums alternative ways of fixing the starting point(s) used for fixing geodeticco-ordinates in a network solution. It also makes it desirable to discuss absolute (singlepoint) and relative (between pairs of points) co-ordinates and their follow up by discussing some more esoteric concepts; these sections are provided asreading material, but we will discuss them only conceptually, to get the ideas across. Thelearning objective of this is, that you will not be completely surprised if in future work,these concepts turn up; you will be somewhat prepared to read up on these subjects anduse them in your work.
2 The subjects are (in blue):1. Criterion matrices for modelling spatial variance structures2. Stochastic processes as a means of modelling time series behaviour statistically;signal and noise processes3. Some even more esoteric subjects:a) Statistical testing and its philosophical backgroundsb) Bayesian inferencec) Inter-model comparisons by information theoretic methods, theAkaikein-formation expressed in variances and .. and covariances .. propagation .. , point errors, error ellipses .. and relative variances .. point variances from a big variance matrix112 Relative location error by error .. (1) .. (2): uncertainty in surface area fromuncertain edge co-ordinates ..123Co-ordinate uncertainty , datum and is a datum? Levelling network example .. transformations.
3 Example.. is an S-transformation?..174 Modelling location uncertainty by criterion absolute and relative precision .. in ppm and inmm/ km.. spatial uncertainty behaviour by crite-rion matrices ..185 The modelling of .. processes .. function .. collocation .. and kriging .. processes ..246 Modern approaches in testing .. background of Statistical testing . inference .. theoretical methods ..3231 uncertainty expressed in variances and covariancesIn this text we discuss uncertainty as approached by physical geoscientists, which differssomewhat from approaches more commonly found in geoinformatics [Devillers and Jeansoulin, 2006, ]. Central concepts are variances and covariances the variance-covariance matrix especially of location information in the form of co-ordinates.
4 We shall elaborate in thechapters that DefinitionsWe can describe Statistical uncertainty about the value of a quantity, its random vari-ation when it is observed again and again, by , we can describethe tendency of two quantities to vary randomly in somewhat the same way, as quantity(random variate) is a method to producerealizationsof a phys-ical quantity. The number of realizations is in principle unlimited. , throwing adie produces one realization of a stochastic process defined on the discrete domain{1,2,3,4,5,6}. Throwing a coin similarly is defined on the discrete domain{0,1}, whereheads is 0, tails spatial information, stochastic quantities typically exist on acontinuous domain, ,the real numbersR. measure a distanced R. The measurements ared1,d2,d3.
5 And the stochastic quantity is calledd(underline).A stochastic quantity has one more property: aprobability(density)distibution. Whendoing a finite set of measurements, one can construct ahistogramfrom those limitE(x) + Few when the number of measurements increases, the histogram will become more andmore detailed, and in the limit become a smooth function. This theoretical limit1iscalled the stochastic quantityd sprobability density distribution function, distribution1We cannot everdeterminethis function from the observations, only obtainapproximationsto it, whichwill become better, the more observations we have at our disposal. In practice wepostulatesomeform for the distribution function, , the Gaussian or normal distribution, in which there are twofree parameters: the expectancy and the mean error or standard deviation.
6 These parametersare thenestimatedfrom our for short. It is written asp(x), wherexis an element of the domain ofd( ,in this case, a real number, a possible measurement value).Just like the total probability of all possible discrete outcomes, mi=1pi= 1, so is alsothe integral + p(x)dx= integral bap(x)dxagain describes the probability ofxlying within the interval [a,b]. If this probability is0, we say that such an outcome isimpossible; if it is 1, we say that it very common density distribution often found2when measured quantities containa large number of small, independent error contributions, is thenormalorGaussiandistribution ( bell curve ). It is the one depicted above titled theoretical limit . Thecentral axis of the curve corrsponds to the expectationE(x), the two inflection pointson the left and right slope (see Fig.)
7 3) are at locationsE(x) , where is called themean errororstandard deviationassociated with this distribution curve. The broaderthe curve, the larger the mean Variances and covariancesJust like distribution function is the theoretical idealization of histogram , we alsodefineexpectancyas the theoretical counterpart of mean/average value. If we have a setof measurementsxi,i= 1,..,n, the mean is computed asx=1nn i= this set, 1/nis the empirical probability,pi, for measurement valuexito occurif picked at random out of the set ofnmeasurements (and note that ni=1pi= 1). Sowe may writex=n i= the theoretical idealization of this is an integral:E{x}= xp(x)dx.(1)This is theexpectancy, or expected value, ofx; the centre of gravity of the distributionfunction. It is the value to which the average will tend for larger and larger numbers fact, often when the amount of data is too small to clearly establish what the distribution functionis, a normal distribution is routinely assumed, because it occurs so commonly (and is thus likely tobe a correct guess), and because it has nice mathematical (a) Small random errors(b) Large random errors(c) Correlated errorsFigure 1: Different types of random error in a stochastic variable onR2.
8 Left, smallrandom error ; middle, large random error . The picture on the right showscorrelationbetween the horizontal and vertical random variables, both defining variances and covariances is easy. Let us have a real stochastic variablexwith distributionp(x). Then itsvarianceisVar{x}=E{(x E{x})2}.It describes, to first order, the amount of spread , ordispersion, of the quantity aroundits own expected value, how much it is expected to deviate from this defined fortwostochastic variables,xandy:Cov(x,y)=E{(x E{x})(y E{y})}.It describes to what extent variablesxandy co-vary randomly, in other words, howlikely it is, whenxis bigger (or smaller) than its expected value, that then also thecorresponding realization ofywill we have the covariance, we can also define thecorrelation:Corr{x,y}=Cov{x,y} Var{x}Var{y}Correlation is covariancescaledrelative to the variances of the two stochastic variablesconsidered.
9 It is always in the range of [ 1,1], or [ 100%,+100%]. If the correlation isnegative, we say :correlation doesn t prove causation!Or more precisely, correlation doesn ttell us anything about what is the cause and what is the effect. Typically, whentwo stochastioc quantities correlate, it may be that one causes the other, that thesecond causes the first, or that both have a common third cause. The correlationas such proves nothing about this. However, sufficiently strong correlation (asestablished by Statistical testing) is accepted in science a proof of theexistenceofa causal (a) None(b) Weak(c) Strong(d) 100%(e) Anticorrelation(f) -100%Figure 2: Examples of correlations and error propagationThe expectancy operatorE[ ] islinear. This means that, ifu=ax, thenE{u}=aE{x}.
10 This follows directly from the definition:E{u}= up(u)du= + axp(x)dx=aE{x},becausep(u)du=p(x)dx=dpre fers to the same infinitesimal propagation law can be extended to variances and covariances (withv=by):Var{u}=a2 Var{x},Cov{u,v}=abCov{x,y}.(2)If a stochastic variable is a linear combination of two variables, say,u=ax+by,we getsimilarly (we also give the matric version):Var{u}=a2 Var{x}+b2 Var{y}+ 2abCov{x,y}=7=[a b][Var{x}Cov{x,y}Cov{x,y}Var{y}][ab],Cov {u,v}=aCov{x,v}+bCov{y,v}==[a b][Cov{x,v}Cov{y,v}].(3) Co-ordinates, point errors, error ellipsesIn spatial data, the quantities most often studied areco-ordinatesof points on theEarth s surface. Spatial co-ordinates are typically three-dimensional; there is howevera large body of theory and geodetic practice connected with treating location or mapco-ordinates, which are two-dimensional.