Transcription of Introduction to Nonlinear Regression - ETH Z
1 Introduction to Nonlinear RegressionAndreas RuckstuhlIDP Institut f r Datenanalyse und ProzessdesignZHAW Z rcher Hochschule f r Angewandte WissenschaftenOctober 2010 Contents1. The Nonlinear Regression Model12. Methodology for Parameter Estimation53. Approximate Tests and Confidence Intervals84. More Precise Tests and Confidence Intervals135. Profile t-Plot and Profile Traces156. Parameter Transformations177. Forecasts and Calibration238. Closing Comments27A. Gauss-Newton Method28 The author thanks Werner Stahel for his valuable comments. E-Mail Address: Internet: The Nonlinear Regression Model1 GoalsThenonlinear Regression modelblock in the Weiterbildungslehrgang (WBL) in ange-wandter Statistik at the ETH Zurich should1.
2 Introduce problems that are relevant to the fitting of Nonlinear Regression func-tions,2. present graphical representations for assessing the quality of approximate confi-dence intervals, and3. introduce some parts of the statistics softwareRthat can help with solvingconcrete The Nonlinear Regression ModelaThe Regression studies the relationship between avariable ofinterestYand one or moreexplanatory or predictor variablesx(j). The generalmodel isYi=hhx(1)i, x(2)i, .. , x(m)i; 1, 2, .. , pi+ ,his an appropriate function that depends on the explanatory variables andparameters, that we want to summarize with vectorsx= [x(1)i, x(2)i, .. , x(m)i]Tand = [ 1, 2, .. , p]T. The unstructured deviations from the functionhare describedvia the random errorsEi.
3 The normal distribution is assumed for the distribution ofthis random error, soEi ND0, 2E, The Linear Regression (multiple) linear Regression , functionshare con-sidered that are linear in the parameters j,hhx(1)i, x(2)i, .. , x(m)i; 1, 2, .. , pi= 1ex(1)i+ 2ex(2)i+..+ pex(p)i,where theex(j)can be arbitrary functions of the original explanatory variablesx(j).(Here the parameters are usually denoted as jinstead of j.)c The Nonlinear Regression ModelIn Nonlinear Regression , functionshare consideredthat can not be written as linear in the parameters. Often such a function is derivedfrom theory. In principle, there are unlimited possibilities for describing the determin-istic part of the model. As we will see, this flexibility oftenmeans a greater effort tomake statistical d speed with which an enzymatic reaction occurs depends ontheconcentration of a substrate.
4 According to the informationfrom Bates and Watts(1988), it was examined how a treatment of the enzyme with an additional substancecalled Puromycin influences this reaction speed. The initial speed of the reaction ischosen as the variable of interest, which is measured via radioactivity. (The unit of thevariable of interest is count/min2; the number of registrations on a Geiger counter pertime period measures the quantity of the substance present,and the reaction speed isproportional to the change per time unit.)A. Ruckstuhl, ZHAW21. The Nonlinear Regression :Puromycin Example. (a) Data ( treated enzyme; untreated enzyme) and (b)typical course of the Regression relationship of the variable of interest with the substrate concentrationx(in ppm)is described via the Michaelis-Menten functionhhx; i= 1x 2+ infinitely large substrate concentration (x ) results in the asymptotic speed 1.
5 It has been suggested that this variable is influenced by theaddition of experiment is therefore carried out once with the enzymetreated with Puromycinand once with the untreated enzyme. Figure shows the result. In this section thedata of the treated enzyme is e Oxygen determine the biochemical oxygen consumption , river watersamples were enriched with dissolved organic nutrients, with inorganic materials, andwith dissolved oxygen, and were bottled in different bottles. (Marske, 1967, see Batesand Watts (1988)). Each bottle was then inoculated with a mixed culture of microor-ganisms and then sealed in a climate chamber with constant temperature. The bottleswere periodically opened and their dissolved oxygen content was analyzed.
6 From thisthe biochemical oxygen consumption [mg/l] was model used to con-nect the cumulative biochemical oxygen consumptionYwith the incubation timex,is based on exponential growth decay, which leads tohhx, i= 1 1 e 2x . Figure shows the data and the Regression function to be f From Membrane Separation Technology(Rapold-Nydegger (1994)). The ratio ofprotonated to deprotonated carboxyl groups in the pores of cellulose membranes isdependent on the pH valuexof the outer solution. The protonation of the carboxylcarbon atoms can be captured with13C-NMR. We assume that the relationship canbe written with the extended Henderson-Hasselbach Equation for polyelectrolyteslog10 1 yy 2 = 3+ 4x ,WBL Applied Statistics Nonlinear Regression1.
7 The Nonlinear Regression Model312345678101214161820 DaysOxygen DemandDaysOxygen DemandFigure :Oxygen consumption example. (a) Data and (b) typical shape of the (=pH)y (= chem. shift)(a)xy(b)Figure :Membrane Separation Technology.(a) Data and (b) a typical shape of the the unknown parameters are 1, 2and 3>0 and 4<0. Solving foryleadsto the modelYi=hhxi; i+Ei= 1+ 210 3+ 4xi1 + 10 3+ 4xi+ Regression funtionhhxi, ifor a reasonably chosen is shown in Figure nextto the A Few Further Examples of Nonlinear Regression Functions: Hill Model (Enzyme Kinetics):hhxi, i= 1x 3i/( 2+x 3i)For 3= 1 this is also known as the Michaelis-Menten Model ( ). Mitscherlich Function (Growth Analysis):hhxi, i= 1+ 2exph 3xii.
8 From kinetics (chemistry) we get the functionhhx(1)i, x(2)i; i= exph 1x(1)iexph 2/x(2) Ruckstuhl, ZHAW41. The Nonlinear Regression Model Cobbs-Douglas Production FunctionhDx(1)i, x(2)i; E= 1 x(1)i 2 x(2)i useful Regression functions are often derived from the theory of the applicationarea in question, a general overview of Nonlinear Regression functions is of limitedbenefit. A compilation of functions from publications can befound in Appendix 7 ofBates and Watts (1988).hLinearizable Regression Nonlinear Regression functions can belin-earizedthrough transformation of the variable of interest and the explanatory example, a power functionhhx; i= 1x 2can be transformed for a linear (in the parameters) functionlnhhhx; ii= lnh 1i+ 2lnhxi= 0+ 1ex ,where 0= lnh 1i, 1= 2andex= lnhxi.
9 We call the Regression functionhlin-earizable, if we can transform it into a function linear in the (unknown) parametersvia transformations of the arguments and a monotone transformation of the are some more linearizable functions (also see Daniel and Wood, 1980):hhx, i= 1/( 1+ 2exph xi) 1/hhx, i= 1+ 2exph xihhx, i= 1x/( 2+x) 1/hhx, i= 1/ 1+ 2/ 11xhhx, i= 1x 2 lnhhhx, ii= lnh 1i+ 2lnhxihhx, i= 1exph 2ghxii lnhhhx, ii= lnh 1i+ 2ghxihhx, i= exph 1x(1)exph 2/x(2)ii lnhlnhhhx, iii= lnh 1i+ lnhx(1)i 2/x(2)hhx, i= 1 x(1) 2 x(2) 3 lnhhhx, ii= lnh 1i+ 2lnhx(1)i+ 3lnhx(2) last one is the Cobbs-Douglas Model from The Statistically Complete linear Regression with the linearized regressionfunction in the referred-to example is based on the modellnhYii= 0+ 1exi+Ei,where the random errorsEiall have the same normal distribution.
10 We back transformthis model and thus getYi= 1 x 2 eEiwitheEi= exphEii. The errorseEi,i= 1, .. , nnow contribute multiplicatively andare lognormal distributed! The assumptions about the random deviations are thusnow drastically different than for a model that is based directily onh,Yi= 1 x 2+E iwith random deviationsE ithat, as usual, contribute additively and have a specificnormal Applied Statistics Nonlinear Regression2. Methodology for Parameter Estimation5A linearization of the Regression function is therefore advisable only if the assumptionsabout the random deviations can be better satisfied - in our example, if the errorsactually act multiplicatively rather than additively and are lognormal rather thannormally distributed.