Example: barber

Statistical Methods Principles - Department of Statistics ...

Statistical MethodsPrinciplesModel CheckingDr Eleni Ramsey and Schafer The Statistical Sleuth Davison Statistical Models Faraway Linear Models with R Faraway Extending the Linear model with R A Gelman, Carlin, Stern and Rubin Bayesian Data Analysis 1 AssumptionsModel inference, prediction, selection etc. usually rely on the assumptions are violated the results can be seriously , andchecking, the model assumptions is vital for anyvalid example, you have learned that in the normal linear model we assumethat the errors are N(0, 2) and that the model allthe necessary variables have been state any assumptions you makeWill the assumptions holdexactly?Probably not.

Statistical Methods Principles Model Checking Dr Eleni Matechou matechou@stats.ox.ac.uk References: F.L. Ramsey and D.W. Schafer \The Statistical Sleuth" A.C. Davison \Statistical Models" ... Understanding, and checking, the model assumptions is vital for any valid analysis.

Tags:

  Principles, Model, Methods, Statistical, Checking, Statistical methods principles, Statistical methods principles model checking

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Statistical Methods Principles - Department of Statistics ...

1 Statistical MethodsPrinciplesModel CheckingDr Eleni Ramsey and Schafer The Statistical Sleuth Davison Statistical Models Faraway Linear Models with R Faraway Extending the Linear model with R A Gelman, Carlin, Stern and Rubin Bayesian Data Analysis 1 AssumptionsModel inference, prediction, selection etc. usually rely on the assumptions are violated the results can be seriously , andchecking, the model assumptions is vital for anyvalid example, you have learned that in the normal linear model we assumethat the errors are N(0, 2) and that the model allthe necessary variables have been state any assumptions you makeWill the assumptions holdexactly?Probably not.

2 Many distributional results rely onasymptoticswhichmeans they hold for large sample if the asymptotics hold, real data will not be exactly how theassumptions is acceptable?3 We do not like to ask: Is the model true or false? since probability models in mostdata analyses will not be perfectly more relevant question is: Do the model s deficiencies have a noticeable effect on thesubstantive inferences? . Gelman et al. chapter 6 DiagnosticsHow do we check if the assumptions of the model hold?We performdiagnostic We may divide diagnostic Methods into two types. Some Methods are designed to detect single case orsmall groups of cases that do not fit the pattern of therest of the data. Outlier detection is an example of this.

3 Other Methods are designed to check the assumptions ofthe model , such as the choice and transformation of thepredictors, and those that check the stochastic part ofthe model , such as the nature of the variance about themean response .Faraway (2) section the model fit well? This can be difficult to the observed to thefittedvalues should give an certain data types, eg. contingency tables or binomial data, thereexist goodness-of-fit tests, such as the residual , using these tests can only tell you if the model fits well or notand cannot suggest ways to improve the fit, something which is possibleusing, less formal but often more revealing, diagnostic If the model fits, then replicated data generated underthe model should look similar to observed data.

4 Gelman et al. chapter 6 Influential observationsAre there observations which control/influencethe fit more than wewould like to?This could lead to erroneous results which are driven by one or a smallgroup of much do our conclusions change if these observations are removed?6 Influential points/outliers can mask other influential points/outliers,which is why leave-one-out Methods do not always spot the Yes Yes No Yes No 7 Added variable plotsAdded variable plotsreduce the higher-dimensional regression problemto a series of two-dimensional plots and show leverage and influence ofthe observations on each coefficient of the can also indicate whether a variable should be added to the model ,afterthe other variables have been , they can prove misleading when diagnosing other sorts ofproblems, such as there observations that are not fitted by the model well?

5 An observation can be outlying for one model but not for There are two ways to deal with excessively influential observations: one is to use procedures that are robust/resistant to theseobservations the other is to examine them closely to see whether they areindeed influential, why they are influential and whether theyprovide some interesting extra information about the processunder study. Ramsey and Schafer chapter of model -fittingDetailed model -fitting should be performed after the model assumptionsand influential/outlying observations have been Often unexpected discrepancies between a fitted model anddata will lead to further thought, and then to more cyclesof model -fitting, checking and interpretation, iterated until abroadly satisfactory model has been found.

6 Davison section variations of the model improve the fit?There are cases where transforming the variables leads to a better fittingmodel which complies with the include the log, square root, square transformations several transformations result in a similar fit, then the transformationwhich makes interpretation of the results more straightforward should set inlibrary(MASS) response variable is the time it took to complete the race, inminutes, and the two potential explanatory variables are the total heightgained during the route, in feet, and the distance on the map, in a linear model make sense?llllllllllllllllllllllllllllllllll l051015202530050100150200250disttimeBens of JuraLairig GhruTwo BreweriesMoffat ChaseKnock Hilllllllllllllllllllllllllllllllllllll0 2000400060008000050100150200250climbtime Bens of JuraBen NevisTwo BreweriesMoffat ChaseLairig GhruSeven HillsKnock Hill12 Added variable plotslllllllllllllllllllllllllllllllllll 10 505101520050100residuals of dist~climbresiduals of time~climbBens of JuraKnock Hilllllllllllllllllllllllllllllllllllll 4000 2000020004000 40 20020406080100residuals of climb~distresiduals of time~distBens of JuraKnock Hill13 The slopes of these two simple linear regressionsare equal to the coefficients in the multiple linearregression for the

7 Corresponding predictor observations withhi>2 (3/35)are:Bens of Jura Lairig Ghru Two Breweries Moffat observations with studentised residuals>3are:Bens of Jura Knock numberCook's distancelm(time ~ dist + climb)Cook's distanceBens of JuraKnock HillLairig Ghru14 Bens of Jura is highly influential, muchmore than Knock Hill although the latterwas further from the fitted line. Therefore,although Knock Hill is an outlier, it doesnot have the ability of Bens of Jura to pullthe line towards these observations are removed andthe model is , this is donefor demonstration purposes only and greatcare should be taken when data points areremoved from the model assumptionslllllllllllllllllllllllllllll llll50100150200 2 101234 Fitted valuesStudentised residualsTwo Brewerieslllllllllllllllllllllllllllllll ll 2 1012 2 101234 Theoretical QuantilesSample Quantiles15 Two Breweries has appeared now as a possible outlier!

8 Is this observation influential? Check it on your own.


Related search queries