Example: quiz answers

The normal distribution assumption and other assumptions.

The normal distribution assumption and other t- tests assume that the data from the population are distributed how do we know if a population has a normal distribution ?As usual, we use the sample and use this as and estimate (sort of).If you take a sample from a normally distributed population, it makes sense that your sample will also be normally re more likely to pick items near the average, because more items will be NEAR the average!As your sample size goes up, you start picking up more extreme values (just as in a normal distribution )So how do you figure out if a population is actually normally distributed?

A lot of people don’t like “tests” of normality. In all cases you're trying to prove the H0, which is impossible! We won’t be covering them in this class, but if you ever do need to use ... QQ plots (sometimes called normal probability plots) - probably the best method for figuring out if something's normal. II. Making Q-Q plots ...

Tags:

  Tests, Distribution, Normal, Probability, Plot, Normal distribution, Normality, Normal probability plots

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of The normal distribution assumption and other assumptions.

1 The normal distribution assumption and other t- tests assume that the data from the population are distributed how do we know if a population has a normal distribution ?As usual, we use the sample and use this as and estimate (sort of).If you take a sample from a normally distributed population, it makes sense that your sample will also be normally re more likely to pick items near the average, because more items will be NEAR the average!As your sample size goes up, you start picking up more extreme values (just as in a normal distribution )So how do you figure out if a population is actually normally distributed?

2 There are two types of methods:1) statistical tests :A lot of people don t like tests of all cases you're trying to prove the H0, which is impossible!We won t be covering them in this class, but if you ever do need to use them ( , someone tells you to, even when they shouldn't) here are some that work:Shapiro-Wilks test - better for smaller test - better for larger samples.(The division is not clear cut, and there's lots of overlap).NEVER, never use a goodness of fit test to test for normality (You can literally do anything you want with this one!

3 2) graphical methods:Histograms - these will tell you very quickly when something's not normal , but it's more difficult to use these to figure out if something is plots - are similar to histograms - they'll let you know when something's not normal , but aren't so good to figure out if something's actually plots (sometimes called normal probability plots) - probably the best method for figuring out if something's Making Q-Q plots (= quantile - quantile plots)In general, we ll let the computer do this. But let's explain how it works:1) Determine your sample ) The calculate what you would expect from this sample size (more below).

4 3) plot the expected values (x-axis) vs. the actual values of your data (y-axis).4) If the points are more or less on a straight line, then your sample is probably normal . The details:Suppose you took a sample of size would you expect the smallest value to be?At the 10th percentile (= 1/10).The second smallest value should be at the 20th percentile (= 2/10) and so you could simply get the z-score for the 10th percentile and plot that vs. the smallest value in your data.[You could convert your z-score into a y-value (the 4th edition does this), but it really isn't necessary - the graph will look the same:This does, however, plot the actual vs.]

5 Expected values using the same scale, which might be a little easier to understand.]However, we can't do this. If we just plot the percentiles like this, eventually we'd get to the 100th is no z-score for 100% of the area (it's at infinity!)In other words, we can't do:in to calculate our percentiles because this will eventually give us 1 or 100%.Notice that we use i, not the actual values of our we have to tweak our percentiles a bit and we use:i 12n This way we can't get 100%, and the difference between this and the actual quantiles (percentiles) is minimal, at least as far as the plot is !

6 You still have to convert these to a z-score!Remember that you need to use i, not the actual value of your we get really confused, let's just do an example: While IQ is a suspect measure (there's a lot of debate about how well it actually works), it was designed to have a normal distribution . Let's examine this. Suppose we have a sample of IQ's for 12 people:IQ: 101 102 106 120 110 107 119 94 117 100 86 95 Let's find out if these IQ scores have a normal note that n = we calculate the z-scores we would expect for the smallest height, the second smallest height, and so the lowest IQ score we have:quantile = percentile = 1 1212= then look up the z-score for this let's try the second smallest value:quantile=2 1212= we look , you re finding the quantile (or percentile), then doing a reverse lookup for your z-value.

7 The other 10 numbers would be calculated the same way - aren t you glad the computer can do all this?Once you have all 12 z-scores you then plot everythingYou plot your smallest observed value against the 1st z-scoreYou plot the second smallest observed value against the 2nd z-scoreAnd so on, until you plot all 12 values:**(You could plot all 12 numbers against the expected IQ scores. To do this, you would convert the z-scores back to IQ scores (using y = and s = (it doesn't make much difference here that we're not using and )):lowest IQ score = ( ) + = lowest IQ score = ( ) + = so on.)

8 The plot will be identical to the plot of all 12 numbers vs. the z-scores, so we save ourselves a step. The only advantage is that here you can actually see the expected values instead of z-scores which might be a little hard to understand until you've done this a few times).**III. Interpreting Q-Q what should a q-q plot look like?If things are perfectly normal , then all the dots should line up on a perfectly straight of this a little like plotting x = y. If your data points are exactly normal , all your z-scores will match your observed data perfectly, and everything will be on a straight line).

9 In other words, the straighter, the plot of IQ's is reasonably straight and is a good example of a q-q plot that shows normally distributed are some guidelines to interpreting q-q plots:1) Don t worry about every little bump. There ll be lots of is every perfectly straight (unless the data are made up).2) Worry if you see a strong curve of some you see a backwards S (if the ends of the curve point towards the vertical) this indicates long tails which is BAD. Long tails can be very difficult to work takes a long time for the central limit theorem to work.

10 If you see a regular S that indicates short 's not so bad, and even a moderate sample size means things are probably sufficiently practice, you can tell a lot about the distribution of your data from q-q plots. Here are some examples:Caution: some software ( , Minitab) reverse the axes on q-q plotsIf the axes are reversed, then the above comments need to be reversed as well (a regular S becomes long tailed).Always make sure you look at the labels of the axes before you interpret a probability 'll get a homework exercise where you will need to interpret some of comments on the normal on how the data are not normal , the t-test might still do pretty good, though there are often more powerful tests (we'll learn one soon).


Related search queries