Transcription of Chapter 8 Describing Data: Measures of Central …
1 100 Part 2 / Basic Tools of Research: Sampling, Measurement, Distributions, and Descriptive StatisticsChapter 8 Describing data : Measures ofCentral tendency and DispersionIn the previous Chapter we discussed measurement and the various levels at which we can usemeasurement to describe the extent to which an individual observation possesses a particulartheoretical construct. Such a description is referred to as a datum. An example of a datum couldbe how many conversations a person initiates in a given day, or how many minutes per day a personspends watching television, or how many column inches of coverage are devoted to labor issues inThe Wall Street Journal. Multiple observations of a particular characteristic in a population or in asample are referred to as we collect a set of data , we are usually interested in making some statistical summarystatements about this large and complex set of individual values for a variable.
2 That is, we want todescribe a collective such as a sample or a population in its entirety. This description is the first stepin bridging the gap between the measurement world of our limited number of observations, andthe real world complexity. We refer to this process as Describing the distribution of a are a number of basic ways to describe collections of 8: Describing data : Measures of Central tendency and Dispersion101 Part 2 / Basic Tools of Research: Sampling, Measurement, Distributions, and Descriptive StatisticsDescribing DistributionsDescription by EnumerationOne way we can describe the distribution of a variable is by enumeration, that is, by simplylisting all the values of the variable. But if the data set or distribution contains more than just a fewcases, the list is going to be too complex to be understood or to be communicated effectively.
3 Imag-ine trying to describe the distribution of a sample of 300 observations by listing all 300 by Visual PresentationAnother alternative that is frequently used is to present the data in some visual manner, suchas with a bar chart, a histogram, a frequency polygon, or a pie chart. Figures 8-1 through 8-5 giveexamples of each of these, and the examples suggest some limitations that apply to the use of thesegraphic first limitation that can be seen in Figure 8-1 is that the data for bar charts should consist ofa relatively small number of response categories in order to make the visual presentation is, the variable should consist of only a small number of classes or categories. The variable CDPlayer Ownership is a good example of such a variable. Its two classes ( Owns a CD Player and Does not own a CD Player ) lend themselves readily to presentation via a bar 8-2 gives an example of the presentation of data in a histogram.
4 In a histogram thehorizontal axis shows the values of the variable (in this case the number of CD discs a person reportshaving purchased in the previous year) and the vertical axis shows the frequencies associated withChapter 8: Describing data : Measures of Central tendency and Dispersion102 Part 2 / Basic Tools of Research: Sampling, Measurement, Distributions, and Descriptive StatisticsChapter 8: Describing data : Measures of Central tendency and Dispersion103 Part 2 / Basic Tools of Research: Sampling, Measurement, Distributions, and Descriptive Statisticsthese values, that is, how many persons stated that they purchased, for instance, 8 histograms or bar charts, the shape of the distribution can convey a significant amount ofinformation. This is another reason why it is desirable to conduct measurement at an ordinal orinterval level, as this allows you to organize the values of a variable in some meaningful that the values on the horizontal axis of the histogram are ordered from lowest to highest, ina natural sequence of increasing levels of the theoretical concept ( Compact Disc Purchasing ).
5 Ifthe variable to be graphed is nominal, then the various classes could be arranged visually in any oneof a large number of sequences. Each of these sequences would be equally natural , since nominalcategories contain no ranking or ordering information, and each sequence would convey differentand conflicting information about the distribution of the variable. The shape of the distributionwould convey no useful information at all. Bar charts and histograms can be used to compare therelative sizes of nominal categories, but they are more useful when the data graphed are at theordinal or higher level of 8-3 gives an alternative to presenting data in a histogram. This method is called a fre-quency polygon, and it is constructed by connecting the points which have heights correspondingwith the frequencies on the vertical axis.
6 Another way of thinking of a frequency polygon is as a linewhich connects the midpoints of the tops of the bars in the that the number of response categories that can be represented in the histogram orfrequency polygon is limited. It would be very difficult to accommodate a variable with many moreclasses. If we want to describe a variable with a large number of classes using a histogram or afrequency polygon, we would have to collapse categories, that is, combine a number of previouslydistinct classes, such as the classes 0, 1, 2, etc. into a new aggregate category, such as 0 through 4, 5through 9, 10 through 14, etc. Although this process would reduce the number of categories andincrease the ease of presentation in graphical form, it also results in a loss of information.
7 For in-stance, a person who purchased 0 CDs would be lumped together with a person who purchased asmany as 4 CDs in the 0-4 class, thereby losing an important distinction between these two individu- Chapter 8: Describing data : Measures of Central tendency and Dispersion104 Part 2 / Basic Tools of Research: Sampling, Measurement, Distributions, and Descriptive Statisticsals. Figure 8-4 illustrates the results of such a reclassification or recoding of the original data fromFigure another way of presenting data visually is in the form of a pie chart. Figure shows a piechart which presents the average weekly television network ratings during prime charts are appropriate for presenting the distributions of nominal variables, since the or-der in which the values of the variable are introduced is immaterial.
8 The four classes of the variableas presented in this chart are: tuned to ABC, tuned to NBC, tuned to CBS and, finally, tuned toanything else or not turned on. There is no one way in which these levels of the variable can orshould be ordered. The sequence in which these shares are listed really does not matter. All we needto consider is the size of the slice associated with each class of the StatisticsAnother way of Describing a distribution of data is by reducing the data to some essentialindicator that, in a single value, expresses information about the aggregate of all observations. De-scriptive statistics do exactly that. They represent or epitomize some facet of a distribution. Notethat the name descriptive statistics is actually a misnomer we do not limit ourselves to descriptive statistics allow us to go beyond the mere description of a distribution.
9 Theycan also be used for statistical inference, which permits generalizing from the limited number ofobservations in a sample to the whole population. We explained in Chapter 5 that this is a majorgoal of scientific endeavors. This fact alone makes descriptive statistics preferable to either enu-meration or visual presentation. However, descriptive statistics are often used in conjunction withvisual statistics can be divided into two major categories: Measures of Central tendency ;and Measures of Dispersion or Variability. Both kinds of Measures focus on different essential char-acteristics of distributions. A very complete description of a distribution can be obtained from arelatively small set of Central tendency and dispersion Measures from the two of Central TendencyThe Measures of Central tendency describe a distribution in terms of its most frequent , typi-cal or average data value.
10 But there are different ways of representing or expressing the idea of typicality . The descriptive statistics most often used for this purpose are the Mean (the average),the Mode (the most frequently occurring score), and the Median (the middle score). Chapter 8: Describing data : Measures of Central tendency and Dispersion105 Part 2 / Basic Tools of Research: Sampling, Measurement, Distributions, and Descriptive StatisticsThe MeanThe mean is defined as the arithmetic average of a set of numerical scores, that is, the sum of allthe numbers divided by the number of observations contributing to that = = sum of all data values / number of data valuesor, more formally,which is the formula to be used when the data are in an array, which is simply a listing of a setof observations, organized by observation number.