4.3 Measures of central tendency and dispersionIB Maths: Analysis and Approaches HL: Revision notes
Section 1
Mean, median and mode
For a list of values the mean is and for a frequency table it is . At SL the data set is treated as the whole population, so the formula booklet writes the mean as .
The median is the middle value once the data are in order: it is the th value. With an even number of values, take the mean of the two middle values. The mode is the most frequent value; a data set can have more than one mode or none.
If you know the mean of a group, you know its total: total . This is the quickest way to handle a value being added or removed.
Finding the 'middle' of an unordered list. Always put the data in order before looking for the median.
A new value that makes the mean of 9 values equal to 7 must bring the total to ; subtract the old total to find it.
Section 2
Grouped data: estimated mean and modal class
When data are grouped into classes such as , the exact values are lost. To estimate the mean, replace every value in a class by its mid-interval value (here 27.5) and use .
The modal class is the class with the highest frequency. On this course the modal class is only used with equal class widths.
The answer is an estimate because it assumes the values in each class are centred on the mid-interval value.
Using the class width or an upper boundary instead of the mid-interval value, or dividing by the number of classes instead of the total frequency.
Section 3
Measures of dispersion
Dispersion measures how spread out the data are.
- Range largest smallest: simple, but depends only on the two most extreme values.
- Interquartile range : the spread of the middle 50% of the data, so it is not affected by extreme values.
- Standard deviation : a typical distance of the values from the mean. The variance is .
On this course you find the standard deviation and variance with technology only: enter the data (with frequencies, or mid-interval values for grouped data) into your GDC and read off . The standard deviation has the same units as the data; the variance has squared units.
At SL use the population standard deviation, labelled on most calculators, not .
Squaring the variance to get the standard deviation. It is the other way round: variance , so .
Section 4
Quartiles and technology
The lower quartile and upper quartile split the ordered data into quarters. For discrete data you find them with technology. Be aware that different methods exist: some calculators exclude the median and take the median of each half, others interpolate. For 38, 39, 41, 42, 44, 45, 46, 47, 48, 51, 120, one method gives and , another gives and . IB markschemes accept values from standard methods.
Section 5
Effect of constant changes
If every value is transformed by :
- mean: (the median and mode transform the same way);
- standard deviation: , because adding shifts the data but does not change the spread;
- variance: .
Example: temperatures with mean and converted by have mean and .
Adding the constant to the standard deviation. measures distances between values, and a shift moves every value by the same amount.
Choosing measures, and must know
An extreme value pulls the mean and the standard deviation towards it but barely moves the median and IQR. For skewed data, or data with an extreme value, the median and IQR usually describe a typical value and the spread better; for roughly symmetric data the mean and standard deviation use all of the data.
Must know
- Mean ; median is the middle of the ordered data; mode is the most frequent value.
- Grouped data: use mid-interval values, and the result is an estimate. Modal class = highest frequency.
- IQR ; and by technology.
- : mean becomes , standard deviation becomes .
That's the notes covered.
Carry on to the next subtopic.