All revision notes topics

Measures of dispersionEdexcel International A Level Maths: Revision notes

Section 1

Range and quartiles

A measure of dispersion (or spread) describes how widely data vary. The range is the largest value minus the smallest, which is simple but depends on two values only. The quartiles split ordered data into four equal parts. For nn values, Q1Q_1 is the n4\frac n4th value, the median Q2Q_2 is the n2\frac n2th and Q3Q_3 is the 3n4\frac{3n}{4}th. If the position is not a whole number, round up to the next whole position; if it is a whole number, take the mean of that value and the next one. The interquartile range is IQR=Q3−Q1IQR=Q_3-Q_1, the spread of the middle 50% of the data, so it is not affected by extreme values. Example: 2, 4, 4, 4, 5, 5, 7, 9 has n=8n=8, so Q1=4+42=4Q_1=\frac{4+4}{2}=4, Q3=5+72=6Q_3=\frac{5+7}{2}=6 and IQR=2IQR=2.

Key termsrangequartileinterquartile range
Common mistake

Using the range to describe consistency when there is an extreme value. The IQR or standard deviation is a better choice.

Section 2

Percentiles and interpolation for grouped data

An interpercentile range is the difference between two percentiles, such as the 10th to 90th, P90−P10P_{90}-P_{10}, which ignores the most extreme 10% at each end. For grouped continuous data, estimate any percentile by linear interpolation assuming even spread in each class: locate the class containing the required position (for example n4\frac{n}{4} for Q1Q_1), then value ≈L+position−Ff×w\approx L+\frac{\text{position}-F}{f}\times w. Example: 80 waiting times with cumulative frequencies 6, 30, 58, 74, 80 at 10, 20, 30, 40, 60 minutes. Q1Q_1 is the 20th value, in the 10 to 20 class: 10+20−624×10=15.8310+\frac{20-6}{24}\times10=15.83. Q3Q_3 is the 60th value, in the 30 to 40 class: 30+60−5816×10=31.2530+\frac{60-58}{16}\times10=31.25. So IQR≈15.4IQR\approx15.4.

Key termspercentileinterpercentile range
Exam tip

Write down the position first (n4\frac{n}{4}, n2\frac{n}{2}, 3n4\frac{3n}{4}), then find the class using cumulative frequencies.

Section 3

Variance and standard deviation

The variance measures the mean squared distance from the mean: σ2=∑(x−xˉ)2n=∑x2n−xˉ2\sigma^2=\frac{\sum(x-\bar x)^2}{n}=\frac{\sum x^2}{n}-\bar x^2, and the standard deviation is its square root, σ=variance\sigma=\sqrt{\text{variance}}. For a frequency table or grouped data (using midpoints), σ2=∑fx2∑f−xˉ2\sigma^2=\frac{\sum fx^2}{\sum f}-\bar x^2. Use the second form, which is quicker. Example: n=20n=20, ∑x=140\sum x=140, ∑x2=1160\sum x^2=1160: xˉ=7\bar x=7, variance =58−49=9=58-49=9, σ=3\sigma=3. The variance is never negative; if you get a negative value you have made an error. Standard deviation has the same units as the data, whereas variance has the units squared.

Key termsvariancestandard deviation
Common mistake

Forgetting to subtract xˉ2\bar x^2, or forgetting to take the square root when the standard deviation is asked for.

Common mistake

Dividing by n−1n-1. On this specification the divisor is nn.

Section 4

Coding and combining data

If y=ax+by=ax+b, then yˉ=axˉ+b\bar y=a\bar x+b but the standard deviation becomes σy=∣a∣σx\sigma_y=|a|\sigma_x and the variance becomes a2σx2a^2\sigma_x^2. Adding or subtracting a constant does not change the spread. Example: σx=2\sigma_x=2 and y=3x+2y=3x+2 give σy=6\sigma_y=6. To add or remove a value, update nn, ∑x\sum x and ∑x2\sum x^2 and recalculate. Adding a value equal to the mean leaves the mean unchanged but reduces the standard deviation. For example, adding 7 to the sample above gives 120921−49=8.57\frac{1209}{21}-49=8.57, a smaller variance than 9.

Key termscoding
Common mistake

Applying the added constant to the standard deviation. Only the multiplier affects spread.

Section 5

Interpreting and comparing spread

Compare data sets using a measure of location and a measure of spread, in context. A smaller standard deviation (or IQR) means the data are more consistent, a larger one means they are more spread out. Use the median and IQR when data are skewed or contain extreme values, and the mean and standard deviation otherwise. Example: two machines have target 500 ml. A has mean 501.5 and σ=1.89\sigma=1.89; B has mean 500.6 and σ=0.8\sigma=0.8. B is better as its mean is closer to the target and it is more consistent. A value far from the mean increases the standard deviation a lot, because deviations are squared.

Key termsconsistent
Exam tip

Never say a data set is better just because its mean is larger. Link the measures to what the question asks about, such as consistency or closeness to a target.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Measures of dispersion

  1. The reaction times, in tenths of a second, of eight athletes were 2, 4, 4, 4, 5, 5, 7 and 9.
    Find the interquartile range of the eight values.2 marks
  2. A sample of 20 delivery times, xx minutes, has ∑x=140\sum x=140 and ∑x2=1160\sum x^2=1160. A second sample of delivery times, from another firm, has the same mean but a standard deviation of 5 minutes.
    A 21st delivery time of 7 minutes is added to the first sample. Find the new standard deviation.2 marks
  3. The waiting times, in minutes, of 80 patients at a clinic are summarised as follows: under 10 minutes, 6 patients; 10 to under 20 minutes, 24 patients; 20 to under 30 minutes, 28 patients; 30 to under 40 minutes, 16 patients; 40 to under 60 minutes, 6 patients. Assume that times are spread evenly within each class.
    Estimate the interquartile range of the waiting times.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).