All revision notes topics

Measures of location and spreadEdexcel A-Level Maths: Revision notes

Section 1

Measures of central tendency

The mean is xˉ=∑xn\bar x=\frac{\sum x}{n}, or ∑fx∑f\frac{\sum fx}{\sum f} from a frequency table. The median is the middle value of the ordered data; the mode is the most frequent value (or modal class for grouped data). The mean uses every value but is affected by outliers; the median is not affected by extreme values; the mode is the only measure for non-numerical data. For grouped data use class midpoints, so the mean is an estimate.

Key termsmeanmedianmodemidpoint
Common mistake

Using the class boundaries rather than midpoints when estimating the mean from grouped data.

Section 2

Measures of variation

The range is the largest minus the smallest value. A percentile splits ordered data: the ppth percentile has p%p\% of values below it. The interquartile range is Q3−Q1Q_3-Q_1, and an interpercentile range such as the 10th to 90th percentile covers the middle 80%80\%. For grouped data, estimate a percentile by linear interpolation: find the class containing position p100n\frac{p}{100}n and use L+position−Fbelowf×w.L+\frac{\text{position}-F_{\text{below}}}{f}\times w. Ranges based on percentiles ignore extreme values; the range does not.

Key termsrangepercentileinterpercentile rangelinear interpolation
Exam tip

Position for the 10th percentile of n=31n=31 is 0.1×31=3.10.1\times31=3.1. Find the class containing that position, not the 3.1th value of the data.

Section 3

Variance and standard deviation

The standard deviation measures the typical distance of values from the mean. Define Sxx=∑x2−(∑x)2n.S_{xx}=\sum x^2-\frac{(\sum x)^2}{n}. Then variance =Sxxn=∑x2n−xˉ2=\frac{S_{xx}}{n}=\frac{\sum x^2}{n}-\bar x^2 and standard deviation =Sxxn=\sqrt{\frac{S_{xx}}{n}}. The version with Sxxn−1\sqrt{\frac{S_{xx}}{n-1}} is also accepted. For frequency tables, use ∑fx2\sum fx^2 and ∑f\sum f in place of ∑x2\sum x^2 and nn. Example: ∑x=240\sum x=240, ∑x2=6100\sum x^2=6100, n=10n=10 gives Sxx=340S_{xx}=340 and SD =34=5.83=\sqrt{34}=5.83.

Key termsvariancestandard deviation$S_{xx}$
Common mistake

Writing (∑x)2\left(\sum x\right)^2 as ∑x2\sum x^2. They are different: the first sums then squares; the second squares then sums.

Section 4

Grouped data

From a grouped frequency table, estimate the mean using midpoints: xˉ≈∑fx∑f\bar x\approx\frac{\sum fx}{\sum f}, and the standard deviation from ∑fx2\sum fx^2. Both are estimates because the real values within classes are unknown. Example: classes with midpoints 5,15,25,405,15,25,40 and frequencies 4,12,20,144,12,20,14 give ∑fx=1260\sum fx=1260, xˉ=25.2\bar x=25.2; ∑fx2=37700\sum fx^2=37700, variance =754−635.04=118.96=754-635.04=118.96, SD =10.9=10.9.

Key termsestimatefrequency table
Exam tip

Say 'estimate' in your answer, since midpoints are used.

Section 5

Coding

Coding simplifies data: y=x−aby=\frac{x-a}{b}, so x=by+ax=by+a. Then

  • xˉ=byˉ+a\bar x=b\bar y+a (the mean is multiplied by bb and shifted by aa),
  • standard deviation of x=b×x=b\times standard deviation of yy (shifting does not change spread),
  • variance of x=b2×x=b^2\times variance of yy. Example: y=x−40010y=\frac{x-400}{10} with yˉ=3\bar y=3 and variance 7.257.25 gives xˉ=430\bar x=430 and variance 725725.
Key termscoding
Common mistake

Multiplying the variance by bb instead of b2b^2.

Section 6

Large data set and interpretation

Pearson's large data set gives weather data. Questions may use its terminology (daily mean temperature, rainfall) but need no knowledge of the actual data. To compare two data sets, make one comment on a measure of location (mean or median) and one on a measure of spread (SD, IQR), each in context: 'a higher mean means warmer on average' and 'a smaller standard deviation means more consistent temperatures'.

Key termsmeasure of locationmeasure of spread
Exam tip

When asked to compare, mention both average and spread, and use the words of the context.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Measures of location and spread

  1. The delivery times, xx minutes, of 10 parcels have ∑x=240\sum x=240 and ∑x2=6100\sum x^2=6100.
    Calculate the standard deviation of the delivery times.2 marks
  2. A sample of 8 masses, xx grams, is coded using y=x−40010y=\frac{x-400}{10}. The coded data have ∑y=24\sum y=24 and ∑y2=130\sum y^2=130.
    A second sample is formed by adding 5 g to every mass in this sample. State the effect on (i) the mean and (ii) the standard deviation.2 marks
  3. A clinic records the waiting times of 50 patients, in minutes: 4 patients waited from 0 up to 10, 12 from 10 up to 20, 20 from 20 up to 30 and 14 from 30 up to 50.
    Estimate the mean waiting time.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).