All revision notes topics

Outliers and data cleaningAQA A-Level Maths: Revision notes

Section 1

What is an outlier?

An outlier is a value that lies unusually far from the rest of the data. There is no single universal rule, so a question will normally tell you which rule to use. Two common ones:

  • Quartile rule: a value is an outlier if it is more than 1.5×IQR1.5\times\text{IQR} above Q3Q_3 or more than 1.5×IQR1.5\times\text{IQR} below Q1Q_1.
  • Standard deviation rule: a value is an outlier if it is more than kk standard deviations (often 2 or 3) from the mean. A value that meets the rule is only a possible outlier. It might be an error, or it might be genuine.
Key termsoutlierinterquartile range
Common mistake

Using 'greater than' as 'greater than or equal to'. The rule says 'more than', so a value exactly on the boundary is not an outlier.

Section 2

Testing for outliers

Quartile rule worked example: Q1=20Q_1=20, Q3=27Q_3=27. IQR =7=7. Upper boundary 27+1.5×7=37.527+1.5\times7=37.5; lower boundary 20−10.5=9.520-10.5=9.5. An age of 71 is above 37.5, so it is an outlier, while 17 is within the boundaries. Standard deviation rule: mean 50, standard deviation 6, 2-standard-deviation rule gives boundaries 50±12=3850\pm12=38 and 6262. A reading of 70 is an outlier; 38.5 is not. Always calculate both boundaries and compare, then state your conclusion.

Exam tip

Write the two boundaries clearly, then list which values lie outside them.

Section 3

Outliers in diagrams

  • Box plot: outliers are marked individually (usually with a cross) beyond the whiskers.
  • Scatter diagram: an outlier is a point far from the general pattern of the others.
  • Histogram: an isolated bar a long way from the rest, or a very long tail.
  • Stem-and-leaf: a value separated from the others by a gap. When you interpret a diagram, say what the outlier is and suggest a possible cause in context, such as a measurement error, a different group, or a rare but genuine event.
Key termsbox plot

Section 4

Cleaning data

Data cleaning means dealing with problems before analysis.

  • Missing data (blank cells): exclude them from calculations and say how many; do not treat blanks as 0.
  • Errors: impossible values (such as a negative spend or an age of 200) or codes like −1-1. Correct them if the true value can be found (for example from the source or an obvious typing slip), otherwise remove them.
  • Outliers: check them against the source. If they are errors, correct or remove them. If they are genuine, keep them and comment on their effect, for example by quoting results with and without the outlier. Example: a total of 25 777 over 194 entries includes three entries of −1-1 and one of 4500 (really 450). The cleaned total is 25777+3−4500+450=2173025777+3-4500+450=21730 over 191 entries, mean 113.77113.77.
Key termsdata cleaningmissing data
Common mistake

Deleting an outlier just because it is large. Remove a value only when there is a reason to think it is wrong.

Section 5

Choosing and critiquing presentation

Match the diagram to the data: histograms for continuous grouped data, bar charts for categories, scatter diagrams for two variables, box plots to compare groups. When critiquing, look for:

  • Unequal class widths shown with frequency instead of frequency density.
  • Misleading or truncated axes, or missing labels and units.
  • Outliers hidden or exaggerated, or data that has not been cleaned. Measures of average are affected by outliers: the mean is pulled towards them, while the median and IQR resist them. With an outlier present, the median is usually the better measure of a typical value.
Exam tip

Always say which measure you would choose and why, in context.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Outliers and data cleaning

  1. The daily numbers of visitors to a museum over a period have lower quartile 24 and upper quartile 36. A value is treated as an outlier if it is more than 1.5×1.5\times the interquartile range above the upper quartile or more than 1.5×1.5\times the interquartile range below the lower quartile.
    The value 63 is found to be a typing error caused by swapping two digits. State the corrected value and explain whether it is still an outlier.2 marks
  2. Ten temperature readings, in ∘^\circC, from a sensor have mean 50 and standard deviation 6. One of the readings is 70. A reading is treated as an outlier if it lies more than 2 standard deviations from the mean.
    The reading of 70 is confirmed as a sensor fault and removed. Find the mean of the remaining 9 readings.2 marks
  3. A supermarket stores the weekly grocery spend, in pounds, of 200 households in a spreadsheet. Six entries are blank, three entries are −1-1 (the till uses this to mean 'card declined') and one entry is 4500. All the other values are between 15 and 310.
    For each of the three kinds of problem entry (blank, −1-1, 4500), state how it should be dealt with before the mean spend is calculated, giving a reason.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).