Outliers and cleaning dataEdexcel A-Level Maths: Revision notes
Section 1
What is an outlier?
An outlier is a value that lies a long way from the rest of the data. There is no single definition, so the question will state a rule. Two common rules:
- more than above or below , so the limits are and ;
- more than standard deviations from the mean (often ), so the limits are . A value exactly on a limit is not outside it. On a box plot, outliers are plotted individually and the whiskers stop at the most extreme values that are not outliers.
Use the rule exactly as the question states it. Do not assume 1.5 IQR if the question says otherwise.
Section 2
Applying the IQR rule
Worked example: , . , so . Lower limit ; upper limit . A value of 50 is above 34, so it is an outlier. A value of 3 is above 2, so it is not. Always show the calculation of both limits, then compare each suspect value with them.
Adding to instead of IQR, or forgetting to check both ends.
Section 3
Applying the mean and standard deviation rule
For mean g and standard deviation g, 3 standard deviations is g, so the limits are and . A mass of is an outlier; is not. The mean and standard deviation are themselves affected by the outlier: an extreme value increases both, which widens the limits and can hide other unusual values. It can therefore be better to remove the outlier and recalculate. To find a mean after removing a value: new mean .
Dividing by rather than after one value is removed.
Section 4
Dealing with outliers
First find the cause.
- A recording or measurement error (for example a misplaced decimal point): correct it if the true value is known, otherwise remove it.
- A genuine but unusual value: keep it, as it is real data, but consider reporting measures not affected by it, such as the median and IQR. Outliers strongly affect the mean, standard deviation and range; they have little effect on the median and IQR. Always justify what you do with an outlier.
Say whether the outlier is an error or genuine before deciding to keep or remove it.
Section 5
Cleaning data
Cleaning data means removing or correcting values that would mislead the analysis.
- Errors: values that are impossible or implausible (such as 128 cm for a plant that is about 13 cm tall); correct from the source if possible, otherwise remove.
- Missing data: either omit the case (this reduces the sample size) or replace it with an estimate such as the mean of the others (which is not a real measurement and reduces spread), or re-collect it.
- Duplicates or inconsistent entries: remove or standardise. Record what you changed so that others can see it.
Replacing a missing value with the mean and then reporting the standard deviation as if it were fully real data.
Section 6
Selecting and critiquing presentation
Choose a diagram for the question:
- Box plot: shows median, quartiles and outliers, good for comparing groups.
- Histogram: shows shape (skew, peaks) of continuous data.
- Cumulative frequency: for medians and percentiles. Then interpret: comment on a measure of location and a measure of spread, in context. A diagram can mislead if outliers are left in, axes are cut off or classes are badly chosen.
When critiquing, say what the diagram hides: a box plot hides the shape of the distribution inside the box.
That's the notes covered.
Carry on to the next subtopic.
Exam questions on Outliers and cleaning data
- The numbers of text messages sent in one day by 40 students have minimum 3, lower quartile 14, upper quartile 22 and maximum 50. A value is an outlier if it is more than IQR above the upper quartile or below the lower quartile.The maximum value of 50 was recorded in error and should have been 15. State what should be done with this value and find the effect on the mean.2 marks
- The masses of 200 loaves from a bakery have mean 805 g and standard deviation 6 g. A loaf is classed as an outlier if its mass is more than 3 standard deviations from the mean.A loaf of mass 842 g is an outlier because the scale was faulty. Find the mean mass of the other 199 loaves.2 marks
- A student measures the heights, in cm, of 8 plants. The values entered in a spreadsheet are 12.5, 13.1, 12.8, 128, 13.4, 12.9 and 13.0, and the cell for the eighth plant was left blank.Identify the likely error in the data, explain how you would deal with it, and calculate the mean height of the plants measured, after dealing with it.3 marks
Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).