All revision notes topics

The large data setAQA A-Level Maths: Revision notes

Section 1

What the large data set is

The AQA course expects you to be familiar with real data from the large data set, published by AQA in advance of the exam. Real data are rich enough to explore data presentation and interpretation: many records (rows, one per observation) and many variables (columns). Know, for each variable, its type, its units, how it was collected, and how missing readings are shown. Variables are categorical (labels, such as fuel type) or quantitative: discrete (counts) or continuous (measured on a scale, such as temperature). The type decides the diagram: bar chart or pie chart for categorical data, histogram, box plot or cumulative frequency graph for continuous data, and scatter diagram for two quantitative variables. Exam questions refer to the data set, so use the AQA files and their guidance notes to explore it yourself beforehand.

Key termsrecordsvariablescategoricaldiscretecontinuous
Exam tip

Always note the units of each variable. A correct calculation with the wrong units loses the interpretation mark.

Section 2

Using technology

Use a spreadsheet or statistics package to explore the data: sort, filter and select a subset (for example one location, one year or one fuel type), then calculate summary statistics and draw diagrams. Useful features are sorting, filtering by a condition, and built-in functions for the mean, median, quartiles and standard deviation. You can also analyse a subset using a calculator with standard statistical functions: enter the data (or a frequency table) in statistics mode and read off nn, ∑x\sum x, ∑x2\sum x^2, the mean xˉ\bar x, the standard deviation σ\sigma, the median and the quartiles. Standard deviation is σ=∑x2n−xˉ2\sigma=\sqrt{\frac{\sum x^2}{n}-\bar x^2}.

Key termssubsetspreadsheet
Common mistake

Using the wrong standard deviation key on a calculator. Check which key gives σ\sigma (divides by nn) and which gives ss (divides by n−1n-1), and use the one your course requires.

Section 3

Cleaning the data

Real data are untidy. Before analysing, look for missing values (blank cells or a code such as −99-99), impossible values, inconsistent units and duplicates. Missing codes must be removed, not treated as numbers: including a −99-99 would drag the mean down and inflate the standard deviation. Replacing blanks by 0 has the same effect, because the 0 adds nothing to the total but counts as an observation. A value that is unusual but possible is an outlier. Do not delete it automatically: check whether it is an error, then decide, and say what you did and why.

Key termsmissing valueoutliercleaning
Exam tip

Say why a value is removed. 'It is a code for missing data, not a real reading' earns the mark.

Section 4

Outliers and summary statistics

An outlier is usually defined by a rule that the question will give you:

  • more than 1.5×1.5\times IQR above Q3Q_3 or below Q1Q_1; or
  • more than kk standard deviations from the mean. Example. For a sample with n=40n=40, ∑x=612\sum x=612 and ∑x2=9774\sum x^2=9774: xˉ=15.3\bar x=15.3 and σ=244.35−234.09=3.20\sigma=\sqrt{244.35-234.09}=3.20. Is 25.5 an outlier by the 2 standard deviation rule? The upper limit is 15.3+2(3.20)=21.715.3+2(3.20)=21.7, and 25.5>21.725.5>21.7, so yes. Compare the median and IQR (resistant to outliers) with the mean and standard deviation (affected by them). If the mean is well above the median, the distribution is positively skewed.
Key termsinterquartile rangestandard deviationskew
Common mistake

Applying two different rules and not commenting when they disagree. Say which is more reliable and why.

Section 5

Interpreting data and drawing conclusions

Interpret real data in summary or graphical form and use it to investigate a question in context. Compare two samples using one measure of location (median or mean) and one of spread (IQR or standard deviation), in the context of the question, and note differences in skew. A sample may not represent the whole data set. State the limitations: how the sample was chosen, whether the data are only from some places or times, and whether the conclusion applies more widely. Correlation in a scatter diagram does not prove that one variable causes the other.

Key termslocationspreadlimitation
Exam tip

Write the comparison in words that refer to the context, such as 'a typical day is windier at A', not just 'A is bigger'.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on The large data set

  1. A large data set records, for each new car sold in a year, the fuel type (petrol, diesel or hybrid), the number of doors, the engine size in litres and the CO₂ emissions in g/km. Some records have the CO₂ emissions blank.
    Name a suitable diagram for displaying the fuel type of the cars, and explain why a histogram would not be suitable.2 marks
  2. A student takes a random sample of 11 petrol cars from a large data set and records the CO₂ emissions in g/km: 110, 114, 118, 121, 124, 126, 129, 131, 134, 138, 205. Quartiles are the medians of the lower and upper halves of the ordered data, leaving out the median. A value is an outlier if it is more than 1.5×1.5\times IQR above the upper quartile or below the lower quartile.
    Show that 205 g/km is an outlier.2 marks
  3. A student uses a large data set of daily weather records from one station. Missing readings are shown by the value −99-99. A random sample of 40 valid days gives ∑x=612\sum x=612 and ∑x2=9774\sum x^2=9774 for the daily mean temperature xx in °C. A value is an outlier if it lies more than 2 standard deviations from the mean.
    Explain what would go wrong if the readings of −99-99 were included when calculating the mean and standard deviation, and what the student should do.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).