The large data setAQA A-Level Maths: Revision notes
Section 1
What the large data set is
The AQA course expects you to be familiar with real data from the large data set, published by AQA in advance of the exam. Real data are rich enough to explore data presentation and interpretation: many records (rows, one per observation) and many variables (columns). Know, for each variable, its type, its units, how it was collected, and how missing readings are shown. Variables are categorical (labels, such as fuel type) or quantitative: discrete (counts) or continuous (measured on a scale, such as temperature). The type decides the diagram: bar chart or pie chart for categorical data, histogram, box plot or cumulative frequency graph for continuous data, and scatter diagram for two quantitative variables. Exam questions refer to the data set, so use the AQA files and their guidance notes to explore it yourself beforehand.
Always note the units of each variable. A correct calculation with the wrong units loses the interpretation mark.
Section 2
Using technology
Use a spreadsheet or statistics package to explore the data: sort, filter and select a subset (for example one location, one year or one fuel type), then calculate summary statistics and draw diagrams. Useful features are sorting, filtering by a condition, and built-in functions for the mean, median, quartiles and standard deviation. You can also analyse a subset using a calculator with standard statistical functions: enter the data (or a frequency table) in statistics mode and read off , , , the mean , the standard deviation , the median and the quartiles. Standard deviation is .
Using the wrong standard deviation key on a calculator. Check which key gives (divides by ) and which gives (divides by ), and use the one your course requires.
Section 3
Cleaning the data
Real data are untidy. Before analysing, look for missing values (blank cells or a code such as ), impossible values, inconsistent units and duplicates. Missing codes must be removed, not treated as numbers: including a would drag the mean down and inflate the standard deviation. Replacing blanks by 0 has the same effect, because the 0 adds nothing to the total but counts as an observation. A value that is unusual but possible is an outlier. Do not delete it automatically: check whether it is an error, then decide, and say what you did and why.
Say why a value is removed. 'It is a code for missing data, not a real reading' earns the mark.
Section 4
Outliers and summary statistics
An outlier is usually defined by a rule that the question will give you:
- more than IQR above or below ; or
- more than standard deviations from the mean. Example. For a sample with , and : and . Is 25.5 an outlier by the 2 standard deviation rule? The upper limit is , and , so yes. Compare the median and IQR (resistant to outliers) with the mean and standard deviation (affected by them). If the mean is well above the median, the distribution is positively skewed.
Applying two different rules and not commenting when they disagree. Say which is more reliable and why.
Section 5
Interpreting data and drawing conclusions
Interpret real data in summary or graphical form and use it to investigate a question in context. Compare two samples using one measure of location (median or mean) and one of spread (IQR or standard deviation), in the context of the question, and note differences in skew. A sample may not represent the whole data set. State the limitations: how the sample was chosen, whether the data are only from some places or times, and whether the conclusion applies more widely. Correlation in a scatter diagram does not prove that one variable causes the other.
Write the comparison in words that refer to the context, such as 'a typical day is windier at A', not just 'A is bigger'.
That's the notes covered.
Carry on to the next subtopic.
Exam questions on The large data set
- A large data set records, for each new car sold in a year, the fuel type (petrol, diesel or hybrid), the number of doors, the engine size in litres and the CO₂ emissions in g/km. Some records have the CO₂ emissions blank.Name a suitable diagram for displaying the fuel type of the cars, and explain why a histogram would not be suitable.2 marks
- A student takes a random sample of 11 petrol cars from a large data set and records the CO₂ emissions in g/km: 110, 114, 118, 121, 124, 126, 129, 131, 134, 138, 205. Quartiles are the medians of the lower and upper halves of the ordered data, leaving out the median. A value is an outlier if it is more than IQR above the upper quartile or below the lower quartile.Show that 205 g/km is an outlier.2 marks
- A student uses a large data set of daily weather records from one station. Missing readings are shown by the value . A random sample of 40 valid days gives and for the daily mean temperature in °C. A value is an outlier if it lies more than 2 standard deviations from the mean.Explain what would go wrong if the readings of were included when calculating the mean and standard deviation, and what the student should do.3 marks
Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).