All revision notes topics

Analysing DataAQA GCSE Maths: Revision notes

Section 1

How Do We Measure Central Tendency and Spread?

To summarise and compare data sets, use appropriate measures of central tendency and measures of spread:

Central tendencySpread
MeanRange
MedianInterquartile range (IQR)
Mode / modal classConsideration of outliers

The interquartile range (Higher) is the difference between the upper quartile and lower quartile, and is less affected by outliers than the range. These measures let you apply statistics to describe a population as a whole, rather than looking at every individual value.

Key termsmeanmedianinterquartile rangeoutlier
Exam tip

The median and interquartile range are better measures to use than the mean and range when a data set contains outliers.

Section 2

How Do We Construct and Interpret Histograms? (Higher)

Histograms are used for grouped continuous (or grouped discrete) data, especially when class intervals are unequal. Unlike a bar chart, the height of each bar represents frequency density, not frequency directly:

frequency density=frequencyclass width\text{frequency density} = \frac{\text{frequency}}{\text{class width}}

This means the area of each bar (not the height) represents the frequency for that class.

Key termshistogramfrequency density
Common mistake

Do not read frequency directly from the height of a histogram bar when class widths are unequal — you must calculate frequency = frequency density × class width.

Section 3

How Do We Use Cumulative Frequency Graphs and Box Plots? (Higher)

A cumulative frequency graph plots the running total of frequencies against the upper class boundary, and is used to estimate the median and quartiles from grouped data.

A box plot displays the minimum, lower quartile, median, upper quartile and maximum of a data set, and clearly shows the interquartile range. Comparing distributions using box plots or cumulative frequency diagrams lets you judge which of two data sets has a higher average or greater spread.

Key termscumulative frequencybox plot
Example

If Class A's box plot shows a higher median but a wider IQR than Class B's, Class A performed better on average but with more variability.

Section 4

How Do We Interpret Scatter Graphs and Correlation?

A scatter graph displays bivariate data (two related variables) as points, used to identify correlation:

  • Positive correlation: as one variable increases, so does the other
  • Negative correlation: as one variable increases, the other decreases
  • No correlation: no clear relationship
  • Correlation can also be described as weak or strong, depending on how closely the points follow a pattern

A line of best fit can be drawn to make predictions. Interpolating (predicting within the range of the data) is more reliable than extrapolating (predicting beyond the data), and correlation does not imply causation.

Key termscorrelationline of best fitinterpolationextrapolation
Common mistake

Just because two variables are correlated does not mean one causes the other — correlation does not indicate causation.

Must Know

  • Measures of central tendency: mean, median, mode/modal class; measures of spread: range and interquartile range
  • Frequency density = frequency ÷ class width; histogram bar area (not height) represents frequency
  • Cumulative frequency graphs estimate the median and quartiles from grouped data
  • Box plots show minimum, lower quartile, median, upper quartile and maximum, and are used to compare distributions
  • Scatter graphs show correlation (positive, negative, none, weak or strong) between two variables
  • Correlation does not indicate causation; interpolation is more reliable than extrapolation

That's the notes covered.

Carry on to the next subtopic.