All revision notes topics

Scatter diagrams, correlation and regressionAQA A-Level Maths: Revision notes

Section 1

Scatter diagrams and correlation

A scatter diagram plots paired values of two variables, one point for each item. Correlation describes the pattern:

  • Positive: as one variable increases the other tends to increase.
  • Negative: as one variable increases the other tends to decrease.
  • No correlation: no clear pattern. The closer the points lie to a straight line, the stronger the correlation. Describe both the direction and the strength, in context: 'the more hours of revision, the higher the test score tends to be'.
Key termsscatter diagramcorrelationbivariate data
Common mistake

Describing correlation as 'x causes y'. Correlation only describes the pattern in the data.

Section 2

Regression lines

A regression line is the straight line that best fits the data, written y=a+bxy=a+bx (or y=bx+ay=bx+a). Here yy is the response variable and xx the explanatory variable. You will not be asked to calculate the line in this course; you must be able to interpret it.

  • The gradient bb is the average change in yy for each increase of 1 in xx. For s=34+4.5hs=34+4.5h, each extra hour of revision is associated with 4.5 more marks on average.
  • The intercept aa is the predicted yy when x=0x=0. It is only meaningful if x=0x=0 is in the range of the data and makes sense in context. Substitute a value of xx into the line to get an estimate of yy: s=34+4.5×8=70s=34+4.5\times8=70.
Key termsregression lineexplanatory variableresponse variable
Exam tip

Always interpret the gradient with units, in context, using the words 'on average' and 'associated with'.

Section 3

Interpolation and extrapolation

Interpolation is estimating yy for a value of xx within the range of the data. It is usually reasonably reliable if the correlation is strong. Extrapolation is estimating outside the range. It is unreliable because the relationship may not continue in the same way. For s=34+4.5hs=34+4.5h with hh between 1 and 12, using h=30h=30 gives 169169, impossible for a test marked out of 100. When asked whether an estimate is reliable, state the range of the data, say whether the value is inside or outside it, and comment on the strength of the correlation.

Key termsinterpolationextrapolation
Common mistake

Using a regression line for values far outside the data and treating the answer as reliable, or ignoring impossible results.

Section 4

Distinct sections of the population

Sometimes the points form separate clusters, because the data contains different groups, for example women and men, or weekdays and weekends. A single regression line through all the points can be misleading: the overall correlation may come mainly from the difference between the groups, and the line may fit neither group well. If the clusters correspond to groups you can identify, it is better to analyse each group separately, with its own line. Comment on how the groups differ, for example by comparing gradients: a larger gradient means a steeper rise in yy per unit of xx for that group.

Key termscluster
Exam tip

When you see two clusters, say what the groups might be and suggest fitting a line to each.

Section 5

Correlation does not imply causation

Two variables can be correlated without one causing the other. Possible reasons: a third (confounding) variable drives both, the link is a coincidence, or the direction of causation is the other way round. Example: ice-cream sales and beach rescues are positively correlated because hot weather increases both. Banning ice cream would not reduce rescues. To comment on a claim of causation, say the data show an association only, suggest another factor that could explain it, and describe what extra evidence (such as data on the other factor, or a wider range of data) would help.

Key termscausationconfounding variable
Common mistake

Writing 'there is correlation, so it must be a coincidence'. A strong correlation may reflect a real link through a third variable.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Scatter diagrams, correlation and regression

  1. A teacher records, for 25 students, the hours per week hh spent revising and the test score ss out of 100. The scatter diagram shows positive correlation and the regression line of ss on hh is s=34+4.5hs=34+4.5h. The values of hh range from 1 to 12.
    A student claims that revising for 30 hours a week would give a score of 34+4.5×3034+4.5\times30. Explain why this prediction is unreliable.2 marks
  2. A scatter diagram shows the monthly sales of ice cream and the monthly number of beach rescues for 24 months in a seaside town. There is strong positive correlation between the two variables.
    A councillor proposes banning ice-cream sales to reduce the number of rescues. Comment on this proposal.2 marks
  3. A scatter diagram shows the mass (kg) against the height (cm) of 60 adults. There is strong positive correlation overall, but the points form two distinct clusters: a lower cluster of 28 points and an upper cluster of 32 points. Further inspection shows that the lower cluster is the women and the upper cluster is the men.
    Explain why a single regression line fitted to all 60 adults could be misleading.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).