All revision notes topics

4.4 Correlation and linear regressionIB Maths: Applications and Interpretation SL: Revision notes

Section 1

Scatter diagrams and correlation

Bivariate data are pairs of values (x,y)(x,y). A scatter diagram plots them as points. Correlation describes the linear association between the variables: positive (as xx increases, yy tends to increase), negative (yy tends to decrease) or zero (no linear pattern). It is strong if the points lie close to a straight line and weak if they are widely scattered; with no pattern there is no correlation. A curved pattern can be strong without being linear. Example: hours of revision against test score usually shows positive correlation; the age of a car against its value usually shows negative correlation.

Key termsbivariate datascatter diagrampositivenegativezero correlation

Section 2

Pearson's product-moment correlation coefficient

Pearson's product-moment correlation coefficient rr measures the strength and direction of linear correlation, with −1≤r≤1-1\le r\le1. r=1r=1 is perfect positive, r=−1r=-1 perfect negative and r=0r=0 no linear correlation. A common guide: ∣r∣|r| close to 11 is strong, around 0.50.5 moderate, close to 00 weak. Use your GDC to calculate rr (hand calculation can help understanding). Critical values of rr will be given where appropriate: if ∣r∣|r| is greater than the critical value for the sample size, the linear correlation is significant. Note that rr is only meaningful for linear relationships: a strong curved pattern can give an rr close to 00. Example: x=1,…,6x=1,\dots,6 and y=3,5,4,7,8,9y=3,5,4,7,8,9 give r=0.949r=0.949, strong positive linear correlation.

Key termsPearson's rlinearcritical value
Common mistake

Treating rr as a percentage or as the gradient. It is a measure of how close the points lie to a line, not how steep the line is.

Section 3

Correlation and causation

Correlation does not imply causation. Two variables may move together because of a third variable (ice-cream sales and sunburn both rise in hot weather), or by coincidence. A causal claim needs more evidence than a high rr. When asked to comment, say that the data show an association and name a plausible third variable if you can.

Key termscausationthird variable
Common mistake

Writing 'a high correlation shows that xx causes yy'. State only that they are associated.

Section 4

Line of best fit by eye

On a scatter diagram, draw a straight line of best fit by eye so that about half the points lie on each side and the line follows the trend. It should pass through the mean point (xˉ,yˉ)(\bar{x},\bar{y}). You can then read approximate values, or find its equation from two points on it. Example: for x=1,…,6x=1,\dots,6, y=3,5,4,7,8,9y=3,5,4,7,8,9 the mean point is (3.5,6)(3.5,6); any good line passes through it.

Key termsline of best fitmean point

Section 5

Regression line of y on x and its parameters

The regression line of yy on xx, y=ax+by=ax+b, is found with your GDC (linear regression); it minimises the squares of the vertical distances from the points to the line and passes through (xˉ,yˉ)(\bar{x},\bar{y}). Interpret the parameters in context: aa (gradient) is the average change in yy for each one-unit increase in xx; bb (intercept) is the predicted value of yy when x=0x=0 (which may not make sense if 00 is outside the data). Example: y=1.2x+1.8y=1.2x+1.8 for the data above. a=1.2a=1.2: each extra unit of xx goes with 1.2 more units of yy; mean point check 1.2(3.5)+1.8=61.2(3.5)+1.8=6.

Key termsregression linegradientintercept
Exam tip

State the interpretation of aa and bb using the variable names and units from the question.

Section 6

Prediction, interpolation and extrapolation

Use the regression line to predict yy for a given xx by substituting. A prediction within the range of the data is interpolation and is reasonably reliable if ∣r∣|r| is high. A prediction outside the range is extrapolation and is unreliable: the trend may not continue (a predicted negative value of a car, for example, is impossible). The line of yy on xx should only be used to predict yy from xx. Predicting xx from a given yy by rearranging the equation is not always reliable; use the regression line of xx on yy instead.

Key termsinterpolationextrapolation
Common mistake

Using the yy on xx line to find xx for a given yy. Use the xx on yy regression instead.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on 4.4 Correlation and linear regression

  1. Over eight days a café records the maximum daily temperature, T ∘T\,^\circC, and the number of cold drinks sold, nn. TT: 18, 21, 24, 26, 29, 31, 34, 36. nn: 42, 55, 61, 70, 82, 85, 96, 104. The regression line of nn on TT has equation n=aT+bn=aT+b. Use your GDC.
    Use your regression line to estimate the number of cold drinks sold on a day when the maximum temperature is 28 ∘28\,^\circC.2 marks
  2. Over 12 months, a coastal town records its monthly ice-cream sales and the number of sunburn cases treated at its clinic. Pearson's product-moment correlation coefficient for the data is r=0.91r=0.91.
    Suggest a third variable that could explain the correlation between ice-cream sales and sunburn cases.2 marks
  3. A teacher records the number of hours, xx, that eight students revised for a test and their test score, yy (out of 100). xx: 2, 4, 5, 6, 8, 9, 10, 12. yy: 45, 52, 50, 61, 63, 70, 66, 78. Use your GDC.
    Find the value of Pearson's product-moment correlation coefficient rr, and the equation of the regression line of yy on xx.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).