All revision notes topics

Correlation and regressionIB MYP Maths Extended: Revision notes

Section 1

Scatter graphs and types of correlation

A scatter graph plots two variables for the same items, with one point for each item. Correlation describes the relationship between the variables.

  • Positive correlation: as one variable increases, the other tends to increase (points rise from left to right).
  • Negative correlation: as one variable increases, the other tends to decrease (points fall from left to right).
  • No correlation: no clear pattern. Correlation describes a pattern; it does not prove that one variable causes the other.
Key termsscatter graphcorrelation
Common mistake

Saying that a strong correlation proves one variable causes the other. It only shows that they tend to change together.

Section 2

Strength of correlation and the value of r

The correlation coefficient rr is found with technology. It is always between −1-1 and 11.

  • The sign gives the direction: r>0r>0 is positive, r<0r<0 is negative.
  • The size ∣r∣|r| gives the strength: close to 11 is strong, around 0.50.5 is moderate, close to 00 is weak or none.
  • r=1r=1 or r=−1r=-1 means all points lie exactly on a straight line. For example, r=−0.82r=-0.82 is a strong negative correlation and r=0.1r=0.1 is a very weak correlation. The strength depends on how close the points are to a line, not on how steep the line is.
Key termscorrelation coefficientstrongweak
Exam tip

Describe correlation with two words: the strength (strong, moderate, weak) and the direction (positive, negative).

Section 3

Line of best fit using technology

The line of best fit is a straight line that follows the trend of the points. Using technology (a graphic display calculator or spreadsheet) enter the data in two lists and ask for linear regression. This gives the equation y=mx+cy=mx+c and rr. Example: temperature TT and drinks sold nn give n=3.02T−43.2n=3.02T-43.2 with r=0.990r=0.990. Interpret the numbers in context. The gradient mm is the change in yy for each increase of 11 in xx: here about 3 more drinks for every extra degree. The intercept cc is the value of yy when x=0x=0, which may have no meaning in practice (here −43.2-43.2 drinks).

Key termsline of best fitregressiongradient
Exam tip

Use the variable names in the equation, such as n=3.02T−43.2n=3.02T-43.2, rather than just yy and xx.

Section 4

Interpolation and extrapolation

You can use the equation to predict a value.

  • Interpolation: predicting a value inside the range of the data you collected. This is usually reliable, especially when the correlation is strong.
  • Extrapolation: predicting a value outside the range of the data. This is usually unreliable, because the pattern may not continue. Example: data for TT from 2424 to 3636. At T=30T=30 the equation gives 47.547.5 (reliable). At T=45T=45 it gives 92.892.8 (unreliable, outside the data). Check for impossible values too: a car price of −4.6-4.6 thousand AED shows that a model has failed.
Key termsinterpolationextrapolation
Common mistake

Trusting a prediction far outside the data. Always compare the xx-value with the range of the data.

Section 5

Commenting on reliability

When you comment on a prediction, mention these things:

  1. Strength of correlation: a strong correlation (close to ±1\pm1) makes predictions more reliable.
  2. Interpolation or extrapolation: inside the data is reliable, outside is not.
  3. Sensible answer: is the predicted value possible? Prices and counts cannot be negative.
  4. Amount of data: the more points, the better the line. Write a short conclusion, for example: 'The prediction for 30 °C is reliable because it is within the data and r=0.990r=0.990.'
Key termsreliablesensible
Exam tip

Give a reason with every comment: 'unreliable because it is outside the range of the data'.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Correlation and regression

  1. A researcher in Toronto collects data from 40 teenagers and finds that the correlation coefficient rr between their daily screen time (hours) and their sleep (hours) is r=−0.82r=-0.82.
    The researcher also finds r=0.1r=0.1 between screen time and test scores. Describe this correlation and say what its scatter graph would look like.2 marks
  2. A student uses technology to investigate four sets of data. Set 1: line of best fit y=2x+3y=2x+3, points close to the line, r=0.96r=0.96. Set 2: line of best fit y=−0.5x+10y=-0.5x+10, points close to the line, r=−0.93r=-0.93. Set 3: line of best fit y=0.3x+1y=0.3x+1, points widely scattered about the line, r=0.28r=0.28. Set 4: line of best fit y=−4x+7y=-4x+7, points widely scattered about the line, r=−0.31r=-0.31.
    A sixth set has line of best fit y=−3x+20y=-3x+20 and the points are widely scattered about the line. Use the pattern to describe its correlation.2 marks
  3. A kiosk in Mumbai records the maximum daily temperature TT (∘^\circC) and the number nn of cold drinks sold on six days. The data are: (24,31)(24,31), (27,38)(27,38), (29,41)(29,41), (31,52)(31,52), (33,57)(33,57), (36,66)(36,66). Use technology.
    Find the equation of the line of best fit for nn against TT, and the value of rr. Describe the correlation.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).