4.10 Regression line of x on yIB Maths: Analysis and Approaches HL: Revision notes
Section 1
Why are there two regression lines?
The regression line of on is the least squares line that minimises the squared vertical distances from the points to the line. It is built to predict from a given .
The regression line of on , written , minimises the squared horizontal distances. It is built to predict from a given .
The two lines are different unless the correlation is perfect (). The weaker the correlation, the further apart they are. So an on line cannot reliably predict from , and a on line cannot reliably predict from .
Rearranging to make the subject and calling it the on line. It is a different line unless .
Ask 'which variable am I predicting?' The variable you want goes on the left-hand side of the line you use.
Section 2
Finding the x on y line with technology
On a GDC, enter the data in two lists and run linear regression with the lists swapped: as the independent (explanatory) list and as the dependent list. The calculator returns an equation in its own letters. Rewrite it as .
For example, for sleep hours and reaction time ms, the calculator gives , while the on line is . Give coefficients to 3 significant figures, but keep full accuracy in the calculator for predictions.
Reporting the equation as '' because the calculator uses for its output. The line must be written as .
Data : and .
Section 3
Both lines pass through the mean point
Every least squares regression line passes through the mean point . So:
- the on and on lines intersect at the mean point, and solving them simultaneously gives and ;
- if you know one mean and one line, substitute to find the other mean;
- an unknown data value can be found from a given mean, e.g. gives .
With and : , so , giving and . The mean point is .
To verify that a line passes through the mean point, substitute one mean and show that you get the other. Keep full calculator accuracy so the check comes out exactly.
Section 4
Prediction and reliability
A prediction from a regression line is more reliable when:
- the correct line is used ( on to predict );
- the correlation is strong ( close to 1);
- the given value lies within the range of the data (interpolation).
Predicting outside the data range is extrapolation. The linear pattern may not continue there, so the estimate is unreliable. For example, the reaction times in the data run from 248 ms to 305 ms, so estimating sleep for a reaction time of 340 ms is extrapolation. In an exam comment, refer to both the value of and the range of the data.
Checking the range of the wrong variable. For the on line, check that the given lies within the range of the -data.
A strong correlation does not rescue an extrapolation. Both conditions matter.
Must know
- on predicts from ; on predicts from . Never swap them.
- Find on the GDC with as the independent list; do not rearrange the on line.
- Both lines pass through , so they intersect there.
- The lines are the same only when .
- Predictions are reliable only with strong correlation and interpolation.
That's the notes covered.
Carry on to the next subtopic.