All revision notes topics

4.13 Non-linear regressionIB Maths: Applications and Interpretation HL: Revision notes

Section 1

Fitting curves by least squares

In regression we fit a model to data so that the vertical distances between the points and the curve are as small as possible. The least squares regression curve minimises the sum of the squares of these distances. The distance between an observed value yiy_i and the model's value y^i\hat y_i is the residual. In examinations you may be asked about linear, quadratic, cubic, exponential, power and sine regression, and you evaluate the curve using technology: enter the data in lists, choose the regression type and read off the parameters (to 3 significant figures unless told otherwise). Always state which type of regression you used and write the full equation.

Key termsleast squaresresidualregression curve
Exam tip

Write the full equation with parameters and the type of regression, e.g. exponential regression gives N=3.50×1.38tN=3.50\times1.38^t.

Section 2

The models and their forms

  • Quadratic: y=ax2+bx+cy=ax^2+bx+c. Cubic: y=ax3+bx2+cx+dy=ax^3+bx^2+cx+d.
  • Exponential: y=abxy=ab^x (or aekxae^{kx}). If xx takes equally spaced values 0,1,2,…0,1,2,\ldots then the yy values form a geometric sequence with first term aa and common ratio bb (link to SL 1.3). A value b=1.38b=1.38 means an increase of 38%38\% per unit.
  • Power: y=axby=ax^b. A power model is natural where one quantity scales with another, such as period and distance in orbits.
  • Sine: y=asin⁡(bx+c)+dy=a\sin(bx+c)+d; amplitude ∣a∣|a|, period 2πb\frac{2\pi}{b} (with xx in radians), principal axis y=dy=d. Choose a model from the shape of the data and the context.
Key termsquadraticexponentialpowersine regressiongeometric sequence
Common mistake

Confusing exponential abxab^x with power axbax^b. In exponential models the variable is in the exponent.

Section 3

Sum of squared residuals

The sum of squared residuals, SSres=∑(yi−y^i)2SS_{res}=\sum(y_i-\hat y_i)^2, measures how closely the model fits the data. A smaller SSresSS_{res} means a closer fit. It depends on the units and the size of the data, so use it to compare models fitted to the same data. The least squares curve is the one with the smallest SSresSS_{res} of its type. Example: the data (v,d)(v,d) in the exam question has SSres=79.0SS_{res}=79.0 for the linear model and 4.674.67 for the quadratic model, so the quadratic fits much more closely.

Key termssum of squared residuals$SS_{res}$fit
Exam tip

Compare SSresSS_{res} values only for models fitted to the same data.

Section 4

The coefficient of determination R2R^2

The coefficient of determination R2R^2 is evaluated using technology. It gives the proportion of the variability in the second variable (yy) accounted for by the chosen model. For example R2=0.998R^2=0.998 means 99.8%99.8\% of the variation in braking distance is accounted for by the model. R2R^2 lies between 00 and 11; closer to 1 means a better fit. (Awareness that R2=1−SSresSStotR^2=1-\frac{SS_{res}}{SS_{tot}}, and so is 11 when SSres=0SS_{res}=0, may help understanding but is not examined.) For a linear model, R2=r2R^2=r^2, where rr is Pearson's product-moment correlation coefficient. If r=0.982r=0.982 then R2=0.964R^2=0.964.

Key termscoefficient of determination$R^2$Pearson's $r$
Common mistake

Saying R2=0.998R^2=0.998 means the model is 99.8%99.8\% likely to be correct or that predictions are 99.8%99.8\% accurate. It is the proportion of variability explained.

Section 5

Validity of models: why R2R^2 is not enough

Many factors affect the validity of a model, and R2R^2 alone is not a good way to decide between models. Think about:

  • Context: does the model make sense? A linear model for braking distance gives −15.9-15.9 m at 00 km/h, which is impossible.
  • Range: using the model outside the data range is extrapolation, which is less reliable than interpolation.
  • Complexity: a model with more parameters (a cubic compared with a quadratic) often has a higher R2R^2 without being better.
  • Residual pattern and sample size: small samples can fit by chance. In a comment, give a reason that refers to the context or the data, not just "R2R^2 is high".
Key termsextrapolationinterpolationvalidity of a model
Exam tip

In a "comment" question, give a reason and relate it to the context, e.g. "the model is used outside the range of the data, so it may not hold".

Section 6

Worked example: choosing and using a model

Data: t=0t=0 to 66 hours, N=3.5,4.8,6.7,9.2,12.7,17.5,24.2N=3.5, 4.8, 6.7, 9.2, 12.7, 17.5, 24.2 thousand. The ratios of successive values are about 1.381.38, so an exponential model is suitable.

  1. Exponential regression on the GDC: N=3.50×1.38tN=3.50\times1.38^t.
  2. Predict at t=8t=8: 3.50×1.388=46.03.50\times1.38^8=46.0 thousand (extrapolation, so treat with care).
  3. Interpret: the culture grows by about 38%38\% each hour, and the hourly values form a geometric sequence. For a periodic tide, use sine regression: with h=1.30sin⁡(0.508t−1.02)+2.60h=1.30\sin(0.508t-1.02)+2.60 the amplitude is 1.301.30, the period is 2π0.508=12.4\frac{2\pi}{0.508}=12.4 hours and the mean depth is 2.602.60.
Key termsgrowth factoramplitudeperiod
Exam tip

Round only the final answer; keep full GDC values in your working.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on 4.13 Non-linear regression

  1. The number of bacteria NN (in thousands) in a culture is counted every hour for six hours. At times t=0,1,2,3,4,5,6t=0, 1, 2, 3, 4, 5, 6 hours the counts are 3.5,4.8,6.7,9.2,12.7,17.5,24.23.5, 4.8, 6.7, 9.2, 12.7, 17.5, 24.2. Use your GDC.
    Interpret the value 1.381.38 in the model N=3.50×1.38tN=3.50\times1.38^{t}, and state the link between this model and a geometric sequence.2 marks
  2. The braking distance dd metres of a car was measured at speeds vv km/h. The results for (v,d)(v, d) were (20,6.4)(20, 6.4), (30,9.5)(30, 9.5), (40,16.9)(40, 16.9), (50,23.0)(50, 23.0), (60,34.6)(60, 34.6), (70,43.8)(70, 43.8), (80,58.7)(80, 58.7). Using your GDC, linear regression gives r=0.982r=0.982, and quadratic regression gives d=0.00940v2−0.0719v+3.88d=0.00940v^2-0.0719v+3.88 with sum of squared residuals SSres=4.67SS_{res}=4.67 and R2=0.998R^2=0.998.
    A student says: "The quadratic model has a larger R2R^2 than the linear model (0.998>0.9640.998>0.964), so it will give a reliable braking distance at 150150 km/h." Give two reasons why this conclusion is not justified.2 marks
  3. Astronomers record the mean distance dd of each planet from the Sun, in astronomical units (AU), and its orbital period TT in years: Mercury (0.387,0.241)(0.387, 0.241), Venus (0.723,0.615)(0.723, 0.615), Earth (1.000,1.000)(1.000, 1.000), Mars (1.524,1.881)(1.524, 1.881), Jupiter (5.203,11.86)(5.203, 11.86), Saturn (9.537,29.46)(9.537, 29.46). A power model T=a×d bT=a\times d^{\,b} is to be fitted. Use your GDC.
    Use your GDC to find the power regression model T=a×d bT=a\times d^{\,b}, giving aa and bb to 3 significant figures.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).