All revision notes topics

Estimators, bias and the sampling distribution of the meanEdexcel International A Level Further Maths: Revision notes

Section 1

Parameters, statistics and estimators

A parameter is a fixed but usually unknown number describing a population, such as the mean μ\mu or the variance σ2\sigma^2. A statistic is a quantity calculated from a random sample that does not involve any unknown parameter, for example Xˉ=1n∑Xi\bar X=\frac{1}{n}\sum X_i. Because the sample is random, a statistic varies from sample to sample and is itself a random variable with its own distribution. An estimator is a statistic used to estimate a parameter. An estimate is the numerical value the estimator takes for one particular sample. For instance Xˉ\bar X is an estimator of μ\mu, and xˉ=12.5\bar x=12.5 is an estimate. Capital letters are used for random variables (estimators) and lower case for the observed values (estimates).

Key termsparameterstatisticestimatorestimate
Common mistake

Mixing up an estimator (random variable, Xˉ\bar X) with an estimate (a number, xˉ\bar x). Say 'estimate' when you give a value.

Section 2

Bias and unbiased estimators

For an estimator TT of a parameter θ\theta, the bias is E(T)−θE(T)-\theta. If E(T)=θE(T)=\theta, then TT is an unbiased estimator: on average, over many samples, it gives the true value. A biased estimator systematically over- or under-estimates θ\theta. To test a linear estimator, find its expected value using E(aX1+bX2)=aE(X1)+bE(X2)E(aX_1+bX_2)=aE(X_1)+bE(X_2). For observations with mean μ\mu:

  • T=X1+2X23T=\frac{X_1+2X_2}{3} has E(T)=μ+2μ3=μE(T)=\frac{\mu+2\mu}{3}=\mu, so it is unbiased;
  • T=X1+X2+X32T=\frac{X_1+X_2+X_3}{2} has E(T)=3μ2E(T)=\frac{3\mu}{2}, so it is biased, with bias μ2\frac{\mu}{2} (it overestimates). An estimator is not 'better' just because it is unbiased. The sample size matters too, through the standard error below.
Key termsbiasunbiased estimator
Exam tip

To show bias, work out E(T)E(T) and compare it with θ\theta. State clearly whether it is above or below.

Section 3

Unbiased estimates of the mean and variance

From a random sample x1,…,xnx_1,\dots,x_n of a population with mean μ\mu and variance σ2\sigma^2: xˉ=∑xn  is an unbiased estimate of μ,s2=1n−1∑(xi−xˉ)2  is an unbiased estimate of σ2.\bar x=\frac{\sum x}{n}\ \text{ is an unbiased estimate of }\mu,\qquad s^2=\frac{1}{n-1}\sum(x_i-\bar x)^2\ \text{ is an unbiased estimate of }\sigma^2. For calculation, use s2=1n−1(∑x2−nxˉ2)=1n−1(∑x2−(∑x)2n)s^2=\frac{1}{n-1}\left(\sum x^2-n\bar x^2\right)=\frac{1}{n-1}\left(\sum x^2-\frac{(\sum x)^2}{n}\right). Worked example: n=8n=8, ∑x=100\sum x=100, ∑x2=1292\sum x^2=1292. Then xˉ=12.5\bar x=12.5 and s2=1292−8(12.5)27=427=6s^2=\frac{1292-8(12.5)^2}{7}=\frac{42}{7}=6. The alternative 1n∑(xi−xˉ)2\frac1n\sum(x_i-\bar x)^2 is biased: its expected value is n−1nσ2\frac{n-1}{n}\sigma^2, so on average it underestimates σ2\sigma^2. This is why the divisor is n−1n-1. (No proofs are required.)

Key termssample varianceunbiased estimate
Common mistake

Dividing by nn instead of n−1n-1 when the question asks for an unbiased estimate of the population variance.

Section 4

The sampling distribution of the sample mean

If X1,…,XnX_1,\dots,X_n is a random sample from a population with mean μ\mu and variance σ2\sigma^2, the sample mean Xˉ\bar X has E(Xˉ)=μ,Var⁡(Xˉ)=σ2n.E(\bar X)=\mu,\qquad \operatorname{Var}(\bar X)=\frac{\sigma^2}{n}. Its standard deviation, σn\frac{\sigma}{\sqrt n}, is called the standard error of the mean. If the population is itself Normal, X∼N(μ,σ2)X\sim N(\mu,\sigma^2), then exactly Xˉ∼N(μ,σ2n).\bar X\sim N\left(\mu,\frac{\sigma^2}{n}\right). The distribution of Xˉ\bar X is called a sampling distribution. It is centred on μ\mu (so Xˉ\bar X is unbiased), and it is narrower than the distribution of a single observation. When σ\sigma is not known, the standard error is estimated by sn\frac{s}{\sqrt n}.

Key termsstandard errorsampling distribution
Common mistake

Using σ2\sigma^2 rather than σ2n\frac{\sigma^2}{n} as the variance when standardising a sample mean.

Section 5

Using the distribution of the sample mean

To find probabilities about Xˉ\bar X, state its distribution and standardise with the standard error: Z=Xˉ−μσ/n∼N(0,1).Z=\frac{\bar X-\mu}{\sigma/\sqrt n}\sim N(0,1). Worked example: apple masses are N(150,202)N(150,20^2) and n=16n=16. Then Xˉ∼N(150,25)\bar X\sim N(150,25) with standard error 55, so P(Xˉ>156)=P(Z>156−1505)=P(Z>1.2)=0.115.P(\bar X>156)=P\left(Z>\frac{156-150}{5}\right)=P(Z>1.2)=0.115. A single apple within 44 g of 150150 g has probability P(∣Z∣<0.2)=0.159P(|Z|<0.2)=0.159, but the sample mean has P(∣Z∣<0.8)=0.576P(|Z|<0.8)=0.576: the mean of a sample is much less variable. Because the standard error is σn\frac{\sigma}{\sqrt n}, it halves when the sample size is multiplied by 4, and larger samples give more reliable estimates.

Key termsstandardise
Exam tip

Write the distribution Xˉ∼N(μ,σ2/n)\bar X\sim N(\mu,\sigma^2/n) first, as the mark for it is easy to earn.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Estimators, bias and the sampling distribution of the mean

  1. A random sample of 8 observations of a variable XX gives ∑x=100\sum x=100 and ∑x2=1292\sum x^2=1292.
    Calculate an estimate of the standard error of the sample mean.2 marks
  2. X1X_1, X2X_2 and X3X_3 are independent observations from a population with mean μ\mu and variance σ2\sigma^2.
    Show that T=X1+X2+X32T=\frac{X_1+X_2+X_3}{2} is a biased estimator of μ\mu, and state the bias.2 marks
  3. The masses of apples from an orchard are Normally distributed with mean 150150 g and standard deviation 2020 g. Random samples of 1616 apples are taken, and the sample mean Xˉ\bar X is calculated for each sample.
    State the distribution of Xˉ\bar X, and find P(Xˉ>156)P(\bar X>156).3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).