All revision notes topics

4.18 Hypothesis testing: means, proportions, correlation and errorsIB Maths: Applications and Interpretation HL: Revision notes

Section 1

The language of a test

A hypothesis test decides whether sample data give enough evidence against a claim. The null hypothesis H0\mathrm{H}_0 is the claim to be tested (it states a parameter equals a value, for example μ=500\mu=500 or p=0.3p=0.3). The alternative hypothesis H1\mathrm{H}_1 says what we suspect: ≠\ne (two-tailed) or << / >> (one-tailed). Hypotheses are about the population parameter, never the sample statistic. The significance level α\alpha (for example 5%5\%) is the probability of rejecting H0\mathrm{H}_0 when it is true. The pp-value is the probability, assuming H0\mathrm{H}_0 is true, of a result at least as extreme as the one observed. If p<αp<\alpha, reject H0\mathrm{H}_0; otherwise do not reject. For a two-tailed test compare pp with α\alpha after doubling the tail probability (or compare each tail with α2\frac{\alpha}{2}). The critical region is the set of values of the test statistic that lead to rejecting H0\mathrm{H}_0; its boundary is the critical value. Conclusions must be in context and cautious: 'sufficient evidence to reject H0\mathrm{H}_0' or 'insufficient evidence to reject H0\mathrm{H}_0', never 'proved'.

Key termsnull hypothesisalternative hypothesissignificance levelp-valuecritical region
Common mistake

Writing hypotheses about xˉ\bar{x} or the sample. They are about the population parameter (μ\mu, pp, λ\lambda, ρ\rho).

Common mistake

Concluding that the null hypothesis is 'true' when p>αp>\alpha. You only have insufficient evidence to reject it.

Section 2

Tests for a population mean

σ\sigma known (normal population, or large nn): under H0\mathrm{H}_0, Xˉ∼N(μ0,σ2n)\bar{X}\sim N\left(\mu_0,\frac{\sigma^2}{n}\right), so use the zz-test; the GDC gives the pp-value. Example: μ0=500\mu_0=500, σ=8\sigma=8, n=36n=36, xˉ=497.4\bar{x}=497.4, H1:μ<500\mathrm{H}_1:\mu<500: z=497.4−5008/6=−1.95z=\frac{497.4-500}{8/6}=-1.95 and p=0.0256<0.05p=0.0256<0.05, so reject H0\mathrm{H}_0. σ\sigma unknown: use the one-sample tt-test with n−1n-1 degrees of freedom, whatever the sample size. The GDC tt-test takes the data (or xˉ\bar{x}, sn−1s_{n-1}, nn) and gives tt and the pp-value; you are not expected to calculate critical regions for tt-tests. Paired data: samples may be paired (matched pairs). Take the differences d=x1−x2d=x_1-x_2 and treat them as a single sample, testing H0:μd=0\mathrm{H}_0:\mu_d=0. Unpaired samples are two independent samples.

Key terms$z$-test$t$-testmatched pairs
Exam tip

Say which direction the differences are taken (before minus after) so the alternative hypothesis has the right sign.

Section 3

Tests for a proportion and for a Poisson mean

Proportion (binomial): H0:p=p0\mathrm{H}_0:p=p_0. With nn trials, X∼B(n,p0)X\sim\mathrm{B}(n,p_0) under H0\mathrm{H}_0. For H1:p>p0\mathrm{H}_1:p>p_0 the pp-value for observing xx is P(X≥x)=1−P(X≤x−1)\mathrm{P}(X\ge x)=1-\mathrm{P}(X\le x-1). Example: n=25n=25, p0=0.3p_0=0.3, x=12x=12: p=0.0442<0.05p=0.0442<0.05, so reject H0\mathrm{H}_0. Poisson mean: H0:λ=λ0\mathrm{H}_0:\lambda=\lambda_0, and X∼Po(λ0)X\sim\mathrm{Po}(\lambda_0) under H0\mathrm{H}_0. Rescale λ\lambda to the period observed. Binomial and Poisson tests are one-tailed only in IB. The critical region for a discrete variable is found by testing values: choose the region with the largest probability that is still less than the significance level. Example: X∼B(20,0.7)X\sim\mathrm{B}(20,0.7), H1:p<0.7\mathrm{H}_1:p<0.7: P(X≤10)=0.0480<0.05\mathrm{P}(X\le10)=0.0480<0.05 but P(X≤11)=0.113>0.05\mathrm{P}(X\le11)=0.113>0.05, so the critical region is X≤10X\le10.

Key termsone-tailedcritical value
Common mistake

For H1:p>p0\mathrm{H}_1:p>p_0 using P(X>x)\mathrm{P}(X>x) instead of P(X≥x)\mathrm{P}(X\ge x).

Section 4

Type I and Type II errors

A Type I error is rejecting H0\mathrm{H}_0 when it is true; P(Type I)\mathrm{P}(\text{Type I}) is the probability that the test statistic lies in the critical region under H0\mathrm{H}_0. For continuous tests this equals α\alpha; for discrete tests it is the actual probability of the critical region, which may be below α\alpha (for example 0.04800.0480). A Type II error is not rejecting H0\mathrm{H}_0 when it is false. Its probability is found under a specific true value of the parameter: P(test statistic not in the critical region∣true value)\mathrm{P}(\text{test statistic not in the critical region}\mid\text{true value}). Examples: Normal, σ=8\sigma=8, n=36n=36, H0:μ=500\mathrm{H}_0:\mu=500, H1:μ<500\mathrm{H}_1:\mu<500 at 5%5\%: critical region xˉ<497.8\bar{x}<497.8. If the true μ=497\mu=497, P(Type II)=P(Xˉ≥497.8∣μ=497)=0.273\mathrm{P}(\text{Type II})=\mathrm{P}(\bar{X}\ge497.8\mid\mu=497)=0.273. Binomial: B(20,0.7)\mathrm{B}(20,0.7), critical region X≤10X\le10; if p=0.5p=0.5, P(Type II)=P(X≥11)=0.412\mathrm{P}(\text{Type II})=\mathrm{P}(X\ge11)=0.412. Poisson: Po(6)\mathrm{Po}(6), critical region X≥11X\ge11; if λ=9\lambda=9, P(Type II)=P(X≤10)=0.706\mathrm{P}(\text{Type II})=\mathrm{P}(X\le10)=0.706.

Key termsType I errorType II error
Exam tip

A Type I error uses the null value of the parameter; a Type II error uses the true (alternative) value given in the question.

Section 5

Testing for correlation

For bivariate data from a bivariate normal population, the sample product-moment correlation coefficient rr estimates the population coefficient ρ\rho. To test whether there is a linear correlation: H0:ρ=0\mathrm{H}_0:\rho=0 against H1:ρ≠0\mathrm{H}_1:\rho\ne0 (or ρ>0\rho>0, ρ<0\rho<0). Use technology; in examinations the data (or the output) will be given. If the pp-value is less than the significance level, reject H0\mathrm{H}_0: there is sufficient evidence of a linear correlation. Example: r=0.78r=0.78 with p=0.004p=0.004 at 5%5\% gives evidence that ρ≠0\rho\ne0. Correlation does not prove causation.

Key termsproduct-moment correlation coefficient
Common mistake

Saying a significant correlation means one variable causes the other.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on 4.18 Hypothesis testing: means, proportions, correlation and errors

  1. A machine is set to fill bags of sugar with a mean mass of 500500 g. The mass of a bag is normally distributed with a known standard deviation of 88 g. An inspector suspects that the machine is under-filling the bags. She takes a random sample of 3636 bags, which has a mean mass of 497.4497.4 g, and tests the claim at the 5%5\% significance level.
    State the conclusion of the test, giving a reason for your answer.2 marks
  2. A trainer records the 100100 m sprint times of 88 athletes before and after a training programme. The differences in time, in seconds, calculated as before minus after, are 0.12, 0.05, −0.03, 0.20, 0.08, 0.15, 0.02, 0.100.12,\ 0.05,\ -0.03,\ 0.20,\ 0.08,\ 0.15,\ 0.02,\ 0.10. The differences are normally distributed. The trainer tests, at the 5%5\% significance level, whether the programme reduces the mean sprint time. Let μd\mu_d be the population mean difference. Use your GDC where helpful.
    State the conclusion of the test in context, giving a reason for your answer.2 marks
  3. A manufacturer claims that 30%30\% of its customers choose the premium model of a phone. A retailer believes that the proportion is higher. She asks 2525 randomly chosen customers and finds that 1212 of them choose the premium model. She tests the manufacturer's claim at the 5%5\% significance level.
    Let pp be the proportion of customers who choose the premium model. Write down the null and alternative hypotheses, and state the distribution of XX, the number of customers in the sample who choose the premium model, if the null hypothesis is true.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).