Hypothesis tests for the difference between two meansEdexcel International A Level Further Maths: Revision notes
Section 1
Comparing two population means
Two independent random samples are taken, one from each of two populations, to test whether their means differ. Let have mean and variance , with sample size and sample mean , and similarly with , , and . The hypotheses are (or ) against (two-tailed), or or (one-tailed). As always, the null hypothesis holds the equality and the hypotheses concern population means, not sample means. The statistic that is tested is the difference of the sample means, .
Writing the hypotheses with and . They must use and .
Section 2
The distribution of
If and are independent, then and , so The means subtract but the variances add, because subtracting independent random variables does not remove any of the variability. The standard error of the difference is . Example: , , , gives variance .
Subtracting the variances. Variances always add, even for a difference.
Section 3
Carrying out the test (variances known)
Under (), the test statistic is Compare with the critical value for the significance level and the tail(s): (two-tailed ), or (one-tailed ), (one-tailed ), (two-tailed ). Worked example: fertilisers and , , , , , , , . The variance of the difference is , so . Do not reject : insufficient evidence at the level that gives taller seedlings.
Keep the order of the means the same as in : for use , which gives a negative when the result supports .
Section 4
Large samples with unknown variances
When the population variances are not known, use the unbiased estimates and . If both samples are large, the Central Limit Theorem makes and approximately Normal (even if the populations are not Normal), and , are close enough to , that Example: stores and with , , , , , . The variance is and , so there is insufficient evidence of a difference at the level. For small samples this approximation is not reliable, and a different distribution would be needed, which is not required here. Always state that the samples are large.
Using and with small samples as though the result were exact. Large samples are needed for this method.
Section 5
Conclusions and assumptions
Give the conclusion in context and at the stated level: 'there is evidence that method gives a shorter mean time than method ', or 'there is insufficient evidence that the mean masses differ'. Never claim that is proved. Assumptions to state or check: the samples are random and independent; for small samples the populations are Normal (and the variances known); for large samples the Central Limit Theorem is used. A test about means says nothing about every individual: some members of the population with the smaller mean may still exceed some members of the population with the larger mean. A result can be significant at but not at , so the level matters. If , the test is significant in a one-tailed test at () and at ().
When asked to evaluate a conclusion, comment on the population versus individuals, the assumptions, and the strength of the evidence.
That's the notes covered.
Carry on to the next subtopic.
Exam questions on Hypothesis tests for the difference between two means
- Two machines, and , fill bags with sugar. The masses, in grams, are Normally distributed, with known standard deviations of for machine and for machine . A random sample of bags from has a mean mass of and a random sample of bags from has a mean mass of . A test of against is carried out at the significance level.Complete the test and state your conclusion in context.2 marks
- Seedlings are grown using either fertiliser or fertiliser . The heights are Normally distributed, with known standard deviations of cm for and cm for . A random sample of seedlings grown with has a mean height of cm and a random sample of seedlings grown with has a mean height of cm. A gardener wants to test, at the significance level, whether fertiliser gives a greater mean height than fertiliser .Carry out the test and state your conclusion in context.2 marks
- A supermarket chain compares customer spending at two stores, and . A random sample of customers at spent a mean of £42.50 with sample variance , and a random sample of customers at spent a mean of £39.80 with sample variance . The distributions of spending are not assumed to be Normal. The chain tests against at the significance level.State the approximate distribution of under , and justify why this distribution may be used.3 marks
Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).