4.18 Hypothesis testing: means, proportions, correlation and errorsIB Maths: Applications and Interpretation HL: Revision notes
Section 1
The language of a test
A hypothesis test decides whether sample data give enough evidence against a claim. The null hypothesis is the claim to be tested (it states a parameter equals a value, for example or ). The alternative hypothesis says what we suspect: (two-tailed) or / (one-tailed). Hypotheses are about the population parameter, never the sample statistic. The significance level (for example ) is the probability of rejecting when it is true. The -value is the probability, assuming is true, of a result at least as extreme as the one observed. If , reject ; otherwise do not reject. For a two-tailed test compare with after doubling the tail probability (or compare each tail with ). The critical region is the set of values of the test statistic that lead to rejecting ; its boundary is the critical value. Conclusions must be in context and cautious: 'sufficient evidence to reject ' or 'insufficient evidence to reject ', never 'proved'.
Writing hypotheses about or the sample. They are about the population parameter (, , , ).
Concluding that the null hypothesis is 'true' when . You only have insufficient evidence to reject it.
Section 2
Tests for a population mean
known (normal population, or large ): under , , so use the -test; the GDC gives the -value. Example: , , , , : and , so reject . unknown: use the one-sample -test with degrees of freedom, whatever the sample size. The GDC -test takes the data (or , , ) and gives and the -value; you are not expected to calculate critical regions for -tests. Paired data: samples may be paired (matched pairs). Take the differences and treat them as a single sample, testing . Unpaired samples are two independent samples.
Say which direction the differences are taken (before minus after) so the alternative hypothesis has the right sign.
Section 3
Tests for a proportion and for a Poisson mean
Proportion (binomial): . With trials, under . For the -value for observing is . Example: , , : , so reject . Poisson mean: , and under . Rescale to the period observed. Binomial and Poisson tests are one-tailed only in IB. The critical region for a discrete variable is found by testing values: choose the region with the largest probability that is still less than the significance level. Example: , : but , so the critical region is .
For using instead of .
Section 4
Type I and Type II errors
A Type I error is rejecting when it is true; is the probability that the test statistic lies in the critical region under . For continuous tests this equals ; for discrete tests it is the actual probability of the critical region, which may be below (for example ). A Type II error is not rejecting when it is false. Its probability is found under a specific true value of the parameter: . Examples: Normal, , , , at : critical region . If the true , . Binomial: , critical region ; if , . Poisson: , critical region ; if , .
A Type I error uses the null value of the parameter; a Type II error uses the true (alternative) value given in the question.
Section 5
Testing for correlation
For bivariate data from a bivariate normal population, the sample product-moment correlation coefficient estimates the population coefficient . To test whether there is a linear correlation: against (or , ). Use technology; in examinations the data (or the output) will be given. If the -value is less than the significance level, reject : there is sufficient evidence of a linear correlation. Example: with at gives evidence that . Correlation does not prove causation.
Saying a significant correlation means one variable causes the other.
That's the notes covered.
Carry on to the next subtopic.
Exam questions on 4.18 Hypothesis testing: means, proportions, correlation and errors
- A machine is set to fill bags of sugar with a mean mass of g. The mass of a bag is normally distributed with a known standard deviation of g. An inspector suspects that the machine is under-filling the bags. She takes a random sample of bags, which has a mean mass of g, and tests the claim at the significance level.State the conclusion of the test, giving a reason for your answer.2 marks
- A trainer records the m sprint times of athletes before and after a training programme. The differences in time, in seconds, calculated as before minus after, are . The differences are normally distributed. The trainer tests, at the significance level, whether the programme reduces the mean sprint time. Let be the population mean difference. Use your GDC where helpful.State the conclusion of the test in context, giving a reason for your answer.2 marks
- A manufacturer claims that of its customers choose the premium model of a phone. A retailer believes that the proportion is higher. She asks randomly chosen customers and finds that of them choose the premium model. She tests the manufacturer's claim at the significance level.Let be the proportion of customers who choose the premium model. Write down the null and alternative hypotheses, and state the distribution of , the number of customers in the sample who choose the premium model, if the null hypothesis is true.3 marks
Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).