4.11 Hypothesis testing: chi-squared and t-testsIB Maths: Applications and Interpretation SL: Revision notes
Section 1
Hypotheses, significance levels and p-values
A hypothesis test decides whether sample data give enough evidence against a claim.
- The null hypothesis is the claim being tested, written as an equation, such as , or in words (for example 'the variables are independent').
- The alternative hypothesis is what we accept if there is enough evidence against . For a mean it may be (two-tailed) or or (one-tailed).
- The significance level is the probability of rejecting when it is true. Common levels are , and .
- The -value is the probability of getting a result at least as extreme as the one observed, if is true. Decision rule: if significance level, reject ; if significance level, do not reject . State the conclusion in context. Say 'there is evidence that…' or 'there is not enough evidence that…'. Never say that has been proved.
Writing 'accept ' or 'H0 is true'. Say you do not reject and that there is not enough evidence.
Section 2
The chi-squared test for independence
The test for independence tests whether two categorical variables are associated. : the variables are independent. : they are not independent (associated). Data are in a contingency table of observed frequencies . The expected frequency in each cell, if is true, is The statistic is and the degrees of freedom are . In examinations tables have at most rows or columns, , and you use your GDC for and the -value. Only upper-tail tests are used: reject if is greater than the critical value (given), or if the significance level. Example: Year travel and Year travel (walk, bus, car). The expected values are in both rows, , and , so there is no evidence of an association.
Check that every expected frequency is greater than before using the test.
Section 3
The chi-squared goodness of fit test
The goodness of fit test tests whether data fit a stated distribution. : the data fit the distribution (for example, equally likely categories). : they do not. Find expected frequencies from the distribution: total probability of each category. For categories, at SL the degrees of freedom are . Then and compare with the critical value, or compare with the significance level. Example: customers choose among desserts: . If the choices are equally likely, each expected frequency is . and . At the critical value is . Since (and ), reject : there is evidence that the desserts are not chosen equally often.
Using for goodness of fit. At SL it is the number of categories minus .
Section 4
The t-test for comparing two means
The -test uses the sample means to test whether the means of two populations differ. In examinations the samples are unpaired, the population variance is unknown and you assume equal variances, so the pooled two-sample -test is used on your GDC. The test needs the data in both populations to be normally distributed. Hypotheses: and (two-tailed), or (one-tailed). Enter the sample means, sample standard deviations and sample sizes (or the raw data) into the GDC and choose the correct alternative. The GDC returns and the -value for the alternative chosen. Compare with the significance level. Example: class A ( students, mean , standard deviation ) and class B ( students, mean , standard deviation ). One-tailed test, : , reject . Two-tailed: , do not reject . The one-tailed -value is half the two-tailed value when the sample difference is in the direction of .
Choose one-tailed or two-tailed from the wording: 'greater than' is one-tailed, 'different from' is two-tailed.
Section 5
Interpreting results and limitations
Write conclusions in context and compare with the given significance level: the same can be significant at but not at (for example ).
- A significant result is evidence of an effect, not proof, and it does not say how large or important the effect is.
- For tests, expected frequencies of or less make the test unreliable, so categories may need combining.
- The -test needs the populations to be normal with equal variances; if not, the conclusion may be unreliable.
- A larger sample makes it easier to detect a real difference. Do not use a two-tailed test when only one direction was predicted in advance, and do not change the significance level after seeing the result.
Choosing the significance level to get the conclusion you want. Use the level given in the question.
That's the notes covered.
Carry on to the next subtopic.
Exam questions on 4.11 Hypothesis testing: chi-squared and t-tests
- One hundred and fifty students in Year and Year were each asked how they usually travel to school. Year (75 students): walk , bus , car . Year (75 students): walk , bus , car . A test for independence is to be carried out between year group and method of travel.Use your GDC to find the -value of the test, and state your conclusion at the significance level.2 marks
- A café sells four desserts. The manager claims that customers choose the four desserts equally often. In one week, customers chose cheesecake , brownie , sorbet and tart times. A goodness of fit test is used to test the claim.Write down the null and alternative hypotheses, and the number of degrees of freedom.2 marks
- A teacher compares the scores of two classes in the same test. Class A has students with mean score and sample standard deviation . Class B has students with mean score and sample standard deviation . The scores in both classes are normally distributed with equal variances. Use a pooled two-sample -test on your GDC.Test, at the significance level, whether the mean score of class A is greater than that of class B. State the hypotheses, the -value and your conclusion.3 marks
Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).