All revision notes topics

Type I and Type II errors and powerEdexcel A-Level Further Maths: Revision notes

Section 1

Type I and Type II errors

A hypothesis test can go wrong in two ways. A Type I error is rejecting H0\mathrm{H}_0 when it is true. A Type II error is failing to reject H0\mathrm{H}_0 when it is false (so H1\mathrm{H}_1 is actually true). P(Type I error)=P(reject H0∣H0 true),P(Type II error)=P(do not reject H0∣H1 true).\mathrm{P}(\text{Type I error})=\mathrm{P}(\text{reject }\mathrm{H}_0\mid\mathrm{H}_0\text{ true}),\qquad\mathrm{P}(\text{Type II error})=\mathrm{P}(\text{do not reject }\mathrm{H}_0\mid\mathrm{H}_1\text{ true}). Always say what the errors mean in the context of the question: for a coin, a Type I error is deciding it is biased when it is fair.

Key termsType I errorType II error
Common mistake

Describing a Type II error as 'accepting H1\mathrm{H}_1 wrongly'. It is failing to reject H0\mathrm{H}_0 when H1\mathrm{H}_1 is true.

Section 2

Size of a test

The size of a test is the probability of a Type I error, which is the probability that the test statistic falls in the critical region when H0\mathrm{H}_0 is true. For a discrete distribution the actual size is usually below the nominal significance level because the critical region has to be made of whole values. Example: X∼B(10,p)X\sim\mathrm{B}(10,p), H0:p=0.5\mathrm{H}_0:p=0.5, critical region X≥8X\geq8. Size =P(X≥8∣p=0.5)=45+10+11024=0.0547=\mathrm{P}(X\geq8\mid p=0.5)=\frac{45+10+1}{1024}=0.0547. For a normal mean, Xˉ∼N(μ,σ2n)\bar X\sim\mathrm{N}\left(\mu,\frac{\sigma^2}{n}\right): reject if xˉ<499\bar x<499 with μ=500\mu=500, σ=2\sigma=2, n=16n=16 gives P(Z<−2)=0.0228\mathrm{P}(Z<-2)=0.0228.

Key termssizecritical region

Section 3

Probability of a Type II error

To find P(Type II error)\mathrm{P}(\text{Type II error}) you need a specific value of the parameter from H1\mathrm{H}_1. Use the same distribution family but with that value, and find the probability of landing in the acceptance region (the values where H0\mathrm{H}_0 is not rejected). Example: critical region X≥8X\geq8 for B(10,p)\mathrm{B}(10,p) and true p=0.7p=0.7: P(Type II)=P(X≤7∣p=0.7)=0.617\mathrm{P}(\text{Type II})=\mathrm{P}(X\leq7\mid p=0.7)=0.617. Different true values of pp give different Type II probabilities, which is why a single value for the error needs a stated alternative.

Key termsacceptance region
Common mistake

Calculating a Type II probability using the value from H0\mathrm{H}_0. You must use a value from H1\mathrm{H}_1.

Section 4

Power and the power function

The power of a test is the probability of rejecting H0\mathrm{H}_0 when it is false: power=P(reject H0∣H1 true)=1−P(Type II error).\text{power}=\mathrm{P}(\text{reject }\mathrm{H}_0\mid\mathrm{H}_1\text{ true})=1-\mathrm{P}(\text{Type II error}). The power function gives the power as a function of the parameter: power(p)=P(X∈critical region∣p)\text{power}(p)=\mathrm{P}(X\in\text{critical region}\mid p). At the value in H0\mathrm{H}_0 it equals the size of the test, and it rises as the true parameter moves further from the H0\mathrm{H}_0 value. For B(10,p)\mathrm{B}(10,p) with critical region X≥8X\geq8: power at p=0.6p=0.6 is 0.1670.167 and at p=0.7p=0.7 is 0.3830.383.

Key termspowerpower function
Exam tip

Power is always 1 minus the Type II error probability, at the same true parameter value.

Section 5

Effectiveness and trade-offs

A good test has small size and large power. Widening the critical region (for example X≥4X\geq4 instead of X≥5X\geq5) increases the power but also the size; narrowing it does the opposite. You cannot reduce both errors just by moving the boundary. The usual way to improve both is to increase the sample size. When you compare tests, compute the size of each and the power of each at the same alternative value, then judge against what is wanted. The same ideas work for the binomial, Poisson and normal distributions from A level Mathematics and Further Statistics 1.

Key termseffectiveness of a test
Exam tip

For 'evaluate' questions, quote both errors for each test and give a recommendation.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Type I and Type II errors and power

  1. A coin is suspected of being biased towards heads. It is tossed 10 times and XX is the number of heads, where X∼B(10,p)X\sim\mathrm{B}(10,p). The hypotheses are H0:p=0.5\mathrm{H}_0:p=0.5 and H1:p>0.5\mathrm{H}_1:p>0.5, and H0\mathrm{H}_0 is rejected if X≥8X\geq8.
    Find the power of the test when p=0.6p=0.6.2 marks
  2. The number of defects on a sheet of metal is modelled by Po(λ)\mathrm{Po}(\lambda). One sheet is inspected to test H0:λ=2\mathrm{H}_0:\lambda=2 against H1:λ>2\mathrm{H}_1:\lambda>2. The test rejects H0\mathrm{H}_0 if the sheet has 5 or more defects.
    The critical region is changed to X≥6X\geq6. State, with a supporting calculation, the effect on the size of the test and on the power of the test when λ=4\lambda=4.2 marks
  3. A machine fills cereal boxes, and the mass of cereal in a box, in grams, is normally distributed with standard deviation 22. The mean mass should be 500500 but the manager suspects it is lower. The manager tests H0:μ=500\mathrm{H}_0:\mu=500 against H1:μ<500\mathrm{H}_1:\mu<500 using the mean xˉ\bar x of a random sample of 16 boxes, and rejects H0\mathrm{H}_0 if xˉ<499\bar x<499.
    Find the size of the test.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).