All revision notes topics

4.1 Sampling, data quality and outliersIB Maths: Applications and Interpretation SL: Revision notes

Section 1

Population, sample and types of data

A population is the whole group being studied; a sample is a smaller part of it used to draw conclusions about the whole. A random sample gives every member of the population an equal chance of selection. At SL, a data set given in a question is treated as the population unless stated otherwise. Discrete data can take only separate values, usually counted (number of siblings, goals scored). Continuous data can take any value in a range, usually measured (height, time, mass). Example: the number of cars passing a school in an hour is discrete; the time a car takes to pass a point is continuous.

Key termspopulationsamplerandom samplediscrete datacontinuous data
Exam tip

Counting gives discrete data; measuring gives continuous data.

Section 2

Reliability and bias

A data source is reliable if it is trustworthy: it comes from a reputable organisation, uses a clear method and a suitably large sample, and gives similar results when repeated. Bias is a systematic tendency for a sample to differ from the population, for example when some members have no chance of being chosen, when a survey question is leading, or when only some people reply (non-response bias). A larger sample does not remove bias: a large sample from the wrong group is still biased.

Key termsreliablebiasnon-response bias
Common mistake

Saying a sample is biased only because it is small. Small samples are unreliable; bias comes from how members are chosen.

Section 3

Sampling techniques (1): random, systematic and stratified

Simple random sampling: number every member, then use random numbers (GDC) to pick distinct members. Fair, but needs a full list and may miss small groups. Systematic sampling: choose every kkth member from an ordered list, starting at a random place, where k=populationsample sizek=\frac{\text{population}}{\text{sample size}}. Quick, but biased if the list has a repeating pattern that matches kk. Stratified sampling: split the population into groups (strata) and sample each in proportion to its size: sample from a group=group sizepopulation×sample size.\text{sample from a group}=\frac{\text{group size}}{\text{population}}\times\text{sample size}. Example: 600 members (240 juniors, 210 adults, 150 seniors), sample of 80. Juniors =240×80600=32=240\times\frac{80}{600}=32, adults =28=28, seniors =20=20. Check 32+28+20=8032+28+20=80.

Key termssimple randomsystematicstratifiedstrata
Common mistake

Rounding each stratum separately and not checking the total: the stratum sizes must add up to the sample size.

Section 4

Sampling techniques (2): convenience and quota

Convenience sampling: choose whoever is easiest to reach (e.g. the first people to walk past). Cheap and fast, but often biased and not representative. Quota sampling: the interviewer must find a set number of people in each group (e.g. 20 men and 20 women) but chooses them non-randomly. It reflects the group structure like stratified sampling, but selection within groups is not random, so bias is still possible. No sampling frame is needed. Effectiveness: random and stratified samples are the most representative; convenience samples are the least.

Key termsconveniencequotasampling frame
Exam tip

Quota = stratified structure with non-random selection. Mention both parts if asked to compare.

Section 5

Errors and missing data

Data may contain recording errors (typing mistakes, impossible values such as 400 hours in a week, a stopwatch left running) or missing data. Sensible responses: check the source and correct the value; remove an item that cannot be corrected; follow up non-respondents or replace them; or report the issue and treat conclusions with caution. Never silently keep an impossible value. Missing data matter most when the missing items are different from the rest, because then the remaining data are biased.

Key termsrecording errormissing data

Section 6

Outliers

An outlier is a data item more than 1.5×IQR1.5\times\text{IQR} from the nearest quartile: a value is an outlier if it is below Q1−1.5×IQRQ_1-1.5\times\text{IQR} or above Q3+1.5×IQRQ_3+1.5\times\text{IQR}, where IQR=Q3−Q1\text{IQR}=Q_3-Q_1. Example: 4,5,5,6,7,8,9,254,5,5,6,7,8,9,25. Then Q1=5Q_1=5, Q3=8.5Q_3=8.5, IQR=3.5\text{IQR}=3.5, upper boundary 8.5+1.5(3.5)=13.758.5+1.5(3.5)=13.75. Since 25>13.7525>13.75, 2525 is an outlier. Lower boundary 5−5.25<05-5.25<0, so there are no low outliers. In context, an outlier may be a valid part of the sample (an exceptionally fast runner) or an error. Check before removing. Outliers are shown as crosses on box and whisker diagrams (4.2) and they inflate the standard deviation and range (4.3).

Key termsoutlierinterquartile range
Common mistake

Measuring 1.5×IQR1.5\times\text{IQR} from the median or mean. It is measured from the nearest quartile.

Exam tip

Quartiles from a GDC and by hand can differ slightly for odd sample sizes; use the GDC value in exams.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on 4.1 Sampling, data quality and outliers

  1. A council wants to find out how its residents travel to work. A researcher stands outside the main railway station one morning and questions the first 80 people who walk past. The council has 12 000 residents on its electoral register.
    Describe how the council could take a simple random sample of 80 residents from the electoral register.2 marks
  2. A sports club has 600 members: 240 juniors, 210 adults and 150 seniors. The club takes a stratified sample of 80 members for a survey about its facilities.
    Explain why a stratified sample is more suitable than a simple random sample for this survey.2 marks
  3. Eleven students record the time, in minutes, they take to finish a puzzle: 12, 14, 15, 15, 17, 18, 19, 21, 22, 24, 41. The student who took 41 minutes says she forgot to stop her timer. Use your GDC to find quartiles.
    Find the interquartile range (IQR) of the times.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).