Collecting and classifying dataIB MYP Maths Extended: Revision notes
Section 1
Types of data
Categorical data are described in words or categories, such as bag colour or favourite sport. Numerical data are numbers and are either discrete or continuous.
- Discrete data can only take separate values, usually found by counting: number of siblings, goals scored, shoe size.
- Continuous data can take any value in a range, usually found by measuring: height, mass, time. Continuous data are always rounded when recorded, but the true value could be anywhere in a range.
Calling shoe size continuous because it is measured. It only takes listed values such as and , so it is discrete.
Ask: can it take a value between two neighbouring values? If yes, it is continuous.
Section 2
Populations and samples
The population is every person or item you want to know about. A sample is a smaller group chosen from the population. A census asks the whole population, but this is often too slow or expensive, so we use a sample. A good sample is representative, meaning it has the same mix of people as the population. A larger sample is usually more reliable.
Section 3
Sampling methods
- Simple random: every member has an equal chance of being chosen (random numbers or a draw).
- Systematic: choose every th member from a list, starting at a random point. For students and a sample of , choose every th.
- Stratified: split the population into groups (strata) and take a random sample from each in proportion to its size.
- Convenience: choose people who are easy to reach. This is quick but often biased. Stratified example: from members, of whom are adults: adults.
Choosing stratified numbers without using the proportion of each group in the population.
Section 4
Bias
Bias means the results are pushed in one direction, so they do not truthfully represent the population. It can come from:
- a sample that is not representative (only asking people at one place);
- self-selection, where volunteers take part, such as an online poll;
- non-response, where many people chosen do not reply, so the replies may be different from the rest;
- leading questions that suggest a particular answer. To reduce bias, use a random or stratified sample, a larger sample and neutral questions, and follow up non-responders.
In a bias question say who is left out and which way it changes the results.
Section 5
Designing questionnaires and collecting data
A good question is clear, neutral and gives a time period ("each school day"). Answer boxes should cover every possible answer, not overlap, and be easy to tick: None; Less than 1 hour; 1 to less than 2 hours; 2 hours or more. Include "none" or "other" where it is needed. Do not ask two questions in one. Primary data are collected by you (survey, observation, experiment); secondary data were collected by someone else (a database or website). Use a tally chart to record counts as you go. A short pilot survey with a few people shows up unclear questions before the real survey.
Overlapping boxes such as to hours and to hours. Which box does hours go in? Use to less than , then to less than .
That's the notes covered.
Carry on to the next subtopic.
Exam questions on Collecting and classifying data
- A school survey records four pieces of information about each student: shoe size, the colour of their bag, their height in centimetres and the number of siblings.Explain why shoe size is discrete data.2 marks
- A town council wants to find out how residents travel to work. It has a list of all residents, of whom live in the north of the town and live in the south.A councillor suggests surveying only people who are at the railway station at 8 am on a weekday. Explain why this sample would be biased.2 marks
- A student wants to find out how long pupils in her school spend on social media each day. She asks: "You spend too much time on social media, don't you?" The answer boxes are: "A lot", "Some" and "Hardly any".Identify three faults with the question and the answer boxes.3 marks
Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).