4.1 Populations, samples and data collectionIB Maths: Analysis and Approaches SL: Revision notes
Section 1
Populations, samples and types of data
The population is the whole group you want to know about; a sample is the part you actually collect data from. A random sample is one in which every member of the population has an equal chance of being chosen, and choices are independent.
Data are discrete if they are counted and can only take particular values (number of people in a household, shoe sizes), and continuous if they are measured and can take any value in a range (time, mass, electricity use). Continuous data are always recorded to some accuracy, but that does not make them discrete.
Calling age discrete because it is written as a whole number of years. Age is measured, so it is continuous (it has been rounded down).
Section 2
Sampling techniques
- Simple random: number every member of the population and use a random number generator to pick distinct numbers. Needs a complete list (a sampling frame).
- Systematic: from an ordered list, pick every th member after a random start, where . Can be biased if the list has a repeating pattern.
- Stratified: split the population into groups (strata) and randomly sample from each in proportion to its size: number from a stratum .
- Quota: fill a set number from each group, but choose people non-randomly (e.g. whoever passes). Quick, but open to bias.
- Convenience: use whoever is easiest to reach. Cheap and quick, but usually biased.
Example: 330 DP1 and 270 DP2 students, sample of 40 stratified by year: from DP1 and 18 from DP2.
Confusing stratified and quota sampling: both use groups, but only stratified sampling chooses randomly within each group.
When stratified numbers are not whole, round sensibly and check the total still equals the sample size.
Section 3
Bias and reliability
A sampling method is biased if it tends to over- or under-represent some part of the population, so the results are systematically wrong. When explaining bias, always name who is left out or over-represented and why that matters for the variable being measured; 'it is not random' alone earns little credit.
The reliability of a data source depends on who collected it, how, and why: a sample from a railway station at 8 a.m. says little about people who drive or work nights. Stratified sampling improves the estimate when the variable differs between groups (e.g. commuting time by district), because each group is represented in the right proportion.
Link every comment to the context: which people, which variable, which direction the estimate is pushed.
Section 4
Outliers, errors and missing data
An outlier is a value more than below the lower quartile or above the upper quartile: Being an outlier does not automatically mean a value is wrong. Decide in context:
- A genuine extreme (a very tall player in a basketball squad) should be kept.
- An error (a negative electricity reading, a distracted participant, a typing mistake) should be corrected if possible or removed.
Missing data should be left out of calculations, not replaced by 0, which would distort the mean. Report how many values were excluded.
Removing a value just because it is an outlier. You must give a reason in context why it is an error.
, : IQR , so outliers are below 160 or above 320.
Must know
- Population versus sample; random means equal chance for every member.
- Discrete = counted; continuous = measured.
- Five techniques: simple random, systematic, stratified, quota, convenience; know how each works and one strength and weakness of each.
- Stratified numbers: .
- Outlier boundaries: and .
- Outliers can be genuine or errors; missing values are excluded, never set to 0.
That's the notes covered.
Carry on to the next subtopic.