All revision notes topics

4.12 Data collection, reliability and validityIB Maths: Applications and Interpretation HL: Revision notes

Section 1

Designing a valid data collection method

A survey or questionnaire collects data from a sample. A valid method begins with the question to be answered: choose the relevant variables from the many that could be recorded, and collect relevant and appropriate data. Collecting everything makes the analysis messy; collecting too little makes a conclusion impossible. The sample must represent the population; asking only people who are easy to reach (for example students leaving the canteen) can introduce bias. Check that the data type suits the analysis: numerical data for means and regression, categorical data for χ2\chi^2 tests.

Key termssurveyquestionnairerelevant variablesbias
Common mistake

Choosing a convenient sample and assuming it represents everyone. Say who is left out and how that could change the results.

Section 2

Writing good questions

Questions can be biased or unbiased. A biased (leading or loaded) question pushes the answer, for example "Don't you agree that the new menu is far better?". An unbiased question is neutral: "How satisfied are you with the new menu?". Personal questions (age, income, health) may get untruthful answers, so ask only what is needed and consider ranges. Unstructured (open) questions allow free answers; they give rich detail but are hard to analyse. Structured questions give consistent answer choices, for example a five-point scale from "very satisfied" to "very dissatisfied", which are easy to count and compare. The choices should be balanced, cover all possible answers and not overlap. Precise questioning means exact wording and units, for example "How many hours did you sleep last night?" rather than "Do you sleep well?".

Key termsbiased questionunstructuredstructuredprecise
Exam tip

To check a question, ask: does it hint at the answer? could two people read it differently? do the choices cover everyone and avoid overlap?

Section 3

Categorising numerical data for a chi-squared test

To use a χ2\chi^2 test on numerical data (for example hours slept), group the values into categories (classes) and count how many fall in each. Choose categories that make sense for the question and justify the choice, for example equal width, or boundaries with a natural meaning. The test requires expected frequencies greater than 5 in every category (a common rule). If a category has a smaller expected frequency, combine it with a neighbouring category and recalculate. In a contingency table the expected frequency is row total×column totalgrand total\frac{\text{row total}\times\text{column total}}{\text{grand total}}. Example: for "under 5 hours" with 6 students, expected alert =6×87155=3.37<5=\frac{6\times87}{155}=3.37<5, so combine "under 5" with "5 to under 6".

Key termscategoriesexpected frequencycombine categories
Common mistake

Combining categories using the observed frequencies instead of the expected ones. The rule of 5 applies to the expected values.

Section 4

Degrees of freedom when parameters are estimated

For a χ2\chi^2 goodness of fit test, with kk categories (after combining) the degrees of freedom are ν=k−1−m,\nu=k-1-m, where mm is the number of parameters estimated from the data (for example pp for a binomial, or the mean for a Poisson). If the parameters are given, m=0m=0. For a contingency table, ν=(rows−1)(columns−1)\nu=(\text{rows}-1)(\text{columns}-1). Example: a binomial B(5,p)B(5,p) with pp estimated, and 3 categories after combining, gives ν=3−1−1=1\nu=3-1-1=1. Use your GDC for χ2\chi^2 and the pp-value: reject H0H_0 if p<p< the significance level.

Key termsdegrees of freedomparameters estimatedgoodness of fit
Exam tip

Count categories after combining, not before. Then subtract 1 for the total and 1 for every parameter you estimated.

Section 5

Reliability

A test is reliable if it gives consistent results when repeated under the same conditions. Reliability does not depend on whether the test measures the right thing. Test-retest: give the same test to the same people on two occasions and find the correlation between the scores; a high rr (close to 1) means high reliability. Parallel forms: give two different but equivalent versions of the test to the same people and correlate the scores.

Key termsreliabletest-retestparallel forms
Common mistake

Saying a high test-retest correlation proves the test is valid. It only shows consistency.

Section 6

Validity, and how it differs from reliability

A test is valid if it measures what it is intended to measure. Content validity: the questions cover all the relevant content (for example an algebra test that includes all algebra topics in the syllabus), usually judged by experts. Criterion-related validity: scores agree with an independent, accepted measure of the same thing (the criterion), for example a questionnaire compared with a counsellor's diagnosis. The two ideas are different: a bathroom scale that is always 3 kg too heavy is reliable (consistent) but not valid (accurate). A test can be reliable without being valid, but a test that is not reliable cannot be valid.

Key termsvalidcontent validitycriterion-related validity
Exam tip

Reliable means "repeatable", valid means "on target". In an answer, name the type of test and say what it shows.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on 4.12 Data collection, reliability and validity

  1. A school council wants to find out how students feel about the new canteen menu. One member suggests asking students as they leave the canteen at lunchtime: "Don't you agree that the new menu is far healthier and tastier than the old boring one?"
    Explain why asking only students who are leaving the canteen at lunchtime may give biased results.2 marks
  2. A psychologist designs a 20-question questionnaire to measure exam anxiety. She gives it to 40 students and then gives it to the same students again two weeks later. The Pearson's product-moment correlation coefficient between the two sets of scores is r=0.91r=0.91. She also compares the scores with a diagnosis of exam anxiety made independently by a school counsellor.
    Explain what is meant by validity, and state the type of validity test carried out when the scores are compared with the counsellor's diagnosis.2 marks
  3. A researcher records, for 155 students, the number of hours slept the previous night and whether the student felt alert in the first lesson. The numbers of students who felt alert and who did not feel alert were: under 5 hours: 1 and 5; 5 to under 6 hours: 9 and 15; 6 to under 7 hours: 25 and 25; 7 to under 8 hours: 32 and 15; 8 hours or more: 20 and 8. In total 87 students felt alert. A χ2\chi^2 test for independence is to be carried out.
    Show that the expected number of students who slept under 5 hours and felt alert is 3.373.37 to 3 significant figures, and explain which categories should be combined, giving a reason.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).