All revision notes topics

4.2 Presentation of dataIB Maths: Analysis and Approaches HL: Revision notes

Section 1

Frequency tables and class intervals

A frequency distribution lists each value or class with its frequency (how many data items it contains). For continuous data, classes are written as inequalities without gaps or overlaps, e.g. 10≤L<1510 \le L < 15, 15≤L<2015 \le L < 20. Every possible value then belongs to exactly one class: 15 cm goes into 15≤L<2015 \le L < 20.

Grouping loses the individual values, so anything calculated from a grouped table (median, quartiles) is an estimate.

Key termsfrequency distributionclass intervalfrequency
Common mistake

Writing classes like 1010–1515, 1616–2020 for continuous data. A value of 15.5 would fit neither class.

Section 2

Frequency histograms

A frequency histogram with equal class widths has touching bars whose heights are the class frequencies. It shows the shape of a distribution: where the data cluster, whether it is symmetric or skewed, and any gaps.

Here everything is described in words, e.g. 'bars of heights 4, 9, 15, 12, 7, 3 for classes of width 5 cm from 10 cm'. Read these exactly like a frequency table. (Frequency density histograms with unequal widths are not required.)

Key termsfrequency histogram
Exam tip

Unlike a bar chart for categories, a histogram has no gaps between bars, because the data are continuous.

Section 3

Cumulative frequency, median, quartiles and percentiles

The cumulative frequency up to a value is the number of data items less than it. From a cumulative frequency graph or table you can estimate:

  • median: the value at position n2\frac{n}{2};
  • lower quartile Q1Q_1 at n4\frac{n}{4} and upper quartile Q3Q_3 at 3n4\frac{3n}{4};
  • the ppth percentile at position p100×n\frac{p}{100} \times n;
  • IQR =Q3−Q1= Q_3 - Q_1 and range == maximum −- minimum.

Without a graph, use linear interpolation: if the position lies rr items into a class of width ww containing ff items starting at aa, the estimate is a+rf×wa + \frac{r}{f} \times w.

Example: 200 apples, 90 below 140 g, 60 in 140≤m<160140 \le m < 160. Median (100th) ≈140+1060×20=143.3\approx 140 + \frac{10}{60} \times 20 = 143.3 g.

Key termscumulative frequencypercentileinterquartile rangelinear interpolation
Common mistake

Using n+12\frac{n + 1}{2} with grouped data. For cumulative frequency, use n2\frac{n}{2}, n4\frac{n}{4} and 3n4\frac{3n}{4}.

Exam tip

Cumulative frequency is plotted at the upper class boundary, e.g. 30 apples at m=120m = 120.

Section 4

Box and whisker diagrams and outliers

A box and whisker diagram displays the five-number summary: minimum, Q1Q_1, median, Q3Q_3, maximum. The box runs from Q1Q_1 to Q3Q_3 with a line at the median. Outliers (more than 1.5×IQR1.5 \times \text{IQR} beyond a quartile) are shown as separate points, and the whisker stops at the most extreme value that is not an outlier.

On this platform the five-number summary is given in words; you give the features (quartiles, IQR, outlier boundaries), not the drawing.

Key termsfive-number summarybox and whisker diagram
Example

Q1=95Q_1 = 95, Q3=145Q_3 = 145: IQR =50= 50, boundaries 95−75=2095 - 75 = 20 and 145+75=220145 + 75 = 220, so 230 is an outlier.

Section 5

Comparing distributions and judging normality

When comparing two distributions, compare one measure of centre (median) and one measure of spread (IQR or range), and interpret each in context: 'the median wait at B is 40 s shorter, so a typical customer is served faster'.

Symmetry: if the median is roughly in the middle of the box and the whiskers are similar lengths, the data are roughly symmetric and may be normally distributed. If the median is much closer to Q1Q_1 (longer right side), the data are positively skewed; closer to Q3Q_3, negatively skewed. Skewed data are not well modelled by a normal distribution.

Key termssymmetricpositively skewednegatively skewed
Common mistake

Comparing only the numbers ('80 < 120') without saying what that means in context.

Exam tip

Use the IQR rather than the range to compare spread when there are outliers, because the range is distorted by extreme values.

Must know

  • Classes as inequalities, no gaps or overlaps.
  • Histograms here have equal class widths; heights are frequencies.
  • Median at n2\frac{n}{2}, quartiles at n4\frac{n}{4} and 3n4\frac{3n}{4}, ppth percentile at p100n\frac{p}{100}n; interpolate within a class.
  • IQR =Q3−Q1= Q_3 - Q_1; outliers beyond 1.5×IQR1.5 \times \text{IQR} from the quartiles.
  • Compare a centre and a spread, in context.
  • Symmetric box and whiskers suggest a normal model may fit; skewed data do not.

That's the notes covered.

Carry on to the next subtopic.