4.2 Presentation of dataIB Maths: Analysis and Approaches HL: Revision notes
Section 1
Frequency tables and class intervals
A frequency distribution lists each value or class with its frequency (how many data items it contains). For continuous data, classes are written as inequalities without gaps or overlaps, e.g. , . Every possible value then belongs to exactly one class: 15 cm goes into .
Grouping loses the individual values, so anything calculated from a grouped table (median, quartiles) is an estimate.
Writing classes like –, – for continuous data. A value of 15.5 would fit neither class.
Section 2
Frequency histograms
A frequency histogram with equal class widths has touching bars whose heights are the class frequencies. It shows the shape of a distribution: where the data cluster, whether it is symmetric or skewed, and any gaps.
Here everything is described in words, e.g. 'bars of heights 4, 9, 15, 12, 7, 3 for classes of width 5 cm from 10 cm'. Read these exactly like a frequency table. (Frequency density histograms with unequal widths are not required.)
Unlike a bar chart for categories, a histogram has no gaps between bars, because the data are continuous.
Section 3
Cumulative frequency, median, quartiles and percentiles
The cumulative frequency up to a value is the number of data items less than it. From a cumulative frequency graph or table you can estimate:
- median: the value at position ;
- lower quartile at and upper quartile at ;
- the th percentile at position ;
- IQR and range maximum minimum.
Without a graph, use linear interpolation: if the position lies items into a class of width containing items starting at , the estimate is .
Example: 200 apples, 90 below 140 g, 60 in . Median (100th) g.
Using with grouped data. For cumulative frequency, use , and .
Cumulative frequency is plotted at the upper class boundary, e.g. 30 apples at .
Section 4
Box and whisker diagrams and outliers
A box and whisker diagram displays the five-number summary: minimum, , median, , maximum. The box runs from to with a line at the median. Outliers (more than beyond a quartile) are shown as separate points, and the whisker stops at the most extreme value that is not an outlier.
On this platform the five-number summary is given in words; you give the features (quartiles, IQR, outlier boundaries), not the drawing.
, : IQR , boundaries and , so 230 is an outlier.
Section 5
Comparing distributions and judging normality
When comparing two distributions, compare one measure of centre (median) and one measure of spread (IQR or range), and interpret each in context: 'the median wait at B is 40 s shorter, so a typical customer is served faster'.
Symmetry: if the median is roughly in the middle of the box and the whiskers are similar lengths, the data are roughly symmetric and may be normally distributed. If the median is much closer to (longer right side), the data are positively skewed; closer to , negatively skewed. Skewed data are not well modelled by a normal distribution.
Comparing only the numbers ('80 < 120') without saying what that means in context.
Use the IQR rather than the range to compare spread when there are outliers, because the range is distorted by extreme values.
Must know
- Classes as inequalities, no gaps or overlaps.
- Histograms here have equal class widths; heights are frequencies.
- Median at , quartiles at and , th percentile at ; interpolate within a class.
- IQR ; outliers beyond from the quartiles.
- Compare a centre and a spread, in context.
- Symmetric box and whiskers suggest a normal model may fit; skewed data do not.
That's the notes covered.
Carry on to the next subtopic.