All revision notes topics

Single-variable data: diagrams and histogramsEdexcel A-Level Maths: Revision notes

Section 1

Histograms and frequency density

A histogram is used for continuous data in classes, which may have different widths. The area of each bar represents the frequency, so the vertical axis shows frequency density: frequency density=frequencyclass width,frequency=frequency density×class width.\text{frequency density}=\frac{\text{frequency}}{\text{class width}},\qquad\text{frequency}=\text{frequency density}\times\text{class width}. There are no gaps between bars. To estimate the frequency in part of a class, assume the data are spread evenly: for example 8 to 12 minutes uses 4.5×24.5\times2 from the bar up to 10 and 1.2×21.2\times2 from the bar after. If one bar's height and frequency are known, the scale for all other bars follows from the ratio of areas.

Key termshistogramfrequency densityclass width
Common mistake

Plotting frequency, not frequency density, when the class widths differ. This makes wide classes look too big.

Section 2

Frequency polygons

A frequency polygon is drawn by plotting the frequency (or frequency density, if the histogram is being described) at the midpoint of each class and joining the points with straight lines. It shows the shape of the distribution and is useful for comparing two data sets on the same axes. The midpoint of the class 20≤x<3020\le x<30 is 2525.

Key termsfrequency polygonmidpoint
Exam tip

Plot at the class midpoint, not at the class boundary.

Section 3

Cumulative frequency diagrams

The cumulative frequency is the running total of frequencies. Plot each total at the upper class boundary and join with a smooth curve (or straight lines for a polygon). For nn values, read off or interpolate to estimate the median at n2\frac{n}{2}, the lower quartile Q1Q_1 at n4\frac{n}{4}, and the upper quartile Q3Q_3 at 3n4\frac{3n}{4}. Linear interpolation within a class: estimate=L+position−Fbelowf×w,\text{estimate}=L+\frac{\text{position}-F_{\text{below}}}{f}\times w, where LL is the lower boundary of the class, ff its frequency and ww its width. The interquartile range is Q3−Q1Q_3-Q_1. Example: 8 plants under 10 cm and 30 under 20 cm gives Q1≈10+25−822×10=17.7Q_1\approx10+\frac{25-8}{22}\times10=17.7.

Key termscumulative frequencyinterquartile rangelinear interpolation
Common mistake

Plotting cumulative frequency at the class midpoint or lower boundary. It must be at the upper boundary.

Section 4

Box and whisker plots and outliers

A box and whisker plot shows the minimum, Q1Q_1, median, Q3Q_3 and maximum. The box covers the interquartile range. Where a rule is given, values beyond the fences are outliers and are plotted separately, with the whiskers stopping at the most extreme values that are not outliers. A common rule: outliers lie below Q1−1.5×IQRQ_1-1.5\times\text{IQR} or above Q3+1.5×IQRQ_3+1.5\times\text{IQR}. For marks with Q1=17Q_1=17, Q3=26Q_3=26: IQR =9=9 and the upper fence is 26+13.5=39.526+13.5=39.5, so 45 is an outlier. To compare two box plots, compare a measure of average (median) and a measure of spread (IQR or range), in context.

Key termsbox and whisker plotquartileoutlier
Exam tip

When comparing two data sets, always give one comment on average and one on spread, each referring to the context.

Section 5

Connection to probability distributions

If you divide each frequency by the total, you get relative frequency, and dividing by class width gives relative frequency density. A histogram drawn like this has total area 1, and the area of a bar is the relative frequency, an estimate of the probability that a randomly chosen value lies in that class. As classes get narrower and the data set larger, the histogram's outline approaches a smooth curve, the probability density function of a continuous distribution. Probability is then the area under the curve. Example: 60 of 200 eggs in a class gives area 0.30.3 for that bar.

Key termsrelative frequencyprobability density
Common mistake

Reading the height of a probability histogram as the probability. The probability is the area.

That's the notes covered.

Carry on to the next subtopic.

Exam questions on Single-variable data: diagrams and histograms

  1. The durations of phone calls made from an office are shown in a histogram. The bar for calls from 2 to 6 minutes has frequency density 7.5, the bar for calls from 6 to 10 minutes represents 18 calls, and the bar for calls from 10 to 20 minutes has frequency density 1.2.
    Estimate the number of calls that lasted between 8 and 12 minutes.2 marks
  2. The marks, out of 50, scored by 11 students in a test were 12, 15, 17, 18, 20, 21, 23, 25, 26, 28 and 45. For this data take the lower quartile to be the median of the lowest five marks and the upper quartile to be the median of the highest five marks. A value is an outlier if it is more than 1.5×1.5\timesIQR above the upper quartile or below the lower quartile.
    Show that the mark of 45 is an outlier.2 marks
  3. The heights of 100 plants are summarised as follows: 8 plants are under 10 cm, 30 are under 20 cm, 65 are under 30 cm, 90 are under 40 cm and all 100 are under 60 cm. Assume the heights are spread evenly within each class.
    Use linear interpolation to estimate the median height of the plants.3 marks
See the full worksheet

Written by the Exaim team, led by Shaun Daswani (Head of Upper Secondary, Improve ME Institute; MSc Financial Mathematics, Imperial College London; BSc, UCL) and Jason Daswani (operational lead, Improve ME Institute; LSE).