Summarise and compare data using measures of centre, spread and box-and-whisker plots.
Practise Statistics in the app →
What gets asked
- Calculate the mean, median or mode of a list of numbers.
- Estimate the mean of grouped data using interval midpoints.
- Find the range, the quartiles, or the interquartile range of a data set.
- Build the five-number summary and draw a box-and-whisker diagram.
You must be able to
- Order data from smallest to largest before finding the median or the quartiles.
- Use the midpoint of each interval when the raw data values are grouped.
- Subtract the lower extreme from the upper extreme to get the range.
- Give the five-number summary in order: minimum, lower quartile, median, upper quartile, maximum.
Traps that cost marks
- Placing the value 20 in the interval , when it belongs in the next interval. Read the inequality carefully. stops just before 20, so 20 goes in the next one.
- Reading as '0 to 19' and getting the wrong midpoint. The midpoint of is , not .
- Halving the upper quartile instead of halving the interquartile range. The semi-interquartile range is half of (upper quartile minus lower quartile).
Worked example
- The data is already ordered, with 6 values.
- The median is the average of the 3rd and 4th values: 9 and 9.
- Median .
Mean, median and mode
Where the middle of a data set sits, which measure to trust, and what one odd value does to it.
- mean
- Add every value, then divide by how many values there are.
- median
- The middle value, once the data is put in order from smallest to largest.
- mode
- The value that appears most often. A set can have none, or more than one.
- outlier
- A value sitting far away from all the others. Not simply the biggest one.
- frequency table
- A table listing each value once, with a count of how often it appears.
The mean shares the total out evenly. Every value pulls on it, so one very large value drags it upwards. The median only counts positions, so it barely moves. That difference decides which one you should report.
A frequency table writes each value once, with its count beside it. Multiply each value by its frequency, add the products, then divide by the total frequency. The number of rows is not the number of values.
Choose the measure that fits the data. When a set holds an outlier, the median describes the group better than the mean. If every value appears exactly once, there is no useful mode. Always say why you chose it.
Adding a value changes the total and the count. Work out the new total, then divide by the new count. An outlier can move the mean a long way and the median hardly at all.
Rules to remember
- mean
Examples
Worked answer
- Add them: .
- Mean: .
- The list is in order already, and there are five values.
- The middle one is the third, so the median is .
- Only repeats, so the mode is .
Answer: mean R, median R, mode R
The single R buy pulls the mean well above the median.
Worked answer
- People in the four homes: .
- Then and .
- Total people: .
- Total homes: .
- Mean: people per home.
Answer: people per home
You divide by the homes, not by the three different values.
Worked answer
- Add the first five: .
- That is , so the mean is .
- The median is the third value, .
- Now the total is , over days.
- New mean: . New median: .
Answer: the mean rises to R, the median to R
The mean moved R and the median only R.
Traps
- Finding the median without putting the values in order first. Order the list from smallest to largest, then count in to the middle.
- Dividing by the number of rows in a frequency table instead of the total frequency. Add the frequencies to see how many values there are. Above that is , not .
- Reporting only the mean for a set that holds one far-off outlier. Give the median too, and say that the outlier pulled the mean up.
- Expecting an outlier to shift the median as far as it shifts the mean. The median only slides along the ordered list, so it shifts very little.
Grouped data
When the data comes in bands instead of single values, you can still get at the middle of it.
- grouped data
- Data sorted into bands, with only a count kept for each band.
- interval
- One band of values, written . The lower end is in, the upper end is not.
- midpoint
- The middle of an interval. Add the two ends and halve the answer.
- modal interval
- The interval carrying the largest frequency.
- cumulative frequency
- A running total of the frequencies, built up down the table.
Grouping keeps a long list readable. A grouped frequency table gives only a count for each interval. The intervals are written as inequalities so that every value has exactly one home. In the sign lets in and the sign keeps out.
Once data is grouped, the raw values are gone. The midpoint stands in for every value in its interval. The midpoint of is . Reading that band as to is what gives the wrong midpoint.
To estimate the mean, multiply each midpoint by its frequency. Add all those products. Then divide by the total frequency, never by the number of intervals.
The answer is only an estimate. Rounding is not what makes it one. The real values were thrown away when the data was grouped. Two different sets can give the same table, and the same estimate.
The modal interval is an interval, not a count. Cumulative frequency is a running total. Add each frequency to the one before it. Read down that column to reach the median, then name the interval it lands in.
Rules to remember
- midpoint
- median position
Examples
Worked answer
- Midpoints: , then , and .
- Multiply by the frequencies: and .
- Then and .
- Add: . Count: .
- Estimated mean: .
Answer: about R
The underneath is the total frequency, not the four intervals.
Worked answer
- The largest frequency is , so the modal interval is .
- Cumulative frequency: , then , then , then .
- Median position: .
- Positions to fill the first two intervals.
- So position lands in the third, .
Answer: modal interval , median in
Grouped data cannot give one median value, so you name its interval.
Worked answer
- The second interval is , and keeps out.
- The third is , and lets in.
- So R is counted in .
Answer:
A value on a boundary joins the interval that starts there.
Traps
- Placing the value 20 in the interval , when it belongs in the next interval. Read the inequality carefully. stops just before 20, so 20 goes in the next one.
- Reading as '0 to 19' and getting the wrong midpoint. The midpoint of is , not .
- Dividing the total by the number of intervals instead of by the total frequency. Four intervals held values, so you divide by .
- Giving the highest frequency as the mode instead of the interval that carries it. The modal interval is , not the count .
Range, quartiles, percentiles
Cutting an ordered data set into quarters, and measuring how far apart the values lie.
- range
- The largest value minus the smallest. One number, not two.
- quartile
- One of the three cuts that split ordered data into four equal-sized parts.
- interquartile range
- The upper quartile minus the lower quartile. It measures the middle half.
- semi-interquartile range
- Half of the interquartile range.
- percentile
- The value that a given percentage of the ordered data sits below.
The range uses only the two end values. Subtract the smallest from the largest, in that order, and give one number. Because it looks at nothing in between, a single far-off value stretches it wide.
The median cuts ordered data in half. The lower quartile is the median of the lower half. The upper quartile is the median of the upper half. If the count is odd, the median joins neither half. Give the value, not its position number.
The interquartile range is the upper quartile minus the lower quartile. It measures the middle half only, so a far-off value cannot stretch it. Halve it and you have the semi-interquartile range.
A percentile works the same way, in hundredths. Take that percentage of the count to find the position. Between two positions, move up to the next one. Landing exactly ON a position, average that value and the next. At this rule gives the median. A question stating a DIFFERENT formula means: use its one.
Grouped data cannot give you one value. Read down the cumulative frequency column until you pass the position you want. Name the interval you stopped in. You are not asked for a value inside it.
Rules to remember
- range largest smallest
- interquartile range
- semi-interquartile range
Examples
Worked answer
- Range: .
- There are nine values, so the median is the fifth, .
- Lower half: ; ; ; . So .
- Upper half: ; ; ; . So .
- The median joined neither half, because the count is odd.
Answer: range ; ; median ;
Each half held four values, so each quartile is the mean of two of them.
Worked answer
- Interquartile range: .
- Semi-interquartile range: .
- The range was , more than twice as wide.
- The single value of stretched the range and left the middle half alone.
Answer: and
Only the middle half is measured, so one far-off value cannot reach it.
Worked answer
- Position: .
- That sits between two positions, so move up to the rd.
- The rd value is .
Answer: marks
A percentile is a mark taken from the list, not the number .
Traps
- Giving the two end values as the range. The range is a single number, the difference: .
- Counting the median into both halves when the number of values is odd. Leave it out of both. The lower half above is ; ; ; .
- Halving the upper quartile instead of halving the interquartile range. The semi-interquartile range is half of (upper quartile minus lower quartile).
- Counting down the frequency column instead of the cumulative frequency column. Only the cumulative column gives positions. Read down that one.
Learn and practise “Range, quartiles, percentiles” in the app →
Comparing data sets
Five numbers that describe a whole data set, and how to use them to compare two of them fairly.
- dispersion
- How spread out the values are. Papers also call this the spread.
- five-number summary
- The minimum, the three quartiles and the maximum, written in that order.
- box-and-whisker diagram
- A picture of the five-number summary drawn above a number line.
- whisker
- A line from one end of the box out to an extreme value.
Two sets can share a centre and still be nothing alike. Dispersion is the difference between them, and there are two ways to measure it. The range uses only the two end values. The interquartile range describes the middle half. Look at both before you decide.
So a bigger range on its own settles nothing. One far-off value stretches a range without touching the middle half. Never set a range against an interquartile range either, because they are not measuring the same thing.
A box-and-whisker diagram is drawn above a number line. The box runs from the lower quartile across to the upper quartile, with a line inside it at the median. One whisker runs from each end of the box out to the smallest and the largest value.
That splits the picture into four parts, and each part holds one quarter of the data. A long whisker means that quarter is spread wide, not that more learners sit in it. Two diagrams can only be compared when the same number line is used for both.
Working out the summaries is half the job. Say what they show about the group. And if the data came to you already grouped, the mean is only an estimate, so write "about" in front of it.
Rules to remember
- five-number summary: minimum, , median, , maximum
- the box holds the middle half, from to
Examples
Worked answer
- There are eight values, and they are already in order.
- Minimum and maximum .
- Median: .
- Lower half ; ; ; , so .
- Upper half ; ; ; , so .
Answer: ; ; ; ;
In that order the five numbers can be read straight off, left to right.
Worked answer
- Both medians are , so the typical wait matches.
- Range A: . Range B: .
- The ranges are equal, so they settle nothing.
- Interquartile range A: . Interquartile range B: .
- B's box is far narrower, so B's waits vary less.
Answer: Route B
The equal ranges hid the difference and the middle half showed it.
Traps
- Calling the set with the bigger range more spread out, without looking at the middle half. One far-off value stretches a range. Check the interquartile range as well.
- Writing the five-number summary out of order, so nobody can tell which value is which. It always runs minimum, lower quartile, median, upper quartile, maximum.
- Reading a long whisker as holding more of the data. Every whisker holds a quarter of the data. A long one is spread wider.
- Using the estimated mean of grouped data as though it were exact. Say "about". The real values were lost when the data was grouped.