Where we left off
Part 3 settled the question of which centre to trust: the mean, the median or the mode, depending on how the data is shaped. Part 4 added the question of spread: how far values typically sit from that centre, using the range, variance, standard deviation and the interquartile range (IQR). The IQR was described there as the width of the middle chunk of the data, built from two values called the first and third quartiles. This part explains exactly what a quartile is, generalises the idea to percentiles, and shows how those numbers turn into the box plot, one of the most common charts in data work.
What a percentile really means
Imagine a class of 20 students takes a test and you are told you scored in the 90th percentile. That does not mean you got 90 percent of the questions right. It means that out of everyone who took the test, 90 percent scored at or below you, and only 10 percent scored higher. Your actual mark might have been 65 out of 100 if the test was hard, or 98 out of 100 if it was easy. The percentile tells you your position relative to other people, not your raw score.
This is the core idea behind every percentile: it is a position marker. The pth percentile of a dataset is the value below which roughly p percent of the data falls. The median, which you already met in part 3, is nothing more than the 50th percentile: half the data lies below it, half above. So percentiles are not a new topic bolted onto what came before, they are a generalisation of the median to any cut point you like, not just the halfway one.
Defining a percentile precisely
To find any percentile you first sort the data from smallest to largest. Position only makes sense once the data is in order. Let's work with a small, concrete dataset: the commute times, in minutes, of 11 colleagues on a given morning.
Sorted commute times: 12, 15, 15, 18, 20, 22, 25, 28, 30, 45, 60.
Two common ways to compute a percentile
Here is something that surprises a lot of people the first time they check their work against a spreadsheet or a library: there is more than one accepted way to calculate a percentile, and they can give slightly different answers on the same data. Neither is wrong, they are just different conventions for handling the case where the target position falls between two data points.
The simplest approach, sometimes called the nearest-rank method, rounds the target position to the closest actual data point and reads off its value. A more common approach in statistical software, called linear interpolation, allows the percentile to fall between two data points and blends them proportionally. For our 11 commute times, the position for the 75th percentile under the interpolation method works out to be between the 8th and 9th sorted values, three-quarters of the way from one to the other, giving a value that is not one of the original numbers at all.
You rarely need to do this arithmetic by hand once you are working with real datasets, but it helps to see it once so the numbers a library prints are not mysterious.
import numpy as np
commute_times = np.array([12, 15, 15, 18, 20, 22, 25, 28, 30, 45, 60])
q1 = np.percentile(commute_times, 25)
median = np.percentile(commute_times, 50)
q3 = np.percentile(commute_times, 75)
print(q1, median, q3)
# 16.5 22.0 29.0
This uses numpy's default interpolation setting, which has long produced linear interpolation between neighbouring points. In recent numpy versions the keyword controlling this changed name from interpolation to method, but the default behaviour and the numbers above are unaffected. pandas, when you call a Series's quantile method, uses the same default, so a pandas and a numpy calculation on the same data will normally agree.
Notice that the median computed this way, 22, lands exactly on one of the original data points, because with 11 values the middle one is unambiguous. The first and third quartiles, 16.5 and 29, do not land on original data points at all, because the interpolation method is free to invent a value between two neighbours. That is completely normal and not an error.
Quartiles: the percentiles we use most
A quartile is simply a percentile at one of three specific, commonly used cut points. The first quartile, written Q1, is the 25th percentile: a quarter of the data lies below it. The second quartile, Q2, is the 50th percentile, which is the median. The third quartile, Q3, is the 75th percentile: three quarters of the data lies below it, one quarter above.
Part 4 introduced the IQR as Q3 minus Q1, the width of the middle half of the data, and used a method you will also see in many textbooks: split the sorted data in half at the median, then take the median of the lower half as Q1 and the median of the upper half as Q3. Applying that method to our commute times: the lower half is 12, 15, 15, 18, 20, whose median is 15, so Q1 equals 15. The upper half is 25, 28, 30, 45, 60, whose median is 30, so Q3 equals 30. That gives an IQR of 30 minus 15, which is 15.
Compare that with the numpy calculation above, which gave Q1 as 16.5 and Q3 as 29, an IQR of 12.5. Same data, two different legitimate methods, two different answers. This is exactly the kind of convention gap you need to know about rather than be surprised by. If a teammate's spreadsheet disagrees with your Python output on a quartile, the first question to ask is not who made an arithmetic mistake, it is which method each tool used. The conclusions you draw from the numbers are usually unaffected even when the exact values shift a little, as you will see below.
The five-number summary
Put the minimum, Q1, the median, Q3 and the maximum together and you get what is called the five-number summary, a compact description of a dataset's centre, spread and extremes without needing the mean or the standard deviation at all. Using the median-of-halves method on our commute times:
- Minimum: 12 minutes
- Q1 (first quartile): 15 minutes
- Median (second quartile): 22 minutes
- Q3 (third quartile): 30 minutes
- Maximum: 60 minutes
From these five numbers alone you can already say a lot: half the colleagues commute for 22 minutes or less, the middle 50 percent of commutes fall between 15 and 30 minutes, and someone has a commute of 60 minutes, which is twice the median and well outside where most of the group sits. This is where the box plot comes in, it turns exactly these five numbers into a picture.
From five numbers to a picture: the box plot
A box plot, sometimes called a box-and-whisker plot, draws the five-number summary as a simple shape. A box stretches from Q1 to Q3, so its height (or width, if drawn sideways) is the IQR. A line inside the box marks the median. Two lines called whiskers extend outward from the box toward the minimum and maximum, but with one important twist: the whiskers usually do not reach all the way to the true minimum and maximum. They stop at a boundary, and anything beyond that boundary is drawn as an individual dot rather than being absorbed into the whisker.
The 1.5 times IQR rule for whiskers
The most common convention, used by default in matplotlib, pandas and many other tools, defines two fences based on the IQR. The lower fence sits at Q1 minus 1.5 times the IQR, and the upper fence sits at Q3 plus 1.5 times the IQR. Any data point beyond either fence is flagged individually as a dot on the plot. The whisker itself then stretches only to the most extreme data point that still falls within the fences, not to the fence itself.
Let's work this out for our commute times using the median-of-halves quartiles, Q1 equal to 15 and Q3 equal to 30, so the IQR is 15. The lower fence is 15 minus 1.5 times 15, which is 15 minus 22.5, giving minus 7.5. The upper fence is 30 plus 22.5, giving 52.5. Our smallest value, 12, is comfortably above the lower fence, so the lower whisker simply reaches down to 12, the true minimum. Our largest value, 60, is above the upper fence of 52.5, so it gets drawn as a separate dot rather than being included in the whisker. The upper whisker instead reaches to the largest value that is still within the fence, which is 45.
So the finished box plot for this dataset would show a box from 15 to 30 with a line at 22, a lower whisker down to 12, an upper whisker up to 45, and a single dot at 60 marking that one colleague's unusually long commute. If you instead used the interpolated quartiles from numpy, Q1 equal to 16.5 and Q3 equal to 29, the fences would land at slightly different numbers, roughly minus 2.25 and 47.75, but the conclusion is the same: 60 still sits beyond the upper fence and still gets drawn as a separate point. This is a good general lesson: the exact fence numbers depend on which quartile convention a tool uses, but whether a point counts as unusually high or low is usually stable across the common conventions.
import matplotlib.pyplot as plt
commute_times = [12, 15, 15, 18, 20, 22, 25, 28, 30, 45, 60]
plt.boxplot(commute_times)
plt.ylabel("Commute time (minutes)")
plt.title("Commute times for 11 colleagues")
plt.show()
Running this with matplotlib (any recent 3.x version behaves the same way here) draws one vertical box with the median line inside it, whiskers above and below, and a single marker above the upper whisker for the 60-minute commute. That marker is matplotlib telling you, visually, that this value is unusually far from the rest of the group, using the same 1.5 times IQR logic described above.
Reading a box plot
Once you know how a box plot is built, reading one is mostly about noticing a few basic shapes.
- A short box means the middle 50 percent of the data is tightly clustered; a tall box means it is spread out.
- The median line's position inside the box tells you about lopsidedness: if it sits closer to Q1 than to Q3, more of the middle data is bunched toward the lower end, and vice versa.
- A long whisker on one side with a short one on the other suggests the data stretches out further in that direction, though a careful read of the actual shape of the distribution is better left to the next part in this path.
- Dots beyond the whiskers mark points the 1.5 times IQR rule treats as unusual. They deserve a closer look, not automatic removal; part 8 of this path is entirely about deciding what to do with such points.
Comparing groups side by side
Box plots are especially useful when you want to compare several groups at a glance, because each group's five-number summary collapses into one compact shape that can sit next to the others on the same axis. Suppose a second team's commute times, also for 11 people, sorted, are: 10, 12, 14, 15, 16, 18, 20, 22, 25, 28, 33.
Following the same median-of-halves method: the median is 18, the lower half (10, 12, 14, 15, 16) has a median of 14, so Q1 is 14, and the upper half (20, 22, 25, 28, 33) has a median of 25, so Q3 is 25. The IQR is 11. The fences are 14 minus 16.5, which is minus 2.5, and 25 plus 16.5, which is 41.5. The maximum, 33, sits comfortably inside the upper fence, so this team has no points flagged as unusual.
Placed next to the first team's box plot, the picture tells a clear story without any further statistics: team two has a lower typical commute (median 18 versus 22), a slightly tighter box, and no outlying points, while team one has a longer typical commute and one colleague whose 60-minute commute stands well apart from everyone else. A table of numbers can say the same thing, but the two boxes side by side make the comparison instant, which is exactly why box plots are so common in reports and dashboards.
Common mistakes
- Reading a percentile as a percentage score. The 90th percentile on a test is about rank among test takers, not about getting 90 percent of questions right; the two numbers are unrelated unless you check.
- Trusting extreme percentiles on small samples. With only 11 data points, the 95th or 99th percentile is so close to the maximum that it barely means anything distinct from just reporting the largest value; percentiles need enough data behind them to be stable.
- Assuming every tool computes quartiles the same way. As shown above, the median-of-halves method and linear interpolation can disagree on Q1 and Q3 for the same dataset. Always check which convention a library or colleague is using before comparing numbers across tools.
- Treating every dot beyond a box plot's whisker as an error to delete. The 1.5 times IQR rule is a flag for a closer look, not a verdict; deciding what an unusual value actually means is the subject of part 8.
- Assuming a box plot shows the full shape of a distribution. Two datasets can share an identical five-number summary while looking completely different when you plot every individual point, for example one being smoothly spread out and another having two separate clusters. A box plot summarises; it does not replace looking at the raw shape, which is where part 6 of this path picks up.
Where percentiles show up in data work
- Service level reporting: engineering teams routinely report the 50th, 90th and 99th percentile of response times (often written p50, p90, p99) rather than just the average, because a slow average can hide a small fraction of very slow requests that still frustrate real users.
- Growth and reference charts: fields like pediatrics plot a child's measurement against percentile curves built from large reference populations, so a reading is interpreted relative to peers rather than as an absolute pass or fail number.
- Robust preprocessing: when a dataset has extreme values, some scaling tools centre and scale features using the median and IQR instead of the mean and standard deviation, precisely because quartiles are less disturbed by a handful of extreme points, an idea that connects directly back to the robustness discussion in part 4.
- Percentile-rank features: converting a raw measurement into "what percentile does this fall in" is a common feature engineering step when the raw scale is hard to compare across different groups or time periods.
- Dashboards and A/B testing: many analytics dashboards report several percentiles side by side rather than a single average, because stakeholders often care as much about the tail of the distribution as about the typical case.
Summary and what comes next
A percentile marks a position in sorted data: the pth percentile has roughly p percent of the data at or below it. Quartiles are just the 25th, 50th and 75th percentiles, with the median as the special case you already met in part 3. The five-number summary, minimum, Q1, median, Q3 and maximum, condenses a dataset into a shape you can draw: the box plot, with its box spanning the IQR, a line at the median, whiskers reaching to the most extreme non-flagged points, and individual dots marking values beyond the 1.5 times IQR fences. Different tools can compute quartiles slightly differently, and that is a convention issue to check for, not a mistake to hunt down. Next, part 6 moves from summarising a dataset with a handful of numbers to describing its overall shape in words: whether it leans to one side, how heavy or light its tails are, and what that tells you before you ever assume a dataset looks like a bell curve.
Comments (0)
No comments yet. Be the first to share your thoughts.