Why the centre is not the whole story

In the previous article you learned how to pick a trustworthy centre for a dataset: the mean, the median, or the mode, depending on shape and outliers. A centre answers the question "where is the middle of this data?" It does not answer a second, equally important question: "how spread out is everything around that middle?" Two datasets can share the exact same mean and look completely different in practice.

Take two small teams of five employees, each measured on minutes spent commuting to work. Team Alpha: 20, 22, 24, 26, 28. Team Beta: 4, 12, 24, 36, 44. Add up Team Alpha's numbers and you get 120, so the mean is 120 divided by 5, which is 24 minutes. Add up Team Beta's numbers and you also get 120, so the mean is also 24 minutes. If you only reported the mean, you would describe both teams identically: "average commute of 24 minutes." But anyone managing these teams would experience them very differently. Alpha's commutes are all clustered close to 24 minutes. Beta has someone who walks four minutes and someone who travels 44, nearly eleven times longer, with the same average sitting in between.

This is what spread, also called dispersion or variability, captures: how far the individual values tend to sit from the centre. A dataset with small spread is predictable and consistent. A dataset with large spread has the same centre but a wider range of real experiences hiding behind it. This article walks through the four most common ways to measure that spread: the range, the variance, the standard deviation, and the interquartile range (IQR). Each one answers the same basic question with a different trade-off between simplicity and robustness.

Range: the simplest measure of spread

The range is the distance between the largest and the smallest value in a dataset. In formula form: range equals maximum minus minimum. For Team Alpha, the maximum is 28 and the minimum is 20, so the range is 8 minutes. For Team Beta, the maximum is 44 and the minimum is 4, so the range is 40 minutes. The range immediately confirms the intuition from the commute example: Beta's spread is five times larger than Alpha's, even though both teams have the same average commute.

The range is appealing because it needs almost no calculation and is easy to explain to anyone. It also has two serious limitations. First, it only looks at two data points, the extremes, and ignores everyone in between. A dataset of a thousand values could have 998 of them clustered tightly together with just two extreme outliers, and the range would describe only those two outliers, saying nothing about the other 998. Second, because it depends entirely on the most extreme values, a single unusual measurement, whether a genuine rare event or a data entry error, can inflate or deflate the range dramatically. You will see this effect clearly later in this article with a worked example.

The range is a reasonable first glance at a new dataset, but it is rarely the number you would report in a serious analysis. It is a starting point, not a conclusion.

Measuring spread around the mean: deviations

To build a better measure of spread, start with a simple idea: for each value in the dataset, measure how far it sits from the mean. This distance is called a deviation, and it is calculated as value minus mean. A value above the mean has a positive deviation; a value below the mean has a negative deviation.

Take Team Alpha again: 20, 22, 24, 26, 28, with mean 24. The deviations are:

  • 20 minus 24 equals -4
  • 22 minus 24 equals -2
  • 24 minus 24 equals 0
  • 26 minus 24 equals 2
  • 28 minus 24 equals 4

Notice something: if you add these deviations together, -4 + -2 + 0 + 2 + 4, you get exactly zero. This is not a coincidence specific to this dataset. The deviations from the mean always sum to zero, for any dataset, because the mean is defined as the balance point of the data. This is a useful mathematical fact, but it is also a practical nuisance: if you tried to average the deviations directly to get a single number representing "typical distance from the mean," you would always get zero, which tells you nothing.

There are two common fixes to this problem. One is to take the absolute value of each deviation before averaging, which produces a measure called the mean absolute deviation. It is intuitive and occasionally used, but it has awkward mathematical properties that make it harder to work with in further statistical calculations, so you will not encounter it often in practice. The second fix, and the one that dominates statistics, is to square each deviation before averaging. Squaring has a convenient side effect: it always produces a positive number, regardless of whether the original deviation was positive or negative, so the squared deviations no longer cancel out to zero. This is the idea behind variance.

Variance: the average squared deviation

Variance is defined as the average of the squared deviations from the mean. Working through Team Alpha's squared deviations:

  • (-4) squared = 16
  • (-2) squared = 4
  • 0 squared = 0
  • 2 squared = 4
  • 4 squared = 16

Adding these up gives 16 + 4 + 0 + 4 + 16 = 40. There are 5 values, so the variance is 40 divided by 5, which equals 8. This number, 8, is in "minutes squared," which is an odd unit that nobody thinks in naturally. Hold that thought; it is exactly why standard deviation exists, covered in the next section.

Now do the same for Team Beta: 4, 12, 24, 36, 44, mean 24. The deviations are -20, -12, 0, 12, 20. Squared, these are 400, 144, 0, 144, 400, which sum to 1088. Dividing by 5 gives a variance of 217.6. Compare the two variances directly: Alpha's variance is 8, Beta's is 217.6, more than 27 times larger. Variance amplifies large deviations because of the squaring step, so a team with a few far-flung commute times gets a dramatically larger variance than a team whose commutes cluster tightly, even when both teams share the same mean.

Population variance versus sample variance

There is a detail in the variance formula that trips up almost everyone the first time they meet it: whether you divide by the number of values, n, or by n minus 1. Both versions exist, and the one you should use depends on whether your data is the entire population you care about or only a sample drawn from a larger population.

If Team Alpha and Team Beta are the only two teams you will ever analyse, and your question is only about these two specific teams, you have the full population of interest. In that case you divide by n, the number of values, exactly as done above. This is called the population variance.

But imagine instead that Team Alpha's five commute times are a sample drawn from a much larger company of hundreds of employees, and your real goal is to estimate the variance of commute times across the whole company, not just these five people. In that case, dividing by n tends to produce an estimate that is slightly too small, on average, compared to the true variance of the full company. The reason is subtle but worth understanding: the sample mean is calculated from the very same data whose spread you are measuring, so the data fits its own sample mean a little too snugly. The deviations you calculate are, on average, a bit smaller than the deviations you would get from the true population mean, which you do not actually know. Dividing by n-1 instead of n corrects for this by making the result a little larger, which compensates for that snugness. This adjustment is called Bessel's correction, and the result is called the sample variance.

In practice, the rule of thumb is simple: if your data is the entire group you want to describe, use n (population variance). If your data is a sample standing in for a bigger group you are trying to draw conclusions about, use n-1 (sample variance). Most of the time in data work you are dealing with a sample, even if it does not feel that way, so n-1 is the more commonly used version, and it is the default behaviour in many statistical software packages, as you will see shortly. With a small sample, the difference between n and n-1 can matter quite a bit; with a large sample, say a few thousand values, dividing by n versus n-1 barely changes the answer.

For Team Alpha treated as a sample, the sum of squared deviations was 40, and n-1 is 4, so the sample variance is 40 divided by 4, which equals 10, compared to the population variance of 8 calculated earlier. The sample version is always a bit larger, as expected.

Standard deviation: back to meaningful units

Variance solved the problem of deviations cancelling to zero, but it introduced a new problem: the result is expressed in squared units. Minutes squared, dollars squared, kilograms squared, these are not quantities anyone reasons about naturally. The fix is simple: take the square root of the variance. This brings the number back into the original units of the data, and the result is called the standard deviation.

For Team Alpha, the population variance was 8, so the population standard deviation is the square root of 8, which is approximately 2.83 minutes. For Team Beta, the population variance was 217.6, so the population standard deviation is the square root of 217.6, approximately 14.75 minutes. Now the comparison reads naturally: Alpha's commute times typically sit about 2.83 minutes from the mean of 24, while Beta's typically sit about 14.75 minutes from the same mean of 24. The standard deviation gives you a single number, in the original units, that represents a kind of typical distance from the centre.

It is worth being precise about what "typical distance" means here, because standard deviation is not literally the average of the absolute deviations (that would be the mean absolute deviation mentioned earlier, which gives a slightly different number). Standard deviation is the square root of the average squared deviation. Because of the squaring and unsquaring, it tends to be pulled upward a little by any large deviations in the data, giving more weight to values that sit far from the mean. This is a useful property in many statistical methods you will meet later in this path, particularly once you reach the normal distribution, but it also means standard deviation is sensitive to outliers in a way that will matter later in this article.

A full worked example with an outlier

To see how range, variance and standard deviation behave when a dataset has one unusual value, consider nine exam scores out of 100 from a small class: 55, 58, 60, 61, 62, 63, 65, 67, 99. Eight students scored within a fairly tight band between 55 and 67, and one student scored 99, far above everyone else, perhaps a genuinely exceptional result, perhaps a data entry error worth double-checking. Either way, it is a single value that sits well away from the rest.

Adding the nine scores: 55+58+60+61+62+63+65+67+99 = 590. Dividing by 9 gives a mean of about 65.56.

The range is 99 minus 55, which is 44 points, a large spread driven almost entirely by that one score of 99.

Working out the squared deviations from the mean (65.56) for each score and adding them up gives a sum of approximately 1361. Dividing by 9 (treating this class as the full population of interest) gives a population variance of about 151.25, and the square root of that is a standard deviation of about 12.3 points.

Now remove the score of 99 and look at the remaining eight scores: 55, 58, 60, 61, 62, 63, 65, 67. The new mean is (590-99)/8 = 491/8 = 61.4. Recomputing the squared deviations from this new mean and summing them gives approximately 102. Dividing by 8 gives a population variance of about 12.7, and the square root gives a standard deviation of about 3.6 points.

Line the two versions up side by side. With the score of 99 included, the range is 44 and the standard deviation is about 12.3. Remove that single value, and the range drops to 12 (67 minus 55) and the standard deviation drops to about 3.6. One data point, out of nine, more than tripled the standard deviation and nearly quadrupled the range. This is the practical meaning of saying that range and standard deviation are sensitive to outliers: a small number of extreme values can dominate the entire summary, making the reported spread describe the outlier more than it describes the bulk of the data.

The interquartile range: a measure built to resist outliers

The interquartile range, or IQR, takes a different approach. Instead of looking at the two most extreme values (like the range) or squaring every deviation (like variance and standard deviation), it looks only at the middle half of the data and ignores the top and bottom quarters entirely. The full machinery of quartiles and percentiles is the subject of the next article in this path, but here is enough to compute the IQR by hand.

Sort the data from smallest to largest. The median, also called the second quartile or Q2, splits the sorted data into a lower half and an upper half. Q1, the first quartile, is the median of the lower half. Q3, the third quartile, is the median of the upper half. The interquartile range is then Q3 minus Q1: the width of the middle 50 percent of the data.

Return to the nine exam scores, sorted: 55, 58, 60, 61, 62, 63, 65, 67, 99. With nine values, the middle one (the 5th) is the overall median, which is 62. The lower half is the four values before it: 55, 58, 60, 61. Their median is the average of the two middle values, 58 and 60, which is 59. So Q1 is 59. The upper half is the four values after the median: 63, 65, 67, 99. Their median is the average of 65 and 67, which is 66. So Q3 is 66. The IQR is 66 minus 59, which equals 7.

Now recompute without the outlier, using the eight remaining scores: 55, 58, 60, 61, 62, 63, 65, 67. The lower half is 55, 58, 60, 61, with median (58+60)/2 = 59, so Q1 stays at 59. The upper half is 62, 63, 65, 67, with median (63+65)/2 = 64, so Q3 is now 64. The IQR is 64 minus 59, which equals 5.

Compare how much each measure moved when the outlier of 99 was removed. The range fell from 44 to 12, a drop of 32 points, roughly 73 percent. The standard deviation fell from about 12.3 to about 3.6, a drop of roughly 71 percent. The IQR fell from 7 to 5, a drop of only 2 points, roughly 29 percent. The IQR barely reacted, because it only cares about where the middle 50 percent of the data sits, and that middle block of values barely changed when a single extreme score was added or removed. This is exactly what "robust to outliers" means in practice: a robust measure gives a similar answer whether or not a small number of unusual values are present, because it is not built from the extremes in the first place.

Checking the arithmetic with code

It helps to confirm these hand calculations with a tool, and doing so also exposes a common trap: different libraries default to different versions of variance and standard deviation. The example below assumes Python 3 with numpy 1.26 and pandas 2.x; the exact default behaviour described here has been stable across recent versions of both libraries, but it is always worth checking the documentation for the version you have installed.

🐍Python
import numpy as np
import pandas as pd

scores = [55, 58, 60, 61, 62, 63, 65, 67, 99]

# numpy defaults to the POPULATION formula (divides by n)
print("numpy std, default (population):", np.std(scores))
print("numpy std, ddof=1 (sample):     ", np.std(scores, ddof=1))

# pandas defaults to the SAMPLE formula (divides by n-1)
series = pd.Series(scores)
print("pandas std, default (sample):   ", series.std())
print("pandas std, ddof=0 (population):", series.std(ddof=0))

print("range:", max(scores) - min(scores))

Running this prints four standard deviation values and the range. numpy's np.std with no extra argument divides by n, giving the population standard deviation of about 12.30. pandas' Series.std with no extra argument divides by n-1, giving the sample standard deviation, which comes out a little larger, around 13.05. The ddof argument (degrees of freedom) is how both libraries let you switch between the two: ddof=0 forces the population version, ddof=1 forces the sample version. The two libraries disagree on their default, which is precisely why it matters to know which formula you are getting rather than assuming the first number a tool gives you is the right one for your purpose.

Choosing the right spread measure

With four measures now on the table, the natural question is which one to report. The honest answer is that it depends on what the data looks like and what the audience needs, in much the same way part 3 of this path showed that the choice between mean, median and mode depends on the shape of the distribution and the presence of outliers.

  • Use the range for a fast, first glance at how wide a dataset is, but do not rely on it as your final summary, and always check whether one or two extreme values are driving it.
  • Use variance when you need the quantity for further calculations: many statistical formulas, including ones you will meet later in this path when correlation and the normal distribution are introduced, are built directly on variance.
  • Use standard deviation when you want a single number, in the original units, that most people can interpret as a typical distance from the mean. It works best when the data is roughly symmetric and does not contain extreme outliers.
  • Use the IQR together with the median when the data is skewed or contains outliers that you do not want to let dominate the summary. The pairing of median and IQR is the robust counterpart to the pairing of mean and standard deviation.
  • Whichever spread measure you choose, report it alongside a measure of centre, never alone. "Average commute 24 minutes" and "average commute 24 minutes, standard deviation 15 minutes" describe very different realities, even though the first half of each sentence is identical.

A practical habit worth adopting now, before later articles build on it: when you first see a new dataset, compute the mean and the standard deviation, and also compute the median and the IQR. If the mean and median are close together and the standard deviation does not dwarf the IQR by an unusual amount, the data is probably well-behaved and symmetric, and the mean and standard deviation are a safe summary to report. If the mean and median disagree noticeably, or the standard deviation looks inflated compared to what the IQR suggests, that is a signal that outliers or skew are present, and the median and IQR are the safer pair to lead with. The next two articles in this path, on percentiles and quartiles and then on the shape of distributions, give you the tools to make this judgement with more confidence.

Common mistakes

  • Using the wrong denominator. Dividing by n when the data is a sample meant to estimate a larger population's variance, or dividing by n-1 when the data is the full population, will quietly bias the result. The difference shrinks as the sample grows, but with small datasets it can change the answer noticeably, as the Team Alpha example showed (variance 8 versus 10).
  • Reporting variance when standard deviation was meant, or vice versa. Variance is in squared units and is rarely the number an audience wants to see directly; standard deviation is the interpretable one. Mixing them up produces numbers that look wrong by a large factor, since variance is the square of standard deviation.
  • Comparing standard deviations across variables measured on different scales or in different units without adjusting for scale. A standard deviation of 5 means something very different for a variable ranging from 0 to 10 than for one ranging from 0 to 10,000. When that comparison is genuinely needed, analysts often use the coefficient of variation, the standard deviation divided by the mean, which puts different scales on a comparable footing, though that detail goes beyond what this article cove
  • Assuming a small standard deviation means there are no outliers. A single moderate outlier in a large dataset can be swamped by hundreds of ordinary values and barely move the standard deviation, so a modest standard deviation is not proof that the data is clean.
  • Treating the IQR's robustness as if it meant the outliers do not exist. The IQR deliberately ignores the top and bottom quarters of the data, which is useful for summarising the typical spread, but it is not a tool for detecting or investigating extreme values. A dataset can have a small, stable IQR and still contain values worth examining; the IQR simply will not tell you that by itself. Later in this path, a dedicated article on outliers shows how the IQR is actually used as a detection rule,
  • Quoting the range from a small sample as if it described the full population. With only a handful of observations, the minimum and maximum are themselves unstable; draw a new sample of the same size and they can shift substantially, even if the underlying spread of the full population has not changed at all.

Summary and what comes next

A centre alone never fully describes a dataset; two groups with identical means or medians can differ enormously in how tightly their values cluster. The range gives the fastest, roughest sense of spread but depends entirely on the two extreme values and is easily distorted by a single unusual observation. Variance builds a more complete picture by squaring every deviation from the mean, which avoids the deviations cancelling to zero, but leaves the result in awkward squared units and requires a choice between dividing by n (population) or n-1 (sample, with Bessel's correction). Standard deviation takes the square root of variance to return to the original units, giving an interpretable typical distance from the mean, though it remains sensitive to outliers because of the squaring step underneath it. The interquartile range sidesteps that sensitivity entirely by looking only at the middle 50 percent of the sorted data, trading some information for robustness, as the worked exam-score example showed directly: removing one outlier barely moved the IQR while it reshaped the range and the standard deviation dramatically.

The next article in this path, on percentiles, quartiles and box plots, picks up exactly where the IQR section here left off. It gives the full, precise definition of quartiles and percentiles, including how to compute them properly when the number of values makes the split less tidy than the nine-score example here, and it introduces the box plot as a single picture that displays the median, the quartiles, the IQR and potential outliers all at once.