Measures of Variability: Range, Variance and Standard Deviation

By Dr. Zubair Khalid, DVM, MS, PhD ·

Measures of Variability: Range, Variance and Standard Deviation

Variability describes how spread out the values in a dataset are. Two datasets can share the same mean and still behave very differently, so a measure of variability is what tells you whether scores cluster tightly or scatter widely. The three most common measures are the range, the variance and the standard deviation.

Quick Answer

  • Variability is the dispersion or differences among scores in a dataset [1].
  • The range is the difference between the highest and lowest values [2].
  • The variance is roughly the arithmetic average of the squared distance from the mean [3].
  • The standard deviation is the square root of the variance and returns to the original units.
  • A small standard deviation means good precision, a large one means the values are widely spread [4].

What Variability Means

In plain terms, variability is how much the values in a dataset differ from each other and from their center. If every score were identical, there would be no variability at all. The more the scores differ, the larger the variability.

The precise statistical definition is the dispersion or differences among scores or qualitative responses in a data set [1]. Measures of variability bring specificity by summarizing that dispersion, showing whether a group of scores tended to be similar or different from each other [1]. This matters because much of science focuses on means, yet the conclusion drawn from a mean alone can change once variability is considered [5]. For a fuller treatment of the center, see mean and standard deviation.

How It Works

Each measure captures spread in a different way.

Range

$$\text{Range} = \text{maximum} - \text{minimum}$$

The range uses only the two most extreme points in the data [3]. It is fast to compute and easy to explain, but a single outlier can dominate it.

Variance

$$s^2 = \frac{\sum_{i=1}^{N}(Y_i - \bar{Y})^2}{N - 1}$$

Here $Y_i$ is each individual value, $\bar{Y}$ is the sample mean, and $N$ is the number of values. The numerator sums the squared distance of every value from the mean. Squaring gives greater weight to values further from the mean [3]. Dividing by $N - 1$ instead of $N$ corrects for the fact that a sample tends to underestimate the spread of the population it came from.

Standard deviation

$$s = \sqrt{s^2}$$

The standard deviation is the square root of the variance. Because it undoes the squaring, it is expressed in the same units as the original data, which makes it easier to interpret than the variance. Precision is quantified by a standard deviation, so good precision implies a small standard deviation [4].

Worked Example

Two classes of 8 exam scores have the same mean but different spread.

StudentClass AClass B
S17255
S27562
S37870
S48080
S58285
S68592
S78898
S890108

Step 1: Find the mean. Both classes sum to 650 with $n = 8$, so the mean is $650 / 8 = 81.2500$ for each.

Step 2: Find the range. Class A: $90 - 72 = 18$. Class B: $108 - 55 = 53$.

Step 3: Find the deviations from the mean.

  • Class A: -9.2500, -6.2500, -3.2500, -1.2500, 0.7500, 3.7500, 6.7500, 8.7500
  • Class B: -26.2500, -19.2500, -11.2500, -1.2500, 3.7500, 10.7500, 16.7500, 26.7500

Step 4: Square each deviation.

  • Class A: 85.5625, 39.0625, 10.5625, 1.5625, 0.5625, 14.0625, 45.5625, 76.5625
  • Class B: 689.0625, 370.5625, 126.5625, 1.5625, 14.0625, 115.5625, 280.5625, 715.5625

Step 5: Sum the squares. Class A: $SS = 273.5000$. Class B: $SS = 2313.5000$.

Step 6: Divide by $n - 1$ to get the variance. Class A: $273.5000 / 7 = 39.0714$. Class B: $2313.5000 / 7 = 330.5000$.

Step 7: Take the square root for the standard deviation. Class A: $\sqrt{39.0714} = 6.2507$. Class B: $\sqrt{330.5000} = 18.1797$.

The two classes have identical means, yet Class B is far more spread out. You can reproduce this with the code below or check your own numbers with the variance calculator.

import pandas as pd
df = pd.DataFrame({'Class A': [72, 75, 78, 80, 82, 85, 88, 90], 'Class B': [55, 62, 70, 80, 85, 92, 98, 108]})
print(df.agg(['mean','var','std','min','max']))

Output:

Class A: mean=81.2500, range=18, var=39.0714, sd=6.2507
Class B: mean=81.2500, range=53, var=330.5000, sd=18.1797

How to Interpret It

A small standard deviation means the values sit close to the mean. A large one means they are widely dispersed. In the example, Class A has $s = 6.2507$, so a typical score is about 6 points from the mean of 81.25. Class B has $s = 18.1797$, so a typical score is about 18 points away.

The range is the quickest read on total spread. Class B spans 53 points while Class A spans 18, which already signals that Class B is more variable. The variance is harder to interpret directly because its units are squared, but it is the building block for many other statistics, including the correlation and covariance you will meet later.

When central tendency and variability are both known, different sets of data can be compared [2]. Reporting both is necessary to characterize a distribution [2].

When to Use It (and when not to)

Use the range for a fast, rough sense of spread, especially when you want the full span of the data.

Use the variance when you are doing further calculations, such as analysis of variance or building a covariance matrix.

Use the standard deviation when you want to describe spread in the original units, compare two groups, or report precision [4].

Avoid relying on the range alone when outliers are present, because it uses only the two most extreme points [3]. Avoid the standard deviation for heavily skewed data without checking, since squaring gives greater weight to values far from the mean [3]. For skewed data, the interquartile range or median absolute deviation may describe spread better, since those measures do not give undue weight to the tails [3].

Variability vs Central Tendency

Central tendency tells you where the middle of the data sits. Variability tells you how tightly the data cluster around that middle. They answer different questions and belong together.

AspectCentral TendencyVariability
Question answeredWhere is the center?How spread out is the data?
Common measuresMean, median, modeRange, variance, standard deviation
UnitsOriginal unitsOriginal units (SD), squared units (variance)
Effect of outliersMean shifts, median resistsRange and SD react strongly
Typical useDescribe a typical valueDescribe consistency or precision

Two distributions with the same mean may still differ completely in shape, which is why variability matters [2]. For more on choosing a center, see mean vs median.

Common Mistakes

  • Reporting the variance as if it were in original units. The variance is in squared units. Take the square root to get the standard deviation before interpreting it.
  • Dividing by $N$ instead of $N - 1$ for a sample. Using $N$ underestimates the population variance. Use $N - 1$ for sample data.
  • Trusting the range with outliers. One extreme value inflates the range. Check the data or use the interquartile range instead.
  • Comparing standard deviations across different units. A standard deviation of 5 means different things in dollars and in kilograms. Compare only like with like.
  • Assuming a small standard deviation means the data are accurate. Precision and accuracy are different. A precise instrument can still be biased, as covered in accuracy vs precision.
  • Ignoring variability when comparing means. A conclusion based on the mean alone can change once variability is considered [5].

Limitations

These measures summarize spread with a single number, so they hide the shape of the distribution. Two datasets can share the same standard deviation while one is symmetric and the other is skewed. The variance and standard deviation also give greater weight to values far from the mean, so a few extreme points can dominate the result [3].

The range depends entirely on two values and tells you nothing about the middle of the data [3]. None of these measures reveal where the spread comes from. In measurement work, variability can arise from short-term sources tied to instrument precision and from long-term sources tied to changes in environment and handling, and a single number will not separate them [4]. Variability is a human and process reality, and if it is not handled correctly it can lead to incorrect conclusions [6].

Frequently Asked Questions

What is the difference between variance and standard deviation?

The variance is the average squared distance from the mean. The standard deviation is the square root of the variance. The standard deviation is easier to interpret because it is in the same units as the data, while the variance is in squared units.

Why do we divide by n minus 1 for the variance?

Dividing by $N - 1$ corrects for the tendency of a sample to underestimate the spread of the population. Using $N$ would give a biased estimate that is too small on average. The $N - 1$ version is the standard sample variance.

Can two datasets have the same mean but different variability?

Yes. In the worked example, both classes have a mean of 81.2500, but Class A has a standard deviation of 6.2507 and Class B has 18.1797. The mean alone would hide this difference entirely.

Is a higher standard deviation always bad?

Not necessarily. A high standard deviation means the values are widely spread. Whether that is good or bad depends on context. In manufacturing you usually want low variability, but in some research a wide spread is exactly what you are studying.

What is the range used for?

The range gives a quick sense of the total span of the data. It is the difference between the highest and lowest values [2]. It is useful for a fast summary but sensitive to outliers, so pair it with the standard deviation or interquartile range.

References

  1. 4.1: Variability - Statistics LibreTexts/04%3A_Measures_of_Variability/4.01%3A_Variability)
  2. Variability
  3. 1.3.5.6. Measures of Scale
  4. 2.1.1.4. Variability
  5. Wensink MJ, Ahrenfeldt LJ, Möller S. (2020). Variability Matters. International journal of environmental research and public health
  6. Li H, Chen Z, Zhu W. (2019). Variability: Human nature and its impact on measurement and statistical analysis. Journal of sport and health science

Related Articles