Standard Deviation vs Variance vs Standard Error: What Each Measures and When to Report It
By Dr. Zubair Khalid, DVM, MS, PhD ·

Every quantitative result you produce in a lab or a clinical dataset comes with two questions attached: how spread out are the measurements, and how precisely have you pinned down the average? Variance, standard deviation (SD) and standard error of the mean (SEM) answer those questions, but they answer different ones. Mixing them up is one of the most common reporting errors in the biomedical literature, and it changes what your figure or table actually claims.
You will meet these three quantities constantly. Software output gives you a variance by default. Journal figures show error bars that could be SD, SEM or a confidence interval, and the legend often does not say which. Reviewers ask you to justify your choice. This article explains what each number measures, how to calculate it, how to read it, and which one belongs in a given sentence or figure.
Quick Answer
- Variance is the average squared distance of observations from the mean. It is in squared units, so a variance of 2.328 g² is hard to interpret directly [4].
- Standard deviation is the square root of the variance. It is back in the original units and describes how widely the data are scattered [4][6].
- Standard error of the mean is the SD divided by the square root of the sample size: SE = SD / √n. It measures the precision of the sample mean, not the spread of the data [2].
- Use the sample formula (divide by n - 1) when your data are a sample; use the population formula (divide by N) only when you truly have every member of a population [6][7][8].
- To describe how variable your measurements are, report the SD. To describe how precisely you have estimated the mean, report the SEM or a confidence interval [2].
- Always label error bars and state n in the figure legend. A bare "mean ± value" tells the reader nothing about which quantity it is [1][2].
Definitions and Intuition
Start with the mean, the balance point of your data. Variance asks a simple question: on average, how far is each observation from that balance point? Distances are squared before averaging, which does two things. It stops positive and negative deviations from canceling out, and it gives extra weight to points far from the mean. NIST describes the variance as roughly the arithmetic average of the squared distance from the mean [4].
The cost of squaring is that the units change. If your measurements are in grams, the variance is in grams squared. Lee, In and Lee note that this squared unit is exactly what makes variance confusing to interpret, which is why the SD, which shares the units of the mean, is usually the better descriptive number [3].
The standard deviation is the square root of the variance, so it returns to the original units. OpenStax defines it as a number that measures how far data values are from their mean [6]. Altman and Bland call it a measure of variability [2]. If your mice weigh 23.6 g on average with an SD of 1.5 g, the SD is directly comparable to the mean and to the individual weights.
The standard error of the mean is a different animal. It is not a property of the data points; it is a property of your estimate of the mean. Altman and Bland describe it as a measure of the precision of the sample mean [2]. Lee, In and Lee define it as the standard deviation of the sampling distribution of sample means, estimated in practice from a single sample as SD / √n [3]. A small SEM means your mean is well pinned down. It says nothing about how spread out the individual observations are.
How to Calculate Each One
The sample variance uses the n - 1 denominator:
$$s^2 = \frac{\sum_{i=1}^{N}(Y_i - \bar{Y})^2}{N - 1}$$
Here $Y_i$ are the individual observations, $\bar{Y}$ is the sample mean, and $N$ is the sample size [4]. The sample standard deviation is its square root:
$$s = \sqrt{\frac{\sum_{i=1}^{N}(Y_i - \bar{Y})^2}{N - 1}}$$
The population versions divide by N instead of n - 1. OpenStax gives the population standard deviation as $\sigma = \sqrt{\sum (x - \mu)^2 / N}$, where $\mu$ is the population mean and $N$ is the population size [6]. OpenStax explains that the sample variance estimates the population variance, and dividing by n - 1 instead of n gives a better estimate [6].
The standard error follows directly from the SD:
$$SE = \frac{s}{\sqrt{n}}$$
where $s$ is the sample SD and $n$ is the sample size [2]. Because n sits under a square root, the SEM shrinks slowly as you collect more data. Quadrupling the sample size halves the SEM. Holding the SD fixed at 1.526 g, the SEM is 0.623 g at n = 6, 0.311 g at n = 24 and 0.156 g at n = 96.
For a confidence interval around the mean, NIST gives the limits as:
$$\bar{Y} \pm t_{(1-\alpha/2,\ N-1)} \times \frac{s}{\sqrt{N}}$$
where $t$ is the percentile of the t distribution with N - 1 degrees of freedom [5]. For large samples, Altman and Bland note that a 95% confidence interval is approximately the mean ± 1.96 × SE [2]. NIST points out that the interval narrows as sample size grows because of the √N term [5].
If you want to check your arithmetic, use the site's Standard Deviation Calculator.
Worked Example
Six mice are weighed, giving 22.1, 24.5, 23.0, 25.8, 21.9 and 24.3 g. The mean is 141.6 / 6 = 23.60 g. The sum of squared deviations from the mean is 11.64 g².
The sample variance divides by n - 1 = 5:
$$s^2 = \frac{11.64}{5} = 2.328\ \text{g}^2$$
The population variance divides by n = 6, giving 11.64 / 6 = 1.940 g². The sample SD is √2.328 = 1.526 g, and the population SD is √1.940 = 1.393 g. In Excel, STDEV.S returns 1.526 and STDEV.P returns 1.393 [7][8].
The SEM is 1.526 / √6 = 0.623 g. For a 95% confidence interval, use t(0.975, 5) = 2.571: 23.60 ± 2.571 × 0.623 = 23.60 ± 1.60, or 22.00 to 25.20 g.
| Quantity | Value | Units | What it tells you |
|---|---|---|---|
| Mean | 23.60 | g | Center of the data |
| Sample variance | 2.328 | g² | Average squared spread |
| Sample SD | 1.526 | g | Typical distance from the mean |
| SEM | 0.623 | g | Precision of the mean estimate |
| 95% CI | 22.00 to 25.20 | g | Plausible range for the true mean |
Two honest ways to report this: "mean 23.6 g (SD 1.5 g, n = 6)" if you are describing the animals, or "mean 23.6 g (95% CI 22.0 to 25.2 g)" if you are describing how precisely you know the mean. Note that the SD carries units of g while the variance carries g². If the same SD held with n = 24, the SEM would be 0.311 g; with n = 96, 0.156 g.
Reading and Interpreting Each Number
The SD is a descriptive statistic. For normally distributed data, about 68.3% of values lie within 1 SD of the mean, 95.4% within 2 SD and 99.7% within 3 SD. Cumming, Fidler and Vaux classify SD error bars as descriptive because they show how the data are spread, and note that SD bars include about two thirds of the sample while 2 × SD bars encompass roughly 95% [1].
The SEM is an inferential statistic. It tells you how much your sample mean would bounce around if you repeated the experiment many times. It does not tell you where individual observations fall. Because the SEM is always smaller than the SD, Lee, In and Lee warn that authors are tempted to report it when describing samples, which can mislead readers about how variable the data really are [3].
The behavior of the two quantities under increasing sample size is the cleanest way to keep them apart. Altman and Bland state that as sample size increases the standard error falls, while the SD does not systematically change [2]. If you add more mice to a study, the weights stay just as variable, but your estimate of the average weight gets sharper.
SD vs SEM Error Bars: What Your Figure Is Claiming
Error bars are where these distinctions become visible to readers, and where ambiguity causes real damage. Altman and Bland warn that the notation "mean ± value" gives no indication whether the second figure is an SD or an SE, so it must be labeled [2]. Cumming and colleagues make this their first rule: when showing error bars, always describe in the figure legend what they are [1].
Their remaining rules are worth knowing before you build a figure. State n in the legend and distinguish independent experiments from technical replicates [1]. Show error bars and statistics only for independently repeated experiments, never for replicates; a single representative experiment has n = 1 and should carry no error bars or P values [1]. When comparing results with controls, inferential error bars (SE or CI) are usually more appropriate than SD, and if n is very small, such as n = 3, plot the individual data points instead [1].
If you use SE bars, the conversion to a confidence interval depends on n. Cumming and colleagues note that SE bars can be doubled to approximate a 95% CI when n is 10 or more, but at n = 3 the SE must be multiplied by about 4 [1]. The multiplier is the t critical value: t(0.975, 2) = 4.303 for n = 3, 2.262 for n = 10 and 2.045 for n = 30. Their example makes the arithmetic concrete: measurements 28.7, 38.7 and 52.6 give a mean of 40.0, an SD of 12.0 and an SE of 6.93 [1].
Their overlap rules are approximations for two independent groups of similar n and similar SE. With n of 10 or more, a gap of 1 SE between SE bars suggests P around 0.05 and a gap of 2 SE suggests P around 0.01; at n = 3, if doubled SE bars do not overlap, P < 0.05 [1]. With 95% CIs and n = 3, overlap of one full arm suggests P around 0.05 and overlap of half an arm suggests P around 0.01 [1]. These rules do not apply to paired or repeated-measures data, which their Rule 8 addresses directly: for repeated measurements on the same group, SE bars or CIs are irrelevant to within-group comparisons [1].
Common Mistakes
- Reporting SEM while describing variability. The SEM is smaller than the SD and shrinks with n, so it makes scattered data look tidy. If the sentence is about how variable the measurements are, use the SD [2][3].
- Writing "mean ± SEM" without saying so. A reader cannot tell an SD from an SE by looking at the number. Label the statistic in the text and the figure legend [1][2].
- Using the population formula on sample data. Dividing by n instead of n - 1 underestimates the population variance. Use STDEV.S and VAR.S for samples, and reserve STDEV.P and VAR.P for data that truly represent an entire population [7][8][9].
- Treating the variance as directly interpretable. A variance of 2.328 g² does not mean anything in grams. Take the square root before you describe spread to a reader [3][4].
- Putting error bars on technical replicates. Error bars and P values belong to independently repeated experiments. A single representative run has n = 1 [1].
- Assuming a small SEM means the data are homogeneous. A tiny SEM with a large SD just means you collected a lot of variable observations. The two facts coexist.
- Forgetting that the sample SD is still slightly biased. The n - 1 divisor makes the sample variance unbiased for the population variance, but the square root of an unbiased estimator is not itself unbiased. The sample SD tends to run a little low, which matters most at small n.
Limitations
The 68.3, 95.4 and 99.7 percentages assume a normal distribution. For skewed data, the SD alone is a poor summary, and a median with an interquartile range often describes the data better. The SD is also sensitive to outliers, since deviations are squared before averaging.
The n - 1 correction is standard practice, but the claim that it makes the sample variance unbiased for the population variance is standard textbook theory, not something the cited sources state. The sample SD remains slightly biased low even after the correction.
Cumming's overlap rules are approximations for two independent groups with similar n and similar SE. They do not extend to paired designs or repeated measures, and they should not be treated as a substitute for the actual test.
There is genuine disagreement in the literature about when the SEM is acceptable for comparing groups. Lee, In and Lee state that either SEM or SD can be used to compare groups of equal size, but that position is debated and is not settled guidance.
Finally, the choice between the sample and population formula is not always obvious in practice. If your "population" is really a convenience sample, the sample formula is the honest choice even when the dataset feels complete.
Frequently Asked Questions
What is the difference between standard deviation and variance?
Variance is the average squared distance from the mean, and standard deviation is its square root. The SD is in the same units as your data, while the variance is in squared units, which is why the SD is usually easier to report and interpret [3][4].
What is standard deviation vs standard error?
The SD describes how spread out the observations are. The SEM describes how precisely the sample mean estimates the population mean, and it equals SD / √n [2]. As n grows, the SEM falls while the SD stays roughly the same [2].
How do I calculate standard deviation by hand?
Subtract the mean from each observation, square each deviation, sum them, divide by n - 1 for a sample, then take the square root [4][6]. The worked example above walks through each step with real numbers.
When should I report SD vs SEM error bars?
Use SD bars when the point of the figure is to show how variable the data are. Use SEM or confidence interval bars when the point is to show the precision of a mean or to compare groups with a control [1][2]. Either way, name the statistic in the legend and state n [1].
What is the difference between sample and population standard deviation?
The sample SD divides the squared deviations by n - 1, while the population SD divides by N [6]. The n - 1 version gives a better estimate of the population variance when you are working from a sample, and the two converge as n gets large [6][7][8].
References
- Cumming, Fidler & Vaux 2007, Error bars in experimental biology, J Cell Biol 177:7-11
- Altman & Bland 2005, Standard deviations and standard errors, BMJ 331:903
- Lee, In & Lee 2015, Standard deviation and standard error of the mean, Korean J Anesthesiol 68:220-223
- NIST/SEMATECH e-Handbook 1.3.5.6 Measures of Scale
- NIST/SEMATECH e-Handbook 1.3.5.2 Confidence Limits for the Mean
- OpenStax Introductory Statistics 2e, 2.7 Measures of the Spread of the Data
- Microsoft Support, STDEV.S function
- Microsoft Support, STDEV.P function
- Microsoft Support, VAR.S function