# Sample Variance Equation: Formula, Steps and Examples

The sample variance equation measures how far the values in a sample spread around their own mean. You subtract the mean from each value, square those deviations, add them up, and divide by $n - 1$. The result, written $s^2$, is the average squared distance from the sample mean, and it estimates the variance of the wider population the sample came from [1].

## Quick Answer

- Sample variance is $s^2 = \dfrac{\sum (x - \bar{x})^2}{n - 1}$, where $x$ is each value, $\bar{x}$ is the sample mean, and $n$ is the sample size [2].
- The numerator, $\sum (x - \bar{x})^2$, is the sum of squared deviations from the mean.
- The denominator is $n - 1$, not $n$. Dividing by $n - 1$ gives an unbiased estimate of the population variance, while dividing by $n$ tends to underestimate it [3].
- Variance is in squared units. Its square root, the standard deviation, is back in the original units [2].
- For the five reaction times below, $s^2 = 0.2650$ seconds squared and $s = 0.5148$ seconds.

## The Formula

The sample variance equation is:

$$s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n - 1}$$

Each symbol means something specific [2]:

| Symbol | Meaning |
|---|---|
| $x_i$ | Each individual value in the sample |
| $\bar{x}$ | The sample mean, $\frac{\sum x_i}{n}$ |
| $n$ | The number of values in the sample |
| $x_i - \bar{x}$ | The deviation of one value from the mean |
| $(x_i - \bar{x})^2$ | The squared deviation, which removes the sign |
| $\sum$ | Sum across all $n$ values |
| $s^2$ | The sample variance |

The population variance uses a different symbol and denominator: $\sigma^2 = \frac{\sum (x - \mu)^2}{N}$, where $\mu$ is the population mean and $N$ is the population size [2]. The sample version replaces $\mu$ with $\bar{x}$ and $N$ with $n - 1$.

Squaring each deviation matters because raw deviations always sum to zero. Positive and negative deviations cancel, so you square them first to keep the size of the spread and drop the direction.

## How to Calculate It Step by Step

1. Count your values to get $n$.
2. Add all values and divide by $n$ to get the sample mean $\bar{x}$. If you want a refresher on that step, see [sample mean](/blog/data-analysis/sample-mean).
3. Subtract $\bar{x}$ from each value to get its deviation.
4. Square each deviation.
5. Add the squared deviations to get the sum of squares, often written $SS$.
6. Divide $SS$ by $n - 1$. That is your sample variance $s^2$.
7. Take the square root if you want the standard deviation $s$, which is in the original units.

## Worked Example

A lab records five reaction times in seconds: 12.4, 13.1, 11.8, 12.9, 12.3. We want the sample variance of these five measurements.

| Value $x$ | Deviation $x - \bar{x}$ | Squared deviation |
|---|---|---|
| 12.4 | -0.1000 | 0.0100 |
| 13.1 | 0.6000 | 0.3600 |
| 11.8 | -0.7000 | 0.4900 |
| 12.9 | 0.4000 | 0.1600 |
| 12.3 | -0.2000 | 0.0400 |

The sample mean is:

$$\bar{x} = \frac{12.4 + 13.1 + 11.8 + 12.9 + 12.3}{5} = 12.5000$$

The sum of squared deviations is:

$$SS = 0.0100 + 0.3600 + 0.4900 + 0.1600 + 0.0400 = 1.0600$$

The sample variance is:

$$s^2 = \frac{1.0600}{5 - 1} = \frac{1.0600}{4} = 0.2650$$

The sample standard deviation is:

$$s = \sqrt{0.2650} = 0.5148$$

If you divided by $n$ instead, you would get $1.0600 / 5 = 0.2120$. That value underestimates the spread, which is exactly the bias that $n - 1$ corrects [3].

## How to Interpret the Result

The sample variance is 0.2650 seconds squared. Because the units are squared, that number is hard to read on its own. The standard deviation, 0.5148 seconds, tells you that a typical reaction time sits about half a second from the mean of 12.5 seconds.

A larger $s^2$ means the values are more spread out. A value near zero means the observations are tightly clustered. Variance is always zero or positive, since squared deviations cannot be negative.

The sample variance is an estimate, not the population value. Different samples from the same population give different $s^2$ values because of sampling error. Across many samples, the average of the $n - 1$ variances equals the true population variance, while the average of the $n$ variances falls short of it [3]. That is the practical meaning of "unbiased."

Variance also feeds into other work. It appears in standard errors, confidence intervals, hypothesis tests, and sample size planning, where a larger variance means you need more observations to reach the same precision. See [sample size standard deviation formula](/knowledge/diagnostics/research-methods/estimating-variability-how-to-use-standard-deviation-and-variance-in-sample-size-formulas) for how that plays out.

## Doing It in Software

In Excel, `VAR.S` returns the sample variance and `STDEV.S` returns the sample standard deviation. The `VAR.P` and `STDEV.P` functions divide by $N$ and are for population data. For the reaction times above, `VAR.S` gives 0.265 and `STDEV.S` gives 0.5148. There is a full walkthrough in [how to calculate variance in Excel](/blog/data-analysis/how-to-calculate-variance-in-excel).

In Python, NumPy's `np.var` divides by $n$ by default, so you must pass `ddof=1` to get the sample variance. The same applies to `np.std`.

```python
import numpy as np
data = [12.4, 13.1, 11.8, 12.9, 12.3]
s2 = np.var(data, ddof=1)
s  = np.std(data, ddof=1)
print(s2, s)  # 0.2650 0.5148
```

Output:

```text
0.2649999999999997 0.5147815070493497
```

In R, `var(x)` returns the sample variance and `sd(x)` returns the sample standard deviation. Both use $n - 1$ by default, so no extra argument is needed.

If you would rather not write code, the [variance calculator](/tools/variance-calculator) takes a list of values and returns the sample variance and standard deviation.

## Common Mistakes

- **Dividing by $n$ instead of $n - 1$.** This gives the population variance for your sample and underestimates the population spread. Use $n - 1$ whenever your data is a sample from a larger group [3].
- **Forgetting to square the deviations.** Raw deviations sum to zero, so the result would always be zero. Square each one before adding.
- **Using the population mean $\mu$ in the sample formula.** You almost never know $\mu$. Use the sample mean $\bar{x}$ and let $n - 1$ handle the correction.
- **Mixing up variance and standard deviation.** Variance is in squared units. If you need the original units, take the square root. The comparison in [standard deviation vs variance vs standard error](/blog/research-skills/standard-deviation-vs-variance-vs-standard-error) covers when each is reported.
- **Rounding too early.** Keep full precision through the squared deviations. Rounding the mean before squaring can shift the final answer.
- **Treating $n - 1$ as a rule for all denominators.** It applies to the sample variance. Other statistics, such as the population variance, divide by $N$ [2].

## Limitations

The sample variance is sensitive to outliers. One extreme value gets squared, so it can dominate the sum and inflate $s^2$ far beyond what the rest of the data suggests. If your data has heavy tails or clear outliers, the variance may describe the outlier more than the typical spread.

Variance also assumes the values are on an interval or ratio scale, where differences are meaningful. For ordinal data or categories, variance is not a sensible summary. And because it is in squared units, it is rarely reported alone. Most analyses report the standard deviation, which is easier to interpret, or use the variance as an input to a test rather than as a headline number. When the spread changes with the level of the data, a transformation may be needed, as described in [variance stabilizing transformation](/blog/research-skills/variance-stabilizing-transformation-when-and-how).

## Frequently Asked Questions

### Why do we divide by n - 1 in the sample variance equation?

Dividing by $n - 1$ makes the sample variance an unbiased estimate of the population variance. If you divide by $n$, the average of the sample variances across repeated samples falls below the true population variance. With $n - 1$, that average lands exactly on it [3]. The correction is larger for small samples and shrinks as $n$ grows.

### What is the difference between sample variance and population variance?

Population variance uses every member of the group and divides by $N$. Sample variance uses a subset and divides by $n - 1$ [1]. The sample value is an estimate of the population value, and the two formulas differ only in the denominator and the mean they use [2].

### Is sample variance the same as standard deviation?

No. The standard deviation is the square root of the variance. Variance is in squared units, while the standard deviation is in the original units, which makes it easier to interpret [2]. For the reaction times, $s^2 = 0.2650$ and $s = 0.5148$.

### Can sample variance be negative?

No. Every squared deviation is zero or positive, so the sum of squares is never negative. Dividing by $n - 1$ keeps it non-negative. A variance of zero means every value in the sample is identical.

### What does a high sample variance mean?

A high variance means the values are spread far from the mean. In the reaction time example, a variance much larger than 0.2650 would signal less consistent timing. High variance also widens confidence intervals and increases the sample size needed to detect an effect, which is why it matters in planning studies like those in [sample size calculation](/blog/guides/sample-size-calculation-formulas-and-practical-considerations).

## References

1. [Variance - Wikipedia](https://en.wikipedia.org/wiki/Variance)
2. [4.5: Variance - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Introductory_Statistics/Statistics%3A_Open_for_Everyone_(Peter)/04%3A_Measures_of_Variability/4.05%3A_Variance)
3. [4.5 - Why Are the Variance Formulas Different? - Introduction to Statistics and Statistical Thinking](https://open.maricopa.edu/haasstatistics/chapter/4-5/)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)

## Related Articles

- [How to Calculate Variance: Formula, Steps and Examples](/blog/data-analysis/how-to-calculate-variance)
- [Coefficient of Variation: Formula and Examples](/blog/data-analysis/coefficient-of-variation-formula)
- [Covariance Formula: Definition and Calculation Examples](/blog/data-analysis/covariance-formula-definition)
- [Sample Mean: Definition, Formula and Examples](/blog/data-analysis/sample-mean)
- [Standard Deviation of a Binomial Distribution: Formula and Example](/blog/data-analysis/standard-deviation-binomial-distribution)
- [Sample size standard deviation formula](/knowledge/diagnostics/research-methods/estimating-variability-how-to-use-standard-deviation-and-variance-in-sample-size-formulas)
- [Sample Size Calculation: Formulas and Practical Considerations](/blog/guides/sample-size-calculation-formulas-and-practical-considerations)
- [Poisson Distribution: Formula and Examples](/blog/research-skills/poisson-distribution-formula-and-examples)