# Mean vs Median: Differences and When to Use Each

Mean vs median is the choice between two measures of central tendency that answer slightly different questions. The mean is the arithmetic average of all values, while the median is the middle value once the data is sorted. They agree on symmetric data and diverge sharply when a few extreme values pull the mean away from the typical case.

## Quick Answer

- The mean adds every value and divides by the count. It uses all the data and reacts to every value, including outliers.
- The median is the middle value of the sorted data. It ignores how far the extremes sit from the center.
- On symmetric data with no outliers, mean and median are close, so either works.
- On skewed data or data with outliers, the median describes the typical value better, and the mean overstates or understates it.
- Report the mean for roughly symmetric data and for anything you will feed into further arithmetic. Report the median for income, house prices, reaction times and other skewed distributions.

## Key Differences

| Property | Mean | Median |
|---|---|---|
| Definition | Sum of values divided by count | Middle value of sorted data |
| Uses every value | Yes | Only for ordering, not for distance |
| Sensitive to outliers | Yes, strongly | No |
| Works with skewed data | Poorly | Well |
| Typical notation | $\bar{x}$ or $\mu$ | $M$ or $\tilde{x}$ |
| Used in further math | Yes, sums and variances build on it | Rarely |
| Behavior when data is symmetric | Equals the median | Equals the mean |

The core difference is sensitivity. Move one value in a dataset and the mean changes. Move the same value without crossing the middle of the sorted order and the median does not move at all.

## Mean Explained

The mean is the balance point of your data. Add all values, divide by how many you have.

$$\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i$$

For a sample of reaction times, you sum the trials and divide by the number of trials. Every observation contributes, and each contributes in proportion to its size. That is why the mean is the natural input for variance, standard deviation and most hypothesis tests. The [mean and standard deviation](/blog/data-analysis/mean-and-standard-deviation) article walks through how the two are computed together.

The same property that makes the mean useful makes it fragile. A single value far from the rest shifts the balance point toward itself. In a lab with eight trials near 15 seconds and one trial at 95 seconds, the mean lands well above every normal trial.

The mean is also the quantity that most statistical procedures test. Confidence intervals for differences between means, t-tests and ANOVA all target population means [1][2]. If you plan to compare groups formally, you usually need the mean even when the median describes the data better.

## Median Explained

The median is the value that splits the sorted data in half. Half the observations sit at or below it, half at or above it.

For an odd number of values, the median is the single middle value at position $(n+1)/2$. For an even number, it is the average of the two middle values. In R, the `median` function handles this directly and returns a length-one object of the same type as the input [3].

The median depends only on order, not on magnitude. Whether the largest value is 20 or 20,000, the median is unchanged as long as it stays the largest. That property is what statisticians mean when they call the median resistant to outliers. The R documentation's own example makes the point: `median(c(1:3, 100, 1000))` returns 3, because the two extreme values cannot move the middle [3].

The median also has its own inferential tools. SciPy's `median_test` checks whether two or more samples come from populations with the same median, building a contingency table of values above and below the grand median and passing it to a chi-squared test [4]. So you are not giving up statistical testing by choosing the median, you are just testing a different hypothesis.

## Worked Example

A lab records reaction times in seconds across 9 trials. One trial is far slower than the rest.

| Trial | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
|---|---|---|---|---|---|---|---|---|---|
| Reaction time (s) | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 95 |

Step 1. Sum the values.

$$12+13+14+15+16+17+18+19+95 = 219$$

Step 2. Divide by the count to get the mean.

$$219 / 9 = 24.3333$$

Step 3. Sort the data. It is already in order: 12, 13, 14, 15, 16, 17, 18, 19, 95.

Step 4. Find the median position for an odd count.

$$(n+1)/2 = (9+1)/2 = 5$$

Step 5. Read the fifth value: the median is 16.

Step 6. Compare them.

$$24.3333 - 16 = 8.3333$$

The mean sits 8.33 seconds above the median. Eight of the nine trials fall between 12 and 19 seconds, so 16 describes a normal trial and 24.33 describes no trial at all. The outlier at 95 pulled the mean upward.

```python
import statistics
data = [12, 13, 14, 15, 16, 17, 18, 19, 95]
mean = statistics.mean(data)   # 24.3333
median = statistics.median(data)  # 16
```

Output:

```
mean = 24.3333, median = 16
```

You can reproduce this on your own numbers with the [Mean, Median & Mode Calculator](/tools/mean-median-mode-calculator).

## Which One Should You Use?

Ask what shape your data has and what you plan to do with the number.

Use the mean when the distribution is roughly symmetric, when you need the value that variance and standard deviation are built on, and when you will run t-tests, ANOVA or regression [2]. The mean is also the right choice when every observation should count in proportion to its size, such as total revenue per customer.

Use the median when the distribution is skewed, when outliers are real but unrepresentative, or when you want to describe the typical case. Income, house prices, wait times and reaction times are the classic cases. The [measures of central tendency](/blog/data-analysis/measures-of-central-tendency) overview shows how the median fits alongside the mean and mode.

Report both when they disagree. A gap between mean and median is itself information: it tells the reader the distribution is skewed and in which direction. If the mean exceeds the median, the right tail is long. If the median exceeds the mean, the left tail is long.

One more consideration: the level of measurement. The mean requires numeric values you can add. For ordinal data, where the gaps between categories are not equal, the median is the appropriate center. The [nominal vs ordinal variables](/blog/data-analysis/nominal-vs-ordinal-variables) article covers why arithmetic on ordinal codes is misleading.

## Common Mistakes

- **Reporting the mean for skewed data without comment.** A mean income in a town with one billionaire describes nobody. Fix: report the median, or report both and explain the gap.
- **Assuming the median is always safer.** The median throws away information about magnitude. Two datasets with the same median can have wildly different spreads. Fix: pair the median with a spread measure such as the interquartile range, and see [measures of variability](/blog/data-analysis/measures-of-variability).
- **Dropping outliers to protect the mean.** Removing a real observation changes the question you are answering. Fix: keep the value and switch to the median, or report both and state the decision.
- **Confusing "average" with "mean" in a report.** Readers hear "average" and assume the mean. Fix: name the statistic explicitly every time.
- **Using the mean on ordinal survey codes.** Averaging a 1-to-5 satisfaction scale treats the gaps as equal when they may not be. Fix: use the median or report the full distribution.
- **Forgetting the even-count rule.** With an even number of values there is no single middle value. Fix: average the two central values, which is what standard software does [3].

## Limitations

Neither measure tells you about spread, shape or sample size. A mean of 50 from three observations and a mean of 50 from three million look identical in a summary table. Always pair a center with a spread measure and the count.

The median is resistant to outliers but not immune to every problem. It is less efficient than the mean on clean, symmetric data, meaning it needs more observations to estimate the population center with the same precision. It also cannot be decomposed, so you cannot compute a combined median from group medians the way you can combine group means and counts. And when many values tie at the middle, the median can land on a value that is common but not central in any meaningful sense.

## Frequently Asked Questions

### What is the difference between mean and median?

The mean is the sum of all values divided by the count. The median is the middle value of the sorted data. The mean uses the size of every value, so outliers move it. The median uses only position, so outliers do not move it.

### Is median or average better?

It depends on the shape of the data. For symmetric data with no extreme values, the mean is more informative and feeds into more statistical tests. For skewed data, the median better represents the typical value. When they differ a lot, report both.

### Why is the mean higher than the median in income data?

Income distributions have a long right tail. A small number of very high earners raise the sum, and therefore the mean, without changing the middle of the sorted list. The gap between mean and median is a quick signal of right skew.

### Can the mean and median be the same?

Yes. In a perfectly symmetric distribution they coincide. In real samples they are usually close but not identical. A large gap is a warning that the distribution is skewed or contains outliers.

### Which measure should I report in a research paper?

Report the one that matches your analysis. If you ran a t-test or ANOVA, report means with standard deviations [1][2]. If your outcome is skewed or ordinal, report medians with interquartile ranges and use a median-based test [4]. State which you used and why.

## References

1. [7.3.1.2. Confidence intervals for differences between means](https://www.itl.nist.gov/div898/handbook/prc/section3/prc312.htm)
2. [7.4.3. Are the means equal?](https://www.itl.nist.gov/div898/handbook/prc/section4/prc43.htm)
3. [R: Median Value](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/median.html)
4. [median_test, SciPy v1.18.0 Manual](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.median_test.html)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)

## Related Articles

- [Sample vs Population Standard Deviation: When to Use Each](/blog/data-analysis/sample-vs-population-standard-deviation)
- [Mean and Standard Deviation: Definition, Formula and Examples](/blog/data-analysis/mean-and-standard-deviation)
- [Accuracy, Precision and Recall: Differences and When to Use Each](/blog/data-analysis/accuracy-precision-recall-differences)
- [Measures of Variability: Range, Variance and Standard Deviation](/blog/data-analysis/measures-of-variability)
- [Nominal vs Ordinal Variables: Differences and Examples](/blog/data-analysis/nominal-vs-ordinal-variables)
- [Statistical Synonyms: A Guide to Terminology in Statistics](/blog/guides/statistical-synonyms-a-guide-to-terminology-in-statistics)
- [T-Test vs ANOVA: A Decision Framework for Comparing Group Means](/blog/guides/t-test-vs-anova-a-decision-framework-for-comparing-group-means)
- [Standard Deviation vs Variance vs Standard Error: What Each Measures and When to Report It](/blog/research-skills/standard-deviation-vs-variance-vs-standard-error)