# Measures of Central Tendency: Mean, Median and Mode Explained

Measures of central tendency are single values that describe the center of a data set, the point where observations tend to cluster. The three you will use most often are the mean, the median and the mode. Each one answers a slightly different question, and choosing the wrong one can make your summary misleading.

## Quick Answer

- **Mean**: the arithmetic average, found by adding all values and dividing by the count. It uses every observation and is the most common measure of central tendency [1].
- **Median**: the middle value once the data is sorted. It splits the data in half and resists extreme values.
- **Mode**: the most frequently occurring value. It is the only measure that works for categorical data.
- **Choose the mean** for roughly symmetric numeric data with no severe outliers.
- **Choose the median** for skewed data, income, house prices, or anything with extreme values. Use the mode for categories like blood type or product color.

## What Measures of Central Tendency Mean

In plain terms, a measure of central tendency is a summary number that stands in for the whole data set. It tells you the "typical" value around which data points pile up [2]. If someone asks how long a delivery usually takes, you give one number, not a list of 500 times.

The precise statistical definition is a measure of location that represents the center of a frequency distribution. Statisticians also call these measures of central location or averages. The mean, median and mode are the three standard ones, though the geometric mean and trimmed mean are also used in specific situations [2].

Central tendency is only half of a good summary. Two data sets can share the same mean and still look completely different, so you should always pair a measure of central tendency with a measure of spread such as the [range, variance or standard deviation](/blog/data-analysis/measures-of-variability).

## How It Works

Each measure uses a different rule to find the center.

**Mean (arithmetic mean).** Add every value, then divide by the number of observations [1]:

$$\bar{X} = \frac{\sum X}{n}$$

- $\bar{X}$ (X-bar) is the sample mean.
- $\sum$ (uppercase sigma) means "add up."
- $X$ is each individual value.
- $n$ is the number of observations.

**Median.** Sort the values from smallest to largest. If $n$ is odd, the median is the middle value. If $n$ is even, the median is the average of the two middle values.

**Mode.** Count how often each value appears. The mode is the value with the highest frequency. A data set can have no mode, one mode, or several modes.

## Worked Example

A lab recorded 12 reaction times in milliseconds, and one trial was unusually slow at 512 ms.

| Reaction time (ms) |
|---|
| 212 |
| 218 |
| 221 |
| 224 |
| 226 |
| 229 |
| 231 |
| 234 |
| 238 |
| 241 |
| 245 |
| 512 |

**Step 1: Sort the values.**
[212, 218, 221, 224, 226, 229, 231, 234, 238, 241, 245, 512]

**Step 2: Count the observations.** n = 12.

**Step 3: Sum all times.** 3031.

**Step 4: Compute the mean.** 3031 / 12 = 252.5833 ms.

**Step 5: Compute the median.** With n = 12, average the 6th and 7th sorted values: (229 + 231) / 2 = 230.0000 ms.

**Step 6: Compute the mode.** Every value appears once, so all values tie and there is no true mode. Python's `statistics.mode` returns the first value it meets, 212.

**Step 7: Check the outlier's effect.** Remove the 512 ms trial. The mean becomes 2519 / 11 = 229.0000 ms, while the median moves only from 230.0000 to 229.0000 ms. The outlier shifted the mean by 252.5833 - 229.0000 = 23.5833 ms.

Here is the same calculation in Python:

```python
import statistics
times = [212, 218, 221, 224, 226, 229, 231, 234, 238, 241, 245, 512]
mean = statistics.fmean(times)
median = statistics.median(times)
mode = statistics.mode(times)
print(f"mean={mean:.4f}, median={median:.4f}, mode={mode}")
```

Output:

```
mean=252.5833, median=230.0000, mode=212
```

In Excel, `=AVERAGE(...)` returns 252.5833 and `=MEDIAN(...)` returns 230, while `=MODE.SNGL(...)` returns #N/A because no value repeats. You can reproduce all of this with the [Mean, Median & Mode Calculator](/tools/mean-median-mode-calculator).

## How to Interpret It

The mean of 252.6 ms is higher than 11 of the 12 actual times. That is the classic mean problem: it uses every value, so it is a good representative of the data, yet the value itself often never appears in the raw data [1].

The median of 230.0 ms sits right in the middle of the cluster. It tells you that half the trials were faster than 230 ms and half were slower. When the mean and median differ by a lot, that gap is a signal of skew or outliers.

The mode of 212 ms is not very useful here because no value repeats. The mode earns its place with categorical data, where a mean or median makes no sense at all. For numeric data with no repeats, treat the mode as uninformative.

## When to Use It (and when not to)

Use the **mean** when your data is numeric, roughly symmetric, and free of extreme values. It is the measure that best resists fluctuation between samples drawn from the same population, which is why it underpins most statistical tests [1].

Use the **median** when the distribution is skewed or contains outliers. Income, house prices, wait times and reaction times are typical cases. The worked example shows why: one slow trial moved the mean by 23.6 ms and moved the median by only 1 ms.

Use the **mode** for categorical or nominal data such as eye color, blood type or the most common shoe size. It is also handy for spotting the most frequent value in discrete data.

Avoid the mean for open-ended or heavily skewed data. Avoid the median when you need a value that uses all the information in the data set, since the median ignores everything except the middle. Avoid the mode for continuous data measured to many decimal places, where repeats are rare.

## Mean vs Median

The mean and median both describe the center of numeric data, but they respond to extreme values very differently.

| Feature | Mean | Median |
|---|---|---|
| Definition | Sum divided by count | Middle sorted value |
| Uses all values | Yes | No, only position |
| Affected by outliers | Yes, strongly | No |
| Best for | Symmetric data | Skewed data |
| Typical use | Test scores, heights | Income, house prices |
| In the example | 252.5833 ms | 230.0000 ms |

If you want a deeper comparison, see [mean vs median differences and when to use each](/blog/data-analysis/mean-vs-median-differences).

## Common Mistakes

- **Reporting the mean for skewed data.** Income and price data are almost always right-skewed. Fix: report the median, or report both and explain the gap.
- **Ignoring the mode when no value repeats.** A mode of 212 ms in the example is an artifact of the tie-breaking rule, not a real peak. Fix: check the frequencies before quoting a mode.
- **Forgetting to sort before finding the median.** The median is the middle of the ordered list, not the middle of the raw list. Fix: always sort first.
- **Mixing up the sample and population mean symbols.** Both are computed the same way, by dividing the sum by the count, but $\bar{X}$ denotes a sample mean and $\mu$ a population mean. Fix: use the symbol that matches your data.
- **Treating the mean as a value that must appear in the data.** It often does not [1]. Fix: describe it as a balance point, not an observed value.
- **Reporting central tendency alone.** A single center hides the spread. Fix: pair it with a measure of variability.

## Limitations

No measure of central tendency tells you how spread out the data is. A mean of 50 could come from values tightly packed around 50 or from values split between 0 and 100. You need a measure of variability to see the difference.

The mean is also sensitive to sample size. With small samples, one or two extreme values can swing it heavily [2]. The median solves that problem but throws away information from the rest of the data. The mode can be ambiguous or absent entirely. Treat any single summary number as a starting point, not a full description of the data.

## Frequently Asked Questions

### What is the difference between mean, median and mode?

The mean is the arithmetic average of all values. The median is the middle value when the data is sorted. The mode is the most frequent value. They answer different questions about the same data. In a symmetric distribution with a single peak all three coincide, and a large gap between them is a sign of skew.

### Which measure of central tendency is best?

It depends on your data. Use the mean for symmetric numeric data, the median for skewed data or data with outliers, and the mode for categorical data. There is no single best measure, only the one that fits your data type and distribution shape [2].

### Can a data set have more than one mode?

Yes. If two values tie for the highest frequency, the data set is bimodal. If more than two tie, it is multimodal. If every value appears the same number of times, there is no mode at all.

### Why is the median sometimes preferred over the mean?

The median is not affected by extreme values. In the reaction time example, one 512 ms trial pulled the mean up by 23.6 ms while the median moved by only 1 ms, from 229 to 230 ms. For skewed data like income, the median gives a fairer picture of the typical case.

### Does the mean always appear in the data set?

No. The mean uses every value in the data, but the result often does not match any single observation [1]. In the example, no trial took exactly 252.5833 ms.

### What is central tendency in simple terms?

Central tendency is the tendency of data to cluster around a central value. A measure of central tendency is the single number you use to describe that center, whether it is a mean, median or mode [2].

## References

1. [Measures of central tendency: The mean - PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC3127352/)
2. [Measures of Central Tendency](https://www.jmp.com/en/statistics-knowledge-portal/exploratory-data-analysis/descriptive-statistics/measures-of-central-tendency)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)

## Related Articles

- [Mean vs Median: Differences and When to Use Each](/blog/data-analysis/mean-vs-median-differences)
- [Measures of Variability: Range, Variance and Standard Deviation](/blog/data-analysis/measures-of-variability)
- [Confirmatory Factor Analysis: Definition and Example](/blog/data-analysis/confirmatory-factor-analysis)
- [Ranking Correlation Coefficient: Spearman and Kendall](/blog/data-analysis/ranking-correlation-coefficient)
- [Correlation Examples: Positive, Negative and Zero Relationships](/blog/data-analysis/correlation-examples-positive-negative)
- [Statistical Synonyms: A Guide to Terminology in Statistics](/blog/guides/statistical-synonyms-a-guide-to-terminology-in-statistics)
- [Fundamental Statistics: Core Concepts Explained](/blog/research-skills/fundamental-statistics-core-concepts-explained)
- [Trend Analysis in Research: Methods, Applications, and Pitfalls](/blog/guides/trend-analysis-in-research-methods-applications-and-pitfalls)