# How to Find the Median in R (With Examples)

To find the median in R, call `median()` on a numeric vector. It returns the middle value of the sorted data, or the average of the two middle values when the length is even. The r median is a single line of code in most cases, but missing values and grouped data need a little extra care.

## Quick Answer

- `median(x)` returns the middle value of a numeric vector `x` [1].
- For an even number of values, R averages the two middle values [1].
- `median(x, na.rm = TRUE)` strips `NA` values before computing [1].
- With `na.rm = FALSE` (the default) and any `NA` present, the result is `NA` [1].
- For grouped data, use `aggregate()`, `tapply()`, or `dplyr::summarise()`.

## Syntax

The base function signature is `median(x, na.rm = FALSE, ...)` [1].

| Argument | Required? | Meaning |
|---|---|---|
| `x` | Yes | An object with a method defined, or a numeric vector whose median you want [1]. |
| `na.rm` | No | Logical. If `TRUE`, `NA` values are stripped before the computation [1]. Default is `FALSE`. |
| `...` | No | Further arguments for methods. Not used in the default method [1]. |

The default method returns a length-one object of the same type as `x`, except when `x` is logical or integer of even length [1].

## How It Works

R sorts the values, then picks the middle one. If the count $n$ is odd, the median is the value at position $(n+1)/2$. If $n$ is even, the median is the mean of the values at positions $n/2$ and $n/2 + 1$ [1].

$$\text{median} = \begin{cases} x_{(n+1)/2} & n \text{ odd} \\[4pt] \dfrac{x_{n/2} + x_{n/2+1}}{2} & n \text{ even} \end{cases}$$

The default method builds on `is.na`, `sort`, and `mean`, all of which are generic, so `median()` works for most classes where a median makes sense, including `Date` [1]. If there are no values, or if `na.rm = FALSE` and `NA` values are present, the result is `NA` of the same type as `x` [1].

The median is resistant to extreme values. The help page's own example makes the point: `median(c(1:3, 100, 1000))` returns 3, because the two large values do not move the middle [1]. If you want to see how far the mean drifts from the median on skewed data, the [mean vs median comparison](/blog/data-analysis/mean-vs-median-differences) walks through it.

## Worked Example

A lab records 12 reaction times in seconds and wants a typical value that is not dragged around by a slow trial.

| reaction_time_seconds |
|---|
| 1.82 |
| 2.05 |
| 1.97 |
| 2.31 |
| 1.88 |
| 2.14 |
| 2.02 |
| 1.93 |
| 2.27 |
| 1.79 |
| 2.11 |
| 1.96 |

Step 1. Sort the 12 reaction times:

[1.79, 1.82, 1.88, 1.93, 1.96, 1.97, 2.02, 2.05, 2.11, 2.14, 2.27, 2.31]

Step 2. `n` is even, so average the two middle values:

$$(1.97 + 2.02) / 2 = 1.9950$$

Step 3. Run it in R:

```r
x <- c(1.82, 2.05, 1.97, 2.31, 1.88, 2.14, 2.02, 1.93, 2.27, 1.79, 2.11, 1.96)
median(x)
x_na <- c(x, NA)
median(x_na)
median(x_na, na.rm = TRUE)
mean(x)
```

Output:

```
[1] 1.995
[1] NA
[1] 1.995
[1] 2.020833
```

The median is 1.995, and R prints it as 1.995. The mean is 2.0208, so the mean sits 0.0258 seconds above the median. That gap is small here, but it grows fast when a few slow trials pull the tail out. If you want to check the arithmetic by hand or on another platform, the [median calculator](/tools/mean-median-mode-calculator) does the same steps, and the [Excel median guide](/blog/data-analysis/how-to-calculate-median-in-excel) covers the spreadsheet equivalent.

## More Examples

**Odd-length vector.** With five values, R returns the third sorted value directly.

```r
median(c(4, 9, 1, 7, 3))
```

**Missing values.** The default returns `NA` when any value is missing. Add `na.rm = TRUE` to drop them [1].

```r
median(c(10, 20, NA, 40))
median(c(10, 20, NA, 40), na.rm = TRUE)
```

**Dates.** The default method works on `Date` objects because `sort` and `mean` are generic [1].

```r
d <- as.Date(c("2024-01-01", "2024-01-05", "2024-01-09"))
median(d)
```

**Grouped data with `aggregate()`.** Split a numeric column by a grouping column and compute one median per group.

```r
df <- data.frame(
  group = c("A", "A", "A", "B", "B", "B"),
  score = c(10, 12, 14, 20, 22, 30)
)
aggregate(score ~ group, data = df, FUN = median)
```

**Grouped data with `tapply()`.** Same idea, returns a named vector.

```r
tapply(df$score, df$group, median)
```

**Weighted median.** Base R has no weighted median. `DescTools::Median()` accepts a `weights` argument, and `spatstat` provides `weighted.median()` [2][3]. In the weighted case, the median is the value $m$ such that the total weight of data to the left of $m$ equals the total weight to the right, with linear interpolation when no exact value satisfies that [3].

**Grouped frequency data.** For data already tabulated into classes, `DescTools::Median()` estimates the median by linear interpolation inside the class that contains it, using the `Freq` interface [2].

## Errors and How to Fix Them

**`Error in median(x) : object 'x' not found`.** The object does not exist in your environment. Check spelling and confirm you assigned it, for example with `x <- c(1, 2, 3)`. The [c() vector guide](/blog/data-analysis/c-function-in-r-create-vectors) covers building vectors correctly.

**`Error in median.default(x) : need numeric data`.** You passed a factor or a whole data frame. Select a single column, and convert a factor with `as.numeric(as.character(x))` after checking that the values really are numbers.

**`Warning message: In mean.default(X[[i]], ...) : argument is not numeric or logical: returning NA`.** This shows up when you apply `median()` to a column that contains non-numeric entries, often from a bad import. Inspect the column with `str()` and clean it first.

**Result is `NA` and you expected a number.** There is at least one `NA` in the vector and `na.rm` is still `FALSE` [1]. Add `na.rm = TRUE`, or decide whether dropping those rows is defensible.

**Extra arguments are silently ignored.** `median()` passes extra arguments through `...`, and the default method ignores them, so `median(x, weights = w)` returns the unweighted median with no error. Use a package function that supports weights [2][3].

**Median of an ordered factor fails.** Standard R does not implement a median for ordered factors because it is not well defined when the median falls between two levels for even-length factors [2]. Convert to numeric codes if the ordering is genuinely numeric.

## Common Mistakes

- **Forgetting `na.rm = TRUE`.** One missing value turns the whole result into `NA` [1]. Decide up front whether missing values should be dropped, and say so in your write-up.
- **Assuming `na.rm = TRUE` is the default.** It is not. The default is `na.rm = FALSE` [1]. Code that works on clean data silently returns `NA` on real data.
- **Using `median()` on a data frame.** It expects a vector or an object with a method. Select the column first, as in `median(df$score)`.
- **Confusing the median with the mean.** They answer different questions. The median is the middle value, the mean is the arithmetic average, and they diverge as skew grows [1].
- **Treating a grouped median as a plain median.** `median()` on a stacked column ignores the groups. Use `aggregate()`, `tapply()`, or a grouped summary so each group gets its own value.
- **Expecting a weighted median from base R.** It is not there. Reach for a package that implements it [2][3].

## Limitations

`median()` gives you one number and nothing about the shape of the data. Two datasets with the same median can look completely different, so pair it with a spread measure such as the median absolute deviation, which `mad()` computes as the median of the absolute deviations from the median, scaled by 1.4826 by default [4]. For a visual read on where the middle sits, see [finding the median from a histogram](/blog/data-analysis/how-to-find-median-from-histogram).

The function also cannot handle every data type. Ordered factors have no defined median in standard R because it is unclear what to do when the median falls between two levels for an even-length factor [2]. Weighted and grouped-frequency medians require package functions, and those make modeling assumptions, such as linear interpolation within the median class, that you should state when you report the result [2][3].

## Frequently Asked Questions

### How do I find the median in R with missing values?

Add `na.rm = TRUE`, as in `median(x, na.rm = TRUE)`. Without it, any `NA` in the vector makes the result `NA` because the default is `na.rm = FALSE` [1]. Check how many values were dropped before you trust the number.

### What does median() return for an even number of values?

It returns the average of the two middle values after sorting [1]. For the 12 reaction times above, the two middle values are 1.97 and 2.02, so the median is 1.995, which is exactly what R prints.

### How do I calculate a median by group in R?

Use `aggregate(score ~ group, data = df, FUN = median)` for a data frame, or `tapply(df$score, df$group, median)` for a named vector. Both call `median()` once per group. For more complex summaries, a grouped `summarise()` in dplyr works the same way.

### Does R have a weighted median function?

Not in base R. `DescTools::Median()` takes a `weights` argument, and `spatstat::weighted.median()` computes weighted medians and quantiles [2][3]. The weighted median is the value where the total weight on each side balances, with linear interpolation when needed [3].

### Why does median() return NA when my data looks fine?

There is almost certainly an `NA` somewhere in the vector, possibly from a failed numeric conversion during import. Run `sum(is.na(x))` to count them and `which(is.na(x))` to locate them. If the missing values should be excluded, set `na.rm = TRUE` [1].

## References

1. [R: Median Value](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/median.html)
2. [R: (Weighted) Median Value](https://search.r-project.org/CRAN/refmans/DescTools/html/Median.html)
3. [weighted.median function - RDocumentation](https://www.rdocumentation.org/packages/spatstat/versions/1.56-1/topics/weighted.median)
4. [mad function - RDocumentation](https://www.rdocumentation.org/packages/stats/versions/3.6.2/topics/mad)

## Further Reading

- [median function - RDocumentation](https://www.rdocumentation.org/packages/stats/versions/3.6.2/topics/median)
- [Wickham H (2014). Tidy Data. Journal of Statistical Software](https://doi.org/10.18637/jss.v059.i10)
- [Wickham H, Averick M, Bryan J et al. (2019). Welcome to the Tidyverse. Journal of Open Source Software](https://doi.org/10.21105/joss.01686)

## Related Articles

- [How to Find the Median from a Histogram (Step by Step)](/blog/data-analysis/how-to-find-median-from-histogram)
- [What Is a Median? Definition, Formula and Examples](/blog/data-analysis/what-is-a-median)
- [The %in% Operator in R: Syntax and Examples](/blog/data-analysis/in-operator-in-r-syntax-examples)
- [c() in R: How to Create Vectors (With Examples)](/blog/data-analysis/c-function-in-r-create-vectors)
- [F-Test in R: How to Compare Variances (With Example)](/blog/data-analysis/f-test-in-r)