How to Find the Median in R (With Examples)
By Dr. Zubair Khalid, DVM, MS, PhD ·

To find the median in R, call median() on a numeric vector. It returns the middle value of the sorted data, or the average of the two middle values when the length is even. The r median is a single line of code in most cases, but missing values and grouped data need a little extra care.
Quick Answer
median(x)returns the middle value of a numeric vectorx[1].- For an even number of values, R averages the two middle values [1].
median(x, na.rm = TRUE)stripsNAvalues before computing [1].- With
na.rm = FALSE(the default) and anyNApresent, the result isNA[1]. - For grouped data, use
aggregate(),tapply(), ordplyr::summarise().
Syntax
The base function signature is median(x, na.rm = FALSE, ...) [1].
| Argument | Required? | Meaning |
|---|---|---|
x | Yes | An object with a method defined, or a numeric vector whose median you want [1]. |
na.rm | No | Logical. If TRUE, NA values are stripped before the computation [1]. Default is FALSE. |
... | No | Further arguments for methods. Not used in the default method [1]. |
The default method returns a length-one object of the same type as x, except when x is logical or integer of even length [1].
How It Works
R sorts the values, then picks the middle one. If the count $n$ is odd, the median is the value at position $(n+1)/2$. If $n$ is even, the median is the mean of the values at positions $n/2$ and $n/2 + 1$ [1].
$$\text{median} = \begin{cases} x_{(n+1)/2} & n \text{ odd} \\[4pt] \dfrac{x_{n/2} + x_{n/2+1}}{2} & n \text{ even} \end{cases}$$
The default method builds on is.na, sort, and mean, all of which are generic, so median() works for most classes where a median makes sense, including Date [1]. If there are no values, or if na.rm = FALSE and NA values are present, the result is NA of the same type as x [1].
The median is resistant to extreme values. The help page's own example makes the point: median(c(1:3, 100, 1000)) returns 3, because the two large values do not move the middle [1]. If you want to see how far the mean drifts from the median on skewed data, the mean vs median comparison walks through it.
Worked Example
A lab records 12 reaction times in seconds and wants a typical value that is not dragged around by a slow trial.
| reaction_time_seconds |
|---|
| 1.82 |
| 2.05 |
| 1.97 |
| 2.31 |
| 1.88 |
| 2.14 |
| 2.02 |
| 1.93 |
| 2.27 |
| 1.79 |
| 2.11 |
| 1.96 |
Step 1. Sort the 12 reaction times:
[1.79, 1.82, 1.88, 1.93, 1.96, 1.97, 2.02, 2.05, 2.11, 2.14, 2.27, 2.31]
Step 2. n is even, so average the two middle values:
$$(1.97 + 2.02) / 2 = 1.9950$$
Step 3. Run it in R:
x <- c(1.82, 2.05, 1.97, 2.31, 1.88, 2.14, 2.02, 1.93, 2.27, 1.79, 2.11, 1.96)
median(x)
x_na <- c(x, NA)
median(x_na)
median(x_na, na.rm = TRUE)
mean(x)
Output:
[1] 1.995
[1] NA
[1] 1.995
[1] 2.020833
The median is 1.995, and R prints it as 1.995. The mean is 2.0208, so the mean sits 0.0258 seconds above the median. That gap is small here, but it grows fast when a few slow trials pull the tail out. If you want to check the arithmetic by hand or on another platform, the median calculator does the same steps, and the Excel median guide covers the spreadsheet equivalent.
More Examples
Odd-length vector. With five values, R returns the third sorted value directly.
median(c(4, 9, 1, 7, 3))
Missing values. The default returns NA when any value is missing. Add na.rm = TRUE to drop them [1].
median(c(10, 20, NA, 40))
median(c(10, 20, NA, 40), na.rm = TRUE)
Dates. The default method works on Date objects because sort and mean are generic [1].
d <- as.Date(c("2024-01-01", "2024-01-05", "2024-01-09"))
median(d)
Grouped data with aggregate(). Split a numeric column by a grouping column and compute one median per group.
df <- data.frame(
group = c("A", "A", "A", "B", "B", "B"),
score = c(10, 12, 14, 20, 22, 30)
)
aggregate(score ~ group, data = df, FUN = median)
Grouped data with tapply(). Same idea, returns a named vector.
tapply(df$score, df$group, median)
Weighted median. Base R has no weighted median. DescTools::Median() accepts a weights argument, and spatstat provides weighted.median() [2][3]. In the weighted case, the median is the value $m$ such that the total weight of data to the left of $m$ equals the total weight to the right, with linear interpolation when no exact value satisfies that [3].
Grouped frequency data. For data already tabulated into classes, DescTools::Median() estimates the median by linear interpolation inside the class that contains it, using the Freq interface [2].
Errors and How to Fix Them
Error in median(x) : object 'x' not found. The object does not exist in your environment. Check spelling and confirm you assigned it, for example with x <- c(1, 2, 3). The c() vector guide covers building vectors correctly.
Error in median.default(x) : need numeric data. You passed a factor or a whole data frame. Select a single column, and convert a factor with as.numeric(as.character(x)) after checking that the values really are numbers.
Warning message: In mean.default(X[[i]], ...) : argument is not numeric or logical: returning NA. This shows up when you apply median() to a column that contains non-numeric entries, often from a bad import. Inspect the column with str() and clean it first.
Result is NA and you expected a number. There is at least one NA in the vector and na.rm is still FALSE [1]. Add na.rm = TRUE, or decide whether dropping those rows is defensible.
Extra arguments are silently ignored. median() passes extra arguments through ..., and the default method ignores them, so median(x, weights = w) returns the unweighted median with no error. Use a package function that supports weights [2][3].
Median of an ordered factor fails. Standard R does not implement a median for ordered factors because it is not well defined when the median falls between two levels for even-length factors [2]. Convert to numeric codes if the ordering is genuinely numeric.
Common Mistakes
- Forgetting
na.rm = TRUE. One missing value turns the whole result intoNA[1]. Decide up front whether missing values should be dropped, and say so in your write-up. - Assuming
na.rm = TRUEis the default. It is not. The default isna.rm = FALSE[1]. Code that works on clean data silently returnsNAon real data. - Using
median()on a data frame. It expects a vector or an object with a method. Select the column first, as inmedian(df$score). - Confusing the median with the mean. They answer different questions. The median is the middle value, the mean is the arithmetic average, and they diverge as skew grows [1].
- Treating a grouped median as a plain median.
median()on a stacked column ignores the groups. Useaggregate(),tapply(), or a grouped summary so each group gets its own value. - Expecting a weighted median from base R. It is not there. Reach for a package that implements it [2][3].
Limitations
median() gives you one number and nothing about the shape of the data. Two datasets with the same median can look completely different, so pair it with a spread measure such as the median absolute deviation, which mad() computes as the median of the absolute deviations from the median, scaled by 1.4826 by default [4]. For a visual read on where the middle sits, see finding the median from a histogram.
The function also cannot handle every data type. Ordered factors have no defined median in standard R because it is unclear what to do when the median falls between two levels for an even-length factor [2]. Weighted and grouped-frequency medians require package functions, and those make modeling assumptions, such as linear interpolation within the median class, that you should state when you report the result [2][3].
Frequently Asked Questions
How do I find the median in R with missing values?
Add na.rm = TRUE, as in median(x, na.rm = TRUE). Without it, any NA in the vector makes the result NA because the default is na.rm = FALSE [1]. Check how many values were dropped before you trust the number.
What does median() return for an even number of values?
It returns the average of the two middle values after sorting [1]. For the 12 reaction times above, the two middle values are 1.97 and 2.02, so the median is 1.995, which is exactly what R prints.
How do I calculate a median by group in R?
Use aggregate(score ~ group, data = df, FUN = median) for a data frame, or tapply(df$score, df$group, median) for a named vector. Both call median() once per group. For more complex summaries, a grouped summarise() in dplyr works the same way.
Does R have a weighted median function?
Not in base R. DescTools::Median() takes a weights argument, and spatstat::weighted.median() computes weighted medians and quantiles [2][3]. The weighted median is the value where the total weight on each side balances, with linear interpolation when needed [3].
Why does median() return NA when my data looks fine?
There is almost certainly an NA somewhere in the vector, possibly from a failed numeric conversion during import. Run sum(is.na(x)) to count them and which(is.na(x)) to locate them. If the missing values should be excluded, set na.rm = TRUE [1].
References
- R: Median Value
- R: (Weighted) Median Value
- weighted.median function - RDocumentation
- mad function - RDocumentation
Further Reading
- median function - RDocumentation
- Wickham H (2014). Tidy Data. Journal of Statistical Software
- Wickham H, Averick M, Bryan J et al. (2019). Welcome to the Tidyverse. Journal of Open Source Software