# Mann-Whitney U Test: Definition, Formula and Example

The Mann-Whitney U test is a nonparametric test that compares two independent groups without assuming normal distributions. It ranks all observations together and checks whether the ranks in one group tend to be higher than the ranks in the other. This article explains when to use it, how the U statistic is computed, and how to interpret the p-value.

## Quick Answer

- The Mann-Whitney U test compares two independent groups when your outcome is ordinal or numeric but not normally distributed [1].
- It tests whether the two population distributions are the same, not specifically whether the means or medians are equal [1][2].
- You rank all observations from both groups together, then sum the ranks in each group [1].
- The test statistic $U$ is the smaller of $U_1$ and $U_2$, where $U_1 = R_1 - n_1(n_1+1)/2$ and $U_2 = R_2 - n_2(n_2+1)/2$ [3].
- A small p-value (typically below 0.05) means the two distributions differ more than chance would explain.

## What the Mann-Whitney U Test Means

In plain terms, the Mann-Whitney U test asks a simple question: if you pick one observation at random from each group, how often is the value from one group larger than the value from the other? If the two groups behave the same way, that should happen about half the time.

The precise statistical definition is that the test evaluates a null hypothesis that the probability distribution of a randomly drawn observation from one group is the same as the probability distribution of a randomly drawn observation from the other group, against an alternative that those distributions are not equal [4]. This is different from a t-test, which tests a null hypothesis of equal means in two groups against an alternative of unequal means [4].

The test was expanded on Frank Wilcoxon's rank sum test by Henry Mann and Donald Whitney, which is why you will also see it called the Wilcoxon-Mann-Whitney test or the rank-sum test [3][1]. If you are comparing two groups and want the parametric counterpart, see the [two sample t-test formula and example](/blog/data-analysis/two-sample-t-test-formula-example).

## How It Works

The mechanism is ranking. You combine both groups into one list, sort the values from smallest to largest, and assign ranks. Tied values get the average of the ranks they would have occupied. Then you add up the ranks separately for each group.

The two rank sums are $R_1$ and $R_2$. From these you compute:

$$U_1 = R_1 - \frac{n_1(n_1+1)}{2}$$

$$U_2 = R_2 - \frac{n_2(n_2+1)}{2}$$

The test statistic is the smaller of the two:

$$U = \min(U_1, U_2)$$

Each symbol means the following:

- $n_1$ and $n_2$ are the sample sizes of the two groups.
- $R_1$ and $R_2$ are the sums of the ranks for each group.
- $U_1$ and $U_2$ count the number of times a value in one group exceeds a value in the other.
- $U$ is the smaller of those two counts, which is the value reported by most software [3].

For larger samples, the rank sum $R_1$ is approximated by a normal distribution with mean $\mu = \frac{n_1(n_1+n_2+1)}{2}$ [5], which is the same as $U$ having mean $\frac{n_1 n_2}{2}$. Small samples use exact tables of critical values instead. If $U$ is less than or equal to the critical value, you reject the null hypothesis [3].

## Worked Example

Suppose a clinic compares pain scores (0 to 10) for 6 patients on Treatment A and 6 patients on Treatment B. Lower scores mean less pain.

| Treatment | Pain scores |
|---|---|
| A | 3, 5, 6, 7, 8, 9 |
| B | 1, 2, 4, 5, 6, 10 |

**Step 1. Combine and sort all 12 values.**

Combined sorted values: 1, 2, 3, 4, 5, 5, 6, 6, 7, 8, 9, 10.

**Step 2. Assign ranks, averaging ties.**

The value 5 appears twice (positions 5 and 6), so both get rank 5.5. The value 6 appears twice (positions 7 and 8), so both get rank 7.5.

| Value | 1 | 2 | 3 | 4 | 5 | 5 | 6 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rank | 1 | 2 | 3 | 4 | 5.5 | 5.5 | 7.5 | 7.5 | 9 | 10 | 11 | 12 |

**Step 3. Sum the ranks for each group.**

Group A scores are 3, 5, 6, 7, 8, 9 with ranks 3, 5.5, 7.5, 9, 10, 11. So $R_1 = 46$.

Group B scores are 1, 2, 4, 5, 6, 10 with ranks 1, 2, 4, 5.5, 7.5, 12. So $R_2 = 32$.

**Step 4. Compute $U_1$ and $U_2$.**

$$U_1 = 46 - \frac{6 \times 7}{2} = 46 - 21 = 25$$

$$U_2 = 32 - \frac{6 \times 7}{2} = 32 - 21 = 11$$

**Step 5. Take the smaller value.**

$$U = \min(25, 11) = 11$$

**Step 6. Get the p-value and effect size.**

The exact two-sided p-value is 0.3095. The effect size $r = 1 - 2U/(n_1 n_2) = 1 - 22/36 = 0.3889$.

Here is the same calculation in Python. SciPy reports $U_1$ for the first sample (25), so the smaller statistic is $n_1 n_2 - 25 = 11$:

```python
from scipy.stats import mannwhitneyu
A = [3, 5, 6, 7, 8, 9]
B = [1, 2, 4, 5, 6, 10]
U, p = mannwhitneyu(A, B, alternative='two-sided', method='exact')
print(U, p)  # 25.0 0.3095...
```

Output:

```
25.0 0.30952380952380953
```

You can reproduce this with the [Mann-Whitney U test calculator](/tools/mann-whitney-u-test-calculator).

## How to Interpret It

The p-value answers one question: if the two distributions were truly identical, how likely is a difference in ranks at least as extreme as the one you observed? A p-value of 0.3095 means that, if the distributions were identical, about 31 percent of samples would give a U at least this extreme in either direction. That is above the usual 0.05 threshold, so you fail to reject the null hypothesis.

A small p-value (below your chosen alpha, commonly 0.05) is evidence that the two distributions differ. A large p-value is not proof that they are the same, only that your data do not provide strong evidence of a difference.

The effect size $r$ tells you how large the difference is. Values near 0 mean little separation, values near 1 mean strong separation. An $r$ of 0.3889 is a moderate effect. For a fuller treatment of how p-values and distributions relate, see the [Student's t-distribution guide](/blog/data-analysis/students-t-distribution-definition-formula).

## When to Use It (and when not to)

Use the Mann-Whitney U test when:

- Your outcome is numeric or ordinal, and the two groups are independent [1].
- The normality assumption for a t-test fails, or your samples are too small to assess normality [1].
- Your data are ordinal but not interval scaled, so the spacing between adjacent values cannot be assumed constant [4].
- You have outliers that would distort a mean-based test. Because it compares sums of ranks, the Mann-Whitney U test is less likely than the t-test to spuriously indicate significance because of outliers [4].

Do not use it when:

- You want to compare group means specifically. The Mann-Whitney U test and the t-test do not test the same hypotheses, and if a difference of group means is your primary interest, Mann-Whitney is not an appropriate test [4].
- Your observations are paired or matched. Use a test designed for dependent samples instead.
- You have more than two groups. Use a Kruskal-Wallis test instead.

## Mann-Whitney U Test vs the t-Test

The closest related idea is the independent two-sample t-test. Both compare two independent groups, but they target different quantities.

| Feature | Mann-Whitney U test | Independent t-test |
|---|---|---|
| Assumption | No normality requirement [1] | Populations normally distributed [3] |
| What it tests | Whether the two distributions are the same [4] | Whether the two means are equal [4] |
| Data type | Numeric or ordinal [4] | Numeric, interval or ratio |
| Outlier sensitivity | Lower, since it uses ranks [4] | Higher, since it uses means |
| Efficiency when normal | About 0.95 relative to the t-test [4] | Reference standard |

When normality holds, the Mann-Whitney U test has an asymptotic efficiency of about 0.95 compared with the t-test [4]. For distributions far from normal and large samples, it can be considerably more efficient [4]. That comparison should be read with care, because the two tests do not measure the same thing [4].

## Common Mistakes

- **Calling it a test of medians.** The test compares whole distributions, including spread and shape, not only medians [2]. Fix: report the distributions and mean ranks, and only describe a median difference when the shapes are similar [6].
- **Using it automatically whenever data are non-normal.** Non-normality alone does not decide the test. Fix: ask whether you care about means or about the full distribution, and pick accordingly [4].
- **Forgetting to average tied ranks.** Ties change the rank sums. Fix: assign the average rank to every tied value before summing.
- **Reporting U without the sample sizes or p-value.** A bare U value is hard to interpret. Fix: report U, the group sizes, and the p-value together, as in "U = 145, Z = -1.488, p = 0.142" [6].
- **Treating a large p-value as proof of equality.** Failing to reject is not the same as showing the groups are identical. Fix: describe it as insufficient evidence of a difference.
- **Ignoring effect size.** Significance says nothing about magnitude. Fix: report an effect size such as $r$ alongside the p-value.

## Limitations

The Mann-Whitney U test may have worse type I error control when data are both heteroscedastic and non-normal [4]. In other words, if the two groups have very different variances and are not normal, the test can reject too often or too rarely. It also cannot tell you which part of the distribution differs. A significant result could come from a shift in location, a difference in spread, or a difference in shape [2].

The test gives you no estimate of the size of a mean difference, and it does not produce a confidence interval for a difference in means the way a t-test does. If your research question is about means, this test answers a different question, and you should choose a method that matches your estimand [4].

## Frequently Asked Questions

### What is the difference between the Mann-Whitney U test and the Wilcoxon rank sum test?

They are the same test under different names. The test was expanded on Frank Wilcoxon's rank sum test by Henry Mann and Donald Whitney, so it is often called the Wilcoxon-Mann-Whitney test or the rank-sum test [3][1]. Software packages may report either a U statistic or a W statistic, and both describe the same rank comparison [1].

### Does the Mann-Whitney U test compare medians?

Not strictly. It is commonly regarded as a test of population medians, but that is not strictly true, and treating it as such can lead to inadequate analysis [2]. It tests whether one variable tends to have values higher than the other, which reflects both location and shape [2]. Only when the two distributions have the same shape can you describe the result as a difference in medians [2].

### How do I report a Mann-Whitney U test result?

Report the test statistic, the sample sizes, and the p-value. A typical conclusion reads: "There is evidence to suggest that dogs and cats spend different lengths of time in the shelter before getting adopted (W = 45.5, p < 0.05)" [1]. If the shapes are similar, you can add the medians. If not, report mean ranks instead [6].

### What sample size do I need for the Mann-Whitney U test?

There is no single cutoff. For small samples, exact critical values are tabulated, and some combinations are too small to reject the null hypothesis at all [3]. For larger samples, a normal approximation is used, with the rank sum having mean $\mu = \frac{n_1(n_1+n_2+1)}{2}$ [5] and $U$ having mean $\frac{n_1 n_2}{2}$. If your groups are very small, check whether an exact method is available in your software.

### Can I use the Mann-Whitney U test with ordinal data?

Yes. The test is preferable to the t-test when the data are ordinal but not interval scaled, because the spacing between adjacent values cannot be assumed constant [4]. This makes it a natural fit for Likert-type scales, satisfaction ratings, and ranked judgments. For related terminology, see the [statistical synonyms guide](/blog/guides/statistical-synonyms-a-guide-to-terminology-in-statistics).

## References

1. [Mann-Whitney U-Test](https://sites.utexas.edu/sos/guided/inferential/numeric/onecat/2-groups/independent/u-test/)
2. [Mann-Whitney test is not just a test of medians: differences in spread can be important - PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC1120984/)
3. [13.5: Mann-Whitney U Test - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Introductory_Statistics/Mostly_Harmless_Statistics_(Webb)/13%3A_Nonparametric_Tests/13.05%3A__Mann-Whitney_U_Test)
4. [Mann-Whitney U test - Wikipedia](https://en.wikipedia.org/wiki/Mann%E2%80%93Whitney_U_test)
5. [Mann Whitney U Statistic](https://www.itl.nist.gov/div898/software/dataplot/refman2/auxillar/mannwhit.htm)
6. [Mann-Whitney U Test - Introduction to SPSS - ULibraries Research Guides at University of Utah](https://campusguides.lib.utah.edu/c.php?g=160832&p=6707563)

## Further Reading

- [Chicco D, Sichenze A, Jurman G. (2025). A simple guide to the use of Student's t-test, Mann-Whitney U test, Chi-squared test, and Kruskal-Wallis test in biostatistics. BioData mining](https://pmc.ncbi.nlm.nih.gov/articles/PMC12366075/)

## Related Articles

- [Two Sample t-Test: Formula, Calculation and Example](/blog/data-analysis/two-sample-t-test-formula-example)
- [Likelihood Ratio Test: Definition, Formula and Examples](/blog/data-analysis/likelihood-ratio-test)
- [Harmonic Mean: Formula, Examples and When to Use It](/blog/data-analysis/harmonic-mean-formula)
- [Student's t-Distribution: Definition, Formula and Examples](/blog/data-analysis/students-t-distribution-definition-formula)
- [Mean and Standard Deviation: Definition, Formula and Examples](/blog/data-analysis/mean-and-standard-deviation)
- [Mann whitney u test step by step biology](/knowledge/bioinformatics/mann-whitney-u-test-a-step-by-step-guide-for-comparing-two-independent-groups-with-ordinal-or-non-no)
- [One-Sample t-Test: Formula, Calculation, and Interpretation](/blog/guides/one-sample-t-test-formula-calculation-and-interpretation)
- [Understanding the t-Test: Meaning, Assumptions, and Applications](/blog/guides/understanding-the-t-test-meaning-assumptions-and-applications)