# Simpson Diversity Index: Formula, Calculation and Examples

The Simpson diversity index measures the probability that two individuals drawn at random from a sample belong to the same species. It combines species richness and evenness into a single number, which makes it one of the most widely used measures of biodiversity [1]. This article covers the formula, a full hand calculation, code, and how to interpret the result.

## Quick Answer

- The Simpson index $D$ is the sum of squared proportional abundances: $D = \sum p_i^2$ [1].
- $D$ ranges from 0 to 1. A value near 0 means high diversity, and a value near 1 means low diversity [1].
- Because that direction is counterintuitive, the index is usually reported as $1 - D$ (the Gini-Simpson index) or $1/D$ (the inverse Simpson index) [1][2].
- $1 - D$ is the probability that two randomly chosen individuals belong to different species [2].
- Both $D$ and $1 - D$ use the same data. You only need the count of each species and the total count.

## The Formula

The basic form of the index is:

$$D = \sum_{i=1}^{R} p_i^2$$

Each symbol means the following:

- $R$ is the number of species in the sample, also called species richness [1].
- $p_i$ is the proportional abundance of species $i$, calculated as the count of that species divided by the total number of individuals $N$ [1].
- $D$ is the sum of those squared proportions across all species.

Two common transformations are used alongside it:

$$1 - D \quad \text{and} \quad \frac{1}{D}$$

The first is the Gini-Simpson index, and the second is the inverse Simpson index, also written as ${}^{2}D$ and interpreted as the effective number of species [2]. Both increase as diversity increases, which is why they are often preferred for reporting [2].

## How to Calculate It Step by Step

1. Count the individuals of each species in your sample.
2. Add the counts to get the total number of individuals, $N$.
3. Divide each species count by $N$ to get its proportion $p_i$.
4. Square each proportion.
5. Add the squared proportions. The total is $D$.
6. If you want a value that rises with diversity, compute $1 - D$ or $1/D$.

## Worked Example

The dataset below records counts of five species in two quadrats, both with 100 individuals.

| Species | Quadrat 1 | Quadrat 2 |
|---|---|---|
| Sp A | 40 | 90 |
| Sp B | 25 | 5 |
| Sp C | 15 | 3 |
| Sp D | 10 | 1 |
| Sp E | 10 | 1 |

**Quadrat 1**

The total is $N = 40 + 25 + 15 + 10 + 10 = 100$. Each proportion is squared:

- Sp A: $(40/100)^2 = 0.4000^2 = 0.1600$
- Sp B: $(25/100)^2 = 0.2500^2 = 0.0625$
- Sp C: $(15/100)^2 = 0.1500^2 = 0.0225$
- Sp D: $(10/100)^2 = 0.1000^2 = 0.0100$
- Sp E: $(10/100)^2 = 0.1000^2 = 0.0100$

Summing these gives $D = 0.1600 + 0.0625 + 0.0225 + 0.0100 + 0.0100 = 0.2650$. The Gini-Simpson value is $1 - 0.2650 = 0.7350$.

**Quadrat 2**

The total is $N = 90 + 5 + 3 + 1 + 1 = 100$. Squaring each proportion:

- Sp A: $(90/100)^2 = 0.9000^2 = 0.8100$
- Sp B: $(5/100)^2 = 0.0500^2 = 0.0025$
- Sp C: $(3/100)^2 = 0.0300^2 = 0.0009$
- Sp D: $(1/100)^2 = 0.0100^2 = 0.0001$
- Sp E: $(1/100)^2 = 0.0100^2 = 0.0001$

Summing these gives $D = 0.8100 + 0.0025 + 0.0009 + 0.0001 + 0.0001 = 0.8136$. The Gini-Simpson value is $1 - 0.8136 = 0.1864$.

Both quadrats contain the same five species, so richness is identical. Quadrat 1 is far more even, and that is what drives the difference. Quadrat 1 has the higher diversity, with $1 - D = 0.735$ against $0.186$ for Quadrat 2.

## How to Interpret the Result

The raw value $D$ is a probability. It answers the question: if you pick two individuals at random from the sample, with replacement, what is the chance they are the same species [2]? In Quadrat 2 that chance is 0.8136, which is high because Sp A dominates the sample. In Quadrat 1 it is 0.2650, so a random pair is much more likely to be two different species.

The transformed value $1 - D$ flips the direction. It is the probability that two randomly chosen individuals belong to different species [2]. Higher means more diverse. Quadrat 1 at 0.735 is clearly the more diverse community.

The inverse $1/D$ gives the effective number of species, which is easier to talk about than a probability. A community with $1/D = 4$ behaves, in terms of dominance, like a community with four equally common species [2]. This is the same quantity as true diversity of order 2 [2].

A useful comparison point is the minimum possible value of $D$ for a given richness. If all $R$ species are equally abundant, then $D = 1/R$. With five species, the most even community possible has $D = 0.2$ and $1 - D = 0.8$. Quadrat 1 at $D = 0.2650$ is close to that ceiling, while Quadrat 2 is far from it.

## Doing It in Software

In Python, the calculation is a few lines with NumPy. The function below takes an array of counts, converts it to proportions, squares them, and returns both $D$ and $1 - D$.

```python
import numpy as np
q1 = np.array([40, 25, 15, 10, 10])
q2 = np.array([90, 5, 3, 1, 1])
def simpson(c):
    p = c / c.sum()
    D = (p**2).sum()
    return D, 1 - D
D1, inv1 = simpson(q1)  # D=0.2650, 1-D=0.7350
D2, inv2 = simpson(q2)  # D=0.8136, 1-D=0.1864
```

Output:

```text
Quadrat 1: D = 0.2650, 1-D = 0.7350
Quadrat 2: D = 0.8136, 1-D = 0.1864
More diverse: Quadrat 1
```

In Excel, put the counts in a column, compute the total with `SUM`, then add a column of `(count/total)^2` and take its `SUM`. That final sum is $D$. Subtract it from 1 for the Gini-Simpson value.

In R, the `vegan` package provides `diversity()` with `index = "simpson"`, which returns the Gini-Simpson form. If you prefer to avoid dependencies, the manual route is `sum((x / sum(x))^2)` on a vector of counts.

## Common Mistakes

- **Reporting $D$ and calling it diversity without saying which form.** $D$ falls as diversity rises, so a bare number is ambiguous. State whether you are reporting $D$, $1 - D$, or $1/D$ [2].
- **Comparing indices from different sources as if they were the same.** The inverse Simpson and Gini-Simpson indices have both been called "the Simpson index" in the literature, so check the definition before comparing values [2].
- **Using raw counts instead of proportions.** Squaring counts gives values far above 1 and destroys the 0 to 1 scale. Always divide by $N$ first.
- **Ignoring sample size when comparing sites.** A small sample can miss rare species and inflate $D$. Compare samples of similar size, or use rarefaction.
- **Treating equal richness as equal diversity.** The worked example shows two quadrats with the same five species and very different values. Evenness matters as much as richness.
- **Forgetting that $D$ is a probability with replacement.** The interpretation assumes individuals are drawn independently and replaced [2].

## Limitations

The Simpson index is sensitive to the most abundant species and gives relatively little weight to rare ones. Two sites can share the same value even when one contains several uncommon species the other lacks. If rare species matter for your question, pair the Simpson index with a richness measure or with the Shannon index, which weights species more evenly across the abundance range. Choosing among indices is a real methodological decision, and different indices can lead to different conclusions in the same dataset [3].

The index is also a single summary of a community, so it discards information about which species are present. It says nothing about taxonomic or functional differences between species, and it cannot tell you whether a change over time is caused by disturbance, succession, or sampling effort. Treat it as one descriptive number among several, not as a complete description of biodiversity.

## Frequently Asked Questions

### What is the difference between the Simpson index and the Shannon index?

Both combine richness and evenness, but they weight species differently. The Simpson index is dominated by the most abundant species, while the Shannon index gives more weight to rare species. If you want a fuller picture, calculate both and compare. A guide to the [Shannon diversity index formula](/blog/data-analysis/shannon-diversity-index-formula) covers the alternative in detail.

### What does a Simpson index of 0.8 mean?

It depends on which form you are reporting. If $D = 0.8$, diversity is low, since a random pair of individuals has an 80 percent chance of being the same species. If $1 - D = 0.8$, diversity is high, since a random pair has an 80 percent chance of being two different species. Always check the definition used.

### Can the Simpson index be greater than 1?

No. Because it is a sum of squared proportions, $D$ is always between 0 and 1 [1]. The inverse form $1/D$ can exceed 1, and it equals the effective number of species [2]. If you see a value above 1 labeled as $D$, the calculation used counts instead of proportions.

### Does the Simpson index need a minimum sample size?

There is no fixed threshold, but small samples undercount rare species and distort the proportions. A common practice is to compare samples of similar size or to rarefy to a common number of individuals. If your counts come from a designed study, work through [sample size calculation](/blog/guides/sample-size-calculation-formulas-and-practical-considerations) before collecting data.

### How do I report the Simpson index in a paper?

State the form you used, give the value to a consistent number of decimals, and report the sample size alongside it. For example, "Simpson diversity ($1 - D$) was 0.735 in Quadrat 1 and 0.186 in Quadrat 2, based on 100 individuals per quadrat." Naming the form removes the ambiguity that causes most comparison errors [2].

## References

1. [10.1: Introduction, Simpson’s Index and Shannon-Weiner Index - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Applied_Statistics/Natural_Resources_Biometrics_(Kiernan)/10%3A_Quantitative_Measures_of_Diversity_Site_Similarity_and_Habitat_Suitability/10.01%3A_Introduction__Simpsons_Index_and_Shannon-Weiner_Index)
2. [Diversity index - Wikipedia](https://en.wikipedia.org/wiki/Diversity_index)
3. [Morris EK, Caruso T, Buscot F, Fischer M, Hancock C, Maier TS, Meiners T, Müller (2014). Choosing and using diversity indices: insights for ecological applications from the German Biodiversity Exploratories. Ecology and evolution](https://pmc.ncbi.nlm.nih.gov/articles/PMC4224527/)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)

## Related Articles

- [Shannon Diversity Index: Formula and Calculation Examples](/blog/data-analysis/shannon-diversity-index-formula)
- [Measures of Variability: Range, Variance and Standard Deviation](/blog/data-analysis/measures-of-variability)
- [Ranking Correlation Coefficient: Spearman and Kendall](/blog/data-analysis/ranking-correlation-coefficient)
- [Multicollinearity: Definition, Detection and Examples](/blog/data-analysis/multicollinearity-definition-detection)
- [ANOVA Table Explained: Components, Formulas and Example](/blog/data-analysis/anova-table-explained)
- [Statistical Range: Definition, Calculation, and Applications](/blog/guides/statistical-range-definition-calculation-and-applications)
- [Fundamental Statistics: Core Concepts Explained](/blog/research-skills/fundamental-statistics-core-concepts-explained)
- [The Essential Statistics Toolkit for Computational Biologists](/blog/careers/the-essential-statistics-toolkit-for-computational-biologists-from-hypothesis-testing-to-bayesian-me)