# How to Calculate the Gini Coefficient: Formula and Example

To calculate the Gini coefficient, you compare how far a distribution of incomes sits from perfect equality. The most direct formula uses the mean absolute difference between all pairs of incomes, divided by twice the mean. This article gives you the formula, a full step-by-step calculation on a small dataset, and the Excel and Python versions of the same number.

## Quick Answer

- The Gini coefficient measures income or wealth inequality on a scale from 0 (perfect equality) to 1 (one person holds everything).
- The core formula is $G = \dfrac{\text{MAD}}{2\bar{x}}$, where MAD is the mean absolute difference between all pairs of incomes and $\bar{x}$ is the mean income.
- For the 8-household example below, the mean absolute difference is 44.6875 and the mean income is 52.75, so $G = 44.6875 / (2 \times 52.75) = 0.4236$.
- An equivalent Excel formula, $(2\sum i x_i - (n+1)\sum x) / (n \sum x)$, returns the same 0.4236.
- The Gini coefficient is a single summary number, so it hides where in the distribution the inequality sits.

## The Formula

The Gini coefficient has several equivalent forms. The one that is easiest to reason about uses the mean absolute difference:

$$G = \frac{\dfrac{1}{n^2}\sum_{i=1}^{n}\sum_{j=1}^{n}\left|x_i - x_j\right|}{2\bar{x}}$$

Each symbol means:

- $x_i$ is the income of the $i$-th unit (a household or person).
- $n$ is the number of units in the sample.
- $\left|x_i - x_j\right|$ is the absolute difference between two incomes. The double sum adds up every pair, including each unit compared with itself, which contributes 0.
- $\dfrac{1}{n^2}$ averages those differences over all $n^2$ pairs. This is the mean absolute difference, or MAD.
- $\bar{x}$ is the mean income, $\dfrac{1}{n}\sum x_i$.
- Dividing by $2\bar{x}$ rescales the result so the maximum possible value is $(n-1)/n$, which approaches 1 as $n$ grows.

A second form is convenient when your data are already sorted from smallest to largest:

$$G = \frac{2\sum_{i=1}^{n} i\,x_{(i)}}{n\sum_{i=1}^{n} x_{(i)}} - \frac{n+1}{n}$$

Here $x_{(i)}$ is the $i$-th smallest income. This is the version most spreadsheet formulas use, and it produces the same answer as the MAD form.

## How to Calculate It Step by Step

1. **List every income.** Write down the income for each household or person. Keep the raw values, not grouped ranges.
2. **Sort the incomes from smallest to largest.** This matters for the second formula, where each value is multiplied by its rank.
3. **Count the units.** That count is $n$.
4. **Add up all incomes.** This total is $\sum x$.
5. **Divide the total by $n$.** That gives the mean income $\bar{x}$.
6. **Compute the mean absolute difference.** For every pair of incomes, take the absolute difference, add them all, then divide by $n^2$. In code this is a matrix of all pairwise differences.
7. **Divide the MAD by twice the mean.** The result is the Gini coefficient.
8. **Check with the ranked formula.** Multiply each sorted income by its rank, sum those products, and apply the second formula. Both methods should agree.

If you want to review the mean before you start, the steps are the same as in [how to calculate the mean](/blog/data-analysis/how-to-calculate-the-mean).

## Worked Example

The dataset is annual household income in thousands of dollars for 8 households.

| household_id | income_thousands |
|---|---|
| 1 | 12 |
| 2 | 18 |
| 3 | 22 |
| 4 | 30 |
| 5 | 45 |
| 6 | 60 |
| 7 | 85 |
| 8 | 150 |

The incomes are already sorted from smallest to largest.

| Step | Value |
|---|---|
| Sorted incomes $x_{(i)}$ | 12.0, 18.0, 22.0, 30.0, 45.0, 60.0, 85.0, 150.0 |
| $n$ | 8 |
| Total income $\sum x$ | 422.0000 |
| Mean income $\bar{x}$ | 52.7500 |
| Mean absolute difference $\frac{1}{n^2}\sum\left|x_i-x_j\right|$ | 44.6875 |
| Gini $= \text{MAD} / (2\bar{x})$ | 44.6875 / (2 × 52.7500) = 0.4236 |
| Excel check $(2\sum i x_i - (n+1)\sum x)/(n\sum x)$ | (2 × 2614.0000 - 9 × 422.0000) / (8 × 422.0000) = 0.4236 |

The two methods agree at 0.4236. That is a moderately high level of inequality for a small sample, driven mostly by the top household at 150.

Here is the same calculation in Python:

```python
import numpy as np
x = np.array([12, 18, 22, 30, 45, 60, 85, 150], dtype=float)
n = len(x)
mad = np.abs(x[:, None] - x[None, :]).sum() / (n * n)
gini = mad / (2 * x.mean())  # 0.4236
print(round(gini, 4))
```

Output:

```text
0.4236
```

## How to Interpret the Result

The Gini coefficient runs from 0 to 1. A value of 0 means every unit has exactly the same income. A value of 1 means one unit has all the income and everyone else has none. Real countries typically fall somewhere between roughly 0.25 and 0.65 depending on the measure and the data source.

A value of 0.4236 means the average absolute gap between two randomly chosen households is about 42 percent of twice the mean income. In plain terms, this sample is noticeably unequal. The lowest household earns 12 and the highest earns 150, a ratio of 12.5 to 1.

The Gini coefficient is a relative measure. If you doubled every income in the dataset, the Gini would stay at 0.4236 because the shape of the distribution did not change. That property makes it useful for comparing inequality across populations of different sizes and currencies, but it also means the number tells you nothing about how rich or poor the group is in absolute terms.

## Doing It in Software

**Excel.** Sort your incomes in ascending order in a column, say A1:A8. Put the rank 1 to 8 in B1:B8. Then:

```text
=(2*SUMPRODUCT(B1:B8,A1:A8)-(COUNT(A1:A8)+1)*SUM(A1:A8))/(COUNT(A1:A8)*SUM(A1:A8))
```

This returns 0.4236 for the example. SUMPRODUCT multiplies each rank by its income and adds the products. The rest of the formula applies the ranked Gini expression.

**Python.** The snippet above uses NumPy. The pairwise difference matrix is built with broadcasting, and the mean absolute difference is the sum divided by $n^2$. If you already have a sorted array, the ranked formula is faster for large samples.

**R.** With a numeric vector `x`, the ranked formula is a few lines. Sort the vector, multiply by the sequence `1:length(x)`, and apply the same expression. Base R has no built-in Gini function, so you write it out.

If you need to compare spread across datasets with different means, the [coefficient of variation](/blog/data-analysis/coefficient-of-variation-formula) is a related but different measure of relative dispersion.

## Common Mistakes

- **Forgetting to sort the data for the ranked formula.** The $i\,x_{(i)}$ term assumes ascending order. If your data are unsorted, the result will be wrong. Fix: sort first, or use the MAD form, which does not care about order.
- **Dividing by the mean instead of twice the mean.** The MAD form needs $2\bar{x}$ in the denominator. Dropping the 2 doubles your answer. Fix: write the denominator as $2\bar{x}$ and check that a perfectly equal dataset returns 0.
- **Using grouped income brackets as if they were individual values.** Bracket midpoints distort the tails. Fix: use unit-level data where possible, and if you must group, use the bracket boundaries carefully.
- **Comparing Gini values computed on different units.** A Gini on households is not the same as a Gini on individuals. Fix: state the unit and keep it consistent across comparisons.
- **Treating the Gini as a poverty measure.** A low Gini does not mean low poverty. Fix: pair the Gini with mean income or a poverty rate.
- **Rounding too early.** Rounding the MAD or the mean before the final division shifts the last digits. Fix: keep full precision until the final step, then round to 3 or 4 decimals.

## Limitations

The Gini coefficient collapses an entire distribution into one number, so two very different income distributions can share the same Gini. A society with a large middle class and a society split into two extremes can both land near 0.42. The single value cannot tell you which is which, and it is not sensitive to where in the distribution the inequality occurs.

The measure is also sensitive to how income is defined and to data quality at the extremes. Top incomes are often underreported, which pulls the Gini down. Different sources may use disposable income, market income, or consumption, and those choices change the number. Estimation methods that use only a few quantiles, such as the top 10 percent and bottom 40 percent income shares, can approximate the sample Gini closely, with absolute errors within a hundredth in empirical tests [1]. That is useful when full microdata are unavailable, but it is still an estimate.

## Frequently Asked Questions

### What is the formula for the Gini coefficient?

The most direct formula is $G = \text{MAD} / (2\bar{x})$, where MAD is the mean absolute difference between all pairs of incomes and $\bar{x}$ is the mean income. An equivalent ranked formula is $G = \frac{2\sum i x_{(i)}}{n\sum x_{(i)}} - \frac{n+1}{n}$. Both give the same result when the data are sorted for the ranked version.

### Can the Gini coefficient be greater than 1?

No. With non-negative incomes, the maximum value occurs when one unit holds all the income. For a sample of $n$ units that maximum is $(n-1)/n$, which approaches 1 in large populations. If you see a value above 1, you likely divided by the mean instead of twice the mean, or you have negative values in the data, which break the standard interpretation.

### What does a Gini coefficient of 0.42 mean?

It means the distribution is moderately unequal. In the example, the mean absolute gap between two randomly chosen households is 44.6875 thousand dollars, which is 42 percent of twice the mean income of 52.75. The top household earns 12.5 times the bottom household.

### How do I calculate the Gini coefficient in Excel?

Sort the incomes ascending in one column and put the ranks 1 to n in the next column. Then use `=(2*SUMPRODUCT(ranks,incomes)-(n+1)*SUM(incomes))/(n*SUM(incomes))`, replacing the ranges with your cells. For the example data this returns 0.4236.

### Why is my Gini coefficient different from another source?

Differences usually come from the income definition, the unit of analysis, and data coverage. Disposable income, market income, and consumption each give different values. Household-level and individual-level data also differ. Always compare Gini values computed on the same definition and unit before drawing conclusions.

## References

1. [Dai P, Shen S. (2025). Estimation of the Gini coefficient based on two quantiles. PloS one](https://pubmed.ncbi.nlm.nih.gov/39932932/)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)
- [Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician](https://doi.org/10.1080/00031305.2016.1154108)

## Related Articles

- [How to Calculate the Mean: Formula and Step by Step Examples](/blog/data-analysis/how-to-calculate-the-mean)
- [How to Calculate Variance: Formula, Steps and Examples](/blog/data-analysis/how-to-calculate-variance)
- [How to Calculate Correlation Coefficient in Excel (Step by Step)](/blog/data-analysis/calculate-correlation-coefficient-excel)
- [How to Calculate the Correlation Coefficient (Step by Step)](/blog/data-analysis/how-to-calculate-correlation-coefficient)
- [How to Calculate a Percentage: Formula and Examples](/blog/data-analysis/how-to-calculate-a-percentage)
- [How to Calculate Dilution Factor: Formulas, Examples, and Serial Dilutions](/knowledge/diagnostics/molecular/how-to-calculate-dilution-factor)