# Law of Large Numbers: Definition and Examples

The law of large numbers says that as you collect more independent observations from the same distribution, their average gets closer to the true mean of that distribution. Small samples bounce around. Large samples settle down. That single idea is why pollsters, insurers and quality engineers trust averages from big samples.

## Quick Answer

- The law of large numbers (LLN) states that the sample average of independent, identically distributed observations converges to the population mean as the number of observations grows [1].
- The sample average after $n$ observations is $\bar{X}_n = \frac{1}{n}(X_1 + X_2 + \cdots + X_n)$.
- The weak law describes convergence in probability, and the strong law describes convergence with probability 1 [1].
- Convergence is about the average, not about individual values. A single new observation can still be far from the mean.
- The law requires a finite expected value and independent observations from the same distribution. It says nothing about how fast the average settles.

## What the Law of Large Numbers Means

In plain terms, the law of large numbers is a promise about averages. Flip a fair coin ten times and you might get seven heads. Flip it ten thousand times and the proportion of heads will sit very close to one half. The average of many independent draws from one distribution drifts toward the mean of that distribution.

The precise statistical definition comes in two versions. Let $X_1, X_2, \ldots$ be independent and identically distributed random variables with finite expected value $E[X] = \mu$. Define the partial sum $S_n = X_1 + \cdots + X_n$ [1]. The strong law of large numbers states that $S_n / n$ converges to $\mu$ with probability 1 as $n$ goes to infinity [1]. The weak law states the same limit in the weaker sense of convergence in probability, meaning the chance that the average differs from $\mu$ by more than any fixed amount shrinks to zero [2].

The two versions differ only in the type of convergence. For practical data work, both say the same thing: with enough data, the sample mean is a reliable estimate of the population mean.

## How It Works

The mechanism is simple arithmetic plus independence. Write the sample average as

$$\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i$$

Each symbol means the following.

- $X_i$ is the value of the $i$-th observation, drawn from the same distribution.
- $n$ is the number of observations collected so far.
- $\sum_{i=1}^{n} X_i$ is the running total of all observations.
- $\bar{X}_n$ is the running average after $n$ observations.
- $\mu$ is the true population mean, the fixed target the average approaches.

Independence matters because it keeps any single observation from dominating the total. As $n$ grows, each new value contributes a smaller share, so the average smooths out. The variance of the sample mean is $\sigma^2 / n$, where $\sigma^2$ is the population variance. That shrinking variance is the engine behind the convergence.

## Worked Example

The dataset is the first 10 rolls of a fair six-sided die, generated with a fixed random seed so the sequence is reproducible.

| roll_number | value |
| --- | --- |
| 1 | 1 |
| 2 | 5 |
| 3 | 4 |
| 4 | 3 |
| 5 | 3 |
| 6 | 6 |
| 7 | 1 |
| 8 | 5 |
| 9 | 2 |
| 10 | 1 |

The true mean of a fair six-sided die is $(1+2+3+4+5+6)/6 = 3.5$. Now track the running average as rolls accumulate.

- Running average after 1 roll: $\text{sum(first 1)} / 1 = 1.0000$
- Running average after 10 rolls: $\text{sum(first 10)} / 10 = 3.1000$
- Running average after 100 rolls: $\text{sum(first 100)} / 100 = 3.6400$
- Running average after 1000 rolls: $\text{sum(first 1000)} / 1000 = 3.4880$
- Absolute error at $n = 1000$: $|3.4880 - 3.5| = 0.0120$

The first roll gives an average of 1.0000, which is far from 3.5. After 10 rolls the average is 3.1000. After 100 rolls it is 3.6400, and after 1000 rolls it is 3.4880, only 0.0120 away from the true mean. The path is not smooth. It overshoots at 100 rolls and then comes back. That wobble is normal and does not contradict the law.

```python
import numpy as np
rng = np.random.default_rng(42)
rolls = rng.integers(1, 7, size=1000)
running_avg = np.cumsum(rolls) / np.arange(1, 1001)
print(running_avg[999])  # 3.4880
```

Output:

```
3.488
```

## How to Interpret It

Read the law as a statement about long-run behavior, not about any single draw. The average after 1000 rolls landed within 0.0120 of the true mean, but the 1001st roll could still be a 6. The law constrains the aggregate, not the next value.

The convergence is also not monotone. In the example, the average at 100 rolls (3.6400) is farther from 3.5 than the average at 10 rolls (3.1000). Averages can move away from the target before moving back. What shrinks is the typical size of the gap, not the gap at every step.

Sample size is the lever. If you want a tighter estimate of a mean, collect more data. The variance of the average falls as $1/n$, so cutting the typical error in half takes roughly four times the data.

## When to Use It (and when not to)

Use the law of large numbers whenever you average independent observations from a stable process and want to argue that the average estimates a true mean. It underpins polling margins, insurance pricing, casino economics and Monte Carlo simulation. It is also the reason a long-run average cost or defect rate stabilizes.

Do not use it when observations are dependent. If today's value is correlated with yesterday's, the average can converge to something other than the population mean, or fail to converge at all. Do not use it when the distribution changes over time. A process with a drifting mean has no single fixed target. Do not use it when the expected value is infinite, as with some heavy-tailed distributions, because the average has nothing finite to converge to.

## Law of Large Numbers vs Central Limit Theorem

These two results are often confused. The law of large numbers is about where the average goes. The central limit theorem is about how the average is distributed around that target for large but finite samples.

| Feature | Law of Large Numbers | Central Limit Theorem |
| --- | --- | --- |
| Main claim | The sample average converges to the population mean | The standardized sample average approaches a normal distribution |
| Focus | The limit value | The shape of the sampling distribution |
| Sample size | Concerns behavior as $n$ grows without bound | Gives an approximation for large finite $n$ |
| Typical use | Justifying an average as an estimate | Building confidence intervals and tests |

You often use both together. The law tells you the average is aiming at the right target. The central limit theorem tells you how much spread to expect around it.

## Common Mistakes

- Treating the law as a correction mechanism. It does not mean a run of low rolls makes high rolls more likely. Each roll stays independent, and the fix is to stop expecting the past to balance out.
- Applying it to small samples. Ten rolls gave an average of 3.1000, which is 0.4000 off. The law says nothing useful at that size, so check whether your sample is large enough for the precision you need.
- Ignoring dependence. Averages of correlated data can converge to the wrong value. Verify independence before trusting the average.
- Confusing the average with individual outcomes. The average of 1000 rolls was 3.4880, but any single roll is still between 1 and 6. Keep the two levels separate.
- Assuming a fixed rate of convergence. The law gives no speed guarantee. The gap at 100 rolls was larger than at 10 rolls, so do not read a smooth path into the result.
- Using it with a changing distribution. If the underlying process shifts, the target moves too. Confirm the process is stable before averaging.

## Limitations

The law of large numbers says nothing about how many observations you need. Convergence is asymptotic, so it describes the limit as $n$ grows without bound. In practice you never reach that limit, and the gap at any finite sample size can be larger than you expect. The example shows a gap of 0.0120 at 1000 rolls, but a different seed would give a different gap.

The law also assumes independent, identically distributed observations with a finite mean [1]. Real data often violates one of these conditions. Time series, clustered samples and heavy-tailed variables all break the assumptions, and the average may converge slowly or to the wrong value. The law is a guarantee about the ideal case, not a promise about your dataset.

## Frequently Asked Questions

### What is the law of large numbers in simple terms?

It says that the more independent observations you average, the closer that average gets to the true mean of the population. A few rolls of a die can average anything from 1 to 6, but thousands of rolls average close to 3.5. The idea applies to any process where you can repeat independent trials.

### What is the difference between the weak and strong law of large numbers?

Both say the sample average converges to the population mean. The weak law describes convergence in probability, so the chance of a large deviation shrinks toward zero [2]. The strong law describes convergence with probability 1, a stronger statement about the entire sequence of averages [1]. For applied work the practical conclusion is the same.

### Does the law of large numbers mean outcomes even out?

No. It means the average settles near the true mean, not that past deviations get corrected. A run of low die rolls does not make high rolls more likely. The average moves toward 3.5 because new values are averaged in, not because the process remembers earlier results.

### How many samples do I need for the law of large numbers to work?

There is no fixed number. The law is a statement about the limit as sample size grows, so it never names a threshold. How large your sample must be depends on the variability of the data and the precision you want. More variable data needs more observations for the same accuracy.

### Can the law of large numbers fail?

Yes. It fails when observations are dependent, when the distribution changes over time, or when the expected value is infinite. In those cases the average may not converge to a stable value. Checking these assumptions before relying on an average is part of good practice.

If you work with probability distributions and want to compare observed frequencies against expected ones, [Benford's Law: Definition, Formula and Examples](/blog/data-analysis/benfords-law-definition-formula) shows a related convergence idea applied to leading digits.

## References

1. [4.2: The Strong Law of Large Numbers and Convergence WP1 - Engineering LibreTexts](https://eng.libretexts.org/Bookshelves/Electrical_Engineering/Signal_Processing_and_Modeling/Discrete_Stochastic_Processes_(Gallager)/04%3A_Renewal_Processes/4.02%3A_The_Strong_Law_of_Large_Numbers_and_Convergence_WP1)
2. [](https://math.mit.edu/~sheffield/440/Lecture30.pdf)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)
- [Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician](https://doi.org/10.1080/00031305.2016.1154108)

## Related Articles

- [Benford's Law: Definition, Formula and Examples](/blog/data-analysis/benfords-law-definition-formula)
- [How to Calculate a Percentage: Formula and Examples](/blog/data-analysis/how-to-calculate-a-percentage)
- [What Is an Independent Variable? Definition and Examples](/blog/data-analysis/what-is-an-independent-variable)
- [What Is a Dichotomous Variable? Definition and Examples](/blog/data-analysis/dichotomous-variable-definition-examples)
- [How to Calculate the Mean: Formula and Step by Step Examples](/blog/data-analysis/how-to-calculate-the-mean)