# Monte Carlo Simulation: Definition, Steps and Example

Monte Carlo techniques estimate a numerical answer by running the same random experiment many times and averaging the results. You use them when a probability or a distribution is easy to simulate but hard to compute with a formula. This article defines the method, lists the steps and walks through a full simulation of two dice.

## Quick Answer

- A Monte Carlo method is an algorithm that obtains a numerical approximation using repeated random trials [1].
- You simulate the random process $N$ times, record the outcome of each trial, then average or count the results.
- The estimate improves as $N$ grows, because the average of independent trials converges to the true expected value by the law of large numbers [1].
- The method is most useful when the random variable is easy to simulate but hard to compute analytically [1].
- It gives an approximation, not an exact answer, so you always report the error or the uncertainty around it.

## What Monte Carlo Simulation Means

In plain terms, Monte Carlo simulation is guessing an answer by rolling dice many times instead of solving an equation. If you want to know the chance that a sum of two dice is 7 or 11, you can either count the 36 equally likely outcomes or simulate thousands of rolls and count how often the event happens. The second approach is the Monte Carlo method.

The precise statistical definition is narrower. Let $X$ be a random variable whose expected value $\mathbb{E}[X]$ you want. Draw $N$ independent copies $X_1, \dots, X_N$ of $X$ and take their average:

$$\mathbb{E}[X] = \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N X_i$$

This follows from the law of large numbers [1]. The method dates back to the eighteenth century but took its modern form during the push to develop nuclear weapons in World War II [2]. Stanislaw Ulam coined the name in reference to an uncle who loved playing the odds at the Monte Carlo casino [3].

## How It Works

The mechanism has three parts: a random input, a deterministic rule, and an aggregation step.

1. **Define the random input.** Choose the distribution you will sample from. This could be a uniform draw between 1 and 6 for a die, a normal draw for measurement error, or an exponential draw for waiting times. To use Monte Carlo methods, you need to be able to sample from a given distribution [1].
2. **Apply the deterministic rule.** Feed each random draw through the model. For two dice, the rule is simply $s = d_1 + d_2$. For a financial model, it might be a pricing formula.
3. **Aggregate the results.** Count how often an event occurred, or average the outputs. The count divided by $N$ estimates the probability.

The symbols in the formula above are:

| Symbol | Meaning |
|---|---|
| $X$ | The random variable you want to study |
| $X_i$ | The $i$-th independent copy of $X$ |
| $N$ | The number of simulated trials |
| $\frac{1}{N}\sum X_i$ | The sample mean, your Monte Carlo estimate |

In a stochastic simulation, the answer differs from run to run because there is an element of randomness in it, unlike a deterministic simulation that returns the same result every time [3].

## Worked Example

The dataset below shows the first 10 of 10,000 simulated rolls of two fair six-sided dice, generated with seed 42.

| die1 | die2 | sum |
|---|---|---|
| 1 | 3 | 4 |
| 5 | 2 | 7 |
| 4 | 4 | 8 |
| 3 | 4 | 7 |
| 3 | 2 | 5 |
| 6 | 5 | 11 |
| 1 | 4 | 5 |
| 5 | 1 | 6 |
| 2 | 6 | 8 |
| 1 | 1 | 2 |

The goal is to estimate the probability that the sum is 7 or 11. The exact answer comes from counting outcomes. There are 36 equally likely pairs. Six of them sum to 7, and two sum to 11, so the exact probability is $6/36 + 2/36 = 0.1667 + 0.0556 = 0.2222$.

Now run the simulation. Draw two dice 10,000 times, add them, and count how often the sum is 7 or 11.

```python
import numpy as np
rng = np.random.default_rng(42)
n = 10000
d1 = rng.integers(1, 7, size=n)
d2 = rng.integers(1, 7, size=n)
s = d1 + d2
p_sim = np.mean((s == 7) | (s == 11))
p_exact = 8 / 36
print(f"Simulated P(sum=7 or 11) = {p_sim:.4f}")
print(f"Exact P(sum=7 or 11)     = {p_exact:.4f}")
print(f"Absolute error           = {abs(p_sim - p_exact):.4f}")
```

Output:

```
Simulated P(sum=7 or 11) = 0.2225
Exact P(sum=7 or 11)     = 0.2222
Absolute error           = 0.0003
```

The simulated count of rolls where the sum was 7 or 11 was 2225, so the simulated probability is $2225 / 10000 = 0.2225$. The absolute error is $|0.2225 - 0.2222| = 0.0003$.

The full count table shows how closely the simulation tracks the exact distribution.

| Sum | Simulated count | Exact probability | Exact expected count |
|---|---|---|---|
| 2 | 294 | 0.0278 | 278 |
| 3 | 538 | 0.0556 | 556 |
| 4 | 883 | 0.0833 | 833 |
| 5 | 1107 | 0.1111 | 1111 |
| 6 | 1470 | 0.1389 | 1389 |
| 7 | 1656 | 0.1667 | 1667 |
| 8 | 1319 | 0.1389 | 1389 |
| 9 | 1064 | 0.1111 | 1111 |
| 10 | 834 | 0.0833 | 833 |
| 11 | 569 | 0.0556 | 556 |
| 12 | 266 | 0.0278 | 278 |

Every simulated count sits close to its exact expected value. The histogram of the 10,000 sums would show the familiar triangular shape of two-dice totals, with the orange points marking the exact expected counts.

## How to Interpret It

The simulated probability of 0.2225 is an estimate of the true value 0.2222. The gap of 0.0003 is sampling error, not a mistake. If you reran the code with a different seed, you would get a slightly different number, and the average of many such runs would center on 0.2222.

Read the output as a point estimate plus a sense of its precision. With $N = 10{,}000$ trials and an event probability near 0.22, the standard error of the estimate is roughly $\sqrt{p(1-p)/N} \approx 0.004$. So an estimate of 0.2225 is well within one standard error of the truth.

When you report a Monte Carlo result, state $N$, the seed if reproducibility matters, and the uncertainty. A bare number like "0.2225" hides how much of it is noise.

## When to Use It (and when not to)

Use Monte Carlo simulation when the problem is analytically intractable, meaning no closed-form formula gives the answer [2]. It fits well when you can describe the random inputs but cannot integrate or solve the model by hand. Common uses include estimating cost or time overruns, forecasting commodity prices, and modeling resource exploration [4]. It also appears in clinical research for problems that defy solutions using mathematical theory alone [5].

Do not use it when a formula is available and cheap. For two dice, counting 36 outcomes takes less effort than writing simulation code. Do not use it when your inputs are wrong either, because a Monte Carlo simulation is only as good as its inputs, and accurate empirical data is needed to produce realistic results [3].

Skip it when the answer must be exact, when the model has rare events that $N$ trials will miss, or when each trial is so expensive that you cannot afford enough of them.

## Monte Carlo Simulation vs Analytic Calculation

The closest related idea is the analytic calculation, where you solve the problem with algebra or calculus. The table below contrasts them.

| Aspect | Monte Carlo simulation | Analytic calculation |
|---|---|---|
| Method | Repeated random trials [1] | Closed-form formula or exact counting |
| Output | Approximation with sampling error | Exact value |
| Cost | Grows with the number of trials | Fixed, but can be impossible to derive |
| Best for | Complex or intractable models [2] | Simple, well-understood models |
| Reproducibility | Depends on the seed | Always identical |

A related family of methods, Markov Chain Monte Carlo, samples from distributions that are hard to draw from directly. If your problem involves sampling from a posterior or a complex target distribution, see [Markov Chain Monte Carlo: Definition and Examples](/blog/data-analysis/markov-chain-monte-carlo-definition-examples).

## Common Mistakes

- **Using too few trials.** A run of 100 rolls can be off by several percentage points. Increase $N$ until the standard error is small enough for your decision.
- **Forgetting the seed.** Without a fixed seed, you cannot reproduce your own numbers. Set the seed and record it.
- **Confusing the estimate with the truth.** The simulated 0.2225 is not the exact 0.2222. Report the error or a confidence interval.
- **Sampling from the wrong distribution.** If your input distribution does not match reality, the output will not either [3]. Check the distribution against data first.
- **Ignoring dependence between inputs.** Independent draws are assumed in the basic formula [1]. If inputs are correlated, model the correlation explicitly.
- **Reporting too many digits.** An estimate of 0.2225 does not justify four decimal places of confidence. Round to match the precision.

## Limitations

Monte Carlo simulation cannot give you an exact answer, and it cannot fix a bad model. If the input distributions are wrong, the output is wrong in a way that more trials will not correct. The method also struggles with rare events, because a probability of one in a million needs millions of trials before you see even one occurrence.

The convergence rate is slow. Error shrinks roughly with $1/\sqrt{N}$, so cutting the error in half requires four times the trials. For high-dimensional problems this can make the method expensive, and variance reduction techniques are often needed to make it practical. The method also depends on a good random number generator, since poor randomness biases every result [2].

## Frequently Asked Questions

### What is Monte Carlo simulation in simple terms?

It is a way to estimate an answer by running a random experiment many times and averaging the results. Instead of solving an equation, you simulate the process and count what happens. The more trials you run, the closer the estimate gets to the true value.

### How many trials do I need?

It depends on the precision you want. Error falls with the square root of the number of trials, so to halve the error you need four times as many runs. Start with a pilot run, estimate the standard error, then scale up until it is small enough.

### Why is it called Monte Carlo?

The name comes from the Monte Carlo casino in Monaco. Stanislaw Ulam coined the term in reference to an uncle who loved playing the odds there [3]. The gambling reference fits because the method relies on repeated random draws.

### Is Monte Carlo simulation the same as a random number generator?

No. A random number generator produces the random draws, but the simulation is the whole process of sampling, applying a model and aggregating results. The generator is one component, not the method itself [2].

### Can Monte Carlo simulation give a wrong answer?

Yes. It can be wrong because of sampling error, which shrinks with more trials, or because of a bad model, which does not. If the inputs are unrealistic, the simulation will produce realistic-looking numbers that are still incorrect [3].

If you need exact probabilities for a counting process instead of a simulation, a tool like the [Poisson Distribution Calculator](/tools/poisson-distribution-calculator) can give you the closed-form value directly. For the underlying theory of the distributions you sample from, see [Probability Distributions: Definition, Types and Examples](/blog/data-analysis/probability-distributions-explained).

## References

1. [Monte Carlo Methods, and why they are useful](https://www.math.cmu.edu/~gautam/c/2024-387/notes/01-intro.html)
2. [Introduction To Monte Carlo Simulation - PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC2924739/)
3. [Explained: Monte Carlo simulations | MIT News | Massachusetts Institute of Technology](https://news.mit.edu/2010/exp-monte-carlo-0517)
4. [Introduction to Monte Carlo Methods - Physics 132 Lab Manual](https://openbooks.library.umass.edu/p132-lab-manual/chapter/introduction-to-mc/)
5. [Concato J, Feinstein AR. (1997). Monte Carlo methods in clinical research: applications in multivariable analysis. Journal of investigative medicine : the official publication of the American Federation for Clinical Research](https://pubmed.ncbi.nlm.nih.gov/9291696/)

## Further Reading

- [The Monte Carlo Simulation Method - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Computing_and_Modeling/RTG%3A_Simulating_High_Dimensional_Data/The_Monte_Carlo_Simulation_Method)
- [Monte Carlo Methods | Our Pattern Language](https://patterns.eecs.berkeley.edu/?page_id=186)

## Related Articles

- [Markov Chain Monte Carlo: Definition and Examples](/blog/data-analysis/markov-chain-monte-carlo-definition-examples)
- [Probability Distributions: Definition, Types and Examples](/blog/data-analysis/probability-distributions-explained)
- [Probability Density Function: Definition, Formula and Examples](/blog/data-analysis/probability-density-function-definition-formula)
- [Weibull Distribution: Formula, Parameters and Examples](/blog/data-analysis/weibull-distribution-formula-examples)
- [Bayes' Theorem: Definition, Formula and Examples](/blog/data-analysis/bayes-theorem-definition-formula-examples)
- [Markov Chain Monte Carlo (MCMC) for Biologists](/knowledge/bioinformatics/markov-chain-monte-carlo-mcmc-for-biologists-a-non-technical-introduction-to-how-it-works-and-how-to)
- [Poisson Distribution: Formula and Examples](/blog/research-skills/poisson-distribution-formula-and-examples)
- [Molecular Dynamics Simulations for Beginners: A Step-by-Step Guide to Setting Up and Running Your First Simulation](/knowledge/bioinformatics/molecular-dynamics-simulations-for-beginners-a-step-by-step-guide-to-setting-up-and-running-your-fir)