# What Is Bootstrapping? Resampling Explained with Examples

Bootstrapping is a resampling method that estimates the uncertainty of a statistic by repeatedly drawing samples from your own data, with replacement. Instead of assuming a theoretical distribution, you let the data speak for itself. This article explains the mechanism, walks through a full worked example, and shows when bootstrapping helps and when it misleads.

## Quick Answer

- Bootstrapping builds an empirical sampling distribution by drawing many resamples of size $n$ from your observed data, with replacement [1].
- Each resample can contain the same original observation more than once, and some observations not at all. That repetition is the whole point [1].
- You compute your statistic (mean, median, correlation, whatever) on every resample, then summarize the spread of those values.
- The standard deviation of the bootstrap statistics estimates the standard error, and percentiles give a confidence interval [2].
- It works for statistics where no clean formula exists, and it needs no normality assumption [2].

## What Bootstrapping Means

In plain terms, bootstrapping treats your sample as a stand-in for the population. You pretend the sample is the whole world, then simulate the act of sampling from it over and over.

The precise definition: the bootstrap approximates the sampling distribution of a statistic by resampling observations from the original sample with replacement, computing the statistic on each resample, and using the resulting distribution to estimate variability and construct intervals [1][2].

The name comes from the phrase "pulling yourself up by your own bootstraps." You get an estimate of sampling variability without collecting new data and without a distributional assumption.

## How It Works

The mechanism has four steps.

1. Start with your observed sample $x_1, x_2, \ldots, x_n$ of size $n$.
2. Draw a bootstrap resample $x^*_1, x^*_2, \ldots, x^*_n$ by sampling $n$ values from the original data with replacement.
3. Compute the statistic of interest on the resample, call it $\hat{\theta}^*_b$.
4. Repeat steps 2 and 3 for $B$ resamples, giving $\hat{\theta}^*_1, \ldots, \hat{\theta}^*_B$.

The bootstrap standard error is the standard deviation of those $B$ values:

$$\widehat{SE}_{boot} = \sqrt{\frac{1}{B-1}\sum_{b=1}^{B}\left(\hat{\theta}^*_b - \bar{\theta}^*\right)^2}$$

where $\hat{\theta}^*_b$ is the statistic from resample $b$, $\bar{\theta}^*$ is the mean of all bootstrap statistics, and $B$ is the number of resamples.

A percentile confidence interval takes the $\alpha/2$ and $1-\alpha/2$ quantiles of the bootstrap statistics directly. For a 95% interval, you take the 2.5th and 97.5th percentiles [2].

Sampling with replacement is what makes each resample differ. The probability that any given observation is left out of a resample of size $n$ approaches $e^{-1} \approx 0.368$ as $n$ grows, so roughly 63% of the original rows appear at least once.

## Worked Example

The dataset is a customer satisfaction survey: 30 customers rated a service from 0 to 100.

| customer_id | satisfaction_score |
|---|---|
| 1 | 72 |
| 2 | 85 |
| 3 | 91 |
| 4 | 68 |
| 5 | 77 |
| 6 | 83 |
| 7 | 95 |
| 8 | 62 |
| 9 | 88 |
| 10 | 74 |
| 11 | 79 |
| 12 | 90 |
| 13 | 66 |
| 14 | 81 |
| 15 | 87 |
| 16 | 70 |
| 17 | 93 |
| 18 | 76 |
| 19 | 84 |
| 20 | 69 |
| 21 | 89 |
| 22 | 73 |
| 23 | 80 |
| 24 | 92 |
| 25 | 65 |
| 26 | 78 |
| 27 | 86 |
| 28 | 71 |
| 29 | 94 |
| 30 | 82 |

**Step 1. Sample size.** $n = 30$.

**Step 2. Sample mean.** The scores sum to 2400, so the mean is $2400/30 = 80.0000$.

**Step 3. Sample standard deviation.** Using the $n-1$ denominator, $s = 9.4868$.

**Step 4. Classic standard error.** $SE = s/\sqrt{n} = 9.4868/\sqrt{30} = 1.7321$.

**Step 5. Bootstrap procedure.** Draw 1000 resamples of size 30 with replacement and compute the mean of each.

**Step 6. Bootstrap mean of means.** The average of the 1000 bootstrap means is 79.9015, close to the sample mean of 80.

**Step 7. Bootstrap standard error.** The standard deviation of the 1000 bootstrap means is 1.6572. Compare that with the classic formula value of 1.7321. They agree closely, which is what you expect for a mean from a roughly symmetric sample.

**Step 8. Percentile interval.** The 2.5th percentile of the bootstrap means is 76.9000 and the 97.5th percentile is 83.1342.

**Step 9. Result.** The 95% bootstrap confidence interval for the mean satisfaction score is [76.9000, 83.1342], a width of 6.2342.

Here is the code that produced those numbers.

```python
import numpy as np
scores = [72, 85, 91, 68, 77, 83, 95, 62, 88, 74,
          79, 90, 66, 81, 87, 70, 93, 76, 84, 69,
          89, 73, 80, 92, 65, 78, 86, 71, 94, 82]
rng = np.random.default_rng(42)
boot_means = np.array([rng.choice(scores, size=len(scores), replace=True).mean()
                       for _ in range(1000)])
ci_low, ci_high = np.percentile(boot_means, [2.5, 97.5])
print(f"{ci_low:.4f} {ci_high:.4f}")  # 76.9000 83.1342
```

Output:

```
76.9000 83.1342
```

The histogram of the 1000 bootstrap means is roughly bell shaped and centered near 80. The orange lines mark the interval endpoints at 76.90 and 83.13, and the grey dashed line marks the sample mean of 80.00.

## How to Interpret It

The bootstrap standard error tells you how much your statistic would bounce around if you repeated the study many times. A value of 1.6572 means that sample means from this population would typically land within about 1.7 points of the true mean.

The percentile interval gives a plausible range for the population value. You can say the mean satisfaction score is likely between 76.90 and 83.13. That is the same style of statement a classic confidence interval supports, but it came from resampling instead of a $t$ distribution.

The bootstrap distribution also shows shape. If the histogram is skewed or has two humps, your statistic behaves in a way a symmetric formula would hide. That visual check is often more informative than the interval alone [3].

## When to Use It (and when not to)

Use bootstrapping when:

- Your statistic has no simple standard error formula, such as a median, a ratio of two means, or a correlation coefficient.
- The sample is small and parametric assumptions are shaky. Small cohorts are exactly where parametric estimates of item difficulty and similar metrics become unreliable [4].
- You want to check whether a standard formula is trustworthy for your data [2].
- You are building ensemble models. Bagging trains models on bootstrap samples and averages them, which is the same resampling idea applied to prediction.

Do not use it when:

- The sample is not representative of the population. Resampling a biased sample just reproduces the bias with tight intervals.
- Observations are strongly dependent, such as time series or clustered data, unless you use a block or cluster bootstrap variant.
- $n$ is extremely small. With 5 observations, the resample space is tiny and the interval is unstable.

If you are still building intuition for sampling and estimation, a general introduction to statistical learning covers the surrounding ideas.

## Bootstrapping vs Permutation Testing

Both methods resample, but they answer different questions. The bootstrap estimates the variability of a statistic. A permutation test asks whether two groups differ by shuffling group labels.

| Feature | Bootstrapping | Permutation test |
|---|---|---|
| Question answered | How uncertain is this estimate? | Is there an effect? |
| Resampling method | With replacement from the data | Without replacement, shuffling labels |
| Typical output | Standard error, confidence interval | p-value |
| Assumption | Sample represents the population | Exchangeability under the null |
| Works for a single sample | Yes | No, needs groups to compare |

## Common Mistakes

- **Sampling without replacement.** If you draw each value once, every resample equals the original sample and the spread collapses to zero. Always sample with replacement [1].
- **Using too few resamples.** A few hundred gives a rough picture. Use at least 1000 for standard errors and 2000 or more when you need stable tail percentiles.
- **Forgetting to fix the random seed.** Without a seed, the interval shifts slightly on every run. Set one so your numbers are reproducible.
- **Applying it to dependent data.** Standard bootstrapping assumes independent observations. Time series and grouped data need block or cluster variants.
- **Reporting the bootstrap mean as the estimate.** The point estimate is the statistic from the original sample, 80.0000 here. The bootstrap mean of 79.9015 is a diagnostic, not a replacement.
- **Ignoring a skewed bootstrap distribution.** Percentile intervals can be poor when the distribution is heavily skewed. Consider a bias-corrected variant or transform the statistic.

## Limitations

The bootstrap cannot create information that is not in your sample. If the sample misses a segment of the population, every resample misses it too, and the interval will be confidently wrong. It also assumes the observations are independent draws from a common distribution, which fails for time series, repeated measures, and clustered designs.

Percentile intervals have known weaknesses for skewed statistics and for small samples, where the tails of the bootstrap distribution are poorly estimated. The method also cannot fix a statistic that is undefined or unstable on the original data, such as a regression coefficient from a model with more predictors than observations. Treat the bootstrap as a way to quantify uncertainty in a reasonable estimate, not as a repair for a bad study design.

## Frequently Asked Questions

### How many bootstrap resamples do I need?

For standard errors, 1000 resamples is a common and adequate choice. For confidence interval endpoints, especially the 2.5th and 97.5th percentiles, use 2000 or more so the tails are stable. The cost is only computation time, so running more is cheap.

### Does bootstrapping assume a normal distribution?

No. That is its main appeal. The bootstrap builds the sampling distribution from the data itself, so it does not require normality [2]. It does still assume your observations are independent and that the sample represents the population.

### Can I bootstrap any statistic?

Almost any statistic you can compute on a sample, including medians, correlations, ratios, and regression coefficients. The exceptions are statistics that are undefined or extremely unstable on small resamples, and statistics computed on dependent data without a suitable variant.

### Why is my bootstrap standard error different from the formula value?

They estimate the same quantity by different routes, so small differences are expected. In the worked example the bootstrap gave 1.6572 and the formula gave 1.7321. With 1000 resamples, sampling noise of a few percent is normal. The bootstrap also uses the plug-in variance (dividing by $n$, not $n-1$), so even with a very large $B$ it settles near $1.7321 \times \sqrt{29/30} \approx 1.70$ rather than exactly 1.7321.

### Is bootstrapping the same as cross-validation?

No. Cross-validation splits the data into folds to estimate prediction error on held-out data. Bootstrapping resamples with replacement to estimate the variability of a statistic. They answer different questions and are often used together.

If you want to see the same resampling idea used for prediction instead of inference, read about bagging. To practice the coding side, a beginner's guide to Python for machine learning covers the tools you need.

## References

1. [14.5: Using Simulation for Statistics- The Bootstrap - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Introductory_Statistics/Statistical_Thinking_for_the_21st_Century_(Poldrack)/14%3A_Resampling_and_Simulation/14.05%3A_Using_Simulation_for_Statistics-_The_Bootstrap)
2. [Curran-Everett D. (2009). Explorations in statistics: the bootstrap. Advances in physiology education](https://pubmed.ncbi.nlm.nih.gov/19948676/)
3. [Kulesa A, Krzywinski M, Blainey P et al. (2015). Sampling distributions and the bootstrap. Nature Methods](https://doi.org/10.1038/nmeth.3414)
4. [Mohiyeddini C. (2025). Evaluation of exam questions using bootstrapping: Practical applications in R and SPSS with a case study. Anatomical sciences education](https://pubmed.ncbi.nlm.nih.gov/40662397/)

## Further Reading

- [Bland JM, Altman DG (2015). Statistics Notes: Bootstrap resampling methods. BMJ](https://doi.org/10.1136/bmj.h2622)
- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)

## Related Articles

- [Residual Plots: How to Interpret Them with Examples](/blog/data-analysis/residual-plots-how-to-interpret)
- [Dataset Examples: Types of Data Sets With Real Samples](/blog/data-analysis/dataset-examples-types-of-data-sets)
- [What Is Data Preprocessing? Steps, Techniques and Examples](/blog/data-analysis/what-is-data-preprocessing)
- [Quadratic Regression Analysis: Equation and Example](/blog/data-analysis/quadratic-regression-analysis)
- [Phylogenetic Bootstrap Interpretation](/blog/guides/phylogenetic-bootstrap-interpretation)
- [What is Regression Analysis? A Practical Introduction](/blog/guides/what-is-regression-analysis-a-practical-introduction)
- [Introduction To Statistical Learning](/blog/guides/introduction-to-statistical-learning)