# Simple Random Sampling: Definition, Steps and Examples

A simple random sample is a subset of a population in which every member has an equal chance of being selected and every possible subset of a given size is equally likely. It is the foundation of probability sampling, and most introductory statistical formulas assume you have one. This article explains the definition, the exact steps for drawing one, and how to check that your sample behaves as expected.

## Quick Answer

- A simple random sample (SRS) of size $n$ is drawn so that every set of $n$ individuals in the population has an equal chance of being the sample selected [1].
- Every member of the population has the same probability of inclusion, and the sample is self-weighting, meaning the inverse of the selection probability is equal for all units [2].
- You need a complete list of the population, called a sampling frame, and a random number generator or random digit table [1].
- Drawing without replacement gives each unit a selection probability of $n/N$, where $N$ is the population size.
- An SRS is not the same as any equal-probability sample. Systematic and multistage samples can give each unit the same inclusion probability while giving different sets of units different probabilities [2].

## What a Simple Random Sample Means

In plain terms, a simple random sample is a selection method where no member of the population is favored and no combination of members is favored. If you have 50 students and you want a sample of 10, an SRS makes all 50 students equally likely to appear, and it makes every possible group of 10 equally likely to be the group you end up with.

The precise statistical definition is stricter. Moore and McCabe define a simple random sample of size $n$ as a sample of $n$ individuals from the population chosen in such a way that every set of $n$ individuals has an equal chance to be the sample actually selected [1]. That second clause matters. Equal chance for each individual is necessary but not sufficient. The whole subset must be equally likely.

This is why an SRS is an epsem sample (equal probability of selection method), but not every epsem sample is an SRS [2]. If a teacher arranges a class in 5 rows of 6 columns and picks one column at random to get 5 students, each student has the same chance of being picked, but only subsets arranged as a single column can ever be selected [2]. The individual probabilities are equal, the subset probabilities are not.

## How It Works

The mechanism is straightforward once you have a sampling frame, which is a numbered list of every unit in the population [1]. Number the units 1 through $N$, then generate $n$ distinct random numbers in that range and take the units they point to.

The selection probability for each unit is:

$$P(\text{unit } i \text{ selected}) = \frac{n}{N}$$

where $N$ is the population size and $n$ is the sample size. This holds when you sample without replacement, so no unit can appear twice.

Two symbols describe the outcome you compare against:

- $\mu$ is the population mean, the average of all $N$ values.
- $\bar{x}$ is the sample mean, the average of the $n$ values you drew.

The sample mean is an estimate of the population mean. Because selection is random, $\bar{x}$ will usually differ from $\mu$, and that difference is sampling error, not a mistake. The distribution of the number of units of a given type in the sample depends on the population composition. With replacement it follows a binomial distribution, and without replacement it follows a hypergeometric distribution [2]. If you want to see how sample averages behave across many draws, the [sample mean](/blog/data-analysis/sample-mean) article covers the mechanics.

## Worked Example

The dataset is 50 exam scores from a class population on a 0 to 100 scale. We draw a simple random sample of 10 students without replacement and compare the sample mean to the population mean.

| student_id | score | student_id | score | student_id | score | student_id | score | student_id | score |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 58 | 11 | 75 | 21 | 81 | 31 | 86 | 41 | 91 |
| 2 | 62 | 12 | 76 | 22 | 82 | 32 | 87 | 42 | 92 |
| 3 | 65 | 13 | 76 | 23 | 82 | 33 | 87 | 43 | 92 |
| 4 | 67 | 14 | 77 | 24 | 83 | 34 | 88 | 44 | 93 |
| 5 | 68 | 15 | 78 | 25 | 83 | 35 | 88 | 45 | 94 |
| 6 | 70 | 16 | 78 | 26 | 84 | 36 | 89 | 46 | 95 |
| 7 | 71 | 17 | 79 | 27 | 84 | 37 | 89 | 47 | 96 |
| 8 | 72 | 18 | 80 | 28 | 85 | 38 | 90 | 48 | 97 |
| 9 | 73 | 19 | 80 | 29 | 85 | 39 | 90 | 49 | 98 |
| 10 | 74 | 20 | 81 | 30 | 86 | 40 | 91 | 50 | 99 |

Step 1. Set the population size and sample size. $N = 50$ and $n = 10$.

Step 2. Compute the population mean. The scores sum to 4127, so $\mu = 4127/50 = 82.5400$. The population standard deviation, dividing by $N$, is $\sigma = 9.5377$.

Step 3. Draw the sample without replacement. Using a seeded random generator, the selected zero-based positions are 4, 32, 3, 28, 19, 44, 33, 49, 9 and 39, which are student IDs 5, 33, 4, 29, 20, 45, 34, 50, 10 and 40.

Step 4. Record the sampled scores. They are 68, 87, 67, 85, 81, 94, 88, 99, 74 and 91.

Step 5. Compute the sample mean. The sampled scores sum to 834, so $\bar{x} = 834/10 = 83.4000$. The sample standard deviation is $s = 10.8033$.

Step 6. Compare. The difference is $\bar{x} - \mu = 83.4000 - 82.5400 = 0.8600$. The sample mean sits 0.86 points above the population mean, which is well within what random variation produces at this sample size.

```python
import numpy as np
scores = np.array([...])  # 50 exam scores
rng = np.random.default_rng(42)
sample = rng.choice(scores, size=10, replace=False)
print(f"Population mean = {scores.mean():.4f}")
print(f"Sample mean     = {sample.mean():.4f}")
print(f"Difference      = {sample.mean() - scores.mean():.4f}")
```

Output:

```text
Population mean = 82.5400
Sample mean     = 83.4000
Difference      = 0.8600
```

## How to Interpret It

The sample mean of 83.40 is your estimate of the population mean of 82.54. The 0.86-point gap is sampling error. It does not mean the draw was flawed, and it does not mean the population mean is wrong. It means a sample of 10 carries limited information about a population of 50.

The right way to read the result is as a range, not a point. With a sample standard deviation of 10.8033 and $n = 10$, the standard error of the mean is about 3.42 points, so the estimate is fairly imprecise. Larger samples shrink that error. The [Student's t-distribution](/blog/data-analysis/students-t-distribution-definition-formula) is the tool for building a confidence interval around $\bar{x}$ when the population standard deviation is unknown.

One draw tells you little about the sampling method itself. To judge whether your procedure is unbiased, you would repeat the draw many times and look at the distribution of the sample means. That distribution should center on $\mu$.

## When to Use It (and when not to)

Use a simple random sample when you have a complete, accurate sampling frame, when the population is reasonably homogeneous, and when you can reach the selected units. It is the cleanest design because the math is simple and the estimator is unbiased.

Do not use it when the population is naturally grouped and travel or cost makes a full random draw impractical. A multistage random sample takes a series of simple random samples in stages, which is often more practical for on-location work such as door-to-door surveys [3]. Do not use it when the population splits into subpopulations that differ on the measurement of interest, because a plain SRS can under-represent a small group. A [stratified sample](/knowledge/diagnostics/research-methods/stratified-sampling-for-heterogeneous-populations-a-practical-guide-for-life-science-studies) draws from each subgroup to keep the sample representative [3]. If your units cluster geographically, [cluster sampling](/blog/data-analysis/what-is-cluster-sampling) is usually cheaper.

## Simple Random Sample vs Stratified Sample

Both are probability methods, and both give every unit a nonzero chance of selection. The difference is how the population is handled before selection.

| Feature | Simple random sample | Stratified sample |
|---|---|---|
| Population treatment | Treated as one pool | Split into strata first |
| Selection | Random draw from the whole list | Random draw within each stratum |
| Best when | Population is homogeneous | Subgroups differ on the outcome |
| Main advantage | Simple, unbiased, easy math | Guarantees representation of each group |
| Main cost | Small groups may be missed | Needs stratum information for every unit |

A stratified sample is obtained by taking samples from each stratum or subgroup of a population [3]. If you expect the measurement to vary across subgroups, stratification controls that variation directly. An SRS does not.

## Common Mistakes

- **Confusing equal individual probability with an SRS.** Picking one column of a grid gives every student the same chance, but only column-shaped subsets can be chosen [2]. Fix: check that every subset of size $n$ is possible, not just that each unit is equally likely.
- **Sampling with replacement by accident.** If you allow repeats, a unit can appear twice and the selection probability changes. Fix: draw distinct indices and confirm the sample has no duplicates.
- **Using a biased or incomplete frame.** If the list omits part of the population, no amount of random drawing fixes it [1]. Fix: audit the frame against the target population before drawing.
- **Sorting the list before drawing.** A list ordered by an outcome-related variable can interact badly with a systematic shortcut. Fix: shuffle or use a proper random number generator.
- **Treating one sample mean as the truth.** A single draw of 10 gave 83.40 against a true 82.54. Fix: report uncertainty, not just the point estimate.
- **Assuming any random-looking method is an SRS.** Systematic and multistage designs can be epsem while giving different sets of units different probabilities [2]. Fix: name your design accurately and use the matching variance formula.

## Limitations

A simple random sample cannot fix a bad sampling frame. If the list of the population is incomplete or out of date, the sample inherits that bias, and the compromises made to build a workable frame can easily produce a sample that is biased or not close enough to random to be suitable [1]. The method also assumes you can actually contact the units you select, which fails for dispersed or hard-to-reach populations.

The method is also inefficient when the population is heterogeneous or geographically spread. You spend effort reaching units that add little information, and small subgroups can be missed entirely by chance. In clinical and survey research, simple random sampling is often replaced by systematic, stratified or cluster designs for exactly these reasons [4]. Finally, an SRS gives no guarantee of representativeness in any single draw. It guarantees fairness of the procedure, not balance in the outcome.

## Frequently Asked Questions

### What is the difference between a random sample and a simple random sample?

A random sample is any sample drawn by a chance mechanism. A simple random sample is the specific case where every subset of size $n$ has an equal chance of being selected [1]. All simple random samples are random samples, but not all random samples are simple random samples.

### How do you draw a simple random sample step by step?

Number every unit in the population from 1 to $N$. Generate $n$ distinct random numbers in that range using a random number generator or random digit table. Select the units matching those numbers [1]. Confirm no unit appears twice and that the sample size matches your plan.

### Does a simple random sample need replacement?

No. Standard practice is to sample without replacement so each unit appears at most once. The two versions differ in their distributions. With replacement the count of a given type follows a binomial distribution, and without replacement it follows a hypergeometric distribution [2].

### Is a systematic sample a simple random sample?

No. Systematic random sampling gives each unit the same probability of inclusion, but different sets of units have different probabilities of being selected [2]. That makes it an epsem sample, not an SRS. It is often more practical, but you must use the correct variance formula for it.

### Why is my sample mean different from the population mean?

Because selection is random. In the worked example the sample mean was 83.40 against a population mean of 82.54, a gap of 0.86 points. That gap is sampling error and shrinks as the sample size grows. It is expected behavior, not evidence of a broken procedure.

## References

1. [Simple Random Samples](https://web.ma.utexas.edu/users/mks/statmistakes/SRS.html)
2. [Simple random sample - Wikipedia](https://en.wikipedia.org/wiki/Simple_random_sample)
3. [Sampling](http://www.stat.yale.edu/Courses/1997-98/101/sample.htm)
4. [Zrineh A, Al-Usta M, Alwawi A. (2026). Sampling Methods and Sample Size Determination in Clinical Research: An Educational Review. Journal of general and family medicine](https://pmc.ncbi.nlm.nih.gov/articles/PMC12897549/)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)

## Related Articles

- [What Is Cluster Sampling? Definition and Examples](/blog/data-analysis/what-is-cluster-sampling)
- [What Is Random Forest? Algorithm and Examples](/blog/data-analysis/what-is-random-forest)
- [Probability Distributions: Definition, Types and Examples](/blog/data-analysis/probability-distributions-explained)
- [Sample Mean: Definition, Formula and Examples](/blog/data-analysis/sample-mean)
- [Student's t-Distribution: Definition, Formula and Examples](/blog/data-analysis/students-t-distribution-definition-formula)
- [Random Sampling in Biological Research](/knowledge/diagnostics/research-methods/random-sampling-in-biological-research-methods-assumptions-and-practical-implementation)
- [Randomized Experiments: Why Randomization Matters](/blog/guides/randomized-experiments-why-randomization-matters)
- [Randomized Experiment Design: A Practical Guide for Reducing Bias](/blog/guides/randomized-experiment-design-a-practical-guide-for-reducing-bias)