# What Is Cluster Sampling? Definition and Examples

A cluster sample is a probability sample in which the population is split into groups called clusters, and you randomly select entire clusters instead of individual units. Everyone inside each selected cluster is measured. This makes data collection cheaper and faster when a complete list of individuals is hard to build but a list of groups is easy to get.

## Quick Answer

- A cluster sample selects whole groups (schools, villages, clinics) at random, then studies every unit inside the chosen groups.
- It is a probability method, so you can still estimate standard errors and confidence intervals.
- It works best when clusters are easy to list and reach, and when travel between units is expensive.
- Cluster sampling usually has a larger standard error than simple random sampling of the same size, because people in one cluster tend to be similar.
- The design effect (deff) measures that penalty. A deff of 1.27 means the cluster design needs about 27% more sample to match simple random sampling.

## What Cluster Sampling Means

In plain terms, you do not pick people one by one. You pick groups, then take everyone in the groups you picked. If you want to survey students in a city, you might randomly select 5 schools and survey every student in those 5 schools. The school is the cluster.

The precise statistical definition: cluster sampling is a probability sampling design in which the population is partitioned into $N$ non-overlapping groups called primary sampling units (PSUs), a probability sample of $n$ clusters is drawn, and all elements within each selected cluster are observed. When every element in a selected cluster is measured, the design is one-stage cluster sampling. If you then subsample within clusters, it becomes two-stage or multistage sampling.

The key feature is that the sampling unit (the cluster) is not the same as the observation unit (the person or item you measure). That mismatch is what drives the math and the extra uncertainty.

## How It Works

The mechanism is a two-step random draw. First you sample clusters. Then you take all units inside them. The estimate of the population mean is the mean of the cluster means when clusters are equal-sized:

$$\bar{y}_{cl} = \frac{1}{n}\sum_{i=1}^{n}\bar{y}_i$$

where:

- $\bar{y}_{cl}$ is the cluster sample estimate of the population mean
- $n$ is the number of clusters selected
- $\bar{y}_i$ is the mean of cluster $i$

The variance of that estimate uses the spread between cluster means, not the spread between individuals:

$$\widehat{Var}(\bar{y}_{cl}) = \left(1 - \frac{n}{N}\right)\frac{s_b^2}{n}$$

where:

- $N$ is the total number of clusters in the population
- $1 - n/N$ is the finite population correction (fpc)
- $s_b^2$ is the sample variance of the cluster means, computed with $n-1$ in the denominator

The standard error is the square root of that variance. Survey software applies the same logic. In SAS, for example, the `n =` option on `proc surveymeans` states how many PSUs were available in the population, and that total is used to calculate the fpc [1].

## Worked Example

The dataset below describes 20 schools, each with a mean test score and an enrollment count. Five schools were selected as clusters, and every student in those schools was measured.

| school_id | mean_score | enrollment |
|---|---|---|
| S01 | 72.4 | 48 |
| S02 | 68.1 | 52 |
| S03 | 75.6 | 45 |
| S04 | 70.2 | 50 |
| S05 | 66.8 | 55 |
| S06 | 73.9 | 47 |
| S07 | 69.5 | 51 |
| S08 | 77.1 | 43 |
| S09 | 71.3 | 49 |
| S10 | 64.7 | 58 |
| S11 | 74.8 | 46 |
| S12 | 67.4 | 53 |
| S13 | 76.2 | 44 |
| S14 | 70.9 | 50 |
| S15 | 65.5 | 56 |
| S16 | 72.8 | 48 |
| S17 | 69.1 | 52 |
| S18 | 78.3 | 42 |
| S19 | 71.7 | 49 |
| S20 | 66.1 | 54 |

Step 1. Compute the population mean, weighted by school size. The result is 70.7858.

Step 2. The selected clusters are S02, S06, S10, S14 and S18.

Step 3. Their cluster means are 68.10, 73.90, 64.70, 70.90 and 78.30.

Step 4. The cluster sample estimate is the mean of those five values:

$$(68.10 + 73.90 + 64.70 + 70.90 + 78.30) / 5 = 71.1800$$

Step 5. The between-cluster variance, using the sample formula with $n-1$, is 27.4120.

Step 6. The finite population correction is $1 - 5/20 = 0.7500$.

Step 7. The variance of the cluster mean is $0.7500 \times 27.4120 / 5 = 4.1118$.

Step 8. The standard error is $\sqrt{4.1118} = 2.0278$.

Step 9. For comparison, simple random sampling on the same data gives a standard error of 1.7962.

Step 10. The design effect is $4.1118 / 3.2264 = 1.2744$.

Step 11. The 95% confidence interval is $71.1800 \pm 1.9600 \times 2.0278$, which is [67.2057, 75.1543].

Here is the same calculation in Python:

```python
import statistics, math
cluster_means = [68.1, 73.9, 64.7, 70.9, 78.3]
est = statistics.mean(cluster_means)
var_b = statistics.variance(cluster_means)
fpc = 1 - 5/20
se = math.sqrt(fpc * var_b / 5)
print(est, se)  # 71.1800 2.0278
```

Output:

```
71.18 2.027757381937001
```

## How to Interpret It

The point estimate of 71.18 sits close to the true population mean of 70.79, so this particular sample performed well. The standard error of 2.03 tells you how much the estimate would move across repeated cluster samples of the same design.

The design effect is the number to watch. A deff of 1.2744 means the cluster design carries about 27% more variance than simple random sampling of the same number of units. If you wanted the same precision as simple random sampling, you would need roughly 27% more observations. That penalty exists because students within a school share teachers, neighborhoods and resources, so their scores are correlated.

The confidence interval [67.21, 75.15] is wider than a simple random sampling interval would be. That width is the honest cost of the cheaper field design. When you report results, always report the design-based standard error, not the one you would get by pretending the data came from a simple random sample. If you want to see how the underlying average is computed, the [sample mean guide](/blog/data-analysis/sample-mean) walks through the formula.

## When to Use It (and when not to)

Use cluster sampling when:

- You cannot build a complete list of individuals, but you can list groups such as schools, clinics or blocks.
- Travel or interviewer time is the main cost, so visiting fewer locations matters more than visiting fewer people.
- The population is naturally grouped and the groups are geographically spread out.
- You need a probability design that field teams can actually execute.

Do not use it when:

- A complete sampling frame of individuals already exists and is cheap to use. Simple random sampling or [stratified sampling](/blog/guides/types-of-sampling-in-biostatistics) will be more precise.
- Clusters are very homogeneous inside. High within-cluster similarity inflates the design effect and wastes budget.
- You need precise estimates for small subgroups. Cluster designs often leave too few clusters per subgroup.
- The rare characteristic you care about sits on clustered units. In that case adaptive cluster sampling, which adds nearby units when the trait appears, can be a better fit [2].

## Cluster Sampling vs Stratified Sampling

These two designs are often confused because both split the population into groups. The difference is what happens next.

| Feature | Cluster sampling | Stratified sampling |
|---|---|---|
| Groups are called | Clusters | Strata |
| How groups are used | Randomly select some groups, measure all units inside | Select units from every group |
| Groups should be | Internally heterogeneous, similar to each other | Internally homogeneous, different from each other |
| Main goal | Lower cost and easier fieldwork | Higher precision |
| Effect on variance | Usually increases it | Usually reduces it |
| Typical example | Randomly pick 5 schools, test all students | Split by grade, sample students within each grade |

The rule of thumb: clusters should look like small versions of the whole population, and strata should be internally alike. Getting this backwards is one of the most common design errors.

## Common Mistakes

- **Treating the cluster estimate like a simple random sample estimate.** The standard error will be too small and your confidence intervals too narrow. Fix: use the between-cluster variance formula or survey software that accounts for the design.
- **Forgetting the finite population correction.** When you sample a meaningful share of clusters, the fpc shrinks the variance. Fix: record $N$, the total number of clusters, and apply $1 - n/N$.
- **Choosing clusters that are too similar to each other.** If all selected schools are in one wealthy district, the estimate is biased for the city. Fix: randomize cluster selection from a complete list, and consider stratifying clusters by region first.
- **Confusing clusters with strata.** Selecting a few groups and measuring everyone is cluster sampling. Sampling from every group is stratification. Fix: decide whether you want to reduce cost or reduce variance, then pick the matching design.
- **Ignoring unequal cluster sizes.** When clusters differ a lot in size, the simple mean of cluster means is no longer the best estimate. Fix: weight each cluster mean by its size, or use a self-weighting design.
- **Reporting the number of people instead of the number of clusters.** Precision depends mainly on how many clusters you sampled, not how many people are inside them. Fix: report both, and base the standard error on the cluster count.

## Limitations

Cluster sampling cannot fix a bad sampling frame. If the list of clusters is incomplete or outdated, the design loses its probability basis and the estimates become biased in ways no formula can repair. It also cannot rescue a small number of clusters. With only a handful of clusters, the between-cluster variance is estimated from very few values, so standard errors and confidence intervals become unstable and may understate the true uncertainty.

The method also tends to be less precise than simple random sampling for the same number of measured units. That is a structural trade-off, not a mistake. You accept wider intervals in exchange for lower field costs. If your clusters are highly internally similar, the penalty can be large enough that the cost savings no longer justify the loss of precision, and a different design will serve you better.

## Frequently Asked Questions

### What is a cluster sample in simple terms?

It is a sample where you randomly pick whole groups and then study everyone in those groups. Instead of selecting 200 students across a city one by one, you select 5 schools and survey all students in them. The group is the sampling unit, and the person is the observation unit.

### What is the difference between a cluster sample and a stratified sample?

In cluster sampling you select some groups and measure everything inside them. In stratified sampling you divide the population into groups and then draw units from every group. Clusters should be internally mixed and similar to each other. Strata should be internally alike and different from each other.

### Why is cluster sampling less precise than simple random sampling?

People in the same cluster tend to be similar, so each new person inside a cluster adds less new information than a person drawn from a different cluster. This correlation inflates the variance, which shows up as a design effect greater than 1. In the worked example the design effect was 1.2744.

### How many clusters should I sample?

More clusters generally beat more people per cluster. Precision depends mainly on the number of clusters, so if your budget is fixed, spread it across more clusters and measure fewer units in each. A pilot study or prior data on the between-cluster variance helps you size the design properly.

### Can I analyze cluster sample data in standard software?

Yes, but you must tell the software about the design. Survey procedures let you declare the cluster variable and the population cluster total so the finite population correction is applied [1]. If you analyze the data as if it were a simple random sample, your standard errors will be too small.

## References

1. [How do I analyze survey data with a one-stage cluster design? | SAS FAQ](https://stats.oarc.ucla.edu/sas/faq/how-do-i-analyze-survey-data-with-a-one-stage-cluster-design/)
2. [Adaptive Cluster Sampling for Forest Inventories | US Forest Service Research and Development](https://research.fs.usda.gov/treesearch/625)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)

## Related Articles

- [Simple Random Sampling: Definition, Steps and Examples](/blog/data-analysis/simple-random-sampling-definition-examples)
- [Sample Mean: Definition, Formula and Examples](/blog/data-analysis/sample-mean)
- [What Is Data Aggregation? Definition and Examples](/blog/data-analysis/what-is-data-aggregation)
- [Dataset Examples: Types of Data Sets With Real Samples](/blog/data-analysis/dataset-examples-types-of-data-sets)
- [K-Means Clustering: How It Works With a Worked Example](/blog/data-analysis/k-means-clustering-how-it-works)
- [Cluster Sampling in Field Biology](/knowledge/diagnostics/research-methods/cluster-sampling-in-field-biology-when-to-use-it-and-how-to-avoid-common-pitfalls)
- [Types Of Sampling In Biostatistics](/blog/guides/types-of-sampling-in-biostatistics)