# What Is a Population in Statistics? Definition and Examples

A **population definition** in statistics is simple: it is the complete set of units you want to draw conclusions about. A sample is a smaller subset of that set, measured because the full population is too large, too costly or impossible to reach. The gap between what the sample shows and what the population actually is drives most of applied statistics.

## Quick Answer

- A population is the entire group you are studying, not the group you happened to measure.
- It is defined by your research question, so the same data can belong to different populations under different questions.
- A sample is a subset of the population, and facts about a sample are not automatically facts about the population [1].
- Population quantities use Greek letters: mean $\mu$, standard deviation $\sigma$, size $N$.
- Sample quantities use Latin letters: mean $\bar{x}$, standard deviation $s$, size $n$.

## What Population Means

In everyday speech, "population" means the number of people living somewhere. In statistics it means something broader. A population is the full collection of individuals, objects, events or measurements that share the characteristic you are investigating.

The precise statistical definition is this: a population is the complete set of units about which you want to make inferences, together with the values of the variable you are measuring on those units. That second half matters. You are not just naming a group of mice, you are naming a group of mice and the weight of each one.

Two features follow from that definition.

First, a population is defined by the question, not by geography or biology. If you want to know the average weight of adult male mice in one breeding colony, that colony is your population. If you want to know the average weight of adult male mice in the species, the colony is a sample. The same 500 mice can be a population in one study and a sample in another.

Second, populations can be finite or effectively infinite. The 500 mice in a colony are finite and countable. All future output of a manufacturing process is conceptually infinite, because you cannot list units that have not been produced yet. In that case the population is defined by the process and the time window you care about.

A related term is the target population, which is the group your conclusions are meant to apply to. The sampled population is the group you could actually draw units from. When these two differ, your conclusions are limited to the sampled population no matter how large your sample is [1].

If you want the vocabulary for the numbers that describe a population, see [parameter definition in statistics](/blog/data-analysis/parameter-definition-statistics).

## How It Works

The mechanism is straightforward. You want a population quantity, you measure a sample, and you use the sample to estimate the population value.

The population mean is

$$\mu = \frac{1}{N}\sum_{i=1}^{N} x_i$$

where $N$ is the population size and $x_i$ is the value for unit $i$. Every unit is included, so there is no uncertainty about $\mu$ once you have the full list.

The population standard deviation is

$$\sigma = \sqrt{\frac{1}{N}\sum_{i=1}^{N}(x_i - \mu)^2}$$

The divisor is $N$, because you are describing the whole population, not estimating from part of it.

The sample mean is

$$\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i$$

where $n$ is the sample size. The sample standard deviation uses $n-1$ as the divisor, which corrects for the fact that $\bar{x}$ is itself estimated from the same data.

The difference between the two means is the sampling error:

$$\bar{x} - \mu$$

That quantity is unknown in real work, because $\mu$ is unknown. What you can compute is the standard error of the mean, which describes how much $\bar{x}$ would vary from sample to sample:

$$SE = \frac{\sigma}{\sqrt{n}}$$

When $\sigma$ is unknown, you substitute $s$. The standard error shrinks as $n$ grows, which is why larger samples give more precise estimates. The NIST handbook puts it plainly: the more observations you take from a population, the more your sample data resembles the population, and a sample is adequate when facts about it are reasonable approximations of facts about the population [1].

## Worked Example

The dataset is the weight in grams of 500 lab mice, with a sample formed by taking every 16th mouse, giving $n = 30$. The first 12 rows look like this.

| mouse_index | weight_g |
|---|---|
| 0 | 23.50 |
| 1 | 24.41 |
| 2 | 24.73 |
| 3 | 24.23 |
| 4 | 23.05 |
| 5 | 21.58 |
| 6 | 20.35 |
| 7 | 19.78 |
| 8 | 19.98 |
| 9 | 20.74 |
| 10 | 21.63 |
| 11 | 22.19 |

Here the 500 mice are treated as the population, so we can compute the true population values and compare them with what the sample would have told us.

**Step 1. Population size.** $N = 500$.

**Step 2. Population mean.** Summing all 500 weights gives 12746.73 g, so

$$\mu = \frac{12746.73}{500} = 25.4935 \text{ g}$$

**Step 3. Population standard deviation.** Using the $N$ divisor,

$$\sigma = \sqrt{\frac{\sum (x_i - \mu)^2}{500}} = 2.5030 \text{ g}$$

**Step 4. Sample size.** $n = 30$.

**Step 5. Sample mean.** The 30 sampled weights sum to 757.41 g, so

$$\bar{x} = \frac{757.41}{30} = 25.2470 \text{ g}$$

**Step 6. Sample standard deviation.** Using the $n-1$ divisor, which is what Excel's `STDEV.S` returns,

$$s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{29}} = 2.2352 \text{ g}$$

**Step 7. Sampling error.** $\bar{x} - \mu = 25.2470 - 25.4935 = -0.2465$ g. The sample understated the population mean by about a quarter of a gram.

**Step 8. Standard error of the mean.** $SE = \sigma / \sqrt{n} = 2.5030 / \sqrt{30} = 0.4570$ g. The observed sampling error of $-0.2465$ g is well inside one standard error, which is what you would expect from a sample of this size.

The code below reproduces every number.

```python
import statistics, math
pop_mean = statistics.fmean(weights)          # population mean
pop_sd   = statistics.pstdev(weights)         # population SD (N)
samp_mean = statistics.fmean(sample)          # sample mean
samp_sd   = statistics.stdev(sample)          # sample SD (n-1)
sampling_error = samp_mean - pop_mean
se = pop_sd / math.sqrt(len(sample))
```

Output:

```
pop_mean=25.4935, pop_sd=2.5030, samp_mean=25.2470, samp_sd=2.2352, sampling_error=-0.2465, se=0.4570
```

## How to Interpret It

Read the population values as facts and the sample values as estimates. In the example, 25.4935 g is the truth for those 500 mice. The value 25.2470 g is what you would have reported if you had only weighed 30 of them, and it is off by 0.2465 g.

The standard error tells you how much that kind of miss typically looks like. With $SE = 0.4570$ g, a sampling error of about half a gram in either direction is ordinary. An error of two grams would be unusual and would suggest either bad luck or a sampling method that is not representative.

This is the core logic of inference. You never observe the sampling error in practice, because you never observe $\mu$. You observe $\bar{x}$, you estimate the spread of $\bar{x}$ around $\mu$, and you report an interval. The interval is your honest statement about where the population value probably sits.

One more point about interpretation. The population standard deviation of 2.5030 g describes how spread out the 500 individual mice are. The standard error of 0.4570 g describes how spread out the sample mean is across repeated samples. These are different quantities with different meanings, and confusing them is one of the most common errors in reading statistical output.

## When to Use It (and when not to)

Use population language when you can state clearly who or what your conclusions apply to. That statement belongs in the methods section of any study, and it should name the units, the variable, the time frame and any restrictions.

Use it when you are deciding between a census and a sample. A census measures every unit in the population. It removes sampling error but is often impractical, and it can introduce its own problems. The [census definition article](/blog/data-analysis/census-definition-meaning-examples) covers when a full count is worth the cost.

Use it when comparing groups. Comparing reliability between two or more populations means asking whether the samples came from populations with the same reliability function, and the analysis is framed entirely in population terms [2].

Do not use population language when your sample was collected in a way that does not map onto any real group. A convenience sample of volunteers is not a random sample from a defined population, and calling the volunteers "the population" hides that problem instead of solving it.

Do not treat a population as fixed when it is not. A population of customers changes as people join and leave. A population of manufactured parts changes as the process drifts. State the time window and stick to it.

## Population vs Sample

| Feature | Population | Sample |
|---|---|---|
| What it is | Every unit of interest | A subset of those units |
| Size symbol | $N$ | $n$ |
| Mean symbol | $\mu$ | $\bar{x}$ |
| SD symbol | $\sigma$ | $s$ |
| SD divisor | $N$ | $n-1$ |
| Status | Usually unknown | Observed |
| Role | The target of inference | The evidence you have |

The practical difference is that the population is what you want to know and the sample is what you have. Every formula in inferential statistics is a bridge between the two columns.

For the standard deviation specifically, the divisor choice is not cosmetic. See [sample vs population standard deviation](/blog/data-analysis/sample-vs-population-standard-deviation) for when each applies.

## Common Mistakes

- **Calling your dataset "the population" when it is a sample.** If your 200 survey responses were drawn from a mailing list of 40,000, the population is the list, not the 200. Fix: write the population definition before you write the analysis.
- **Using $n$ instead of $n-1$ for a sample standard deviation.** This biases $s$ downward. Fix: use the $n-1$ divisor for sample data, and reserve the $N$ divisor for a complete population.
- **Confusing the standard deviation with the standard error.** $\sigma$ describes individual units, $SE$ describes the sample mean. Fix: label each number with the quantity it describes.
- **Assuming a large sample fixes a biased one.** Size does not repair a sampling frame that misses part of the population. Fix: check representativeness before you check sample size [1].
- **Leaving the population vague.** "All users" is not a population definition. Fix: specify units, variable, place and time.
- **Reporting a sample mean as if it were exact.** A sample mean is an estimate with a known spread. Fix: report the standard error or a confidence interval alongside it.

## Limitations

A population definition is a modeling choice, not a discovered fact. You decide where the boundary lies, and different reasonable analysts can draw it differently. That means two studies with identical data can report different population values simply because they defined the group differently. The math cannot settle that disagreement.

The second limitation is that the population is usually unobservable. You can compute $\mu$ and $\sigma$ only when you have measured every unit, which is rare. In practice you estimate them, and every estimate carries sampling error plus whatever bias your collection method introduced. A perfectly executed calculation on a badly defined population produces a confident answer to the wrong question.

## Frequently Asked Questions

### What is a population in statistics in simple terms?

It is the whole group you want to say something about. If you want to know the average height of every student in a school, the students in that school are the population. You rarely measure all of them, so you measure a sample and use it to estimate the population value.

### What is the difference between a population and a sample?

The population is the complete set of units of interest. The sample is the subset you actually measure. The population mean is written $\mu$ and the sample mean is written $\bar{x}$. Facts about a sample are not necessarily facts about the population, which is why inference is needed at all [1].

### Can a population be infinite?

Yes. All future units from a manufacturing process cannot be listed, so the population is treated as infinite or as a process defined over a time window. Statistical methods for infinite populations work with distributions and parameters instead of a countable list of units.

### How do I define the population for my study?

Name four things: the units, the variable, the place and the time period. "The weight in grams of adult male mice in colony A during March" is a usable definition. "Mice" is not. A precise definition tells your reader exactly which conclusions your results support.

### Why is the population standard deviation divided by N and the sample standard deviation by n-1?

The population formula divides by $N$ because you are describing every unit, so there is nothing to correct. The sample formula divides by $n-1$ because $\bar{x}$ is estimated from the same data, which makes the squared deviations slightly too small on average. Dividing by $n-1$ removes that bias.

## References

1. [3.1.3.4. Populations and Sampling](https://www.itl.nist.gov/div898/handbook/ppc/section1/ppc134.htm)
2. [8.4.4. How do you compare reliability between two or more populations?](https://www.itl.nist.gov/div898/handbook/apr/section4/apr44.htm)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)

## Related Articles

- [Census Definition: Meaning, Types and Examples](/blog/data-analysis/census-definition-meaning-examples)
- [What Is Descriptive Statistics? Definition and Examples](/blog/data-analysis/descriptive-statistics)
- [What Is Prevalence? Definition, Formula and Examples](/blog/data-analysis/prevalence-definition-formula-examples)
- [Parameter Definition in Statistics: Meaning and Examples](/blog/data-analysis/parameter-definition-statistics)
- [What Is Data? Definition, Meaning and Examples in Science](/blog/data-analysis/what-is-data-definition-meaning)
- [populations biology](/blog/careers/populations-biology)
- [What Does Statistical Mean? A Clear Explanation](/blog/guides/what-does-statistical-mean-a-clear-explanation)