What Is Prevalence? Definition, Formula and Examples
By Dr. Zubair Khalid, DVM, MS, PhD ·

Prevalence is the proportion of a population that has a condition at a specific point in time or during a defined period. It answers the question "how many people currently have this condition?" rather than "how many people are newly getting it?" This article gives you the prevalence definition, the formula, a worked example, and the key differences from incidence.
Quick Answer
- Prevalence is the number of existing cases divided by the total population at risk: $\text{prevalence} = \text{cases} / N$ [1].
- It is usually reported as a proportion or a percentage, such as 0.15 or 15% [1].
- Point prevalence measures cases at one specific moment. Period prevalence measures cases across a time window [1][2].
- Prevalence counts everyone who has the condition, including old and new cases [3][2].
- Incidence counts only new cases. Prevalence rises when incidence rises and falls when people are cured or die [1].
What Prevalence Means
In plain language, prevalence is how common a condition is in a group of people right now. If you say "15% of this survey group has the condition," you are stating a prevalence. Clinicians and public health researchers describe it as the percentage of the population at risk that has the disease [1].
The precise statistical definition is the proportion of the population with a condition at a specific point in time (point prevalence) or during a period of time (period prevalence) [1]. A proportion is a ratio where the numerator is a subset of the denominator, so prevalence always falls between 0 and 1 when written as a decimal.
Two flavors of prevalence matter in practice:
- Point prevalence focuses on a single moment. It gives a snapshot of the condition and its spread, which is useful early in an outbreak [2].
- Period prevalence uses a broader window, such as a month or a year. It includes old and new cases plus anyone who was cured or died during the period, so it often gives a more complete picture of a chronic disease [1][2].
Prevalence can differ between populations and across time. The same condition may show one prevalence in one group and a very different value in another [2].
How It Works
The formula is simple because prevalence is a proportion:
$$\text{Prevalence} = \frac{\text{Number of existing cases}}{\text{Total population at risk}}$$
Each symbol means the following:
- Number of existing cases: everyone who has the condition at the time or during the period you are measuring. This includes both old and new cases [3][2].
- Total population at risk: the full group you are studying, usually written as $N$. This is the denominator.
- Prevalence: the resulting proportion. Multiply by 100 to express it as a percentage.
The numerator is always a subset of the denominator, so the value cannot exceed 1 (or 100%). If you count 30 people with a condition out of 200 people surveyed, the prevalence is $30 / 200 = 0.15$, or 15%.
Prevalence moves with two forces. It increases when new cases are identified through incidence, and it decreases when a patient is cured or dies [1]. That is why a long-lasting condition can have high prevalence even if few new cases appear each year.
Worked Example
Suppose you survey 200 participants. At baseline, 30 already have the condition, and over the following year, 12 new cases develop.
The dataset looks like this:
| metric | count | denominator |
|---|---|---|
| existing_cases | 30 | 200 |
| new_cases | 12 | 170 |
Here are the steps with the computed values:
- Total participants (N): 200
- Participants with condition (cases): 30
- Prevalence formula: $\text{prevalence} = \text{cases} / N$
- Prevalence calculation: $30 / 200 = 0.1500$
- Prevalence as percent: 15.00%
- New cases during the 1-year period: 12
- Incidence formula (cumulative incidence): $\text{incidence} = \text{new\_cases} / (N - \text{cases})$, because only the 170 people without the condition at baseline are at risk of becoming new cases
- Incidence calculation: $12 / 170 = 0.0706$
- Incidence as percent: 7.06%
- Excel formula in cell B3:
=B2/B1gives 0.1500
You can reproduce this in Python:
N = 200
cases = 30
new_cases = 12
prevalence = cases / N
incidence = new_cases / (N - cases)
print(round(prevalence, 4)) # 0.15
print(round(incidence, 4)) # 0.0706
Output:
0.15
0.0706
So the prevalence is 15.00% and the one-year cumulative incidence is 7.06% in this 200-person survey. The bar chart below compares the two.
A bar chart comparing prevalence (30 existing cases, 15.0%) and incidence (12 new cases among 170 people at risk, 7.1%) in a 200-person survey.
How to Interpret It
Read prevalence as a snapshot of burden. A prevalence of 15% means that 15 out of every 100 people in the group currently have the condition. This is intuitive, which is why you hear statements like "currently, X% of Americans were overweight or obese" [1].
A few interpretation points:
- Higher prevalence means more people living with the condition. This can reflect a fast-spreading disease or a disease that lasts a long time.
- Prevalence is not a risk. It tells you how many people have the condition now, not how likely a healthy person is to develop it. Incidence answers that question [3].
- Compare like with like. Prevalence values are only comparable when the population, the case definition, and the time window match [2].
If you want to understand the denominator concept more deeply, see what a population in statistics means.
When to Use It (and when not to)
Use prevalence when your research question is about how common a condition is or how much of a population is affected. It is the right measure for planning health services, estimating the burden of chronic disease, and shaping public policy. Data on disease prevalence can guide legislation, screening programs, and educational efforts aimed at managing conditions in a population [2].
Prevalence studies are also central during outbreaks of a new disease. The COVID-19 pandemic highlighted how important prevalence studies are and how much their reporting and methodology had been neglected [4].
Do not use prevalence when you need to measure the speed at which new cases appear. If your question is "how fast is this spreading?" or "what is the risk of getting this?" you need incidence, which can be reported as a risk or as an incidence rate [3]. Prevalence and incidence answer different research questions, so choose the one that matches your question [3].
Prevalence vs Incidence
The core difference is timing. Incidence focuses on new cases, while prevalence deals with total cases, including those that are new [2].
| Feature | Prevalence | Incidence |
|---|---|---|
| What it counts | Existing cases (old and new) | New cases only |
| Question answered | How many have it now? | How many are getting it? |
| Typical formula | cases / N | new cases / population at risk |
| Reported as | Proportion or percentage | Risk or rate |
| Changes when | Cases are added, cured, or die | New cases appear |
| Best for | Burden of disease, planning | Risk, speed of spread |
Prevalence increases when new disease cases are identified through incidence, and it decreases when a patient is cured or dies [1]. A condition with low incidence but no cure can still have high prevalence because cases accumulate over time.
Common Mistakes
- Using prevalence as a risk measure. Prevalence tells you who has the condition now, not the chance a healthy person will develop it. Use incidence for risk [3].
- Mixing point and period prevalence. A point prevalence and a period prevalence over the same condition will not match. State which one you used and keep the time window consistent [1][2].
- Forgetting that prevalence includes old cases. If you count only new cases in the numerator, you have computed incidence, not prevalence [3][2].
- Ignoring the denominator. Prevalence is a proportion, so the population at risk must be clearly defined. A vague denominator makes the number meaningless.
- Comparing prevalence across mismatched populations. Prevalence differs between populations and over time, so comparisons need matching case definitions and time frames [2].
- Assuming prevalence only goes up. Prevalence falls when people are cured or die, so a dropping prevalence can reflect recovery or mortality, not just fewer cases [1].
Limitations
Prevalence cannot tell you how fast a disease spreads or how likely someone is to get it. It is a snapshot, and a snapshot of a long-lasting disease can look large even when new cases are rare. It also cannot separate the reasons behind a change. A rising prevalence could come from more new cases, better survival, or improved detection.
Prevalence studies carry real risks of bias. The two main domains are selection bias, related to the study population, and information bias, related to how the condition or risk factor is measured [4]. Selection bias can enter when people are invited to take part and again when you assess who actually responds and provides valid data. Information bias appears when systematic errors affect the accuracy and reproducibility of the measurement, including misclassification [4]. Cross-sectional studies, which are the usual design for measuring prevalence, capture exposure and outcome at the same time, so they describe association rather than prove cause [5].
Frequently Asked Questions
What is the difference between prevalence and incidence?
Prevalence counts all existing cases of a condition, while incidence counts only new cases [3]. Prevalence answers "how many people have it now?" and incidence answers "how many people are newly getting it?" They serve different purposes and answer different research questions [3].
What is the formula for prevalence?
Prevalence equals the number of existing cases divided by the total population at risk: $\text{prevalence} = \text{cases} / N$. Multiply by 100 to report it as a percentage. In the worked example, 30 cases out of 200 people gives a prevalence of 0.15, or 15%.
What is the difference between point prevalence and period prevalence?
Point prevalence measures the proportion with a condition at one specific moment, giving a snapshot of the disease and its spread [2]. Period prevalence measures it across a window of time, such as a month or a year, and includes old and new cases plus anyone cured or died during that period [1][2].
Can prevalence be greater than 100%?
No. Prevalence is a proportion, so the numerator is a subset of the denominator and the value stays between 0 and 1, or 0% and 100%. If your calculation exceeds 100%, you have likely counted cases that are not in the denominator or double-counted people.
Why does prevalence change over time?
Prevalence increases when new cases are identified through incidence and decreases when patients are cured or die [1]. It can also shift when the population changes or when detection improves. This is why prevalence can differ between populations and at various points in time [2].
If you want to build up the underlying proportion skills, see what a proportion is, how to find the sample mean, and the formula for range.
References
- Tenny S, Hoffman MR. (2023). Prevalence
- Epidemiology Incidence vs. Prevalence: Exploring Two of the Most Impactful Concepts in Public Health | School of Public Health
- Noordzij M, Dekker FW, Zoccali C, Jager KJ. (2010). Measures of disease frequency: prevalence and incidence. Nephron. Clinical practice
- Buitrago-Garcia D, Salanti G, Low N. (2022). Studies of prevalence: how a basic epidemiology concept has gained recognition in the COVID-19 pandemic. BMJ open
- Capili B. (2021). Cross-Sectional Studies. The American journal of nursing