Attributable Risk Formula: Definition and Worked Example
By Dr. Zubair Khalid, DVM, MS, PhD ·

The attributable risk formula is $AR = I_e - I_u$, the incidence in the exposed group minus the incidence in the unexposed group. It answers a practical question: how much extra disease risk comes with the exposure? This article defines the attributable risk formula, walks through a 2x2 table calculation step by step, and shows how to interpret the number.
Quick Answer
- Attributable risk (AR) is a risk difference: incidence in exposed minus incidence in unexposed.
- Formula: $AR = I_e - I_u$, where $I_e$ and $I_u$ are incidence proportions.
- In the worked example, $I_e = 0.2000$, $I_u = 0.0500$, so $AR = 0.1500$.
- AR is an absolute measure. It keeps the same units as risk (cases per person), unlike the risk ratio.
- Dividing AR by $I_e$ gives the attributable fraction, here $0.1500 / 0.2000 = 0.7500$.
The Formula
Attributable risk compares the risk of a health event in one group with the risk in a comparison group, and the difference between those two incidence proportions is the attributable risk [1].
$$AR = I_e - I_u$$
Each symbol means the following.
| Symbol | Meaning | Units |
|---|---|---|
| $I_e$ | Incidence proportion in the exposed group (cases ÷ exposed participants) | Cases per person |
| $I_u$ | Incidence proportion in the unexposed group (cases ÷ unexposed participants) | Cases per person |
| $AR$ | Attributable risk, also called the risk difference | Cases per person |
Two related quantities often appear alongside AR.
$$AF = \frac{AR}{I_e} = \frac{I_e - I_u}{I_e}$$
The attributable fraction in the exposed ($AF$) is the share of the exposed group's risk that is linked to the exposure. It is a proportion between 0 and 1 when the exposure raises risk.
$$RR = \frac{I_e}{I_u}$$
The relative risk ($RR$) is the ratio of the two incidences. A value of 1.0 means equal risk in both groups, above 1.0 means higher risk in the exposed group, and below 1.0 means the exposure may be protective [1]. AR and RR answer different questions, so report both when you can.
How to Calculate It Step by Step
- Build a 2x2 table with exposed and unexposed rows and case and non-case columns.
- Compute $I_e$: divide exposed cases by exposed participants.
- Compute $I_u$: divide unexposed cases by unexposed participants.
- Subtract: $AR = I_e - I_u$.
- Optionally divide AR by $I_e$ to get the attributable fraction.
- Optionally divide $I_e$ by $I_u$ to get the relative risk for context.
Worked Example
The dataset is a cohort of 400 participants, 200 exposed and 200 unexposed, with 40 and 10 cases respectively.
| Group | Participants | Cases | Incidence |
|---|---|---|---|
| Exposed | 200 | 40 | 0.2000 |
| Unexposed | 200 | 10 | 0.0500 |
Step 1. Incidence in the exposed group:
$$I_e = \frac{40}{200} = 0.2000$$
Step 2. Incidence in the unexposed group:
$$I_u = \frac{10}{200} = 0.0500$$
Step 3. Attributable risk:
$$AR = 0.2000 - 0.0500 = 0.1500$$
Step 4. Relative risk, for context:
$$RR = \frac{0.2000}{0.0500} = 4.0000$$
Step 5. Attributable fraction in the exposed:
$$AF = \frac{0.1500}{0.2000} = 0.7500$$
Step 6. Population attributable fraction, with an exposure prevalence of 0.5 (200 of 400). The overall incidence is $I_t = 50/400 = 0.1250$, so:
$$PAF = \frac{I_t - I_u}{I_t} = \frac{0.1250 - 0.0500}{0.1250} = 0.6000$$
So the exposed group carries 0.1500 more cases per person than the unexposed group, and 75% of the exposed group's risk is attributable to the exposure under these assumptions.
How to Interpret the Result
An AR of 0.1500 means 15 extra cases per 100 people in the exposed group compared with the unexposed group. If you prefer percentages, that is 15 percentage points of absolute risk. The relative risk of 4.0 tells a different story: exposed people are four times as likely to become cases. Both numbers are true and both are useful. The absolute difference shows how many cases you might prevent by removing the exposure, while the ratio shows how strongly the exposure is associated with the outcome [1].
The attributable fraction of 0.7500 says that three quarters of the cases among exposed people would not have occurred if the exposure had no effect. That phrasing depends on a causal assumption. An observed association between an exposure and a disease may reflect a causal relationship, but it may also come from sampling error or other biases [1]. Treat AF as a description of the data plus a causal claim you must defend separately.
The population attributable fraction extends the idea to a whole population. It is the proportion of disease in a population that is attributable to a given risk factor, and it estimates the reduction in disease you would expect if that risk factor were reduced or eliminated [2]. In the example, the exposure prevalence is 0.5, so the PAF is 0.6000, smaller than the AF of 0.7500. When exposure is rarer, the PAF shrinks even if the AF stays large.
Doing It in Software
Excel has no dedicated attributable risk function, so you build it from cell arithmetic. Put exposed cases in one cell and exposed participants in another, divide, and repeat for the unexposed group. Then subtract the two incidences. The same logic works in any spreadsheet.
Python is a good fit when you already have a row-level dataset. This snippet builds the 2x2 data, groups it, and computes AR and AF.
import pandas as pd
df = pd.DataFrame({'group': ['exposed']*200 + ['unexposed']*200,
'case': [1]*40 + [0]*160 + [1]*10 + [0]*190})
tab = df.groupby('group')['case'].agg(['sum','count'])
tab['incidence'] = tab['sum'] / tab['count']
Ie = tab.loc['exposed','incidence']
Iu = tab.loc['unexposed','incidence']
AR = Ie - Iu
AF = AR / Ie
print(AR) # 0.1500
print(AF) # 0.7500
Output:
0.15000000000000002
0.7500000000000001
In R, the same calculation is a few lines with table() and prop.table(). If you are working from a contingency table instead of row-level data, the arithmetic is identical, only the input shape changes. For a related absolute measure used in trials, see absolute risk reduction vs relative risk reduction, which uses the same subtraction logic on treatment and control groups.
Common Mistakes
- Confusing attributable risk with relative risk. AR is a difference, RR is a ratio. Reporting "risk went up 15%" when AR is 0.1500 mixes the two. Fix: state the absolute difference in cases per person and the ratio separately.
- Using counts instead of incidence proportions. Subtracting 40 from 10 gives 30, which is not a risk. Fix: divide each count by its group total first, then subtract.
- Mixing up the attributable fraction and the population attributable fraction. AF applies to the exposed group, PAF applies to the whole population and depends on exposure prevalence. Fix: label which one you computed and state the prevalence you assumed.
- Reading AR as a causal effect by default. AR describes an association in your data. Fix: treat causality as a separate argument supported by study design and prior evidence.
- Ignoring the direction of the difference. If $I_u$ exceeds $I_e$, AR is negative and the exposure looks protective. Fix: report the sign and interpret it instead of dropping it.
- Reporting AR without a confidence interval. A single point estimate hides sampling uncertainty. Fix: compute an interval for the risk difference, or at least report group sizes so readers can judge precision.
Limitations
Attributable risk is a difference of two proportions, so it inherits every weakness of those proportions. It cannot correct for confounding. If the exposed and unexposed groups differ in age, sex, or another risk factor, part of the 0.1500 difference may belong to those factors. Adjusted estimates require regression models, and applying simple formulas to adjusted relative risks can produce uninterpretable results [3]. AR also says nothing about whether the exposure actually caused the outcome.
The attributable fraction and population attributable fraction carry stronger assumptions than AR itself. They require that the exposure-outcome relationship is causal, that the exposure can be defined and measured consistently, and that removing the exposure is realistic [3]. Broad exposure definitions and unmeasured confounding can make these fractions misleading. Use them as planning estimates, not as precise counts of preventable cases.
Frequently Asked Questions
What is the attributable risk formula?
The attributable risk formula is $AR = I_e - I_u$, where $I_e$ is the incidence proportion in the exposed group and $I_u$ is the incidence proportion in the unexposed group. It gives the excess risk associated with the exposure in absolute terms. Some texts call it the risk difference.
How do you calculate attributable risk from a 2x2 table?
Divide the exposed cases by the exposed total to get $I_e$, and the unexposed cases by the unexposed total to get $I_u$. Subtract the second from the first. In the example, $0.2000 - 0.0500 = 0.1500$. The table only needs two rows and two outcome columns.
What is the difference between attributable risk and relative risk?
Attributable risk is a subtraction and keeps the units of risk, so it tells you how many extra cases the exposure produces. Relative risk is a division and is unitless, so it tells you how many times more likely the outcome is. A large relative risk can come with a tiny attributable risk when the baseline risk is low.
Can attributable risk be negative?
Yes. If the unexposed incidence is higher than the exposed incidence, the difference is negative. That pattern suggests the exposure is associated with lower risk, which may mean it is protective or that the groups differ in ways you have not accounted for. Report the negative value rather than taking an absolute value.
What is the attributable fraction in the exposed?
It is $AF = AR / I_e$, the proportion of risk in the exposed group that is linked to the exposure. In the worked example it is $0.1500 / 0.2000 = 0.7500$, or 75%. It is bounded at 1 when the exposure raises risk, and it assumes the association is causal.
References
- 12.5: Epidemiologic Measures - Medicine LibreTexts/12%3A_Epidemiology_for_Informing_Population__Community_Health_Decisions/12.05%3A_Epidemiologic_Measures)
- Freese KE, Bodnar LM, Brooks MM, McTIGUE K, Himes KP. (2020). Population-attributable fraction of risk factors for severe maternal morbidity. American journal of obstetrics & gynecology MFM
- Erren TC, Morfeld P. (2011). Attributing the burden of cancer at work: three areas of concern when examining the example of shift-work. Epidemiologic perspectives & innovations : EP+I
Further Reading
- NIST/SEMATECH e-Handbook of Statistical Methods
- Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician
- Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods
Related Articles
- Probability of A Given B: Conditional Probability Formula and Examples
- t Statistic Formula: Definition, Calculation and Examples
- Combinations Formula: Definition and Examples
- Expected Value: Definition, Formula and Examples
- Absolute Risk Reduction vs Relative Risk Reduction: Meaning, Formulas and Examples