Intention-to-Treat vs Per-Protocol Analysis: Definitions, Examples and When to Use Each

By Dr. Zubair Khalid, DVM, MS, PhD ·

Intention-to-Treat vs Per-Protocol Analysis: Definitions, Examples and When to Use Each

Intention-to-treat (ITT) analysis compares participants in the groups to which they were randomized, regardless of what treatment they actually received. Per-protocol analysis restricts the comparison to participants who followed the protocol as written. The two approaches answer different questions, and the gap between their results is often the most informative number in a trial report.

You will meet this distinction in journal clubs, in regulatory guidance, in peer review, and in your own analysis planning. Reviewers ask which population was analyzed because the answer changes what the trial can claim. A drug that looks impressive in a per-protocol analysis may show a smaller but more trustworthy effect under ITT, and in non-inferiority trials the direction of that gap can flip the conclusion entirely.

Quick Answer

  • ITT principle (ICH E9): the effect of a treatment policy is best assessed on the basis of the intention to treat a subject, not the treatment actually given [1]. Subjects are followed, assessed and analyzed as members of their randomized group [1].
  • The rule: once randomized, always analyzed. Every randomized participant stays in the group assigned at randomization, whatever happened afterward [5][8].
  • Per-protocol set: the subset of the full analysis set who complied with the protocol, for example completed a minimal exposure, had primary outcome measurements, and had no major protocol violations [1].
  • Why ITT is the default for superiority trials: non-compliers dilute the estimated effect, which avoids over-optimistic estimates [1]. The ITT estimate describes the effect of being assigned the treatment in a population where some patients stop.
  • Why non-inferiority is different: ITT is generally not conservative in equivalence or non-inferiority trials [1]. Non-adherence, endpoint misclassification and attrition can make arms look similar and create apparent non-inferiority when the test drug is in fact inferior [4].
  • The modern framing: ICH E9(R1) asks you to define an estimand, a precise description of the treatment effect that matches the clinical question, and then choose the analysis that estimates it [2].

Definitions and Intuition

Randomization creates groups that are comparable on measured and unmeasured factors. That comparability is the entire basis for causal inference in a trial. Anything that removes participants from their assigned group after randomization can break it.

The ITT principle protects the randomization. ICH E9 states that preservation of the initial randomisation in analysis is important in preventing bias and provides a secure foundation for statistical tests [1]. The full analysis set is the operational version of this ideal: as complete as possible and as close as possible to including all randomized subjects, with only minimal and justified exclusions [1].

ICH E9 permits a limited number of justified exclusions from the full analysis set: failure to satisfy major entry criteria, failure to take at least one dose of trial medication, and lack of any data after randomization [1]. Each has conditions. Excluding subjects who failed an entry criterion is unbiased only if the criterion was measured before randomization, violations are detected objectively, all subjects get equal scrutiny, and all detected violations are excluded [1]. Excluding subjects who took no trial medication preserves the ITT principle only if the decision to start treatment could not be influenced by knowledge of the assigned treatment [1].

The per-protocol set goes further. ICH E9 defines it as the subset of the full analysis set who complied with the protocol, such as completing a pre-specified minimal exposure, having primary outcome measurements, and having no major protocol violations [1]. It is also called the valid cases, efficacy sample or evaluable subjects set [1]. The reasons for excluding subjects from this set should be fully defined and documented before breaking the blind [1].

The intuition is simple. ITT asks: what happens when you offer this treatment to a population? Per-protocol asks: what happens when people take it as prescribed? Both are legitimate questions. They are not interchangeable.

The Estimand Framework in One Paragraph

ICH E9(R1), adopted in November 2019, reframes the debate [2]. An estimand is a precise description of the treatment effect that reflects the clinical question posed by the trial objective [2]. It has five attributes: the treatment, the population, the variable (endpoint), the handling of intercurrent events, and a population-level summary [2].

Intercurrent events are events after treatment initiation, such as discontinuation of assigned treatment or use of an additional or alternative therapy, that affect either the interpretation or the existence of the measurements associated with the clinical question [2]. Five strategies handle them: treatment policy, hypothetical, composite variable, while on treatment, and principal stratum [2].

Under the treatment policy strategy, the outcome is used regardless of whether the intercurrent event occurred [2]. This corresponds to the effect of a treatment policy in the E9 ITT definition [2]. The FDA issued the E9(R1) addendum as final guidance in May 2021 [3]. The practical consequence: instead of arguing ITT versus per-protocol, you state your estimand and then justify the analysis that estimates it. E9(R1) also notes that it may not be possible to construct a relevant estimand to which a per-protocol set analysis is aligned, because the per-protocol set may not compare similar subjects on different treatments [2].

Worked Example

A hypothetical placebo-controlled trial randomizes 400 patients 1:1. The outcome is a bad clinical event, so lower is better. The numbers below are constructed to show the direction of bias.

Drug arm (n = 200): 160 took the drug as planned (24 events, 15%). Forty stopped the drug early, mostly because they were sicker (16 events, 40%). Ten of those 40 never took a single dose (5 events).

Placebo arm (n = 200): 180 took placebo as assigned (54 events, 30%), of whom 4 never started placebo (1 event). Twenty obtained the active drug outside the protocol as crossovers (4 events).

Four analyses follow. Risk difference (RD) is reported in percentage points with a Wald 95% confidence interval, alongside the risk ratio (RR).

AnalysisDrug armPlacebo armRD (95% CI)RR
ITT, everyone as randomized40/200 = 20.0%58/200 = 29.0%-9.0 (-17.4 to -0.6)0.69
mITT, excluding those who never took a dose35/190 = 18.4%57/196 = 29.1%-10.7 (-19.1 to -2.2)0.63
Per-protocol, adherers only24/160 = 15.0%54/180 = 30.0%-15.0 (-23.7 to -6.3)0.50
As-treated, by treatment actually received28/180 = 15.6%70/220 = 31.8%-16.3 (-24.4 to -8.1)0.49

The risk difference is computed as:

$$\text{RD} = p_{\text{drug}} - p_{\text{placebo}}$$

where $p_{\text{drug}}$ and $p_{\text{placebo}}$ are the event proportions in each group. A negative RD means fewer events in the drug arm.

The risk ratio is:

$$\text{RR} = \frac{p_{\text{drug}}}{p_{\text{placebo}}}$$

An RR below 1 means the drug arm had fewer events.

Read the table carefully. The drug-arm non-adherers had a 40% event rate versus 15% in adherers. Dropping them removes high-risk patients from one arm only and breaks the balance created by randomization. The per-protocol and as-treated estimates (RR about 0.5) look roughly twice as impressive as ITT (RR 0.69), but only ITT compares groups that are comparable by design. The ITT estimate describes the effect of being assigned the drug in a population where some patients stop taking it.

In a non-inferiority trial, the same dilution would push the arms together and make an inferior drug look non-inferior. That is why ICH E9 and the FDA ask for both analyses in that setting [1][4].

How to Read and Interpret the Results

Start with the population definition, not the p-value. CONSORT 2010 dropped the specific request for the label "intention-to-treat" in favor of a clear description of exactly who was included in each analysis, because labels do not reliably say which patients were included [5]. CONSORT 2025 continues this with item 21b, a definition of who is included in each analysis and in which group, and item 21c, how missing data were handled [9].

When both a full analysis set and a per-protocol analysis are planned, as ICH E9 recommends for confirmatory trials, agreement between them increases confidence in the result [1]. Disagreement is a signal to investigate why. Look for the mechanism: did non-adherers have worse outcomes? Did crossovers move in one direction? Was the per-protocol exclusion rule defined before unblinding?

The direction of conservatism depends on the trial design. In superiority trials, the full analysis set is used for the primary analysis because non-compliers dilute the estimated effect, avoiding over-optimistic estimates [1]. In equivalence or non-inferiority trials, the full analysis set is generally not conservative [1]. The FDA states this plainly: ITT approaches, although conservative in superiority trials, are not necessarily conservative in a non-inferiority study and can lead to an incorrect finding of non-inferiority [4]. The CONSORT extension for non-inferiority and equivalence trials was updated by Piaggio et al. in JAMA 2012 [10].

A number from an ITT analysis does not mean the treatment works when taken. It means the treatment policy produced that result in a population where adherence was imperfect. A number from a per-protocol analysis does not mean the treatment works in compliant patients in the real world, because compliance itself can be a marker for prognosis. ICH E9 is direct about this: per-protocol analysis may maximise the chance to show efficacy, but the bias, which may be severe, arises because adherence to the protocol may be related to both treatment and outcome [1].

ITT and Its Relatives: mITT, As-Treated, Per-Protocol

These terms get used loosely. The distinctions matter.

Modified intention-to-treat (mITT) has no single definition. Cochrane describes mITT as adhering to ITT principles except that participants with missing outcome data are excluded [6]. CONSORT 2010 notes the term is widely used for analyses that exclude participants who did not adequately adhere, especially those who did not receive a minimum amount of the intervention [5]. Gupta describes mITT as a subset of the ITT population that excludes some randomized subjects in a justified way, and notes that mITT definitions are often irregular and arbitrary [8]. When you see mITT in a paper, find the exact exclusion rule before you interpret anything.

As-treated analysis assigns participants to the group corresponding to the intervention they actually received, even if they were randomized to another group [6]. Cochrane calls as-treated analyses potentially seriously biased [6]. In the worked example, the as-treated comparison produced the most extreme estimate (RR 0.49) because the 40 high-risk drug-arm non-adherers (40% event rate) were moved into the no-drug group, raising its event rate to 31.8%.

Per-protocol analysis restricts to adherers. Cochrane calls naive per-protocol analyses, meaning those restricted to adherers without methods to correct for the selection, potentially seriously biased [6]. CONSORT 2010 goes further and says a per-protocol analysis should be labelled a non-randomised, observational comparison [5].

AnalysisWho is includedWhat it estimatesMain risk
ITT / full analysis setAll randomized, as randomizedEffect of treatment policyDilution; needs complete follow-up or imputation
mITTRandomized minus a defined subsetDepends entirely on the exclusion ruleInconsistent definitions; can drift toward per-protocol
Per-protocolAdherers with no major violationsEffect among those who compliedSelection bias from conditioning on adherence
As-treatedGrouped by treatment receivedEffect of received treatmentBreaks randomization; potentially seriously biased [6]

Common Mistakes

  • Calling any analysis with a small exclusion "ITT." CONSORT 2010 cites a review of 403 RCTs from 2002 in which 62% reported ITT as the primary analysis, but only 39% of those actually analysed all participants as randomised [5]. Check the numbers, not the label.
  • Treating mITT as a synonym for ITT. The three major sources define mITT differently [5][6][8]. State the exclusion rule explicitly.
  • Using per-protocol as the primary analysis in a superiority trial. ICH E9 reserves the full analysis set for the primary analysis in superiority trials precisely because per-protocol estimates can be over-optimistic [1].
  • Assuming ITT is always conservative. In non-inferiority trials, ITT can make an inferior drug look non-inferior [1][4]. Plan both analyses and interpret them together.
  • Ignoring missing outcome data. A strict ITT analysis needs outcomes for every randomized participant [1][6]. When outcomes are missing, ITT relies on imputation assumptions. CONSORT 2010 notes that last observation carried forward is simple but may introduce bias and makes no allowance for imputation uncertainty [5].
  • Defining per-protocol exclusions after unblinding. ICH E9 requires the reasons for exclusion to be fully defined and documented before breaking the blind [1].
  • Reporting only one analysis when the two disagree. ICH E9 recommends planning both so differences can be discussed [1]. A silent gap between them is a red flag.

Limitations

The ITT label is weaker than it sounds in many published trials. Hollis and Campbell surveyed RCTs published in 1997 in BMJ, Lancet, JAMA and NEJM: 119 (48%) mentioned ITT, 12 of those excluded patients who did not start treatment, and 89 (75%) had missing primary outcome data [7]. The pattern has not disappeared.

Terminology remains unsettled. The definition of mITT varies across Cochrane, CONSORT and review articles [5][6][8]. A reader who trusts the label without reading the methods section will sometimes be misled.

Under E9(R1), the debate is reframed around estimands. ITT corresponds most closely to the treatment policy strategy, but E9(R1) does not say ITT is always the right question [2]. Some clinical questions are better served by a hypothetical strategy, a while-on-treatment strategy, or a principal stratum approach [2].

Methods that estimate a valid per-protocol effect, such as instrumental variable or inverse probability weighting approaches, exist but are outside the scope of this article. They address the selection problem that naive per-protocol analysis ignores, and they require assumptions that are not always defensible.

The worked example is hypothetical and constructed to show the direction of bias. Real differences between ITT and per-protocol can go either way. When adherence is unrelated to prognosis, the two analyses may agree closely. When adherence tracks with prognosis, the gap can be large, and the direction depends on which patients stop and why.

Frequently Asked Questions

What is intention to treat analysis in plain terms?

It means every participant is analyzed in the group they were randomized to, whether or not they took the treatment, switched treatments, or stopped early [1][5]. The comparison preserves the balance that randomization created. The estimate describes the effect of the treatment policy, not the effect of the drug in people who take it perfectly.

What is the difference between intention to treat vs per protocol?

ITT includes all randomized participants in their assigned groups. Per-protocol restricts to participants who complied with the protocol, such as completing a minimal exposure and having no major violations [1]. ITT answers what happens when a treatment is offered. Per-protocol answers what happens among people who follow the plan, but the selection of those people can bias the result [1].

What is modified intention to treat and why is it controversial?

mITT applies ITT principles with some exclusions, but the definition varies. Cochrane frames it as excluding participants with missing outcome data [6]. CONSORT 2010 frames it as excluding those who did not receive a minimum amount of the intervention [5]. Gupta notes the definitions are often irregular and arbitrary [8]. The controversy is that mITT can quietly become a per-protocol analysis with a friendlier name.

What is as-treated analysis and how does it differ?

As-treated analysis groups participants by the intervention they actually received, even if they were randomized to another group [6]. Cochrane calls these analyses potentially seriously biased [6]. In the worked example, moving the high-risk drug-arm non-adherers into the no-drug group produced the most extreme estimate of the four analyses.

Why does non-inferiority trial ITT behave differently?

In a superiority trial, non-adherence dilutes the effect and makes the drug look less effective, which is conservative [1]. In a non-inferiority trial, the same dilution pushes the arms together and can make an inferior drug look non-inferior [1][4]. The FDA states that ITT approaches are not necessarily conservative in a non-inferiority study and can lead to an incorrect finding of non-inferiority [4]. Both analyses should be planned and reported [1].

References

  1. ICH E9: Statistical Principles for Clinical Trials (Step 4, 1998)
  2. ICH E9(R1): Addendum on Estimands and Sensitivity Analysis (2019)
  3. FDA: E9(R1) Addendum: Estimands and Sensitivity Analysis in Clinical Trials
  4. FDA: Non-Inferiority Clinical Trials to Establish Effectiveness (2016)
  5. Moher D et al. CONSORT 2010 Explanation and Elaboration. BMJ 2010;340:c869
  6. Cochrane Handbook, Chapter 8: Assessing risk of bias in a randomized trial
  7. Hollis S, Campbell F. What is meant by intention to treat analysis? BMJ 1999;319:670-674
  8. Gupta SK. Intention-to-treat concept: A review. Perspect Clin Res 2011;2:109-112
  9. Hopewell S et al. CONSORT 2025 statement. BMJ 2025;389:e081123
  10. Piaggio G et al. Reporting of noninferiority and equivalence randomized trials (CONSORT extension). JAMA 2012;308:2594

Related Articles