How to Assess and Control Confounding in Observational Life Science Studies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Confounding arises from a third variable associated with both exposure and outcome, distorting the observed relationship; identification requires causal reasoning, often visualized with Directed Acyclic Graphs (DAGs) built from subject-matter knowledge, not statistical analysis alone.
- To control confounding, researchers must identify variables meeting three criteria: association with exposure, independent risk for outcome, and not being on the causal pathway (mediator); DAGs are crucial for distinguishing confounders from mediators and colliders.
- Common control methods include restriction (limiting population to a confounder level), matching (balancing exposed/unexposed groups on confounders), stratification (analyzing within confounder strata), and multivariable adjustment (regression models); propensity scores summarize confounders into a single balancing score.
- Unmeasured confounding remains a critical limitation; sensitivity analyses are essential to quantify the potential impact of unmeasured variables on study conclusions, as no statistical adjustment can fully correct for unmeasured factors.
- Instrumental variable analysis offers a potential solution for unmeasured confounding if a valid instrument (a variable associated with exposure, not confounders, and affecting outcome only through exposure) can be identified and justified.
- Transparent reporting, including the DAG, adjustment set, and limitations like potential residual or unmeasured confounding, is paramount for the credibility and reproducibility of observational life science studies.
Quick Answer
- Confounding occurs when a third variable distorts the observed relationship between an exposure and outcome, and researchers must identify these variables before analysis using causal reasoning.
- Build a directed acyclic graph (DAG) from subject-matter knowledge, then apply stratification or multivariable adjustment to estimate the exposure effect within confounder strata.
- No statistical method can fully correct for unmeasured confounders, so sensitivity analysis and transparent reporting of assumptions are essential for credible conclusions.
Understanding Confounding in Observational Life Science Studies
Observational studies form the backbone of much biological and medical research because they allow investigators to examine relationships in real-world populations without the ethical and practical constraints of randomized experiments. In these studies, researchers observe exposures and outcomes as they naturally occur, instead of assigning treatments or interventions. This design choice introduces a fundamental challenge: the groups being compared may differ in ways beyond the exposure of interest, and these differences can distort the observed relationship.
Confounding is the distortion of an exposure-outcome association by a third variable that is associated with both the exposure and the outcome. This third variable, called a confounder, creates a spurious association or masks a true one. For example, in a study examining whether coffee consumption is associated with heart disease, age could confound the relationship because older individuals may drink more coffee and also have higher heart disease risk. Without accounting for age, the study might incorrectly attribute heart disease risk to coffee consumption.
The problem of confounding is not a statistical nuisance that can be fixed with more data or more sophisticated software. It is a causal inference problem that requires careful thinking about the underlying biological and social mechanisms that connect variables. The validity of an observational study depends on how well the investigators can identify, measure, and adjust for confounders, and on how honestly they report the limitations of their approach.
Why Confounding Matters in Life Science Research
Life science research spans a wide range of study types, from laboratory experiments using cell lines to large epidemiological cohort studies of human populations. Confounding is most problematic in observational designs where the investigator does not control the assignment of exposure. These designs include cross-sectional surveys, case-control studies, and prospective cohort studies.
In molecular biology, confounding can arise when comparing gene expression between groups that differ in age, sex, or tissue collection time. In ecology, confounding can arise when comparing species diversity between habitats that differ in temperature or soil composition. In clinical research, confounding is a central concern when comparing treatment outcomes between patients who differ in disease severity, comorbidities, or socioeconomic status.
The consequences of unaddressed confounding are serious. A study may report a statistically significant association that is entirely due to a confounder, leading to false conclusions and wasted resources. Alternatively, a true association may be obscured by a confounder, causing investigators to miss an important biological relationship. In either case, the scientific literature becomes less reliable, and subsequent research may build on flawed foundations.
The Counterfactual Framework
To understand confounding precisely, researchers use the counterfactual framework, also known as the potential outcomes framework. This framework asks what would happen to a specific individual if they were exposed, compared to what would happen to the same individual if they were not exposed. The causal effect of exposure for that individual is the difference between these two potential outcomes.
In reality, each individual experiences only one of these potential outcomes. The observed outcome for an individual who was exposed is the exposed outcome, and the unobserved outcome is the counterfactual. The fundamental problem of causal inference is that the counterfactual outcome is never observed.
In a randomized experiment, the random assignment of exposure ensures that, on average, the exposed and unexposed groups are comparable with respect to all other factors. This comparability means that the observed outcome in the unexposed group can serve as a valid estimate of the counterfactual outcome for the exposed group. In an observational study, there is no such guarantee. The exposed and unexposed groups may differ systematically in ways that affect the outcome, and these differences create confounding.
The goal of confounding control is to approximate the conditions of a randomized experiment by making the exposed and unexposed groups comparable with respect to all confounders. This comparability can be achieved through design strategies, such as restriction or matching, or through analysis strategies, such as stratification or multivariable adjustment.
Core Principles of Confounding Control
Confounding control requires a systematic approach that begins with causal thinking and ends with transparent reporting. The process involves several distinct steps, each of which requires careful judgment and subject matter expertise.
The Three Criteria for a Confounder
A variable must satisfy three criteria to be considered a confounder in a given study. First, the variable must be associated with the exposure in the study population. Second, the variable must be an independent risk factor for the outcome, meaning it affects the outcome even in the absence of the exposure. Third, the variable must not be on the causal pathway between the exposure and the outcome.
The third criterion is important because variables on the causal pathway, called mediators, should not be adjusted for in the same way as confounders. Adjusting for a mediator can introduce bias by blocking the indirect effect of the exposure on the outcome. For example, if a study examines whether a high-fat diet causes heart disease, blood cholesterol may be a mediator because the diet affects cholesterol, which in turn affects heart disease risk. Adjusting for cholesterol would remove part of the dietary effect and underestimate the total effect of the diet.
Directed Acyclic Graphs
A directed acyclic graph (DAG) is a visual representation of the causal relationships among variables in a study. DAGs use arrows to indicate the direction of causal influence, and the graph must be acyclic, meaning there are no feedback loops. DAGs are constructed from subject matter knowledge, not from statistical analysis of the data.
The DAG serves several purposes in confounding control. It forces investigators to make their assumptions explicit, which allows others to evaluate the validity of those assumptions. It helps identify the minimal set of variables that need to be adjusted for to remove confounding. It also helps identify variables that should not be adjusted for, such as mediators and colliders.
A collider is a variable that is caused by two or more other variables. Adjusting for a collider can introduce bias, known as collider bias, even when the collider is not a confounder. For example, if a study examines the relationship between two genetic variants and a disease, and the disease is a collider because it is caused by both variants, then selecting the study population based on disease status can create a spurious association between the variants.
The Role of Subject Matter Knowledge
The identification of confounders cannot be done by statistical methods alone. Statistical tests can indicate whether a variable is associated with the exposure and the outcome, but they cannot determine whether the variable is a confounder, a mediator, or a collider. This determination requires subject matter knowledge about the biological, behavioral, or social mechanisms that generate the data.
Subject matter knowledge is used to construct the DAG, which then guides the statistical analysis. The DAG should be constructed before the data are analyzed, and it should be based on the existing literature and the investigators understanding of the underlying processes. The DAG should be updated as new knowledge becomes available, but it should not be modified based on the observed data in a way that capitalizes on chance.
At a Glance: Confounding Control Methods
The following table summarizes the main methods for controlling confounding in observational studies, their key features, and the situations in which they are most useful.
| Method | Key Feature | Best Used When | Limitation |
|---|---|---|---|
| Restriction | Limits the study population to a single level of the confounder | The confounder has a small number of categories and the restricted population is still large enough | Reduces generalizability and may leave residual confounding within the restricted category |
| Matching | Selects exposed and unexposed groups with the same confounder distribution | The confounder is strong and the study is small | Requires careful analysis that accounts for the matching design |
| Stratification | Divides the study population into confounder strata and pools the stratum-specific estimates | The confounder has a small number of categories and the sample size is large enough | Becomes impractical with many confounders or continuous confounders |
| Multivariable adjustment | Uses a regression model to estimate the exposure effect while holding confounders constant | The confounder is continuous or there are many confounders | Requires correct model specification and can be sensitive to outliers |
| Propensity score | Uses a model to estimate the probability of exposure and then adjusts or matches on this score | The study has many confounders and a large sample size | Requires the propensity model to be correctly specified and the overlap assumption to hold |
| Instrumental variable | Uses a variable that affects the exposure but not the outcome except through the exposure | There is a valid instrument and the study is large | Valid instruments are rare and the assumptions are difficult to verify |
Building a Directed Acyclic Graph
The DAG is the foundation of confounding control in observational studies. A well-constructed DAG provides a clear and explicit representation of the causal assumptions that underlie the analysis, and it guides the selection of variables for adjustment.
Identifying the Exposure and Outcome
The first step in building a DAG is to define the exposure and the outcome. The exposure is the variable whose effect is being studied, and the outcome is the variable that is being affected. These definitions should be specific and measurable. For example, in a study of the effect of a high-fat diet on body weight, the exposure is the dietary fat intake and the outcome is body weight change.
The exposure and outcome should be defined in a way that is consistent with the research question. If the research question is about the effect of a specific dietary component, the exposure should be that component, not a broader dietary pattern. If the research question is about the effect of a treatment on disease progression, the outcome should be a measure of disease progression, not a surrogate marker.
Adding the Confounders
The next step is to identify the variables that are associated with both the exposure and the outcome and that are not on the causal pathway. These variables are the confounders, and they should be added to the DAG with arrows pointing from the confounder to the exposure and from the confounder to the outcome.
The identification of confounders requires subject matter knowledge. The investigator should consider all variables that are known or suspected to affect the exposure and the outcome. These variables may include demographic factors, such as age and sex, behavioral factors, such as smoking and physical activity, and clinical factors, such as disease severity and comorbidities.
Identifying Mediators and Colliders
The DAG should also include mediators and colliders, even though these variables should not be adjusted for in the same way as confounders. Mediators are variables on the causal pathway between the exposure and the outcome. They should be included in the DAG to make the causal assumptions explicit, but they should not be adjusted for in the analysis.
Colliders are variables that are caused by two or more other variables. They should be included in the DAG because adjusting for them can introduce bias. The DAG helps identify colliders and prevents the investigator from inadvertently adjusting for them.
Evaluating the DAG
Once the DAG is constructed, it should be evaluated for completeness and consistency. The DAG should include all variables that are known or suspected to be relevant to the causal relationship. It should be consistent with the existing literature and with the investigators understanding of the mechanisms.
The DAG should also be evaluated for the presence of backdoor paths. A backdoor path is a path between the exposure and the outcome that does not follow the direction of the arrows and that is not blocked by a collider. Backdoor paths represent confounding, and they must be blocked by adjusting for one or more variables on the path.
The DAG can be used to identify the minimal sufficient adjustment set, which is the smallest set of variables that, when adjusted for, blocks all backdoor paths. This set can be identified using the rules of DAG analysis, which are based on the concept of d-separation.
Practical Workflow for Confounding Control
The following workflow provides a step-by-step approach to controlling confounding in an observational study. This workflow is designed to be practical and to be applied to a wide range of research settings.
Step 1: Define the Research Question
The first step is to define the research question in a way that is specific and testable. The research question should specify the exposure, the outcome, and the population of interest. For example, a research question might be, "What is the effect of a high-fat diet on body weight in adults with prediabetes?"
The research question should be defined before any data are collected or analyzed. This definition helps ensure that the study is focused and that the analysis is guided by the research question instead of by the data.
Step 2: Construct the DAG
The second step is to construct the DAG based on subject matter knowledge. The DAG should include the exposure, the outcome, and all variables that are known or suspected to be relevant to the causal relationship. The DAG should be constructed before the data are examined, and it should be based on the literature and the investigators understanding of the mechanisms.
The DAG should be reviewed by other investigators with subject matter expertise. This review helps ensure that the DAG is complete and that the causal assumptions are reasonable.
Step 3: Identify the Adjustment Set
The third step is to identify the minimal set of variables that need to be adjusted for to block all backdoor paths. This set can be identified using the rules of DAG analysis. The adjustment set should include all confounders that are not on the causal pathway and should not include mediators or colliders.
The adjustment set should be specified before the data are analyzed. This specification helps prevent the analysis from being influenced by the data and helps ensure that the analysis is reproducible.
Step 4: Collect and Prepare the Data
The fourth step is to collect the data and prepare it for analysis. The data should include the exposure, the outcome, and all variables in the adjustment set. The data should be checked for completeness, and missing data should be handled using appropriate methods.
The data should be prepared in a way that is consistent with the analysis plan. For example, continuous variables may need to be transformed, and categorical variables may need to be coded.
Step 5: Perform the Analysis
The fifth step is to perform the analysis using the adjustment set identified in Step 3. The analysis should be performed using appropriate statistical methods, such as stratification or multivariable regression. The analysis should be performed in a way that is consistent with the DAG and the adjustment set.
The analysis should be performed using a statistical software package that is appropriate for the data and the research question. The results of the analysis should be reported in a way that is transparent and reproducible.
Step 6: Report the Results
The sixth step is to report the results of the analysis in a way that is transparent and reproducible. The report should include the DAG, the adjustment set, and the results of the analysis. The report should also include a discussion of the limitations of the study, including the possibility of unmeasured confounding.
The report should follow the reporting guidelines that are appropriate for the study design. The EQUATOR Network provides a comprehensive list of reporting guidelines for different study designs, and the use of these guidelines helps ensure that the report is complete and transparent.
Stratification and Multivariable Adjustment
Stratification and multivariable adjustment are the two most common methods for controlling confounding in observational studies. Both methods aim to estimate the effect of the exposure on the outcome while holding the confounders constant.
Stratification
Stratification involves dividing the study population into strata based on the confounder and then estimating the exposure effect within each stratum. The stratum-specific estimates are then pooled to obtain an overall estimate of the exposure effect.
Stratification is most useful when the confounder has a small number of categories and the sample size is large enough to provide stable estimates within each stratum. For example, if the confounder is sex, the study population can be divided into male and female strata, and the exposure effect can be estimated separately for each sex.
The pooled estimate is a weighted average of the stratum-specific estimates, with the weights proportional to the precision of the estimates. The pooled estimate is valid if the exposure effect is the same across strata, which is an assumption that should be tested.
Stratification becomes impractical when there are many confounders or when the confounders are continuous. With many confounders, the number of strata increases rapidly, and the sample size within each stratum becomes too small to provide stable estimates. With continuous confounders, the confounder must be categorized, which can lead to residual confounding within categories.
Multivariable Adjustment
Multivariable adjustment uses a regression model to estimate the exposure effect while holding the confounders constant. The regression model includes the exposure and the confounders as independent variables, and the outcome as the dependent variable. The coefficient for the exposure variable represents the effect of the exposure on the outcome, adjusted for the confounders.
The most common regression models are linear regression for continuous outcomes, logistic regression for binary outcomes, and Cox proportional hazards regression for time-to-event outcomes. The choice of model depends on the type of outcome and the research question.
Multivariable adjustment is more flexible than stratification because it can handle continuous confounders and many confounders simultaneously. However, it requires the model to be correctly specified. If the model is misspecified, the adjusted estimate can be biased.
Comparing Stratification and Multivariable Adjustment
Stratification and multivariable adjustment are related methods. In fact, multivariable adjustment can be seen as a form of stratification in which the strata are defined by the combination of confounder values. The two methods provide similar results when the model is correctly specified and the sample size is large.
The choice between stratification and multivariable adjustment depends on the research question and the data. Stratification is more transparent and easier to understand, but it is limited in the number of confounders it can handle. Multivariable adjustment is more flexible but requires the model to be correctly specified.
Propensity Score Methods
Propensity score methods are a class of methods that are used to control confounding in observational studies. The propensity score is the probability of being exposed, given the confounders. The propensity score is estimated using a regression model that includes the confounders as independent variables and the exposure as the dependent variable.
Estimating the Propensity Score
The propensity score is estimated using a logistic regression model in which the exposure is the dependent variable and the confounders are the independent variables. The predicted probability of exposure from this model is the propensity score.
The propensity score is a balancing score, which means that, conditional on the propensity score, the distribution of the confounders is the same in the exposed and unexposed groups. This property allows the propensity score to be used to create comparable groups.
Using the Propensity Score
The propensity score can be used in several ways to control confounding. The most common methods are matching, stratification, and adjustment.
In propensity score matching, each exposed participant is matched to one or more unexposed participants with a similar propensity score. The matched sample is then analyzed as if it were a randomized experiment. In propensity score stratification, the study population is divided into strata based on the propensity score, and the exposure effect is estimated within each stratum. In propensity score weighting, each participant is weighted by the inverse of the probability of exposure, which creates a pseudo-population in which the confounders are balanced.
Advantages and Limitations
The propensity score has several advantages over multivariable adjustment. It can handle many confounders, and it can be used to create comparable groups that are easy to understand. It also allows the investigator to check the balance of the confounders between the exposed and unexposed groups.
The propensity score also has limitations. It requires the model to be correctly specified, and it does not address unmeasured confounding. The propensity score can also be sensitive to the choice of the model and the method used to estimate the score.
Instrumental Variable Analysis
Instrumental variable analysis is a method that can be used to control confounding when there is a variable that affects the exposure but does not affect the outcome except through the exposure. This variable is called an instrumental variable.
The Instrumental Variable Assumptions
An instrumental variable must satisfy three assumptions. First, the instrument must be associated with the exposure. Second, the instrument must not be associated with the confounders. Third, the instrument must not affect the outcome except through the exposure.
These assumptions are difficult to verify, and they are often not met in practice. The instrument must be carefully selected, and the assumptions must be justified.
Using an Instrumental Variable
Instrumental variable analysis uses the instrument to estimate the effect of the exposure on the outcome. The analysis is performed using a two-stage least squares regression or a similar method.
Instrumental variable analysis can provide an unbiased estimate of the exposure effect even in the presence of unmeasured confounding, provided the assumptions are met. However, the estimates are often imprecise, and the assumptions are difficult to verify.
Common Failure Patterns in Confounding Control
Confounding control is a complex process, and there are several common failure patterns that can lead to biased results. These failure patterns can be avoided with careful attention to the design and analysis of the study.
Adjusting for Mediators
One common failure pattern is adjusting for a mediator. A mediator is a variable on the causal pathway between the exposure and the outcome. Adjusting for a mediator blocks the causal effect of the exposure on the outcome, which leads to an underestimate of the true effect.
For example, in a study of the effect of a high-fat diet on body weight, blood cholesterol is a mediator. Adjusting for blood cholesterol would remove the effect of the diet on body weight that is mediated by cholesterol, which would underestimate the total effect of the diet.
Adjusting for Colliders
Another common failure pattern is adjusting for a collider. A collider is a variable that is caused by two or more other variables. Adjusting for a collider can introduce bias, known as collider bias, even when the collider is not a confounder.
For example, in a study of the relationship between two genetic variants and a disease, the disease is a collider if it is caused by both variants. Adjusting for the disease, for example by selecting a study population of patients with the disease, can create a spurious association between the variants.
Overadjustment
Overadjustment is the inclusion of too many variables in the adjustment set. Overadjustment can reduce the precision of the estimate and can introduce bias if the variables are mediators or colliders.
The adjustment set should be limited to the variables that are necessary to block all backdoor paths. The DAG can be used to identify the minimal adjustment set, and the analysis should be limited to this set.
Residual Confounding
Residual confounding occurs when the adjustment does not fully control for the confounder. This can happen when the confounder is measured with error, when the confounder is categorized, or when the model is incorrectly specified.
Residual confounding can be reduced by measuring the confounder accurately and by using a flexible model that can capture the relationship between the confounder and the outcome.
Unmeasured Confounding
Unmeasured confounding occurs when a confounder is not measured in the study. Unmeasured confounding cannot be addressed by adjustment, and it can bias the results.
Unmeasured confounding can be addressed by using an instrumental variable or by conducting a sensitivity analysis. A sensitivity analysis can be used to assess how the results would change if an unmeasured confounder were present.
Sensitivity Analysis
Sensitivity analysis is a method for assessing the robustness of the results to violations of the assumptions. Sensitivity analysis can be used to assess the impact of unmeasured confounding, measurement error, and model misspecification.
Assessing Unmeasured Confounding
Sensitivity analysis for unmeasured confounding involves specifying a range of plausible values for the relationship between the unmeasured confounder and the exposure and the outcome. The analysis then determines how the results would change under these assumptions.
The results of a sensitivity analysis can be presented as a range of estimates that are consistent with the data and the assumptions. If the range includes the null value, the results are not robust to unmeasured confounding.
Assessing Model Misspecification
Sensitivity analysis for model misspecification involves fitting a range of models and comparing the results. The analysis can assess the impact of the choice of the model, the categorization of continuous variables, and the inclusion of interaction terms.
The results of a sensitivity analysis can be used to determine the robustness of the results to the model specification. If the results are consistent across a range of models, the results are robust.
Reporting and Reproducibility
Transparent reporting is essential for the credibility of observational studies. The report should include the DAG, the adjustment set, and the results of the analysis. The report should also include the limitations of the study, including the possibility of unmeasured confounding.
Reporting Guidelines
The EQUATOR Network provides a comprehensive list of reporting guidelines for different study designs. The use of these guidelines helps ensure that the report is transparent and complete. The guidelines include the STROBE statement for observational studies, the CONSORT statement for randomized trials, and the PRISMA statement for systematic reviews.
The use of reporting guidelines is important because it helps readers understand the methods and the results. It also helps reviewers assess the quality of the study.
Data Sharing
Data sharing is an important part of reproducibility. The National Institutes of Health (NIH) has a Data Management and Sharing Policy that requires investigators to plan for the management and sharing of data. The policy requires investigators to submit a data management and sharing plan with their application, and the plan should describe how the data will be managed and shared.
Data sharing allows other investigators to reproduce the results and to conduct additional analyses. It also helps to ensure that the results are credible.
Reproducibility
Reproducibility is the ability of another investigator to reproduce the results of a study using the same data and methods. Reproducibility requires that the data and the analysis code be available and that the methods be described in sufficient detail.
The analysis code should be well-documented and should be made available with the data. The methods should be described in sufficient detail to allow another investigator to reproduce the analysis.
A Practical Decision Framework for Selecting a Confounding Control Method
Choosing among stratification, multivariable adjustment, propensity scores, and instrumental variables is a common source of confusion in observational research. Researchers often default to a familiar method without systematically evaluating whether it fits their data structure, sample size, and causal assumptions. A structured decision framework can reduce this problem by forcing explicit consideration of the constraints that matter most in practice.
Step 1: Confirm the Adjustment Set Is Measured
Before selecting any analytic method, verify that every variable in the minimal sufficient adjustment set has been measured with acceptable accuracy. If a required confounder is missing, no adjustment method can recover the causal effect. The choice then narrows to instrumental variable analysis, which can address unmeasured confounding under strong assumptions, or a sensitivity analysis that quantifies how much unmeasured confounding would be needed to change the conclusion.
If all confounders are measured, proceed to the next step. If some are measured with error, consider whether the error is likely to be nondifferential with respect to the outcome. Nondifferential measurement error in a confounder typically leads to residual confounding, meaning the adjusted estimate will still be biased toward the unadjusted association.
Step 2: Assess the Number and Type of Confounders
The number of confounders and their measurement scale determine which methods are feasible. With one categorical confounder and a large sample, stratification is transparent and easy to interpret. With several continuous confounders, multivariable adjustment is more practical because stratification would require creating many strata and would quickly exhaust the sample size.
When the number of confounders is large relative to the number of events or observations, multivariable adjustment can become unstable. In this situation, propensity score methods are attractive because they summarize all confounders into a single score. The propensity score reduces the dimensionality of the adjustment problem and allows the investigator to check balance directly.
Step 3: Check the Overlap Assumption
Every adjustment method requires that the exposure groups overlap in their confounder distributions. If exposed and unexposed participants occupy different regions of the confounder space, no method can credibly estimate the exposure effect without extrapolation.
For propensity score methods, the overlap assumption is assessed by examining the distribution of the propensity score in the exposed and unexposed groups. If the distributions do not overlap substantially, the analysis should be restricted to the region of common support. For multivariable adjustment, overlap is checked by examining the range of confounder values in each exposure group. Poor overlap is a design problem, not an analysis problem, and it should be reported as a limitation.
Step 4: Choose the Primary Method
The primary method should be selected based on the first three steps. The following table summarizes the decision points.
| Decision Point | Condition | Recommended Method |
|---|---|---|
| Confounder measurement | All confounders measured accurately | Stratification, multivariable adjustment, or propensity score |
| Confounder measurement | One or more confounders unmeasured | Instrumental variable analysis or sensitivity analysis |
| Confounder count | One categorical confounder | Stratification |
| Confounder count | Multiple continuous confounders | Multivariable adjustment |
| Confounder count | Many confounders relative to events | Propensity score |
| Overlap | Poor overlap in confounder distributions | Restrict to common support and report limitation |
Step 5: Run a Parallel Method as a Check
A single analysis method can hide specification errors. Running a second method that relies on different assumptions provides a useful consistency check. For example, if the primary analysis uses multivariable adjustment, a propensity score analysis can confirm that the results are not sensitive to the functional form of the adjustment model.
The two estimates should be similar if both models are correctly specified. Large differences between the methods indicate that at least one model is misspecified or that the data do not support the assumptions. This discrepancy should be investigated before reporting results.
Step 6: Document the Decision Trail
The decision framework should be documented in the study protocol or analysis plan before the data are analyzed. The documentation should include the DAG, the adjustment set, the reasons for choosing the primary method, and the planned sensitivity analyses. This documentation supports reproducibility and allows reviewers to evaluate whether the method choice was justified.
Records and Measurements for the Decision Framework
The decision framework requires specific records to be useful. The following records should be maintained for each study.
Confounder Measurement Records
For each confounder in the adjustment set, record the measurement instrument, the timing of measurement, and the proportion of missing values. If a confounder is measured with error, record the validation data that quantify the error. This record supports the assessment of residual confounding.
Overlap Assessment Records
Record the range and distribution of each confounder in the exposed and unexposed groups. For propensity score analyses, record the propensity score distribution and the region of common support. These records allow reviewers to verify that the overlap assumption was checked.
Analysis Log
Maintain a log of all analyses performed, including the primary analysis, the parallel analysis, and any sensitivity analyses. The log should record the software version, the model specification, and the date of each analysis. This log supports reproducibility and helps identify errors in the analysis.
Troubleshooting Common Decision Errors
Several recurring errors appear when researchers apply confounding control methods. Recognizing these patterns can prevent wasted effort and biased conclusions.
Choosing a Method Before Building the DAG
Selecting an adjustment method before constructing the DAG is a common error. The DAG determines which variables need adjustment, and the method should be chosen after the adjustment set is known. Building the DAG first prevents the inclusion of mediators or colliders in the adjustment set.
Using Propensity Scores Without Checking Overlap
Propensity score methods are often applied without checking the overlap assumption. If the propensity score distributions do not overlap, the analysis will rely on extrapolation and produce biased estimates. The overlap check should be a mandatory step in any propensity score analysis.
Adjusting for a Variable That Is Not a Confounder
Including variables that are not confounders in the adjustment model can reduce precision and introduce bias. The DAG should be used to identify the minimal adjustment set, and the analysis should be limited to this set. Variables that are not confounders should not be included.
Ignoring the Parallel Analysis
A parallel analysis is a check, not an optional extra. If the primary and parallel analyses disagree, the discrepancy should be investigated. Ignoring the discrepancy and reporting only the primary analysis undermines the credibility of the results.
Troubleshooting When Methods Disagree
When the primary and parallel analyses produce different estimates, the following steps can identify the cause.
Check the Model Specification
The first step is to check whether the models are correctly specified. For multivariable adjustment, examine whether the functional form of the confounders is appropriate. For propensity score methods, check whether the propensity model includes the correct terms and interactions.
Check the Overlap
The second step is to check whether the overlap assumption holds. If the overlap is poor, the two methods may be estimating effects in different populations. The analysis should be restricted to the common support region.
Check for Influential Observations
The third step is to check for influential observations. A small number of observations with extreme confounder values can drive the difference between methods. The analysis should be repeated after removing these observations to assess their impact.
Report the Discrepancy
If the discrepancy cannot be resolved, it should be reported as a limitation. The report should describe the discrepancy, the possible causes, and the implications for the conclusions. This transparency allows readers to evaluate the robustness of the findings.
Welfare and Safety Context
The decision framework has direct implications for research integrity and participant safety. A biased analysis can lead to incorrect conclusions about the safety or effectiveness of an intervention. The Committee on Publication Ethics core practices emphasize the importance of accurate reporting and the avoidance of misleading results. The decision framework supports these practices by making the method choice transparent and by requiring a parallel analysis that can reveal errors.
The framework also supports the responsible use of research resources. The NIH Grants and Funding process requires that research be conducted with rigor and transparency. The decision framework provides a documented trail that can be reviewed by funders and regulators. The NIH Data Management and Sharing Policy requires that data and analysis plans be shared, and the decision framework provides the documentation needed to meet this requirement.
Frequently Asked Questions
How do I choose between propensity score matching and propensity score weighting?
Propensity score matching creates a sample of exposed and unexposed participants with similar scores, while weighting uses the inverse probability of exposure to create a pseudo-population. Matching is easier to understand and check, but it discards data. Weighting uses all data but can be sensitive to extreme weights. The choice depends on the sample size and the distribution of the scores.
What is the minimum sample size for propensity score methods?
There is no fixed minimum sample size. The requirement is that the propensity score model can be estimated and that the overlap is adequate. With a small sample, the score model may be unstable and the overlap may be poor. In this case, multivariable adjustment may be more appropriate.
Can I use the decision framework for a case-control study?
Yes. The framework applies to case-control studies, but the propensity score must be estimated with the case-control sampling in mind. The score is the probability of exposure given the confounders, and the sampling design must be accounted for in the estimation.
What should I do if a confounder is measured in only a subset of participants?
The first step is to assess the missingness. If the missingness is related to the exposure or outcome, the analysis may be biased. Multiple imputation can be used to fill in the missing values, but the imputation model must include the exposure, outcome, and other confounders.
How do I report the decision framework in a paper?
The decision framework should be described in the methods section. The description should include the DAG, the adjustment set, the reasons for the method choice, and the parallel analysis. The results section should report the primary and parallel estimates and any discrepancy.
What is the role of the DAG in the decision framework?
The DAG is the foundation of the framework. It identifies the adjustment set and the variables that should not be adjusted for. The DAG should be built before the method is selected, and it should be reported in the paper.
Frequently Asked Questions
What is the difference between a confounder and a mediator?
A confounder is a variable that is associated with both the exposure and the outcome and that is not on the causal pathway. A mediator is a variable on the causal pathway between the exposure and the outcome. Confounders should be adjusted for, while mediators should not be adjusted for because doing so blocks the causal effect of the exposure.
How do I know if a variable is a confounder?
A variable is a confounder if it is associated with the exposure, associated with the outcome, and not on the causal pathway. The identification of confounders requires subject matter knowledge and should be based on the existing literature and the causal mechanisms.
What is a directed acyclic graph?
A directed acyclic graph (DAG) is a visual representation of the causal relationships among variables in a study. DAGs use arrows to represent the direction of causal influence, and they are used to identify the variables that need to be adjusted for to control confounding.
What is the difference between stratification and multivariable adjustment?
Stratification divides the study population into strata based on the confounder and estimates the exposure effect within each stratum. Multivariable adjustment uses a regression model to estimate the exposure effect while holding the confounders constant. Both methods aim to control confounding, but multivariable adjustment is more flexible.
What is a propensity score?
A propensity score is the probability of being exposed, given the confounders. The propensity score is used to create comparable groups in observational studies. It can be used for matching, stratification, or weighting.
What is unmeasured confounding?
Unmeasured confounding occurs when a confounder is not measured in the study. Unmeasured confounding cannot be addressed with adjustment, and it can bias the results. Sensitivity analysis can be used to assess the impact of unmeasured confounding.
What are reporting guidelines?
Reporting guidelines are checklists that describe the minimum information that should be included in a report of a study. The EQUATOR Network provides a comprehensive list of reporting guidelines for different study designs. The use of reporting guidelines helps ensure that the report is transparent and complete.
How can I make my analysis reproducible?
To make your analysis reproducible, you should make the data and the analysis code available and describe the methods in sufficient detail. The analysis code should be well-documented, and the methods should be described in a way that allows another investigator to reproduce the analysis.
Related Bioinformatics Guides
- What Is a Data Warehouse? A Practical Guide for Life Science Organizations
- Data Science and AI in Life Sciences: Applications and Emerging Trends
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- FAIR Data Maturity Model: A Practical Assessment Framework for Bioinformatics Workflows
- Metabolomics Data Analysis in R: A Practical Workflow
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Sensitivity Analysis in Observational Research: Introducing the E-Value.. Annals of internal medicine, 2017.
- Accounting for Confounding in Observational Studies.. Annual review of clinical psychology, 2020.
- Confounding adjustment in observational studies on cardiothoracic interventions: a systematic review of methodological practice.. European journal of cardio-thoracic surgery : official journal of the European Association for Cardio-thoracic Surgery, 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.