Handling Missing Data in Longitudinal Studies

By Dr. Zubair Khalid, DVM, MS, PhD ·

Handling Missing Data in Longitudinal Studies

Key Takeaways

  • The mechanism of missing data (Missing Completely at Random - MCAR, Missing at Random - MAR, or Missing Not at Random - MNAR) is paramount and dictates the defensibility of analytical methods; MCAR implies missingness is unrelated to any observed or unobserved data, MAR means missingness depends only on observed data, and MNAR signifies missingness is related to the unobserved value itself.
  • Complete-case analysis, which excludes any subject with missing data, is only valid under MCAR and risks substantial bias if missingness is related to observed or unobserved outcomes, particularly in longitudinal studies where dropout is common.
  • Last Observation Carried Forward (LOCF) is a single imputation method that assumes outcome constancy post-dropout, a biologically implausible assumption for most longitudinal processes, leading to significant bias and underestimated variance, and should be reserved for sensitivity analyses.
  • Multiple Imputation (MI) is the preferred primary analysis method when data are MAR, as it leverages observed data, including auxiliary predictors of dropout, to generate multiple plausible complete datasets, thereby accounting for imputation uncertainty and yielding valid inferences.
  • For MNAR data, where missingness is informative and directly related to the unobserved value (e.g., a worsening biomarker leading to study withdrawal), no standard method can produce unbiased estimates without strong, untestable assumptions; study design must prioritize participant retention and collection of dropout predictors.
  • A structured workflow involving documenting missing data patterns, assessing the missingness mechanism based on observed data and study design, selecting the primary method (MI for MAR), conducting sensitivity analyses, and transparently reporting all steps is crucial for robust longitudinal data analysis.

Quick Answer

  • Missing data in longitudinal studies requires identifying the mechanism (missing completely at random, missing at random, or missing not at random) before selecting an analysis method.
  • Complete-case analysis and last observation carried forward introduce bias under most realistic dropout patterns, while multiple imputation provides valid inference when data are missing at random.
  • No method recovers information lost to missing not at random processes, so study design should prioritize retention and collection of auxiliary predictors of dropout.

At a Glance

MethodAssumption About MissingnessBias RiskPractical Use
Complete-case analysisMissing completely at randomHigh when dropout relates to observed or unobserved valuesOnly defensible when few cases are missing and the mechanism is verifiable
Last observation carried forwardMissing completely at random and constant trajectoryHigh when values change over timeAvoid for primary analysis, may serve as sensitivity check
Multiple imputationMissing at randomLow when auxiliary variables are includedPreferred for primary analysis when missingness depends on observed data

Missing Data Mechanisms in Longitudinal Research

Longitudinal studies collect repeated measurements from the same subjects over time. The defining feature of this design is that each subject contributes multiple observations, and those observations are correlated within the subject. When a subject misses a scheduled visit or withdraws from the study entirely, the resulting gaps in the data create a missing data problem that is distinct from cross-sectional missingness because the missing values are embedded in a sequence of related measurements.

The statistical literature classifies missing data into three mechanisms, and this classification determines which analytical approach is defensible. Data are missing completely at random when the probability of a missing value is unrelated to both observed and unobserved values. Data are missing at random when the probability of missingness depends on observed values but not on the missing values themselves. Data are missing not at random when the probability of missingness depends on the unobserved value, meaning the very measurement that is missing influences whether it was recorded.

These distinctions matter because they determine whether the observed data can stand in for the missing data without introducing bias. When data are missing completely at random, the observed cases are a random subset of the full sample, and complete-case analysis yields unbiased estimates. When data are missing at random, the observed data contain information about the missingness process, and methods such as multiple imputation can recover unbiased estimates. When data are missing not at random, the missingness is informative, and no standard method can produce unbiased estimates without strong and untestable assumptions.

In biological research, the missing at random assumption is often plausible when dropout is related to measured covariates. For example, a subject with a high baseline inflammatory marker may be more likely to miss a follow-up visit because of illness, and if that baseline marker is recorded, the missingness is explained by an observed value. The missing not at random mechanism arises when the missingness is driven by the outcome itself, such as when a subject with a worsening biomarker stops attending visits because of disease progression.

The practical implication is that the investigator must think about why data are missing before choosing an analytical method. This reasoning should be documented in the study protocol and in the final report, because the choice of method rests on assumptions that cannot be fully verified from the data alone. The National Library of Medicine research methods resources provide background on the conceptual foundations of missing data mechanisms and their role in biomedical research design.

Why Dropout Creates Bias in Longitudinal Analysis

Dropout is the most common form of missing data in longitudinal studies, and it differs from intermittent missingness in an important way. When a subject drops out, all subsequent measurements are missing, and the subject contributes only a partial sequence of data. Intermittent missingness, by contrast, occurs when a subject misses one visit but returns for later visits, leaving gaps in an otherwise complete sequence.

The bias introduced by dropout depends on the relationship between the dropout process and the outcome trajectory. If subjects who drop out are systematically different from subjects who remain, the observed data no longer represent the original sample. The mean trajectory estimated from the observed data will be shifted toward the subjects who stayed, and this shift can be large when dropout is common.

Consider a study of cognitive decline in aging adults. If participants with faster cognitive decline are more likely to withdraw because they find the testing burdensome, the observed data will show a flatter decline than the true population trajectory. The complete-case analysis will underestimate the rate of decline, and the magnitude of the bias will grow with the dropout rate.

The direction of bias is not always predictable. In some studies, subjects with worse outcomes may be more likely to remain because they are receiving treatment, while in other studies, subjects with worse outcomes may be more likely to leave because they are dissatisfied. The investigator cannot assume that dropout is benign, and the analysis must account for the possibility that the observed data are a selected subset.

The bias also affects the precision of estimates. When subjects drop out, the effective sample size at later time points is reduced, and the standard errors of the estimated trajectories increase. This loss of precision is separate from the bias and occurs even when the dropout is completely at random.

The severity of the problem depends on the proportion of missing data and the strength of the relationship between missingness and the outcome. A small amount of missing data under a missing completely at random mechanism may have little impact, while a large amount of missing data under a missing not at random mechanism can invalidate the study. The investigator should report the proportion of missing data at each time point and the reasons for missingness when known.

Complete-Case Analysis and Its Limitations

Complete-case analysis restricts the dataset to subjects who have no missing values at any time point. This approach is simple to implement and is often the default in statistical software, but it is only valid under the missing completely at random mechanism. When the missingness depends on observed or unobserved values, complete-case analysis produces biased estimates.

The bias in complete-case analysis arises because the subjects who remain are a selected subset of the original sample. If the selection is related to the outcome, the estimated trajectories are shifted. The direction of the shift depends on the direction of the selection, and the magnitude depends on the proportion of subjects excluded.

Complete-case analysis also discards information. A subject who has complete data at the first three visits but misses the fourth visit is excluded entirely, even though the first three measurements are valid and informative. This loss of information reduces statistical power and can make the study underpowered to detect the effects of interest.

The missing completely at random assumption is rarely testable in practice. The investigator can compare the baseline characteristics of subjects who complete the study with those who drop out, but this comparison only addresses whether missingness is related to observed baseline values. It does not address whether missingness is related to unobserved values, which is the defining feature of the missing not at random mechanism.

Complete-case analysis may be acceptable when the proportion of missing data is small and the investigator has strong evidence that the missingness is unrelated to the outcome. In this situation, the bias is likely to be small, and the simplicity of the approach may outweigh the potential for bias. However, the investigator should report the proportion of missing data and the reasons for missingness so that readers can assess the plausibility of the assumption.

The EQUATOR Network provides reporting guidelines that help investigators document the handling of missing data in a transparent way. Transparent reporting includes a description of the missing data mechanism, the proportion of missing values, and the sensitivity of the results to the analytical method.

Last Observation Carried Forward and Its Risks

Last observation carried forward is a single imputation method that replaces each missing value with the last observed value for that subject. The method is easy to implement and is sometimes used in clinical trials, but it rests on a strong assumption that is rarely met in biological data.

The assumption is that the outcome remains constant after the last observed measurement. This assumption is implausible for most biological processes, which change over time. In a study of tumor growth, a subject who drops out after the third visit is assumed to have the same tumor size at the fourth visit as at the third visit, even though the tumor would likely have grown or shrunk in the interim.

The bias introduced by last observation carried forward depends on the direction of the true change. If the outcome is expected to decline over time, the method overestimates the outcome at later time points. If the outcome is expected to increase, the method underestimates the outcome. The method also underestimates the variance because it treats the imputed values as if they were observed, which leads to confidence intervals that are too narrow.

The method is particularly problematic when dropout is related to the outcome. If subjects with worse outcomes are more likely to drop out, last observation carried forward will carry forward the last observed value, which may be the worst value, and the analysis will underestimate the true decline. This can lead to the conclusion that the treatment is effective when it is not.

Last observation carried forward is not recommended for primary analysis in most longitudinal studies. It may be used as a sensitivity analysis to assess the robustness of the results to the missing data assumption, but it should not be the only method. The investigator should be aware that the method is conservative in some settings and anti-conservative in others, and the direction of the bias depends on the data.

The National Library of Medicine research methods resources describe the limitations of single imputation methods and the rationale for more sophisticated approaches. The key limitation is that single imputation methods do not account for the uncertainty of the imputed values, which leads to underestimated standard errors and inflated type I error rates.

Multiple Imputation for Longitudinal Data

Multiple imputation is a method that creates several complete datasets by filling in the missing values with plausible values drawn from a predictive distribution. Each imputed dataset is analyzed using standard methods, and the results are combined to produce estimates that account for the uncertainty in the imputed values.

The method works in three steps. First, the investigator creates multiple imputed datasets, typically between 5 and 20, by drawing values from a model that predicts the missing values based on the observed data. Second, the investigator analyzes each imputed dataset separately using the same statistical method that would be used for complete data. Third, the investigator combines the estimates and standard errors across the imputed datasets using the rules for combining multiple imputation results.

The imputation model must include all variables that are related to the missingness mechanism and the outcome. These variables are called auxiliary variables, and they are included in the imputation model even if they are not included in the analysis model. The inclusion of auxiliary variables improves the plausibility of the missing at random assumption because it captures more of the information that drives the missingness.

Multiple imputation is valid when the missing at random assumption holds, meaning that the missingness depends on observed values. The method does not require the missing completely at random assumption, which makes it more flexible than complete-case analysis. The method also produces standard errors that reflect the uncertainty in the imputed values, which is a major advantage over single imputation methods.

The implementation of multiple imputation requires careful attention to the imputation model. The model should be specified correctly, and the variables should be measured on the appropriate scale. The imputation model should include the outcome variable, the time variable, and any covariates that are related to the missingness. The model should also include the interactions between time and covariates if the trajectories differ across groups.

The National Library of Medicine research methods resources provide an overview of the principles of multiple imputation and the conditions under which it is valid. The investigator should be aware that multiple imputation is not a solution for missing not at random data, and the results are only as good as the imputation model.

Choosing Between Methods Based on the Missing Data Mechanism

The choice of method depends on the missing data mechanism, which the investigator must reason about from the study design and the observed data. The investigator should document the reasoning in the analysis plan and the final report.

When the missingness is missing completely at random, complete-case analysis is valid and is the simplest approach. The investigator can test the plausibility of this assumption by comparing the baseline characteristics of subjects who complete the study with those who drop out. If the groups are similar, the assumption is more plausible, but the test is not definitive because it does not address unobserved values.

When the missingness is missing at random, multiple imputation is the preferred method. The investigator should include auxiliary variables in the imputation model to capture the information that drives the missingness. The analysis should be repeated with different imputation models to assess the sensitivity of the results to the model specification.

When the missingness is missing not at random, no standard method is valid. The investigator should consider the possibility that the missingness is informative and should conduct a sensitivity analysis to assess the impact of the missingness on the results. The sensitivity analysis can use a pattern-mixture model or a selection model, which explicitly model the relationship between the missingness and the outcome.

The investigator should also consider the proportion of missing data. When the proportion is small, the choice of method may have little impact on the results. When the proportion is large, the choice of method can have a substantial impact, and the investigator should be cautious about the interpretation of the results.

The EQUATOR Network provides reporting guidelines that help investigators describe the missing data mechanism and the methods used to handle it. The guidelines emphasize the importance of transparency in reporting, so that readers can assess the validity of the analysis.

Practical Workflow for Handling Missing Data

The following workflow provides a structured approach to handling missing data in longitudinal studies. The workflow is designed to be implemented at the design stage and the analysis stage.

Step 1: Document the Missing Data Pattern

The first step is to document the missing data pattern in the dataset. The investigator should create a table that shows the number of subjects with complete data, the number with missing values at each time point, and the reasons for the missingness when available. The table should also show the proportion of missing data at each time point.

The documentation should include the reasons for the missingness, such as dropout, missed visits, or technical failures. The reasons are important because they provide information about the missingness mechanism. For example, a subject who misses a visit because of a scheduling conflict is more likely to be missing completely at random than a subject who misses a visit because of illness.

Step 2: Assess the Missingness Mechanism

The second step is to assess the plausibility of the missingness mechanism. The investigator should compare the baseline characteristics of subjects with complete data and subjects with missing data. The comparison should include the outcome variables and the covariates.

The investigator should also examine the relationship between the missingness and the observed values. For example, the investigator can test whether the probability of missingness at a given time point is related to the outcome at the previous time point. This test provides evidence about the missingness mechanism, but it is not conclusive.

Step 3: Choose the Primary Analysis Method

The third step is to choose the primary analysis method based on the missingness mechanism. The investigator should specify the method in the analysis plan before the analysis is conducted. The choice should be based on the reasoning about the missingness mechanism, not on the results of the analysis.

The primary analysis method should be the method that is most likely to be valid under the assumed missingness mechanism. If the missingness is missing at random, multiple imputation is the preferred method. If the missingness is missing completely at random, complete-case analysis is acceptable.

Step 4: Conduct Sensitivity Analyses

The fourth step is to conduct sensitivity analyses to assess the robustness of the results to the missingness assumptions. The sensitivity analyses should include alternative methods, such as last observation carried forward, and alternative imputation models.

The sensitivity analyses should be reported alongside the primary analysis. The investigator should state whether the conclusions are robust to the missingness assumptions. If the conclusions change under different assumptions, the investigator should report this and discuss the implications.

Step 5: Report the Missing Data Handling

The fifth step is to report the missing data handling in the final report. The report should include the number of missing values, the reasons for the missingness, the missingness mechanism, the primary analysis method, and the sensitivity analyses.

The report should follow the reporting guidelines for the study design. The EQUATOR Network provides a list of reporting guidelines for different study designs, and the investigator should select the appropriate guideline for the study.

Records and Measurements for Missing Data

The investigator should maintain records of the missing data throughout the study. The records should include the following information for each subject and each time point:

  • The scheduled visit date and the actual visit date
  • The reason for the missingness, if the visit was missed
  • The outcome measurements at each visit
  • The covariates measured at each visit
  • The date of dropout, if the subject withdrew

The records should be maintained in a structured format, such as a spreadsheet or a database. The records should be checked for accuracy and completeness at regular intervals during the study.

The investigator should also record the number of subjects who were screened, the number who were enrolled, and the number who completed the study. This information is important for assessing the generalizability of the results and the potential for selection bias.

The records should be stored in a secure location and should be accessible to the study team. The records should be retained for the duration of the study and for the period required by the funding agency or the institution.

The National Institutes of Health provides guidance on the management of research data, including the retention and sharing of data. The investigator should be aware of the data management requirements of the funding agency and the institution.

Common Failure Patterns in Missing Data Handling

Several common failure patterns can undermine the validity of the analysis. The investigator should be aware of these patterns and take steps to avoid them.

Failure to Report the Missing Data

The most common failure is the failure to report the missing data. The investigator may not report the number of missing values, the reasons for the missingness, or the method used to handle the missing data. This failure makes it impossible for the reader to assess the validity of the analysis.

Using Last Observation Carried Forward as the Primary Analysis

The use of last observation carried forward as the primary analysis is a common failure. The method is not valid for most longitudinal data because it assumes that the outcome is constant after the last observed value. The method should be used only as a sensitivity analysis.

Ignoring the Missingness Mechanism

The investigator may ignore the missingness mechanism and use a method that is not appropriate for the data. For example, the investigator may use complete-case analysis when the missingness is missing at random, which produces biased estimates.

Overinterpreting the Results

The investigator may overinterpret the results of the analysis when the missing data is substantial. The results should be interpreted with caution, and the uncertainty should be acknowledged.

Failing to Conduct Sensitivity Analyses

The investigator may fail to conduct sensitivity analyses to assess the robustness of the results to the missingness assumption. The sensitivity analyses are important because the missingness mechanism cannot be verified from the data.

Welfare and Safety Context

The handling of missing data has implications for the welfare of the study participants and the safety of the study. The investigator should ensure that the study is conducted in a way that minimizes the burden on the participants and the risk of dropout.

The study design should include strategies to minimize the dropout, such as flexible scheduling, reminders, and incentives. The investigator should also monitor the dropout rate during the study and take action if the dropout rate is higher than expected.

The investigator should also ensure that the study is conducted in accordance with the ethical principles of research. The study should be reviewed by an institutional review board, and the participants should provide informed consent. The investigator should also ensure that the data is handled in a way that protects the privacy of the participants.

The Committee on Publication Ethics provides guidance on the ethical conduct of research, including the handling of data and the reporting of results. The investigator should be aware of the ethical requirements of the study and the publication of the results.

Professional Escalation Criteria

The investigator should seek professional advice when the missing data is complex or when the missingness mechanism is uncertain. The following situations warrant escalation to a statistician or a data scientist:

  • The proportion of missing data is large, such as more than 20 percent of the total data
  • The missingness mechanism is not clear from the observed data
  • The missingness is likely to be missing not at random
  • The analysis requires advanced methods, such as pattern-mixture models or weighting models
  • The results are sensitive to the missingness assumption

The investigator should also escalate when the missing data analysis is not consistent with the reporting guidelines or the requirements of the funding agency.

The National Institutes of Health provides guidance on the statistical analysis of research data, and the investigator should consult the guidance when the analysis is complex. The investigator should also consult the ORCID for Researchers to ensure that the researcher identity is maintained and the research output is properly attributed.

A Decision Framework for Selecting Missing Data Methods in Longitudinal Studies

Choosing a missing data method is not a single decision made once at the analysis stage. It is a sequence of decisions that begins during study design and continues through data collection, analysis, and reporting. A structured decision framework helps investigators avoid the common failure of selecting a method based on convenience or habit instead of on the specific characteristics of their data. The framework below organizes the decision process into a series of checkpoints, each with concrete criteria and actions.

Checkpoint 1: Classify the Missingness Pattern at the Design Stage

The first checkpoint occurs before data collection begins. The investigator should anticipate the types of missingness that are likely to occur and plan the data collection procedures accordingly. This planning includes defining what constitutes a missed visit, how dropout will be recorded, and what auxiliary variables will be collected to support the missingness assumption.

The classification of missingness patterns should distinguish between monotone and non-monotone missingness. Monotone missingness occurs when a subject drops out and never returns, creating a pattern where missingness at one time point implies missingness at all subsequent time points. Non-monotone missingness occurs when a subject misses a visit but returns later, creating gaps in the sequence. The distinction matters because some analytical methods are more straightforward to implement for monotone missingness, while others handle non-monotone patterns more naturally.

The investigator should also plan the collection of auxiliary variables that are likely to be related to the missingness mechanism. These variables might include measures of disease severity, treatment adherence, or participant satisfaction. The inclusion of auxiliary variables in the imputation model improves the plausibility of the missing at random assumption because it captures more of the information that drives the missingness.

The National Institutes of Health provides guidance on the design of longitudinal studies, including the planning of data collection procedures. The investigator should consult this guidance when designing the study and should document the planned missing data handling in the study protocol.

Checkpoint 2: Assess the Missingness Mechanism Using Observed Data

The second checkpoint occurs after data collection is complete and before the primary analysis is conducted. The investigator should assess the plausibility of the missingness mechanism using the observed data. This assessment is not a formal statistical test, but rather a structured examination of the data that informs the choice of method.

The first step in this assessment is to compare the baseline characteristics of subjects with complete data and subjects with missing data. The comparison should include the outcome variables and the covariates. If the groups are similar, the missing completely at random assumption is more plausible. If the groups differ, the missingness is likely to be missing at random or missing not at random.

The second step is to examine the relationship between the missingness and the observed values over time. The investigator can test whether the probability of missingness at a given time point is related to the outcome at the previous time point. This test provides evidence about the missingness mechanism, but it is not conclusive because it does not address unobserved values.

The third step is to examine the pattern of missingness across time points. If the missingness is monotone, the investigator should record the number of subjects who drop out at each time point and the reasons for the dropout. If the missingness is non-monotone, the investigator should record the number of subjects with intermittent missingness and the reasons for the missed visits.

The National Library of Medicine research methods resources provide background on the assessment of missingness mechanisms and the interpretation of the results. The investigator should use these resources to guide the assessment and to document the findings.

Checkpoint 3: Select the Primary Analysis Method Based on the Mechanism

The third checkpoint is the selection of the primary analysis method. The selection should be based on the assessment of the missingness mechanism from Checkpoint 2, not on the results of the analysis. The investigator should specify the method in the analysis plan before the analysis is conducted.

The decision table below summarizes the recommended primary analysis method for each missingness mechanism.

Missingness MechanismRecommended Primary MethodRationale
Missing completely at randomComplete-case analysisThe observed cases are a random subset of the full sample, and the estimates are unbiased
Missing at randomMultiple imputationThe observed data contain information about the missingness process, and the method recovers unbiased estimates
Missing not at randomNo standard methodThe missingness is informative, and the estimates are biased under any standard method

When the missingness is missing completely at random, complete-case analysis is the simplest and most defensible method. The investigator should verify the plausibility of this assumption by comparing the baseline characteristics of subjects with complete and missing data. If the groups are similar, the assumption is more plausible, but the test is not definitive because it does not address unobserved values.

When the missingness is missing at random, multiple imputation is the preferred method. The investigator should include auxiliary variables in the imputation model to capture the information that drives the missingness. The analysis should be repeated with different imputation models to assess the sensitivity of the results to the model specification.

When the missingness is missing not at random, no standard method is valid. The investigator should consider the possibility that the missingness is informative and should conduct a sensitivity analysis to assess the impact of the missingness on the results. The sensitivity analysis can use a pattern-mixture model or a selection model, which explicitly model the relationship between the missingness and the outcome.

Checkpoint 4: Conduct Sensitivity Analyses to Assess Robustness

The fourth checkpoint is the sensitivity analysis. The sensitivity analysis is a set of analyses that assess the robustness of the results to the missingness assumptions. The sensitivity analyses should include alternative methods, such as last observation carried forward, and alternative imputation models.

The sensitivity analysis should be reported alongside the primary analysis. The investigator should state whether the conclusions are robust to the missingness assumptions. If the conclusions change under different assumptions, the investigator should report this and discuss the implications.

The sensitivity analysis is particularly important when the missingness is missing not at random. In this situation, the primary analysis is not valid, and the sensitivity analysis provides the only way to assess the impact of the missingness on the results. The sensitivity analysis can use a pattern-mixture model, which assumes that the missingness is related to the outcome, or a selection model, which assumes that the missingness is related to the outcome and the covariates.

The EQUATOR Network provides reporting guidelines that help investigators document the sensitivity analyses in a transparent way. The guidelines emphasize the importance of reporting the sensitivity analyses so that readers can assess the validity of the analysis.

Checkpoint 5: Document the Decision Process and the Results

The fifth checkpoint is the documentation of the decision process and the results. The documentation should include the missingness pattern, the assessment of the missingness mechanism, the primary analysis method, and the sensitivity analyses. The documentation should be included in the final report.

The documentation should follow the reporting guidelines for the study design. The EQUATOR Network provides a list of reporting guidelines for different study designs, and the investigator should select the appropriate guideline for the study.

The documentation should also include the data management and sharing plan. The National Institutes of Health provides guidance on the data management and sharing expectations for NIH-funded research. The investigator should be aware of the data management requirements of the funding agency and the institution.

Records and Measurements for the Decision Framework

The decision framework requires the collection of specific records and measurements. The investigator should maintain the following records for each subject and each time point:

  • The scheduled visit date and the actual visit date
  • The reason for the missingness, if the visit was missed
  • The outcome measurements at each visit
  • The covariates measured at each visit
  • The date of dropout, if the subject withdrew
  • The auxiliary variables that are related to the missingness

The records should be maintained in a structured format, such as a spreadsheet or a database. The records should be checked for accuracy and completeness at regular intervals during the study.

The investigator should also record the number of subjects who were screened, the number who were enrolled, and the number who completed the study. This information is important for assessing the generalizability of the results and the potential for selection bias.

The records should be stored in a secure location and should be accessible to the study team. The records should be retained for the duration of the study and for the period required by the funding agency or the institution.

Common Failure Patterns in the Decision Framework

Several common failure patterns can undermine the decision framework. The investigator should be aware of these patterns and take steps to avoid them.

Failure to Document the Missingness Pattern

The most common failure is the failure to document the missingness pattern. The investigator may not record the number of missing values, the reasons for the missingness, or the pattern of missingness. This failure makes it impossible to assess the missingness mechanism and to select the appropriate method.

Failure to Assess the Missingness Mechanism

The investigator may fail to assess the missingness mechanism and may select a method based on convenience instead of on the data. For example, the investigator may use complete-case analysis when the missingness is missing at random, which produces biased estimates.

Failure to Conduct Sensitivity Analyses

The investigator may fail to conduct sensitivity analyses to assess the robustness of the results to the missingness assumptions. The sensitivity analyses are important because the missingness mechanism cannot be verified from the data.

Failure to Report the Decision Process

The investigator may fail to report the decision process in the final report. This failure makes it impossible for the reader to assess the validity of the analysis and the plausibility of the missingness assumptions.

Failure to Use the Decision Framework at the Design Stage

The investigator may fail to use the decision framework at the design stage. The framework should be used to plan the data collection procedures and to identify the auxiliary variables that will be collected. The failure to plan for missing data at the design stage can lead to a lack of auxiliary variables and a less plausible missing at random assumption.

Welfare and Safety Context

The decision framework has implications for the welfare of the study participants and the safety of the study. The investigator should ensure that the study is conducted in a way that minimizes the burden on the participants and the risk of dropout.

The study design should include strategies to minimize the dropout, such as flexible scheduling, reminders, and incentives. The investigator should also monitor the dropout rate during the study and take action if the dropout rate is higher than expected.

The investigator should also ensure that the study is conducted in accordance with the ethical principles of research. The study should be reviewed by an institutional review board, and the participants should provide informed consent. The investigator should also ensure that the data is handled in a way that protects the privacy of the participants.

The Committee on Publication Ethics provides guidance on the ethical conduct of research, including the handling of data and the reporting of results. The investigator should be aware of the ethical requirements of the study and the publication of the results.

Professional Escalation Criteria

The investigator should seek professional advice when the missing data is complex or when the missingness mechanism is uncertain. The following situations warrant escalation to a statistician or a data scientist:

  • The proportion of missing data is large, such as more than 20 percent of the total data
  • The missingness mechanism is not clear from the observed data
  • The missingness is likely to be missing not at random
  • The analysis requires advanced methods, such as pattern-mixture models or weighting models
  • The results are sensitive to the missingness assumption

The investigator should also escalate when the missing data analysis is not consistent with the reporting guidelines or the requirements of the funding agency.

The National Institutes of Health provides guidance on the statistical analysis of research data, and the investigator should consult the guidance when the analysis is complex. The investigator should also consult the ORCID for Researchers to ensure that the researcher identity is maintained and the research output is properly attributed.

Frequently Asked Questions

What is the difference between missing completely at random and missing at random?

Missing completely at random means the probability of missingness is unrelated to both observed and unobserved values. Missing at random means the probability of missingness depends on observed values but not on the missing values themselves. The distinction matters because complete-case analysis is valid for missing completely at random, while multiple imputation is valid for missing at random.

Why is last observation carried forward not recommended?

Last observation carried forward assumes the outcome is constant after the last observed measurement, which is rarely true for biological processes. The method produces biased estimates when the outcome changes over time and when dropout is related to the outcome. The method also underestimates the variance because it treats the imputed values as if they were observed.

When is complete-case analysis acceptable?

Complete-case analysis is acceptable when the missingness is missing completely at random. The investigator should verify the plausibility of this assumption by comparing the baseline characteristics of subjects with complete and missing data. The method is also acceptable when the proportion of missing data is small and the missingness is not related to the outcome.

What is multiple imputation and how does it work?

Multiple imputation creates multiple imputed datasets by replacing the missing values with plausible values based on the observed data. Each imputed dataset is analyzed separately, and the results are combined to produce estimates that account for the uncertainty in the imputed values. The method is valid when the missingness is missing at random.

How do I choose the imputation model for multiple imputation?

The imputation model should include all variables that are related to the missingness mechanism and the outcome variables. The model should include the outcome variables, the covariates, and any auxiliary variables that are related to the missingness. The model should be specified to include the interactions between time and the outcome if the trajectories are expected to differ.

What is the difference between missing at random and missing not at random?

Missing at random means the missingness depends on observed values but not on the missing values themselves. Missing not at random means the missingness depends on the unobserved value. The distinction is important because multiple imputation is valid for missing at random but not for missing not at random.

How do I report the missing data handling in my study?

The report should include the number of missing values, the reasons for the missingness, the missingness mechanism, the primary analysis method, and the sensitivity analyses. The report should follow the reporting guidelines for the study design, which are available from the EQUATOR Network.

What should I do if the missingness is missing not at random?

If the missingness is missing not at random, no standard method is valid. The investigator should consider the pattern of the missingness and use a sensitivity analysis to assess the impact of the missingness on the results. The investigator should also consider the use of pattern-mixture models or weighting models, which assume a relationship between the missingness and the outcome.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.