# Handling Missing Data in Biological Research


## Key Takeaways

- The primary decision in handling missing biological data involves identifying the missingness mechanism (MCAR, MAR, or MNAR) before selecting an analysis method, as this dictates the validity of subsequent statistical inferences.
- Multiple imputation is the recommended practical approach for most missing data scenarios, particularly under the Missing at Random (MAR) assumption, as it accounts for imputation uncertainty and yields more robust estimates than single imputation methods.
- Sensitivity analysis is crucial for assessing the robustness of conclusions, especially when missingness is suspected to be Missing Not At Random (MNAR), by evaluating how results change under various plausible missingness assumptions.
- Prevention of missing data during the experimental design and data collection phases remains the most effective strategy, as no statistical method can fully recover information lost due to nonrandom missingness (MNAR).
- Complete case analysis (listwise deletion) is only valid under the stringent Missing Completely At Random (MCAR) assumption, which is rarely met in biological research, and often leads to biased estimates and reduced statistical power.
- Transparent documentation and reporting of missing data patterns, the chosen handling methods, and the results of sensitivity analyses are essential for research reproducibility and adherence to ethical publication standards.

---

## Quick Answer

- Missing data in biological research can bias estimates and reduce statistical power, so the first decision is identifying the missingness mechanism before choosing an analysis method.
- Multiple imputation is the preferred practical approach for most missing data scenarios, and sensitivity analysis is required to assess how conclusions change under different assumptions.
- No method can recover information lost to nonrandom missingness (MNAR), so prevention during data collection remains the most important control.

## Understanding Missing Data in Biological Research

Missing data is a structural feature of nearly every biological study, from genomics and proteomics to clinical trials and ecological field experiments. Instruments fail, samples degrade, subjects withdraw, and recording errors occur. The way researchers handle these gaps determines whether the final analysis reflects biological reality or an artifact of the missingness process.

The core problem is that missing data affects two statistical properties simultaneously. First, it reduces the sample size available for analysis, which lowers statistical power and widens confidence intervals. Second, and more seriously, missing data can introduce bias when the observations that are missing differ systematically from those that are present. A study reporting a treatment effect based on complete cases only may be describing the subset of samples that survived the experiment, not the full population of interest.

The statistical literature distinguishes three missingness mechanisms, and this classification drives every subsequent decision. Missing completely at random (MCAR) means the probability of missingness is unrelated to both observed and unobserved values. Missing at random (MAR) means the probability of missingness depends on observed variables but not on the missing values themselves. Missing not at random (MNAR) means the probability of missingness depends on the unobserved value itself, even after accounting for observed variables.

The distinction matters because each mechanism requires a different analytical approach. MCAR allows simple methods like complete-case analysis without introducing bias, though power is still lost. MAR permits valid inference using likelihood-based methods or multiple imputation. MNAR requires specialized sensitivity analysis because standard methods assume the missingness mechanism is ignorable.

### Why Missing Data Is a Statistical Problem

The statistical consequences of missing data are not abstract concerns. They directly affect the conclusions that researchers draw from biological experiments. Consider a gene expression study where samples with low RNA quality fail sequencing quality control. If the failed samples come from a particular treatment group or tissue type, the remaining data will systematically underrepresent that condition. Any differential expression analysis will then reflect the missingness pattern instead of the true biology.

The same problem appears in longitudinal studies. If animals or patients drop out because of disease progression or adverse events, the remaining data overrepresents healthier individuals. A survival analysis or repeated measures model that ignores this pattern will produce optimistic estimates of treatment effects.

Missing data also affects the reproducibility of research. The National Library of Medicine's Research Methods Resources provides access to authoritative biomedical books and research method references that emphasize the importance of transparent reporting of analytical decisions. When researchers do not report how missing data was handled, other laboratories cannot reproduce the analysis or assess whether the conclusions are robust to alternative approaches.

### The Cost of Ignoring Missing Data

The most common approach to missing data is to ignore it. Many statistical software packages default to complete case analysis, which simply drops any observation with a missing value. This approach is acceptable only when the missingness is MCAR, which is rarely verifiable and rarely true in biological data.

The cost of ignoring missing data appears in several forms. The sample size shrinks, sometimes dramatically. A study with 10 percent missingness on each of five variables can lose more than 40 percent of its observations to complete case analysis. The remaining data is not a random subset of the original sample, so the analysis is biased. The standard errors are also incorrect because the software does not account for the uncertainty introduced by the missing values.

The cost is also scientific. A study that reports a null result when the true effect is real, or a significant result when the true effect is absent, contributes to the broader problem of unreliable research findings. The Committee on Publication Ethics Core Practices emphasizes the importance of data integrity and transparent reporting in the publication process. Researchers who do not address missing data are also making a statistical error but also a publication ethics concern.

## Missing Data Mechanisms

### Missing Completely at Random (MCAR)

MCAR is the strongest assumption. It means the probability of a value being missing is independent of both the observed and unobserved data. In practice, MCAR occurs when the missingness is caused by a random process that is unrelated to the study variables. A laboratory instrument that randomly fails to record a measurement, or a sample that is accidentally dropped during processing, produces MCAR data.

The key property of MCAR is that the observed data is a random sample of the full data. This means that complete case analysis produces unbiased estimates, although the standard errors are larger because the sample size is reduced. MCAR is the only mechanism where listwise deletion is a valid approach.

However, MCAR is rarely verifiable. The researcher cannot test whether the missingness is truly independent of the unobserved values because the unobserved values are, by definition, not available. Statistical tests can compare the observed values of other variables between the missing and nonmissing groups, but these tests cannot rule out dependence on the missing value itself.

### Missing at Random (MAR)

MAR is a less restrictive assumption. The probability of missingness depends on observed variables but not on the unobserved value itself. For example, in a longitudinal study, a patient may be more likely to miss a follow-up visit if they are older, and age is recorded. The missingness is MAR because the probability of missingness is explained by age, which is observed.

MAR is the most common assumption in applied missing data analysis. It is also the assumption that underlies multiple imputation and maximum likelihood methods. These methods use the observed variables to predict the missing values, and they produce valid estimates when the MAR assumption holds.

The practical implication is that the researcher must include in the imputation model all variables that predict missingness. If a variable that predicts missingness is omitted, the MAR assumption is violated and the imputation is biased. This is why the imputation model should include as many relevant variables as possible, including variables that are not part of the primary analysis.

### Missing Not at Random (MNAR)

MNAR is the most difficult case. The probability of missingness depends on the unobserved value itself, even after controlling for observed variables. For example, in a study of depression treatment, patients with the most severe depression may be the most likely to drop out. The missingness is MNAR because the probability of missingness depends on the depression score, which is the outcome of interest.

MNAR cannot be handled by multiple imputation or maximum likelihood methods because these methods assume MAR. The only way to address MNAR is through sensitivity analysis, which examines how the conclusions change under different assumptions about the missingness mechanism.

Sensitivity analysis is not a method for recovering the true values. It is a method for assessing the robustness of the conclusions. The researcher specifies a range of plausible missingness mechanisms and examines whether the conclusions hold across all of them. If the conclusions change under different assumptions, the results are not robust and should be interpreted with caution.

## Methods for Handling Missing Data

### Complete Case Analysis

Complete case analysis, also known as listwise deletion, removes any observation with a missing value on any variable in the analysis. This is the default in many statistical software packages and is the simplest approach.

The method is valid only when the missingness is MCAR. Under MCAR, the complete cases are a random subset of the full data, so the estimates are unbiased. The standard errors are larger because the sample size is reduced, but the estimates are correct.

Under MAR or MNAR, complete case analysis is biased. The bias can be substantial, especially when the missingness is related to the outcome. The method also wastes data, because observations with missing values on one variable are dropped from the entire analysis even if they have complete data on all other variables.

### Available Case Analysis

Available case analysis, also known as pairwise deletion, uses all available data for each calculation. The mean of each variable is calculated from all observations with that variable observed, and the correlation between two variables is calculated from all observations with both variables observed.

The method preserves more data than complete case analysis, but it has a serious problem. The sample size differs across the calculations, and the correlation matrix can be inconsistent. The resulting estimates can be biased, and the standard errors are not correct. The method is not recommended for most applications.

### Single Imputation

Single imputation replaces each missing value with a single estimated value. The most common approaches are mean imputation, regression imputation, and last observation carried forward.

Mean imputation replaces the missing value with the mean of the observed values. This method is simple but it reduces the variance of the variable, because the imputed values are all the same. The standard errors are underestimated, and the correlation with other variables is attenuated.

Regression imputation uses a regression model to predict the missing value from the observed variables. This method preserves the variance better than mean imputation, but it still underestimates the uncertainty because the imputed values are treated as known.

Single imputation methods are generally not recommended because they do not account for the uncertainty in the imputed values. The standard errors are too small, and the confidence intervals are too narrow.

### Maximum Likelihood

Maximum likelihood estimation uses all the observed data to estimate the parameters of the model. The method does not impute missing values. Instead, it estimates the parameters that maximize the likelihood of the observed data, integrating over the missing values.

The method is valid under the MAR assumption and is more efficient than complete case analysis. It also produces correct standard errors because it accounts for the uncertainty in the missing values.

The main limitation is that the method requires a specific model for the data. The researcher must specify the joint distribution of the variables, which is often a multivariate normal distribution. If the model is misspecified, the results can be biased.

### Multiple Imputation

Multiple imputation is the most flexible and widely recommended method for handling missing data. The method creates several complete datasets, each with the missing values filled in with different plausible values. The analysis is run on each dataset, and the results are combined to produce estimates that account for the uncertainty in the imputed values.

The method works in three steps. First, the imputation model is specified. This model predicts the missing values from the observed variables. Second, the imputation is repeated multiple times, typically 5 to 20 times, to create multiple complete datasets. Third, the analysis is run on each dataset, and the results are combined using Rubin's rules.

The key advantage of multiple imputation is that it accounts for the uncertainty in the imputed values. The standard errors are larger than those from single imputation, because the imputation uncertainty is included. The method is valid under the MAR assumption, and it can be extended to handle MNAR through sensitivity analysis.

The main limitation of multiple imputation is that the imputation model must be correctly specified. If the model omits important variables or uses the wrong distribution, the results are biased. The method also requires more computational effort than complete case analysis.

## Practical Workflow for Missing Data

### Step 1: Assess the Missingness Pattern

The first step is to describe the missing data. The researcher should calculate the proportion of missing values for each variable and examine the pattern of missingness across variables. The missingness pattern can be visualized with a matrix plot or a heatmap.

The researcher should also test whether the missingness is related to observed variables. This can be done by comparing the distribution of observed variables between the missing and nonmissing groups. For example, a t test can compare the mean of a continuous variable between the two groups, and a chi square test can compare the proportion of a categorical variable.

These tests are not definitive, because they cannot test the MAR assumption directly. However, they provide evidence about whether the missingness is related to observed variables, which is a necessary condition for the MAR assumption.

### Step 2: Choose the Analysis Method

The choice of method depends on the missingness mechanism and the analysis goal. If the missingness is MCAR and the proportion of missingness is small, complete case analysis may be acceptable. If the missingness is MAR, multiple imputation or maximum likelihood is appropriate. If the missingness is MNAR, sensitivity analysis is required.

The researcher should also consider the proportion of missingness. If the proportion is very small, the choice of method may not matter much. If the proportion is large, the method matters a great deal, and the researcher should use a method that accounts for the uncertainty.

### Step 3: Perform the Imputation

The imputation model should include all variables that are relevant to the analysis, including the outcome variable, the predictors, and any variables that predict the missingness. The model should also include any variables that are correlated with the missing values, because these variables help to predict the missing values.

The number of imputations should be at least 5, and more imputations are better when the proportion of missingness is high. The imputation should be performed with a software that is designed for this purpose, such as the mice package in R or the multiple imputation procedure in SPSS.

### Step4: Run the Analysis and Combine the Results

The analysis is run on each imputed dataset. The results are combined using Rubin's rules, which account for the between imputation variance and the within imputation variance. The combined estimates are the final results.

The combined results include the point estimates, the standard errors, and the confidence intervals. The researcher should report the combined results, not the results from a single imputed dataset.

### Step5: Perform Sensitivity Analysis

The sensitivity analysis examines how the conclusions change under different assumptions about the missingness mechanism. The researcher can specify a range of plausible missingness mechanisms and analyze the data under each assumption. If the conclusions are the same across all assumptions, the results are robust. If the conclusions change, the results are not robust, and the researcher should report the uncertainty.

The sensitivity analysis is particularly important when the missingness is MNAR. The researcher should specify a range of plausible missingness mechanisms and analyze the data under each assumption. The results should be reported for each assumption, and the conclusions should be interpreted in light of the range of results.

## At a Glance

| Missingness Mechanism | Definition | Valid Methods | Invalid Methods |
| --- | --- | --- | --- |
| MCAR | Missingness independent of observed and unobserved values | Complete case analysis, multiple imputation, maximum likelihood | None (all methods are valid) |
| MAR | Missingness depends on observed variables, not on the missing value | Multiple imputation, maximum likelihood | Complete case analysis, single imputation |
| MNAR | Missingness depends on the missing value itself | Sensitivity analysis | Multiple imputation, maximum likelihood, complete case analysis |

## Practical Implementation in R

### The mice Package

The mice package in R is the most widely used tool for multiple imputation. The package provides functions for creating the imputation model, performing the imputation, and combining the results.

The basic workflow is to create the imputation model with the mice function, perform the imputation with the mice function, and combine the results with the pool function. The imputation model is specified with a formula that includes all variables that are used in the analysis.

The mice package also provides diagnostic tools for checking the imputation. The researcher can plot the imputed values against the observed values to check whether the imputed values are plausible. The researcher can also compare the distribution of the imputed values with the distribution of the observed values.

### Example Code

The following code shows a basic multiple imputation workflow in R. The code assumes that the data is in a data frame called data, and the analysis is a linear regression of y on x1 and x2.

```r
library(mice)

## Create the imputation model
imputed_data <- mice(data, m = 5, method = "pmm", seed = 123)

## Run the analysis on each imputed dataset
analysis <- with(imputed_data, lm(y ~ x1 + x2))

## Combine the results
combined_results <- pool(analysis)

## Print the summary
summary(combined_results)
```

The code creates 5 imputed datasets, runs a linear regression on each, and combines the results. The summary shows the combined estimates, the standard errors, and the confidence intervals.

### The SPSS Procedure

SPSS provides a multiple imputation procedure that is similar to the mice package in R. The procedure is accessed through the Analyze menu, then Multiple Imputation, then Impute Missing Data Values.

The procedure requires the researcher to specify the variables that are used in the imputation model. The procedure also requires the researcher to specify the number of imputations and the method for the imputation.

The output of the procedure includes the imputed datasets and the combined results. The combined results are reported in the output, and the researcher can use the results in the analysis.

## Practical Implementation in SPSS

### The Multiple Imputation Procedure

The multiple imputation procedure in SPSS is accessed through the Analyze menu, then Multiple Imputation, then Impute Missing Data. The procedure creates a new dataset with the imputed values, and the analysis is run on the imputed dataset.

The procedure requires the researcher to specify the variables that are used in the imputation model. The researcher can also specify the number of imputed values and the method for the imputation.

The output of the procedure includes the imputed datasets and the combined results. The combined results are available in the analysis, and the researcher can use the combined results in the analysis.

### The Analysis of the Imputed Data

The analysis of the imputed data is run on the imputed dataset. The analysis is the same as the analysis of the complete data, but the results are combined across the imputed datasets.

The combined results are available in the output of the analysis. The combined results include the point estimates, the standard errors, and the confidence intervals. The researcher should report the combined results, not the results from a single imputed dataset.

## Common Failure Patterns

### Ignoring the Missingness Mechanism

The most common failure is to ignore the missingness mechanism and use complete case analysis without checking the assumption. This approach is valid only when the missingness is MCAR, which is rare in practice. The result is biased estimates and incorrect standard errors.

The researcher should always examine the missingness pattern and test whether the missingness is related to observed variables. If the missingness is related to observed variables, the MAR assumption is more plausible, and multiple imputation or maximum likelihood is appropriate.

### Using Single Imputation

The second common failure is to use single imputation, such as mean imputation or last observation carried forward. These methods do not account for the uncertainty in the imputed values, and the standard errors are too small. The result is confidence intervals that are too narrow and p values that are too small.

The researcher should use multiple imputation instead of single imputation. Multiple imputation accounts for the uncertainty in the imputed values, and the standard errors are correct.

### Omitting Important Variables from the Imputation Model

The third common failure is to omit important variables from the imputation model. The imputation model should include all variables that are used in the analysis, including the variables that predict the missingness. If a variable that predicts the missingness is omitted, the imputation is biased.

The researcher should include all variables that are relevant to the analysis in the imputation model. The imputation model should also include variables that are not used in the analysis but predict the missingness.

### Not Performing Sensitivity Analysis

The fourth common failure is to not perform sensitivity analysis. The sensitivity analysis is required when the missingness is MNAR, but it is also useful when the missingness is MAR. The sensitivity analysis examines how the conclusions change under different assumptions about the missingness mechanism.

The researcher should perform a sensitivity analysis for all analyses with missing data. The sensitivity analysis should be reported in the results, and the conclusions should be interpreted with respect to the sensitivity.

## Records and Measurements

### Documenting the Missingness

The researcher should document the missingness pattern in the study. The documentation should include the proportion of missingness for each variable, the pattern of missingness across variables, and the results of the tests for the missingness mechanism.

The documentation should be included in the methods section of the report. The documentation should be detailed enough that the reader can assess the validity of the analysis.

### Reporting the Imputation

The researcher should report the imputation in the methods section of the report. The report should include the number of imputed datasets, the method of imputation, and the variables that are included in the imputation model.

The report should also include the results of the sensitivity analysis. The sensitivity analysis should be reported in the results section, and the conclusions should be interpreted with respect to the sensitivity.

### The Data Management and Sharing Policy

The National Institutes of Health Data Management and Sharing Policy requires that researchers plan for the management and sharing of data. The policy applies to all research that is funded by the NIH, and it requires that the data management plan include the handling of missing data.

The data management plan should describe how the missing data is handled in the analysis. The plan should also describe how the missing data is documented and reported. The plan should be included in the grant application, and the plan is reviewed as part of the grant review process.

## Limitations and Interpretation

### The Limits of Multiple Imputation

Multiple imputation is a powerful method, but it has limitations. The method is valid under the MAR assumption, and the imputation model must be correctly specified. If the imputation model is misspecified, the results are biased.

The method also requires the computational effort, and the results are not always reproducible. The imputation is based on a random seed, and the results can vary across different seeds. The researcher should set the seed to ensure that the results are reproducible.

### The Limits of Sensitivity Analysis

Sensitivity analysis is a method for assessing the robustness of the results, but it is not a method for fixing the missingness. The sensitivity analysis can not determine the true missingness mechanism, and the results are only as good as the assumptions that are made.

The researcher should be cautious in the interpretation of the sensitivity analysis. The sensitivity analysis can not prove that the results are correct, and the results are only as valid as the assumptions that are made.

### The Ethical Considerations

The handling of missing data has ethical implications. The researcher has a responsibility to report the missingness and the methods that are used to handle it. The researcher also has a responsibility to not misrepresent the results.

The Committee on Publication Ethics Core Practices emphasizes the importance of data integrity and transparent reporting. The researcher should report the missingness and the analysis in a transparent way, and the researcher should not misrepresent the results.

The researcher should also consider the impact of the missingness on the conclusions. The researcher should not overstate the conclusions when the missingness is high, and the researcher should be cautious in the interpretation of the results.

## Building a Missing Data Decision Log for Longitudinal Biological Studies

Longitudinal biological studies present a distinct missing data problem that cross sectional analyses do not. When the same animal, plant, plot, or patient is measured repeatedly, missingness at one time point affects the interpretation of all subsequent time points. A researcher who handles missing data correctly for a single measurement may still introduce bias by ignoring the temporal structure of the missingness. This section provides a practical decision framework for longitudinal data, a record system for tracking missingness decisions, and a troubleshooting method for common failures in repeated measures designs.

### The Temporal Missingness Framework

The standard MCAR, MAR, and MNAR classification applies to longitudinal data, but the mechanism must be evaluated at each time point and in relation to previous measurements. A dropout that depends on the value of the previous measurement is MAR if that previous value is observed. A dropout that depends on the value of the measurement that would have been taken at the dropout time is MNAR. The distinction is not academic. It determines whether multiple imputation can recover the missing values or whether the analysis must rely on sensitivity analysis.

Consider a plant growth experiment where leaf area is measured weekly. A plant that dies between week 3 and week 4 has a missing week 4 measurement. If the probability of death depends on the week 3 leaf area, which is observed, the missingness is MAR. Multiple imputation that includes week 3 leaf area as a predictor can produce valid estimates. If the probability of death depends on the week 4 leaf area that would have been measured, the missingness is MNAR. No imputation model can recover that value because the value itself is the cause of the missingness.

The decision framework requires the researcher to answer three questions for each missing value. First, is the missingness related to any observed variable? Second, is the missingness related to the missing value itself? Third, does the missingness pattern change across time? The answers determine whether the researcher can use multiple imputation, whether the imputation model must include time varying covariates, and whether sensitivity analysis is mandatory.

### The Missingness Pattern Matrix

A practical tool for longitudinal studies is the missingness pattern matrix. This matrix records, for each subject and each time point, whether the measurement is present or absent. The matrix is not a statistical test. It is a visual and tabular record that allows the researcher to see the structure of the missingness before choosing an analysis method.

The matrix should include the subject identifier, the time point, the missingness indicator, and the values of any time invariant covariates. The researcher should also record the reason for the missingness when the reason is known. Reasons include instrument failure, sample degradation, subject withdrawal, death, and recording error. The reason is not always known, but when it is known, it provides evidence about the missingness mechanism.

The pattern log serves three purposes. First, it forces the researcher to examine the missingness before running the analysis. Second, it provides the documentation that is needed for the methods section of the report. Third, it allows the researcher to identify whether the missingness is monotone or nonmonotone. Monotone missingness means that once a subject has a missing value, all subsequent values are also missing. This pattern is common in survival studies and longitudinal studies with dropout. Nonmonotone missingness means that a subject can have a missing value at one time and an observed value at a later time. This pattern is common in instrument failure and sample processing errors.

The distinction matters because the imputation method can be simpler for monotone missingness. The imputation model can be built sequentially, using the observed values at earlier time points to predict the missing values at later time points. For nonmonotone missingness, the imputation model must account for the possibility that a subject returns to the study after a missing visit.

### The Decision Framework

The decision framework is a sequence of questions that the researcher answers before choosing an analysis method. The framework is designed to prevent the common failure of applying a single method to all missing data without examining the structure.

The first question is whether the missingness is monotone or nonmonotone. If the missingness is monotone, the researcher can use a sequential imputation approach. If the missingness is nonmonotone, the researcher must use a joint imputation approach that models all variables simultaneously.

The second question is whether the missingness is related to the outcome. If the missingness is related to the outcome, the imputation model must include the outcome variable. If the outcome is the variable with missing values, the imputation model must include the predictors of the outcome and the variables that predict the missingness.

The third question is whether the missingness is related to time. If the missingness is more likely at later time points, the imputation model must include the time variable. If the missingness is related to a time varying covariate, the imputation model must include that covariate.

The fourth question is whether the missingness is related to the missing value itself. If the missingness is MNAR, the researcher cannot use multiple imputation as the primary analysis. The researcher must use sensitivity analysis to assess the robustness of the conclusions.

The framework is not a substitute for statistical expertise. It is a decision aid that ensures the researcher has considered the structure of the missingness before choosing a method. The framework is also a record that can be included in the methods section of the report.

### The Record System

The record system for missing data decisions should be maintained throughout the study, also at the analysis stage. The record should include the following items for each variable and each time point.

The first item is the proportion of missing values. This is calculated as the number of missing values divided by the total number of expected values. The proportion should be calculated for each variable and for each time point.

The second item is the pattern of missingness. This is the description of whether the missingness is monotone or nonmonotone, and whether the missingness is concentrated at specific time points.

The third item is the reason for the missingness. This is recorded when the reason is known. The reason can be instrument failure, sample degradation, animal death, withdrawal, or recording error.

The fourth item is the result of the test for the missingness mechanism. The test compares the distribution of observed variables between the missing and nonmissing groups. The test does not prove the mechanism, but it provides evidence about whether the missingness is related to observed variables.

The fifth item is the analysis method that was chosen. The method should be chosen based on the missingness mechanism and the pattern of missingness.

The sixth item is the sensitivity analysis that was performed. The sensitivity analysis should be described in the record, including the range of assumptions that were examined.

The record system is not a statistical analysis. It is a documentation tool that ensures the researcher has examined the missingness and has made a deliberate decision. The record is also the basis for the reporting in the methods section.

### Troubleshooting Common Longitudinal Failures

The first common failure is treating all missing values as the same. A missing value at the first time point is different from a missing value at the last time point. The first time point missingness may be caused by the recruitment process, while the last time point missingness may be caused by dropout. The researcher should examine the missingness pattern by time point before choosing a method.

The second common failure is using the same imputation model for all time points. The imputation model for the first time point may not include the same predictors as the imputation model for the last time point. The researcher should specify the imputation model separately for each time point, or use a joint model that allows the predictors to vary across time.

The third common failure is ignoring the correlation between time points. The imputation model must account for the correlation between the repeated measurements. If the imputation model treats the time points as independent, the imputed values will not reflect the correlation structure of the data.

The fourth common failure is not checking the convergence of the imputation. The imputation algorithm should be run for enough iterations to ensure that the imputed values are stable. The researcher should check the convergence by examining the trace plots of the imputed values.

The fifth common failure is not performing the sensitivity analysis. The sensitivity analysis is required when the missingness is MNAR, but it is also useful when the missingness is MAR. The sensitivity analysis examines how the conclusions change under different assumptions about the missingness mechanism.

### The Role of the Data Management Plan

The National Institutes of Health Data Management and Sharing Policy requires that researchers plan for the management and sharing of data. The policy applies to all research that is funded by the NIH, and it requires that the data management plan include the handling of missing data. The plan should describe how the missing data is handled in the analysis, and the plan should be included in the grant application.

The data management plan should also describe the record system for the missing data. The plan should specify the variables that are tracked, the reasons for the missingness that are recorded, and the methods that are used to handle the missingness. The plan should be updated as the study progresses, and the plan should be reviewed as part of the grant review process.

The data management plan is not a substitute for the analysis. The plan is a planning document that ensures the researcher has considered the missing data before the data is collected. The plan is also a record that can be used to assess the validity of the analysis.

### Reporting the Missing Data Decisions

The reporting of the missing data decisions should follow the guidelines for transparent research reporting. The EQUATOR Network provides a collection of reporting guidelines for different study types. The researcher should select the appropriate guideline for the study design and follow the recommendations for reporting the missing data.

The report should include the proportion of missing values for each variable, the pattern of missingness, the reason for the missingness when known, the method that was used to handle the missingness, and the sensitivity analysis. The report should also include the record of the missingness decisions, so that the reader can assess the validity of the analysis.

The report should not simply state that multiple imputation was used. The report should describe the imputation model, the number of imputed datasets, and the variables that were included in the model. The report should also describe the sensitivity analysis and the range of assumptions that were examined.

The reporting is not a formality. The reporting is the mechanism by which the reader can assess the validity of the analysis. The reader cannot assess the validity of the analysis if the missingness is not reported. The reporting is also the mechanism by which the research can be reproduced. The reader cannot reproduce the analysis if the missingness decisions are not reported.

### The Decision Framework in Practice

The decision framework is applied at the analysis stage, but the record system is maintained throughout the study. The researcher should begin the record at the start of the data collection, and the record should be updated as the missingness occurs. The record should be reviewed at the analysis stage, and the decisions should be documented in the report.

The decision framework is not a substitute for the statistical analysis. The framework is a decision aid that helps the researcher to choose the appropriate method. The framework is also a record that can be used to assess the validity of the analysis.

The researcher should also consider the limitations of the decision framework. The framework cannot determine the true missingness mechanism. The framework can only help the researcher to make a deliberate decision based on the available evidence. The researcher should be cautious in the interpretation of the results when the missingness is high, and the researcher should report the uncertainty in the results.

## Frequently Asked Questions

### What is the difference between MCAR, MAR, and MNAR?

MCAR means the missingness is independent of the data, MAR means the missingness depends on observed variables, and MNAR means the missingness depends on the missing value itself. The distinction is important because it determines the valid methods for the analysis.

### When is complete case analysis acceptable?

Complete case analysis is acceptable when the missingness is MCAR and the proportion of missingness is small. The method is not acceptable when the missingness is MAR or MNAR, because the estimates are biased.

### What is the difference between single imputation and multiple imputation?

Single imputation replaces each missing value with a single estimated value, while multiple imputation replaces each missing value with multiple estimated values. Multiple imputation accounts for the uncertainty in the imputed values, while single imputation does not.

### How many imputed datasets are needed?

The number of imputed datasets should be at least 5, and more imputed datasets are better when the proportion of missingness is high. The number of imputed datasets is a tradeoff between the computational effort and the accuracy of the results.

### What is sensitivity analysis?

Sensitivity analysis is a method for assessing the robustness of the results under different assumptions about the missingness mechanism. The analysis is performed by specifying a range of plausible missingness mechanisms and analyzing the results under each assumption.

### How should the missingness be reported in the analysis?

The missingness should be reported in the methods section of the report. The report should include the proportion of missingness for each variable, the pattern of missingness, and the methods that are used to handle the missingness.

### What is the role of the data management plan in missing data?

The data management plan should describe how the missing data is handled in the analysis. The plan should be included in the grant application, and the plan is reviewed as part of the grant review process.

### What are the ethical considerations in missing data?

The researcher has a responsibility to report the missingness and the analysis that is used to handle it. The researcher also has a responsibility to not misrepresent the results, and the researcher should be cautious in the interpretation of the results when the missingness is high.

## Using the Evidence

| Source | Best use in this topic | Important limitation |
|---|---|---|
| [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) | official guidance | Check the linked page for current local requirements |
| [EQUATOR Network](https://www.equator-network.org/) | official guidance | Check the linked page for current local requirements |
| [Core Practices](https://publicationethics.org/core-practices) | official guidance | Check the linked page for current local requirements |

## Related Bioinformatics Guides

- [Spatial Transcriptomics Data Analysis: A Practical Workflow from Raw Data to Biological Insights](/knowledge/bioinformatics/spatial-transcriptomics-data-analysis-a-practical-workflow-from-raw-data-to-biological-insights)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)
- [Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight](/knowledge/bioinformatics/metabolomics-data-analysis-workflow-from-raw-data-to-biological-insight)
- [Metagenomics Data Analysis: From Raw Reads to Biological Insights](/knowledge/bioinformatics/metagenomics-data-analysis-from-raw-reads-to-biological-insights)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Missing Data in Trial-Based Cost-Effectiveness Analysis: The Journey Continues.](https://doi.org/10.1007/s40273-026-01651-y). 2026.
- [A methodological review finds that the statistical analysis of most comparative diagnostic accuracy studies had shortcomings.](https://doi.org/10.1016/j.jclinepi.2026.112463). 2026.
- [Missing data handling in pediatric appendectomy research: current practices and National Surgical Quality Improvement Program (NSQIP) analysis.](https://doi.org/10.1007/s00464-026-13123-7). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.