# Survival Analysis with Missing Covariate Data


## Key Takeaways

-   Apply multiple imputation with chained equations (MICE) to handle missing covariate data in Cox proportional hazards models, generating 20 to 100 imputed datasets before pooling estimates using Rubin's rules.
-   Crucially, include the event indicator and the cumulative baseline hazard (estimated from complete cases or a preliminary Cox model) as predictors within the imputation model to preserve the association between covariates and survival outcomes.
-   Complete-case analysis, which discards observations with any missing covariate, can introduce significant bias in hazard ratio estimates if missingness is not completely random (MCAR).
-   The imputation model must incorporate all covariates present in the final Cox analysis model, including any interaction terms, to prevent bias in the estimated coefficients.
-   Avoid imputing survival time directly as a linear predictor; instead, use the event indicator and cumulative baseline hazard to accurately reflect the survival outcome's information within the imputation process.

---

## Quick Answer

- Apply multiple imputation with chained equations to handle missing covariate values in Cox proportional hazards models, creating 20 to 100 imputed datasets before pooling estimates.
- Include the event indicator and cumulative baseline hazard in the imputation model to preserve the relationship between covariates and survival outcomes.
- Missing covariate data that is not missing completely at random requires careful modeling of the missingness mechanism, and complete-case analysis may produce biased hazard ratios.

## Understanding Missing Covariate Data in Survival Analysis

Survival analysis examines the time until an event of interest occurs, such as disease recurrence, death, or recovery. The Cox proportional hazards model is the most widely used regression approach in this setting because it estimates the association between covariates and the hazard of the event without requiring specification of the baseline hazard function. Researchers in biology and the life sciences routinely apply this model to data from cohort studies, clinical trials, and laboratory experiments where follow-up times and event indicators are recorded for each subject.

Missing covariate values are a common problem in these datasets. A laboratory may fail to record a biomarker for some samples, a questionnaire item may be left blank, or a measurement instrument may malfunction for a subset of observations. When covariates are missing, the analyst must decide how to proceed. Deleting observations with any missing value, a method called complete-case analysis, reduces the sample size and can introduce bias if the missingness is related to the outcome or to other covariates. The alternative is to impute the missing values using statistical methods that account for the uncertainty of the imputation.

The central challenge in survival analysis with missing covariates is that the imputation procedure must respect the structure of survival data. The event time and the censoring indicator carry information about the outcome, and this information must be incorporated into the imputation model. A naive imputation that ignores the survival outcome will attenuate the association between the covariate and the hazard, producing hazard ratio estimates that are biased toward the null.

The practical goal is to obtain valid estimates of the hazard ratios for the covariates of interest, with confidence intervals that reflect the uncertainty from both the sampling process and the missing data. Multiple imputation is the standard approach for achieving this goal, and it is implemented in major statistical software packages including R, Stata, and SAS.

## At a Glance

The table below summarizes the main approaches to handling missing covariate data in Cox models, the conditions under which each approach is appropriate, and the practical consequences of each choice.

| Method | When to Use | Practical Consequence |
| --- | --- | --- |
| Complete-case analysis | Missingness is completely random and the proportion of missing data is small | Reduced sample size, potential bias if missingness depends on outcome or covariates |
| Single imputation (mean, median, or last observation carried forward) | Exploratory analysis only | Underestimates variance, produces confidence intervals that are too narrow |
| Multiple imputation with chained equations | Standard approach for missing covariates in Cox models | Requires careful specification of the imputation model, including survival outcome and censoring |
| Inverse probability weighting | Missingness depends on observed covariates | Requires correct specification of the missingness model, can be inefficient |

## The Structure of Survival Data and Its Implications for Imputation

Survival data consist of three components for each observation: the follow-up time, the event indicator, and the covariates. The follow-up time is the time from the start of observation to the event or to the last known follow-up. The event indicator records whether the event occurred or whether the observation was censored. Censoring occurs when the event has not happened by the end of the study, when the participant withdraws, or when the participant is lost to follow-up.

The Cox model specifies the hazard function as the product of a baseline hazard and an exponential function of the linear predictor. The linear predictor is a weighted sum of the covariates, and the weights are the regression coefficients that are estimated from the data. The key property of the Cox model is that the baseline hazard cancels out of the partial likelihood, so the regression coefficients can be estimated without specifying the shape of the baseline hazard.

When a covariate is missing, the partial likelihood cannot be computed for the affected observations unless the missing value is filled in. The imputation model must therefore generate plausible values for the missing covariate, and these values must be consistent with the observed survival data. The survival outcome is informative about the missing covariate because the hazard of the event depends on the covariate. If the covariate is associated with the hazard, then observations with longer follow-up times or with the event observed are more likely to have certain covariate values.

The censoring indicator is also informative. A subject who is censored at a particular time has survived up to that time without the event. The imputation model must therefore include both the event indicator and the cumulative baseline hazard, which is a function of the follow-up time. The cumulative baseline hazard can be estimated from the complete-case data or from a model that includes the observed covariates.

## Multiple Imputation for Cox Models

Multiple imputation is a three-step procedure. In the first step, the missing values are filled in multiple times to create several complete datasets. Each imputation draws from the predictive distribution of the missing values given the observed data. In the second step, the Cox model is fitted to each imputed dataset separately, producing a set of regression coefficients and standard errors. In the third step, the estimates are combined using Rubin's rules, which account for the between-imputation variability and the within-imputation variability.

The number of imputations is a practical decision. The statistical literature has traditionally recommended five to ten imputations, but more recent guidance suggests that 20 to 40 imputations are needed to obtain stable estimates of the standard errors, especially when the fraction of missing data is high. The computational cost of additional imputations is modest for most survival datasets, and the benefit is more reliable confidence intervals.

The imputation model must be compatible with the analysis model. For a Cox model with a missing covariate, the imputation model should include the event indicator and the cumulative baseline hazard as predictors. The cumulative baseline hazard can be estimated from the Nelson-Aalen estimator or from a Cox model fitted on the complete-case data. The imputation model should also include all other covariates that are in the analysis model, because omitting a covariate from the imputation model can induce bias in the coefficient for the missing covariate.

The choice of imputation method depends on the type of the missing covariate. For a continuous covariate, linear regression is the default. For a binary covariate, logistic regression is used. For a categorical covariate with more than two levels, a multinomial or ordinal logistic model is used. For a count covariate, Poisson regression is used. The imputation is performed sequentially, with each variable imputed in turn, and the process is repeated for several cycles to allow the imputed values to converge.

## Including the Survival Outcome in the Imputation Model

The most common error in imputing missing covariates for survival data is to impute the covariates using only the other covariates, without including the survival outcome. This approach is equivalent to assuming that the missingness is independent of the survival time and the event indicator, which is rarely true. When the covariate is associated with the hazard, the imputed values will be too similar to the observed values, and the hazard ratio will be attenuated.

The correct approach is to include the survival outcome in the imputation model. The outcome can be represented by the event indicator and the cumulative baseline hazard, or by the event indicator and the survival time. The cumulative baseline hazard is preferred because it captures the information about the survival time in a way that is compatible with the Cox model. The cumulative baseline hazard can be estimated from the complete-case data, and it is included as a predictor in the imputation model.

The event indicator is included as a binary predictor. The survival time is not included directly, because the distribution of the survival time is not linear in the covariates. Instead, the cumulative baseline hazard is used, because it is a function of the survival time that is linear in the Cox model. This approach is known as the event-based imputation, and it is implemented in the R package mice with the event and the cumulative hazard as predictors.

The imputation model should also include any interactions that are in the analysis model. If the analysis model includes an interaction between a covariate and a treatment group, then the imputation model should include the interaction term. Otherwise, the imputed values will not preserve the relationship between the covariates, and the interaction coefficient will be biased.

## Handling Censoring in the Imputation

Censoring is a special feature of survival data that must be handled carefully in the imputation. The censoring indicator is not a covariate, but it is informative about the survival time. A subject who is censored at time t has survived to time t, and this information should be used in the imputation.

The cumulative baseline hazard at the censoring time is the appropriate summary of the censoring information. The cumulative hazard is the integral of the baseline hazard from time zero to the censoring time, and it is a measure of the risk that the subject has accumulated. Including the cumulative hazard in the imputation model allows the imputed values to reflect the fact that a censored subject has survived to the censoring time.

The event indicator is also included in the imputation model. A subject who has experienced the event has a different covariate distribution than a subject who has not, and the imputation model should reflect this difference. The event indicator is a binary variable, and it is included as a predictor in the imputation model.

The imputation model should not include the survival time directly as a linear predictor, because the relationship between the survival time and the covariates is not linear. The cumulative hazard is the appropriate summary of the survival time for the Cox model, and it is the standard approach in the statistical literature.

## Practical Workflow for Imputing Missing Covariates

The practical workflow for imputing missing covariate data in a Cox model follows a sequence of steps that can be implemented in R, Stata, or SAS. The steps are described below in the context of the R software, but the same logic applies to the other packages.

The first step is to examine the missing data pattern. The researcher should identify which covariates have missing values, the proportion of missing values, and the pattern of missingness. The missingness pattern can be visualized with a matrix plot, and the proportion of missing values can be summarized in a table.

The second step is to decide whether the missingness is ignorable. The missingness is ignorable if the probability of missingness depends only on the observed data, not on the unobserved values. This is the missing at random assumption. If the missingness depends on the unobserved values, the missingness is not ignorable, and the imputation model must be extended to include the missingness model.

The third step is to estimate the cumulative baseline hazard from the complete-case data. This is done by fitting a Cox model with the observed covariates and extracting the baseline hazard. The cumulative hazard is then computed for each subject at the observed survival time.

The fourth step is to specify the imputation model. The imputation model includes the covariates with missing values, the covariates with complete values, the event indicator, and the cumulative hazard. The imputation model is specified using the mice package in R, and the method for each variable is chosen based on the variable type.

The fifth step is to generate the imputed datasets. The number of imputations is set to 20 or more, and the imputation is run. The imputed datasets are stored as a list of complete datasets.

The sixth step is to fit the Cox model to each imputed dataset. The Cox model is fitted using the `coxph` function, and the model specification is the same for each imputed dataset.

The seventh step is to pool the estimates. The pooling is done using the `pool` function, which applies the Rubin rules to combine the coefficient estimates and the standard errors. The pooled estimates are the final results, and the confidence intervals are computed from the pooled standard errors.

The eighth step is to check the convergence of the imputation. The imputation is checked by examining the trace plots of the imputed values, and the convergence is assessed by the stability of the imputed values across the cycles.

## Practical Implementation Steps

The following steps provide a practical guide for the researcher who is imputing missing covariate data in a Cox model.

1.  Inspect the data. Create a table of the missing values for each covariate, and a table of the missingness pattern. The proportion of missingness for each covariate should be recorded.

2.  Assess the missingness mechanism. Consider whether the missingness is related to the observed covariates, the survival time, or the event indicator. If the missingness is related to the survival time or the event, the imputation model must include the survival outcome.

3.  Estimate the cumulative baseline hazard. Fit a Cox model to the complete-case data, and extract the cumulative hazard for each observation. The cumulative hazard is the Nelson-Aalen estimate or the Breslow estimate.

4.  Specify the imputation model. Include all covariates in the analysis model, the event indicator, and the cumulative hazard. Include any interactions that are in the analysis model.

5.  Run the imputation. Use the `mice` function in R, or the equivalent in Stata or SAS. Set the number of the imputations to 20 or more.

6.  Fit the Cox model to each imputed dataset. Use the same model specification for each dataset.

7.  Pool the estimates. Use the `pool` function to combine the estimates and the standard errors. Report the pooled hazard ratios and the confidence intervals.

8.  Check the sensitivity. Repeat the imputation with a different number of the imputations, and with a different imputation model, to check the sensitivity of the results.

## Records and Measurements

The imputation procedure should be documented in the analysis records. The documentation should include the following the items.

The proportion of missingness for each covariate should be recorded. This is the number of the missing values divided by the number of the observations. The proportion of missingness is a key measure of the potential for bias.

The pattern of missingness should be recorded. The pattern is the combination of the variables that are missing for each observation. The pattern can be monotone or arbitrary. The monotone pattern is the pattern where the missingness is nested, and the arbitrary pattern is the pattern where the missingness is not nested.

The imputation model should be recorded. The imputation model includes the predictors, the method for each variable, and the number of the cycles. The imputation model should be described in the methods section of the report.

The number of the imputations should be recorded. The number of the imputations is the number of the complete datasets that are generated. The number of the imputations should be reported in the methods.

The pooled estimates should be recorded. The pooled estimates include the hazard ratios, the confidence intervals, and the p-values. The pooled estimates should be reported in the results.

The convergence of the imputation should be recorded. The convergence is assessed by the trace plots of the imputed values. The convergence should be reported in the methods.

## Common Failure Patterns

The following are the common failure patterns in the imputation of the missing covariate data in the survival models.

The first failure pattern is the imputation without the survival outcome. This is the most common error. The imputation model includes only the covariates, and the survival time and the event indicator are ignored. The result is the attenuation of the hazard ratios, and the confidence intervals that are too narrow.

The second failure pattern is the imputation with the survival time as a linear predictor. The survival time is included as a linear predictor in the imputation model, but the relationship between the survival time and the covariates is not linear. The result is the biased imputed values, and the hazard ratios that are not valid.

The third failure pattern is the imputation with the event indicator but without the cumulative hazard. The event indicator is included, but the cumulative hazard is not. The result is the imputed values that do not reflect the survival time, and the hazard ratios that are biased.

The fourth failure pattern is the imputation with the interactions omitted. The analysis model includes an interaction, but the imputation model does not. The result is the interaction coefficient that is biased, and the main effects that are not valid.

The fifth failure pattern is the imputation with the missingness that is not ignorable. The missingness depends on the unobserved values, and the imputation model does not account for this. The result is the imputed values that are biased, and the hazard ratios that are not valid.

The sixth failure pattern is the use of the single imputation. The missing values are replaced by a single value, and the uncertainty of the imputation is not accounted for. The result is the confidence intervals that are too narrow, and the p-values that are too small.

## Limitations and Assumptions

The imputation of the missing covariate data in the survival models relies on the assumptions that should be stated in the report.

The first assumption is the missingness at random. The probability of the missingness depends only on the observed data, not on the unobserved values. This assumption is not testable from the data, and the sensitivity analysis should be conducted to assess the impact of the violation.

The second assumption is the correct specification of the imputation model. The imputation model includes the correct predictors, the correct functional form, and the correct distribution for each variable. The misspecification of the imputation model can induce the bias in the hazard ratios.

The third assumption is the correct specification of the analysis model. The Cox model is correctly specified, with the correct covariates, the correct functional form, and the correct handling of the censoring. The misspecification of the analysis model can induce the bias in the hazard ratios.

The fourth assumption is the proportional hazards. The hazard ratio is constant over the time. The violation of the proportional hazards can induce the bias in the hazard ratios, and the imputation model should be checked for the proportional hazards.

The fifth assumption is the non-informative censoring. The censoring is independent of the event time, given the covariates. The violation of the non-informative censoring can induce the bias in the hazard ratios, and the imputation model should be checked for the censoring.

The limitations of the imputation should be stated in the report. The imputation does not recover the information that is not in the data. The imputation is based on the assumptions, and the assumptions should be stated. The imputation is not a substitute for the collection of the complete data.

## Sensitivity Analysis

The sensitivity analysis is the assessment of the impact of the assumptions on the results. The sensitivity analysis should be conducted to assess the impact of the missingness assumption, the imputation model, and the number of the imputations.

The first sensitivity analysis is the comparison of the complete-case analysis with the imputation analysis. The complete-case analysis is the analysis of the observations with the complete data. The imputation analysis is the analysis of the imputed data. The comparison of the two analyses shows the impact of the missingness on the results.

The second sensitivity analysis is the comparison of the imputation with the different number of the imputations. The imputation is run with the 10, 20, and 40 imputations, and the results are compared. The comparison shows the impact of the number of the imputations on the results.

The third sensitivity analysis is the comparison of the imputation with the different imputation models. The imputation model is the specified with the different predictors, and the results are compared. The comparison shows the impact of the imputation model on the results.

The fourth sensitivity analysis is the comparison of the imputation with the different missingness assumptions. The missingness is the assumed to be at random, and the sensitivity analysis is the conducted with the missingness that is not at random. The comparison shows the impact of the missingness assumption on the results.

The sensitivity analysis should be reported in the report. The sensitivity analysis should be described in the methods, and the results of the sensitivity analysis should be reported in the results.

## Reporting the Imputation in the Manuscript

The imputation of the missing covariate data should be reported in the manuscript according to the reporting guidelines. The reporting guidelines are the set of the recommendations for the reporting of the research, and the guidelines are the available from the EQUATOR Network. The EQUATOR Network is the repository of the reporting guidelines, and the guidelines are the selected based on the study design.

The reporting of the imputation should include the following items.

The proportion of the missingness for each covariate should be reported. The proportion of the missingness is the number of the missing values divided by the total number of the observations.

The pattern of the missingness should be reported. The pattern of the missingness is the set of the covariates that are the missing for each observation.

The imputation model should be reported. The imputation model includes the predictors, the method for each variable, and the number of the iterations.

The number of the imputations should be reported. The number of the imputations is the number of the complete datasets that are the generated.

The pooling method should be reported. The pooling method is the Rubin rules, and the pooling method is the described in the methods.

The sensitivity analysis should be reported. The sensitivity analysis is the assessment of the impact of the assumptions on the results.

The reporting of the imputation is the important for the transparency of the research. The reporting of the imputation allows the reader to assess the validity of the results, and the reporting is the required by the reporting guidelines.

## The Role of the Data Management and Sharing

The data management and the sharing are the important for the reproducibility of the research. The data management is the process of the collection, the storage, and the documentation of the data. The data sharing is the process of the making the data available to the other researchers.

The [NIH Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy) describes the expectations for the data management and the sharing for the NIH-funded research. The policy requires the data management and the sharing plan for the research that is the funded by the NIH. The plan describes the data that will be the collected, the data that will be the shared, and the data that will be the preserved.

The data management and the sharing are the relevant for the imputation of the missing covariate data. The imputation is the based on the data, and the imputation is the reproducible if the data is the shared. The data sharing is the important for the transparency of the research, and the data sharing is the recommended for the research.

The data management of the imputation should be the documented. The data management includes the description of the data, the description of the imputation, and the description of the analysis. The data management is the documented in the data management of the plan.

The data sharing of the imputation should be the considered. The data sharing includes the sharing of the imputed data, the sharing of the imputation code, and the sharing of the analysis code. The data sharing is the important for the reproducibility of the research.

## The Role of the Researcher Identity and the Publication Ethics

The researcher identity and the publication ethics are the relevant for the reporting of the imputation. The researcher identity is the identification of the researcher, and the publication is the process of the reporting of the research.

The [ORCID for the Researchers](https://info.orcid.org/researchers) describes the researcher identity. The ORCID is the identifier for the researcher, and the ORCID is the used to the link the researcher to the research. The ORCID is the recommended for the researcher, and the ORCID is the used to the identify the researcher in the publication.

The [Core Practices](https://publicationethics.org/core-practices) of the Committee on the Publication Ethics describes the publication of the research. The core practices include the authorship, the peer review, the data, the conflicts, the misconduct, and the publication. The core practices are the recommended for the publication of the research.

The publication of the imputation should be the conducted according to the core practices. The authorship should be the assigned to the researchers who the contributed to the research. The peer review should be the conducted by the experts in the field. The data should be the shared to the extent the possible. The conflicts should be the declared. The misconduct should be the avoided.

The researcher identity and the publication are the important for the reporting of the imputation. The researcher identity is the important for the attribution of the research, and the publication is the important for the integrity of the research.

## The Role of the Funding and the Grants

The funding and the grants are the relevant for the research that the involves the imputation. The funding is the support for the research, and the grants are the mechanism of the funding.

The [NIH Grants and the Funding](https://grants.nih.gov/) describes the funding of the research. The NIH is the funder of the research, and the NIH is the provides the grants for the research. The NIH grants are the awarded through the application and the review process.

The funding of the research should be the declared in the publication. The funding is the declared in the funding statement, and the funding statement is the included in the publication. The funding statement is the important for the transparency of the research.

The grants are the relevant for the imputation of the missing survival data. The grants are the support for the research, and the grants are the used for the collection of the data, the analysis of the data, and the reporting of the research. The grants are the important for the research.

The research of the imputation should be the conducted in the accordance with the funding. The research should be the conducted in the accordance with the grant, and the research should be the reported in the accordance with the grant. The research is the important for the funding.

## The Role of the Reporting Guidelines

The reporting guidelines are the important for the reporting of the imputation. The reporting guidelines are the recommendations for the reporting of the research, and the reporting guidelines are the selected based on the study design.

The [EQUATOR Network](https://www.equator-network.org/) is the repository of the reporting guidelines. The EQUATOR Network is the provides the reporting guidelines for the research, and the EQUATOR Network is the selected the reporting guidelines for the study.

The reporting guidelines for the imputation are the selected based on the study design. The reporting guidelines for the observational study are the STROBE, and the reporting guidelines for the clinical trial are the CONSORT. The reporting guidelines for the imputation are the selected based on the study design.

The reporting guidelines are the important for the reporting of the imputation. The reporting guidelines are the recommendations for the reporting, and the reporting guidelines are the important for the transparency of the research. The reporting guidelines are the recommended for the research.

The reporting of the imputation should be the conducted in the accordance with the reporting guidelines. The reporting should be the conducted in the accordance with the reporting guidelines, and the reporting should be the reported in the accordance with the reporting guidelines. The reporting is the important for the transparency.

## The Role of the Research Methods Resources

The research methods resources are the important for the imputation of the missing survival data. The research methods resources are the resources for the research methods, and the research methods resources are the selected based on the research.

The [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) are the resources for the research methods. The Research Methods Resources are the books for the research methods, and the Research Methods Resources are the selected based on the research.

The research methods resources are the important for the imputation. The research methods resources are the important for the methods of the imputation, and the research methods resources are the important for the analysis of the imputation. The research methods resources are the important for the research.

The research methods resources are the selected based on the research. The research methods resources are the selected based on the study design, and the research methods resources are the selected based on the analysis. The research methods resources are the important for the research.

The research of the imputation should be the conducted in the accordance with the research methods resources. The research should be the conducted in the accordance with the research methods resources, and the research should be the reported in the accordance with the research methods resources. The research is the important for the research.

## Frequently Asked Questions

### What is the difference between complete-case analysis and multiple imputation for missing covariates in survival models?

Complete-case analysis deletes observations with any missing covariate and fits the Cox model to the remaining data. Multiple imputation generates several complete datasets by drawing plausible values for the missing covariates from a predictive distribution, fits the Cox model to each dataset, and pools the estimates. Complete-case analysis is valid only when missingness is completely random, while multiple imputation is valid under the weaker assumption that missingness depends on observed data.

### How many imputations should I use for a Cox model with missing covariates?

The statistical literature has traditionally recommended 20 imputations, but more recent guidance suggests using 20 to 40 imputations to obtain stable estimates of the standard errors. The number of imputations should be increased when the fraction of missing data is large. The computational cost of additional imputations is modest for most survival datasets.

### Should I include the survival time in the imputation model?

The survival time should be represented in the imputation model through the event indicator and the cumulative baseline hazard. The survival time should not be included as a linear predictor because the relationship between the survival time and the covariates is not linear. The cumulative baseline hazard is the appropriate summary of the survival time for the Cox model.

### What is the cumulative baseline hazard and why is it important for imputation?

The cumulative baseline hazard is the integral of the baseline hazard from time zero to the observed time. It is a measure of the risk that a subject has accumulated up to the observed time. It is important for imputation because it summarizes the survival time in a way that is compatible with the Cox model, and it allows the imputation model to use the information in the survival time.

### What should I do if the missingness depends on the unobserved values?

If the missingness depends on the unobserved values, the missingness is not ignorable, and the standard multiple imputation approach is not valid. You should consider a sensitivity analysis that models the missingness mechanism, or you should consult a statistician with expertise in missing data methods. The sensitivity analysis should be reported in the results.

### How do I handle interactions in the imputation model?

The imputation model should include any interactions that are in the analysis model. If the analysis model includes an interaction between a covariate and a treatment group, the imputation model must include the interaction term. Otherwise, the imputed values will not preserve the relationship between the covariates, and the interaction coefficient will be biased.

### What should I report in the methods section about the imputation?

The methods section should report the proportion of missingness for each covariate, the pattern of missingness, the imputation model including the predictors and the method for each variable, the number of imputations, and the pooling method. The sensitivity analysis should also be described in the methods section.

### How do I check the convergence of the imputation?

The convergence of the imputation is checked by examining the trace plots of the imputed values. The trace plots should show that the imputed values are stable across the iterations. If the trace plots show a trend or a drift, the imputation has not converged, and the number of iterations should be increased.

## Using the Evidence

| Source | Best use in this topic | Important limitation |
|---|---|---|
| [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) | official guidance | Check the linked page for current local requirements |
| [EQUATOR Network](https://www.equator-network.org/) | official guidance | Check the linked page for current local requirements |
| [Core Practices](https://publicationethics.org/core-practices) | official guidance | Check the linked page for current local requirements |

## Related Bioinformatics Guides

- [Longitudinal Microbiome Data Analysis: Methods and Best Practices](/knowledge/bioinformatics/longitudinal-microbiome-data-analysis-methods-and-best-practices)
- [Spatial Transcriptomics Data Integration: Aligning and Combining Multiple Datasets](/knowledge/bioinformatics/spatial-transcriptomics-data-integration-aligning-and-combining-multiple-datasets)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Predicting Colorectal Cancer Survival Using Time-to-Event Machine Learning: Retrospective Cohort Study.](https://pubmed.ncbi.nlm.nih.gov/37883174). Journal of medical Internet research, 2023.
- [A Frailty Index for UK Biobank Participants.](https://pubmed.ncbi.nlm.nih.gov/29924297). The journals of gerontology. Series A, Biological sciences and medical sciences, 2019.
- [Cox regression analysis with missing covariates via nonparametric multiple imputation.](https://pubmed.ncbi.nlm.nih.gov/29717943). Statistical methods in medical research, 2019.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.