# Kaplan-Meier Survival Curves


## Key Takeaways

- The Kaplan-Meier estimator is a non-parametric method for estimating survival probability over time, crucially accounting for right-censored observations by including censored subjects in the risk set until their censoring time.
- Construction involves ordering event times, calculating the proportion of subjects at risk who survive each interval, and stepping the survival curve downward only at observed events, with censored observations typically marked by tick marks.
- A critical assumption is non-informative censoring, meaning subjects censored have the same future survival prospects as those remaining under observation; violation of this assumption leads to biased estimates.
- The median survival time is defined as the time at which the survival probability drops to 0.5, a key metric for summarizing survival in a population, and is read directly from the curve.
- Comparing survival curves between groups is typically performed using the log-rank test, which assesses differences in observed versus expected events under the null hypothesis of no survival difference.
- The reliability of survival estimates diminishes in later time intervals as the number of subjects at risk decreases, indicated by widening confidence intervals and a less stable curve.

---

## Quick Answer

- Kaplan-Meier curves estimate survival probability over time while correctly accounting for censored observations, where the event of interest has not yet occurred.
- Construct the curve by ordering event times, calculating the proportion surviving at each interval, and stepping the curve downward only at observed events.
- The method assumes censoring is independent of survival prognosis, and the curve becomes less reliable when few subjects remain at risk in later intervals.

## Understanding Survival Data and Censoring

Survival analysis examines the time until a defined event occurs. In biological research, the event might be death, disease recurrence, metastasis, recovery, or any binary outcome with a measurable time component. The distinguishing feature of survival data is that some subjects do not experience the event during the observation period. These subjects are censored.

Censoring occurs when the event time is unknown because the observation ended before the event happened. A subject may be censored because the study concluded, the subject was lost to follow-up, or the subject withdrew from the study. In each case, the researcher knows the subject survived at least until the censoring time, but the exact event time remains unknown.

Right censoring is the most common form in biological research. The subject enters the study at a defined time, is observed for a period, and the event either occurs during observation or does not occur before observation ends. Left censoring occurs when the event has already happened before the subject enters the study. Interval censoring occurs when the event is known to have happened within a time window, but the exact time is unknown.

The Kaplan-Meier method handles right-censored data effectively. It uses the information from censored subjects up to their last known observation time, then removes them from the risk set. This approach avoids discarding valuable follow-up information while preventing the analysis from assuming censored subjects experience the event at the censoring time.

## Core Principles of the Kaplan-Meier Estimator

The Kaplan-Meier estimator, also called the product-limit estimator, calculates the probability of surviving past a given time point. The method works by dividing the observation period into intervals defined by event times. At each event time, the survival probability is multiplied by the proportion of subjects at risk who survive that interval.

The survival function S(t) represents the probability that a subject survives beyond time t. The Kaplan-Meier estimator computes this as a product of conditional probabilities. At each event time, the conditional probability of surviving that interval equals the number of subjects at risk minus the number of events, divided by the number of subjects at risk.

The estimator assumes that censoring is independent of the survival outcome. This means subjects who are censored have the same future survival prospects as subjects who remain under observation. This assumption is called non-informative censoring. When censoring is informative, meaning censored subjects have different event risks than observed subjects, the Kaplan-Meier estimates become biased.

The estimator also assumes that survival probabilities are the same for subjects who enter the study at different calendar times. This assumption can be violated in long-running studies where treatment protocols or patient populations change over time.

## Constructing a Kaplan-Meier Curve

Constructing a Kaplan-Meier curve requires a dataset with two variables for each subject: the time to event or censoring, and an indicator of whether the event occurred or the subject was censored. The time variable must be measured in consistent units across all subjects.

The construction process follows a stepwise procedure. First, sort all event times in ascending order. Second, at each event time, determine the number of subjects at risk, which is the number of subjects who have not yet experienced the event and have not yet been censored. Third, compute the proportion of subjects at risk who experience the event at that time. Fourth, multiply the previous survival probability by the proportion surviving the interval.

The curve is a step function. The survival probability remains constant between event times and drops at each event time. Censored subjects are marked on the curve with small vertical tick marks, though many software packages omit these marks by default.

The curve begins at time zero with a survival probability of 1.0. If the first event occurs at time t1, the survival probability drops to the proportion of subjects surviving past t1. The curve continues to step downward at each subsequent event time.

## Worked Example with R Code

The following worked example demonstrates Kaplan-Meier curve construction using R. The example uses a simulated dataset of 20 subjects with survival times and censoring indicators.

```r
## Load required packages
library(survival)
library(survminer)

## Create example data
## time: time to event or censoring
## event: 1 = event occurred, 0 = censored
time <- c(5, 8, 10, 12, 15, 18, 20, 22, 25, 28,
          30, 33, 35, 38, 40, 42, 45, 48, 50, 55)
event <- c(1, 1, 0, 1, 1, 0, 1, 1, 0, 1,
           1, 0, 1, 1, 0, 1, 1, 0, 1, 0)

## Create survival object
surv_obj <- Surv(time = time, event = event)

## Fit Kaplan-Meier model
km_fit <- survfit(surv_obj ~ 1)

## Display the survival table
summary(km_fit)

## Plot the Kaplan-Meier curve
plot(km_fit,
     xlab = "Time (months)",
     ylab = "Survival Probability",
     main = "Kaplan-Meier Survival Curve",
     conf.int = TRUE)

## Alternative plot using survminer
ggsurvplot(km_fit,
           data = data.frame(time = time, event = event),
           xlab = "Time (months)",
           ylab = "Survival Probability",
           conf.int = TRUE,
           risk.table = TRUE,
           censor = TRUE)
```

The summary output displays the number at risk, number of events, survival probability, and confidence intervals at each event time. The risk table below the plot shows the number of subjects still at risk at selected time points.

## Interpreting the Curve

Reading a Kaplan-Meier curve requires attention to the step pattern, the confidence intervals, and the number at risk. The curve starts at 1.0 and steps downward at each event time. The height of each step reflects the proportion of subjects at risk who experience the event at that time.

The median survival time is the time at which the survival probability drops to 0.5. This value is read directly from the curve by finding the time point where the curve crosses the 0.5 survival probability line. If the curve does not drop below 0.5, the median survival time is not reached within the observation period.

Confidence intervals around the curve indicate the precision of the survival estimates. Wider intervals at later time points reflect the smaller number of subjects at risk. The confidence intervals typically widen as time progresses because fewer subjects remain in the risk set.

The censoring marks on the curve indicate where subjects were censored. These marks do not change the survival probability but reduce the number of subjects at risk for subsequent intervals.

## Comparing Survival Curves Between Groups

Researchers often need to compare survival between two or more groups, such as treatment versus control or genotype A versus genotype B. The Kaplan-Meier method produces separate curves for each group. The log-rank test is the standard statistical test for comparing these curves.

The log-rank test compares the observed number of events in each group with the expected number under the null hypothesis of no difference in survival. The test statistic follows a chi-square distribution. A significant result indicates that the survival distributions differ between groups.

The log-rank test assumes that the survival curves do not cross. When curves cross, the test may fail to detect differences that exist. In such cases, alternative tests such as the Wilcoxon test or the Peto test may be more appropriate.

The R code for comparing groups uses a grouping variable in the survival formula:

```r
## Create data with group variable
group <- c("A", "A", "B", "A", "B", "B", "A", "B", "A", "B",
           "A", "B", "A", "B", "A", "B", "A", "B", "A", "B")

## Fit Kaplan-Meier model by group
km_by_group <- survfit(Surv(time, event) ~ group)

## Plot both curves
ggsurvplot(km_by_group,
           data = data.frame(time = time, event = event, group = group),
           pval = TRUE,
           conf.int = TRUE,
           risk.table = TRUE)

## Perform log-rank test
logrank_test <- survdiff(Surv(time, event) ~ group)
print(logrank_test)
```

The p-value from the log-rank test appears on the plot when the pval argument is set to TRUE. The risk table shows the number at risk for each group at selected time points.

## At a Glance

| Aspect | What to Do | Common Error | Practical Consequence |
|--------|------------|--------------|---------------------|
| Censoring indicator | Code event as 1 and censored as 0 | Reversing the coding | Survival probabilities become reversed and meaningless |
| Time units | Use consistent units across all subjects | Mixing days and months | Incorrect time scale and biased estimates |
| Risk set | Track number at risk at each event time | Ignoring censored subjects | Overestimation of survival probability |
| Confidence intervals | Report with the curve | Omitting intervals | No measure of uncertainty |
| Group comparison | Use log-rank test | Visual inspection only | No statistical evidence for differences |
| Median survival | Read at 0.5 probability | Reporting mean survival | Mean is biased by censoring |

## Data Requirements and Preparation

The Kaplan-Meier method requires a minimum of two variables per subject. The first variable is the time to event or censoring. The second variable is the event indicator. The event indicator must be coded as a binary variable, typically 1 for event and 0 for censoring.

The time variable must be measured on a continuous scale. Common units include days, weeks, months, or years. The choice of units affects the interpretation of the curve but does not change the statistical properties of the estimator.

The data must be checked for errors before analysis. Common data errors include negative times, event indicators with values other than 0 or 1, and missing time values. These errors produce incorrect survival estimates or cause the analysis to fail.

The dataset should include a subject identifier to track individual records. This identifier is essential for data management and for linking survival data to other subject characteristics.

The sample size affects the precision of the survival estimates. Small samples produce wide confidence intervals and unstable estimates at later time points. The number of events, not the total sample size, determines the statistical power of the log-rank test.

## Handling Censored Observations Correctly

Censored observations are the defining feature of survival data. The Kaplan-Meier estimator handles right censoring by including censored subjects in the risk set until their censoring time, then removing them from subsequent intervals.

The key principle is that censored subjects contribute information up to their censoring time. A subject censored at time t provides evidence that the subject survived to time t. This information is used in the survival probability calculation for all intervals up to time t.

The censoring indicator must be coded correctly. A common error is to code censored subjects as events or to exclude censored subjects from the analysis entirely. Excluding censored subjects biases the survival estimates downward because the analysis assumes all excluded subjects experienced the event.

The censoring mechanism must be non-informative. This means the reason for censoring is unrelated to the subject's survival prospects. For example, if subjects who become sick are more likely to withdraw from the study, the censoring is informative and the Kaplan-Meier estimates are biased.

The researcher should examine the pattern of censoring in the data. If censoring is concentrated at specific time points, such as the end of the study period, this is expected and does not indicate a problem. If censoring occurs at random times throughout the study, the pattern should be investigated.

## Practical Workflow for Analysis

The following workflow provides a systematic approach to Kaplan-Meier analysis:

1. Define the event of interest and the time origin for each subject. The time origin is the point at which the subject enters the study and begins being at risk.
2. Prepare the dataset with time and event indicator variables. Verify that all time values are positive and the event indicator contains only 0 and 1.
3. Examine the data for missing values and errors. Correct any identified problems before proceeding.
4. Fit the Kaplan-Meier model using appropriate software. R, SAS, SPSS, and Python all provide functions for this analysis.
5. Generate the survival curve plot with confidence intervals and censoring marks.
6. Examine the curve for the pattern of survival. Note the median survival time and the shape of the curve.
7. If comparing groups, perform the log-rank test and report the p-value.
8. Check the assumptions of the analysis, particularly the non-informative censoring assumption.
9. Report the results with the number at risk at selected time points and the confidence intervals.

## Options and Tradeoffs in Software

Several software packages provide Kaplan-Meier analysis. R is widely used in biological research and offers the survival and survminer packages. The survival package provides the core functions for fitting the model, while the survminer package provides enhanced plotting capabilities.

SAS provides the LIFETEST procedure for survival analysis. The procedure produces the Kaplan-Meier estimates and the log-rank test. SAS is common in clinical research and regulatory settings.

Python provides the lifelines package for survival analysis. The package includes the Kaplan-MeierFitter class for fitting and plotting survival curves. Python is useful for researchers who work primarily in Python for data analysis.

The choice of software depends on the researcher's familiarity and the requirements of the research setting. The statistical results should be identical across software packages when the same data and methods are used.

## Records and Measurements

The analysis should record the following information for each subject: the subject identifier, the time variable, the event indicator, and any group variables. The dataset should be stored in a format that preserves the original data without modification.

The analysis output should include the survival probability estimates at each event time, the number at risk at each time, and the confidence intervals. The log-rank test output should include the chi-square statistic and the p-value.

The analysis code should be saved and documented so that the analysis can be reproduced. The code should include comments explaining each step and the rationale for the analysis choices.

The data management and sharing expectations for the research should be considered. The National Institutes of Health data management and sharing policy describes expectations for data management and sharing in NIH-funded research. The policy requires a data management and sharing plan that describes how data will be preserved and shared.

## Common Failure Patterns

Several common errors occur in the Kaplan-Meier analysis. The first is the exclusion of censored subjects from the analysis. This error reduces the sample size and biases the survival estimates downward.

The second error is the coding of the event indicator. Reversing the event and censoring codes produces a curve that estimates the probability of the event instead of the probability of survival.

The third error is the use of the mean survival time instead of the median. The mean survival time is not well-defined when censoring is present because the mean requires the full survival distribution. The median is the appropriate summary measure.

The fourth error is the failure to report the number at risk. The number at risk at each time point indicates the reliability of the survival estimate. Without this information, the reader cannot assess the stability of the curve at later time points.

The fifth error is the comparison of survival curves without the log-rank test. Visual inspection of the curves is not sufficient to establish a statistically significant difference.

## Limitations and Interpretation Boundaries

The Kaplan-Meier method has several limitations that must be considered in interpretation. The method does not adjust for covariates. The method estimates the survival distribution for the entire sample or for groups defined by a single categorical variable.

The method assumes that censoring is non-informative. When this assumption is violated, the survival estimates are biased. The researcher should examine the reasons for censoring and consider whether the assumption is reasonable.

The method does not provide a direct estimate of the effect of a continuous variable on survival. The Cox proportional hazards model is the appropriate method for this purpose.

The survival estimates at later time points are based on fewer subjects and are less reliable. The confidence intervals reflect this uncertainty, but the researcher should be cautious in interpreting the tail of the curve.

The Kaplan-Meier curve does not provide a direct estimate of the hazard function. The hazard function describes the instantaneous risk of the event at each time point. The hazard can be estimated separately, but it is not directly visible in the Kaplan-Meier curve.

## Reporting Standards and Transparency

The reporting of the Kaplan-Meier analysis should follow established reporting guidelines. The EQUATOR Network provides a collection of reporting guidelines for health research. The guidelines help ensure that the analysis is reported completely and transparently.

The report should include the number of subjects, the number of events, the number of censored subjects, and the median survival time with confidence intervals. The report should also include the number at risk at selected time points.

The report should describe the censoring pattern and the reasons for censoring. The report should state the assumptions of the analysis and any limitations.

The data and code should be made available for the analysis to be reproduced. The National Institutes of Health data management and sharing policy describes the expectations for data sharing in NIH-funded research. The policy requires a data management plan that describes how data will be preserved and shared.

## Professional Escalation Criteria

The researcher should seek additional statistical expertise when the analysis involves complex data structures or when the assumptions of the method are in question. The following situations warrant consultation with a biostatistician:

- The data include informative censoring or competing risks.
- The survival curves cross between groups.
- The analysis requires adjustment for multiple covariates.
- The data include time-varying covariates.
- The study design involves clustered or correlated data.

The researcher should also consult a biostatistician when the results are critical for a regulatory decision or when the analysis will be submitted for publication in a high-impact journal.

## Decision Framework for Choosing Between Kaplan-Meier and Alternative Survival Methods

The Kaplan-Meier estimator is the appropriate starting point for most time-to-event analyses, but it is not always the correct final method. Researchers frequently default to Kaplan-Meier curves without formally assessing whether the data structure and study question require a different approach. This section provides a practical decision framework that helps researchers determine when Kaplan-Meier analysis is sufficient, when it must be supplemented, and when it should be replaced entirely.

### Step 1: Verify the Data Structure Matches Kaplan-Meier Assumptions

Before constructing a Kaplan-Meier curve, verify that the data meet the three structural requirements of the estimator. First, the time variable must represent time from a well-defined origin to either the event or censoring. The time origin must be consistent across all subjects. In clinical studies, the origin is often randomization, diagnosis, or treatment initiation. In agricultural studies, the origin might be birth, weaning, or enrollment in a feeding trial. If subjects enter the study at different stages of their disease or production cycle, the time origin is not comparable and the Kaplan-Meier estimates will be biased.

Second, the event must be a single, well-defined binary outcome. The Kaplan-Meier method cannot handle multiple event types. If subjects can experience competing events, such as death from disease versus death from culling for production reasons, the standard Kaplan-Meier estimator treats all deaths as the same event. This produces misleading survival probabilities because the risk of the event of interest is not independent of the competing event. The competing risks framework, which estimates the cumulative incidence function, is the appropriate method when competing events exist.

Third, the censoring must be non-informative. The researcher must examine the reasons for censoring before proceeding. If subjects are censored because they became too sick to continue, because they were removed from the study due to poor prognosis, or because they received a different treatment after a change in condition, the censoring is informative and the Kaplan-Meier estimates are biased. The researcher should document the reason for each censored observation and assess whether the censoring mechanism is related to the outcome.

### Step 2: Assess Whether the Research Question Requires Covariate Adjustment

The Kaplan-Meier estimator describes the survival distribution for the entire sample or for groups defined by a single categorical variable. It does not adjust for covariates. If the research question asks whether a treatment affects survival after controlling for age, sex, disease stage, or other prognostic factors, the Kaplan-Meier method is insufficient.

The Cox proportional hazards model is the standard extension for covariate adjustment. The Cox model estimates the hazard ratio for each covariate while allowing the baseline hazard to remain unspecified. The model does not require the survival distribution to follow a specific parametric form, which makes it flexible for most biological data.

The decision between Kaplan-Meier and Cox regression depends on the research question. If the question is descriptive, such as what is the survival probability at one year for this population, the Kaplan-Meier curve is appropriate. If the question is comparative and requires adjustment for confounders, the Cox model is necessary. If the question asks about the effect of a continuous variable, such as the dose of a drug or the level of a biomarker, the Cox model is the appropriate method.

### Step 3: Check for Time-Varying Effects and Non-Proportional Hazards

The Cox proportional hazards model assumes that the hazard ratio between groups is constant over time. When this assumption is violated, the Cox model produces an average hazard ratio that may not represent the true relationship at any specific time point. The Kaplan-Meier curve is useful for detecting non-proportional hazards because it shows the survival curves for each group over time.

The researcher should plot the Kaplan-Meier curves for each group and examine whether the curves cross or converge. If the curves cross, the proportional hazards assumption is violated. In this situation, the log-rank test, which is the standard test for comparing Kaplan-Meier curves, may fail to detect a real difference. The log-rank test gives equal weight to all time points, so it can miss differences that occur early or late in the follow-up period.

When the proportional hazards assumption is violated, the researcher should consider alternative methods. The restricted mean survival time compares the average survival time up to a specified time point. This method does not require the proportional hazards assumption and provides a clinically meaningful summary measure. The weighted log-rank test can be used to emphasize early or late differences, but the choice of weights must be specified before the analysis to avoid bias.

### Step 4: Determine Whether the Sample Size Supports the Analysis

The number of events, not the total sample size, determines the precision of the Kaplan-Meier estimates and the power of the log-rank test. A study with 100 subjects and 90 events provides more precise survival estimates than a study with 200 subjects and 30 events. The researcher should count the number of events in each group before deciding whether the Kaplan-Meier analysis is feasible.

The number at risk at each time point is the critical indicator of estimate reliability. The Kaplan-Meier curve becomes unstable when the number at risk drops below 10 or 20 subjects. The confidence intervals widen substantially at these time points, and the survival probability estimates are based on very few observations. The researcher should report the number at risk at each time point and should not overinterpret the tail of the curve when the number at risk is small.

If the sample size is too small for the Kaplan-Meier analysis, the researcher should consider whether the study should be extended, whether additional subjects can be enrolled, or whether the research question can be answered with a different method. The Kaplan-Meier method does not provide a formal sample size calculation, but the researcher can use the number of events to estimate the precision of the survival estimates.

### Step 5: Evaluate the Censoring Pattern

The pattern of censoring affects the interpretation of the Kaplan-Meier curve. The researcher should examine the distribution of censoring times and the reasons for censoring. If censoring is concentrated at the end of the study period, this is expected and reflects the administrative end of follow-up. If censoring occurs throughout the study period, the researcher should investigate whether the censoring is related to the outcome.

The researcher should create a table of censoring reasons and examine whether the reasons differ between groups. If the reasons for censoring differ between groups, the comparison of survival curves may be biased. For example, if subjects in the treatment group are more likely to be censored because they move away, and moving away is related to the outcome, the censoring is informative and the Kaplan-Meier estimates are biased.

The researcher should also check whether the censoring times are independent of the event times. The Kaplan-Meier estimator assumes that the censoring time is independent of the event time. This assumption is violated when the censoring time is related to the event time, such as when subjects who are at higher risk of the event are more likely to be censored.

### Step 6: Compare the Kaplan-Meier Curve with Parametric Alternatives

The Kaplan-Meier estimator is non-parametric, meaning it does not assume a specific distribution for the survival times. This is a strength because the estimator is flexible and can accommodate any survival pattern. However, the non-parametric approach has limitations. The Kaplan-Meier curve is a step function, and the survival probability is constant between event times. This makes it difficult to estimate the survival probability at time points that do not correspond to observed events.

Parametric survival models, such as the exponential, Weibull, or log-normal models, assume a specific distribution for the survival times. These models provide smooth survival curves and allow the researcher to estimate the survival probability at any time point. The parametric models also provide estimates of the hazard function, which describes the instantaneous risk of the event at each time point.

The choice between the Kaplan-Meier estimator and a parametric model depends on the research question and the data. If the researcher is interested in describing the survival pattern without making assumptions about the distribution, the Kaplan-Meier estimator is appropriate. If the researcher is interested in estimating the hazard function or in predicting survival beyond the observed follow-up period, a parametric model is necessary.

The researcher should compare the Kaplan-Meier curve with the parametric model fit to assess whether the parametric model is appropriate. If the parametric model fits the data well, the parametric model provides more efficient estimates. If the parametric model does not fit the data, the Kaplan-Meier estimator is the appropriate choice.

### Step 7: Document the Decision Process

The decision framework should be documented in the analysis plan before the analysis is conducted. The documentation should include the research question, the data structure, the assumptions of the Kaplan-Meier method, and the reasons for choosing the Kaplan-Meier method over alternative methods. The documentation should also include the results of the assumption checks and the sensitivity analyses.

The documentation is important for transparency and reproducibility. The Committee on Publication Ethics core practices emphasize the importance of transparent reporting of research methods. The documentation of the decision process allows other researchers to understand the rationale for the analysis choices and to assess the validity of the results.

The documentation should be included in the data management and sharing plan. The National Institutes of Health data management and sharing policy requires that the data management plan describe how the data will be preserved and shared. The documentation of the analysis decisions should be included in the data sharing plan to ensure that the analysis can be reproduced.

### Step 8: Escalate to a Biostatistician When the Decision Is Unclear

The decision framework provides a systematic approach to choosing the appropriate survival analysis method. However, some situations require the expertise of a biostatistician. The researcher should escalate the decision to a biostatistician when the data structure is complex, when the assumptions of the Kaplan-Meier method are in question, or when the research question requires advanced methods.

The following situations warrant escalation to a biostatistician:

- The data include competing risks, where subjects can experience multiple types of events.
- The data include time-varying covariates, where the value of a covariate changes over time.
- The data include clustered or correlated observations, such as multiple subjects from the same farm or the same litter.
- The data include informative censoring, where the reason for censoring is related to the outcome.
- The research question requires adjustment for multiple covariates.
- The survival curves cross between groups, indicating non-proportional hazards.

The biostatistician can provide guidance on the appropriate methods and can help the researcher interpret the results. The biostatistician can also help the researcher to plan the analysis before the data are collected, which is the most effective way to ensure that the analysis is appropriate.

### Decision Framework Summary Table

| Decision Point | Question to Answer | Method to Use | Escalation Criterion |
|----------------|-------------------|---------------|----------------------|
| Data structure | Is the event a single binary outcome? | Kaplan-Meier | Competing events present |
| Time origin | Is the time origin consistent across subjects? | Kaplan-Meier | Different disease stages at entry |
| Censoring | Is censoring non-informative? | Kaplan-Meier | Censoring related to outcome |
| Covariates | Does the question require adjustment? | Cox model | Multiple confounders |
| Proportional hazards | Are the hazard ratios constant over time? | Log-rank test | Curves cross |
| Sample size | Are there enough events for precision? | Kaplan-Meier | Fewer than 10 events per group |
| Censoring pattern | Is censoring concentrated at the end? | Kaplan-Meier | Censoring throughout follow-up |
| Distribution | Is a parametric model needed? | Parametric model | Prediction beyond follow-up |

### Practical Implementation Steps

The decision framework should be applied before the Kaplan-Meier analysis is conducted. The following steps provide a practical implementation of the framework:

1. Write the research question in a single sentence. Identify the event of interest, the time origin, and the comparison groups.
2. Create a table of the data structure. List the variables, the event indicator, the time variable, and the censoring indicator.
3. Examine the reasons for censoring. Create a table of the reasons for censoring and the number of subjects in each category.
4. Count the number of events in each group. If the number of events is fewer than 10 in any group, consider whether the analysis is feasible.
5. Plot the Kaplan-Meier curves for each group. Examine the curves for crossing or divergence.
6. Perform the log-rank test. If the test is significant, the survival distributions differ between groups.
7. Fit the Cox proportional hazards model. Check the proportional hazards assumption using the Schoenfeld residuals.
8. Document the decision process. Record the research question, the data structure, the assumptions, and the analysis choices.

### Common Failure Patterns in the Decision Process

The decision framework is designed to prevent common errors in survival analysis. The following failure patterns are common when the framework is not applied:

The first failure pattern is the use of the Kaplan-Meier method when the data include competing events. The researcher treats all events as the same outcome, which produces biased estimates of the survival probability. The competing risks analysis should be used instead.

The second failure pattern is the use of the Kaplan-Meier method when the research question requires covariate adjustment. The researcher compares the survival curves between groups without adjusting for confounders, which produces a biased estimate of the treatment effect. The Cox proportional hazards model should be used instead.

The third failure pattern is the use of the log-rank test when the survival curves cross. The log-rank test has low power to detect differences when the curves cross. The restricted log-rank test or the restricted mean survival time should be used instead.

The fourth failure pattern is the interpretation of the tail of the curve when the number at risk is small. The researcher reports the survival probability at the end of the follow-up period without noting that the estimate is based on a few subjects. The number at risk should be reported at each time point, and the tail of the curve should be interpreted with caution.

The fifth failure pattern is the failure to document the decision process. The researcher does not record the reasons for the analysis choices, which makes the analysis difficult to reproduce and the results difficult to verify.

### Records and Measurements for the Decision Framework

The decision framework should be documented in the research records. The following records should be maintained:

- The research question and the analysis plan.
- The data dictionary that describes the time variable, the event indicator, and the censoring indicator.
- The table of censoring reasons and the number of subjects in each category.
- The number of events in each group.
- The Kaplan-Meier curves for each group.
- The log-rank test results.
- The Cox proportional hazards model results and the proportional hazards assumption check.
- The decision documentation that records the analysis choices and the rationale.

The records should be stored in a format that preserves the original data without modification. The analysis code should be saved and documented so that the analysis can be reproduced. The data management and sharing plan should describe how the data and the analysis code will be preserved and shared.

### Professional Escalation Criteria for the Decision Framework

The decision framework includes escalation criteria for situations that require the expertise of a biostatistician. The researcher should escalate the decision to a biostatistician when the data structure is complex, when the assumptions of the Kaplan-Meier method are in question, or when the research question is significant.

The escalation criteria include the presence of competing events, time-varying covariates, clustered data, informative censoring, and non-proportional hazards. The researcher should also escalate the decision when the analysis will be submitted for publication in a high-impact journal or when the results are critical for a regulatory decision.

The biostatistician can provide the appropriate methods and can help the researcher interpret the results. The biostatistician can also help the researcher design the analysis before the data are collected, which is the most effective approach for ensuring that the analysis is appropriate.

## Frequently Asked Questions

### What does a censored observation mean in a Kaplan-Meier analysis?

A censored observation means the event has not occurred by the time the subject is last observed. The subject may have been lost to follow-up, withdrawn from the study, or the study may have ended before the event occurred. The censored subject contributes information to the survival estimate up to the censoring time.

### How does the Kaplan-Meier method handle censored subjects?

The method includes censored subjects in the risk set for all time intervals before their censoring time. At the censoring time, the subject is removed from the risk set without contributing an event. This approach uses the information that the subject survived to the censoring time without assuming the subject experienced the event.

### What is the difference between the median and mean survival time?

The median survival time is the time at which the survival probability drops to 0.5. The median is read directly from the Kaplan-Meier curve. The mean survival time is the average survival time, which is not well-defined when censoring is present because the full distribution of survival times is not observed.

### Why do the confidence intervals widen at later time points?

The confidence intervals widen because the number of subjects at risk decreases over time. With fewer subjects at risk, the survival probability estimate is less precise. The confidence intervals reflect this reduced precision.

### Can the Kaplan-Meier method be used with continuous variables?

The Kaplan-Meier method is designed for categorical variables. For continuous variables, the survival model such as the Cox proportional hazards model is appropriate. The Cox model can include continuous variables as predictors of survival.

### What is the log-rank test used for?

The log-rank test compares the survival distributions of two or more groups. The test determines whether the observed differences in survival between groups are statistically significant. The test is appropriate when the survival curves do not cross.

### How many subjects are needed for a Kaplan-Meier analysis?

The required sample size depends on the number of events and the desired precision of the estimates. The number of events, not the total sample size, determines the statistical power. A biostatistician should be consulted for a formal sample size calculation.

### What should be reported when presenting a Kaplan-Meier curve?

The report should include the number of subjects, the number of events, the number of censored subjects, the median survival time with confidence intervals, and the number at risk at selected time points. The report should also describe the censoring pattern and the assumptions of the analysis.

## Related Bioinformatics Guides

- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Multi-Omics Integration: A Practical Guide to Combining Data Types](/knowledge/bioinformatics/multi-omics-integration-a-practical-guide-to-combining-data-types)
- [Proteomics Data Analysis in R: A Practical Workflow for Differential Expression and Visualization](/knowledge/bioinformatics/proteomics-data-analysis-in-r-a-practical-workflow-for-differential-expression-and-visualization)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [IPDfromKM: reconstruct individual patient data from published Kaplan-Meier survival curves.](https://pubmed.ncbi.nlm.nih.gov/34074267). BMC medical research methodology, 2021.
- [Kaplan-Meier Survival Analysis: Practical Insights for Clinicians.](https://pubmed.ncbi.nlm.nih.gov/38631048). Acta medica portuguesa, 2024.
- [A practical guide to understanding Kaplan-Meier curves.](https://pubmed.ncbi.nlm.nih.gov/20723767). Otolaryngology--head and neck surgery : official journal of American Academy of Otolaryngology-Head and Neck Surgery, 2010.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.