Power Analysis for Longitudinal Studies

By Dr. Zubair Khalid, DVM, MS, PhD ·

Power Analysis for Longitudinal Studies

Key Takeaways

  • Longitudinal study power is critically influenced by the within-subject correlation (ρ), which reduces the effective sample size; higher ρ necessitates larger sample sizes or more time points to achieve adequate power, and should be estimated from pilot data or prior literature.
  • Dropout rates significantly diminish power by reducing the effective sample size at later time points; sample size calculations must inflate the initial recruitment number to account for anticipated attrition, typically 10-20%, to maintain target power.
  • Simulation-based power calculations in statistical environments like R are paramount for complex longitudinal designs, allowing precise incorporation of specific correlation structures (e.g., autoregressive, exchangeable) and missing data patterns, which are often inadequately handled by simpler formulas.
  • The choice of statistical analysis method (e.g., mixed-effects models vs. Generalized Estimating Equations - GEE) dictates power analysis assumptions; mixed models require specification of variance components and random effects, while GEE necessitates defining a working correlation matrix.
  • Effect size, defined as the clinically or biologically meaningful difference, is a primary driver of power; larger effect sizes require fewer subjects, whereas smaller, subtler effects necessitate larger sample sizes or more repeated measures to detect with sufficient statistical power.

Quick Answer

  • Sample size for longitudinal studies depends on the number of subjects, number of repeated measurements, within-subject correlation, expected effect size, and dropout rate, with mixed models and GEE requiring different assumptions.
  • Use simulation-based power calculations in R or dedicated software to account for the correlation structure and missing data patterns before finalizing your study design.
  • Power analysis for longitudinal designs is more complex than cross-sectional calculations, and incorrect assumptions about correlation or dropout can lead to underpowered studies that waste resources.

At a Glance

Design ParameterImpact on PowerPractical Consideration
Number of subjects (N)Increasing N increases power, but with diminishing returnsBalance recruitment cost against the marginal gain in power
Number of repeated measurements (T)More time points can increase power, especially for detecting trendsEach additional measurement adds cost and participant burden
Within-subject correlation (ρ)Higher correlation reduces the effective sample sizeEstimate from pilot data or published studies in your field
Dropout rateReduces the effective sample size at later time pointsPlan for 10-20% attrition and adjust the initial sample accordingly
Effect sizeLarger effects require fewer subjectsUse clinically meaningful differences, beyond statistically detectable ones

Understanding Longitudinal Study Design

Longitudinal studies track the same subjects over time, collecting repeated measurements from each individual. This design is common in biology and life sciences when researchers want to understand how a biological variable changes over time, how a treatment affects a trajectory, or how early measurements predict later outcomes. The defining feature of a longitudinal study is that observations from the same subject are correlated, meaning that a subject who starts high tends to stay high relative to the group average.

This correlation is the central complication for power analysis. In a cross-sectional study, each observation is independent, and the sample size calculation follows the familiar formulas based on the standard error of the mean difference. In a longitudinal study, the repeated measurements from one subject are not independent, and treating them as if they were would overestimate the effective sample size and produce an underpowered study.

The statistical methods used to analyze longitudinal data, mixed-effects models and generalized estimating equations (GEE), both require explicit attention to the correlation structure. Mixed models treat the subject as a random effect and estimate both the fixed effects of interest and the variance components. GEE uses a working correlation matrix to adjust the standard errors for the within-subject correlation. Both approaches require the researcher to specify the correlation structure in advance for the purpose of power analysis.

The National Library of Medicine provides access to biomedical research-method references that describe the principles of longitudinal data analysis and the importance of accounting for repeated measures in study design. These references are useful for researchers who need to understand the statistical foundations before planning their own studies.

Core Principles of Power Analysis

Power is the probability that a statistical test will detect an effect of a given size when that effect truly exists. The power of a study depends on four factors: the significance level (alpha), the effect size, the sample size, and the variability of the measurements. In a longitudinal study, the variability has two components: the between-subject variability and the within-subject variability. The within-subject variability is further divided into the correlation between repeated measurements and the residual variance.

The significance level is the probability of rejecting the null hypothesis when it is true, typically set at 0.05. The effect size is the magnitude of the difference or association that the study aims to detect. The sample size is the number of subjects, and in a longitudinal study, the number of repeated measurements per subject also matters. The variability is the standard deviation of the outcome variable, which must be estimated from prior data or pilot studies.

The relationship between these factors is expressed in the power function, which gives the probability of rejecting the null hypothesis for a given set of parameter values. In a longitudinal study, the power function is more complex because the repeated measurements introduce a correlation structure that affects the effective sample size.

The effective sample size in a longitudinal study is not simply the total number of observations (subjects times time points). Because the repeated measurements are correlated, each new time point adds less information than an independent observation would. The effective sample size depends on the correlation between measurements and the number of time points. Higher correlation means that repeated measurements provide less new information, so the effective sample size is closer to the number of subjects. Lower correlation means the repeated measurements provide more independent information, so the effective sample size is closer to the total number of observations.

This relationship is important for power analysis because it means that adding more time points does not increase power as much as adding more subjects. The marginal gain from an additional time point depends on the correlation. If the correlation is high, the additional time point adds little new information. If the correlation is low, the additional time point adds more information.

Statistical Models for Longitudinal Data

The two most common approaches for analyzing longitudinal data are mixed-effects models and generalized estimating equations. Both approaches account for the within-subject correlation, but they do so in different ways, and the power analysis must match the intended analysis method.

Mixed-Effects Models

Mixed-effects models, also called multilevel models or hierarchical linear models, treat the subject as a random effect. The model includes fixed effects for the predictors of interest, such as treatment group and time, and random effects for the subject-specific deviations from the population average. The random effects capture the fact that some subjects are systematically higher or lower than the group average across all time points.

The mixed model for a simple longitudinal design with two groups and repeated measurements can be written as:

Y_ij = β0 + β1 × Time_j + β2 × Group_i + β3 × Time_j × Group_i + b_i + ε_ij

where Y_ij is the measurement for subject i at time j, β0 is the intercept, β1 is the time effect, β2 is the group effect, β3 is the interaction between time and group, b_i is the random intercept for subject i, and ε_ij is the residual error.

The random intercept b_i is assumed to follow a normal distribution with mean zero and variance σ²_b. The residual error ε_ij is assumed to follow a normal distribution with mean zero and variance σ²_e. The correlation between two measurements from the same subject is then σ²_b / (σ²_b + σ²_e), which is the intraclass correlation coefficient (ICC).

The power analysis for a mixed model requires specifying the variance components, the ICC, the effect size, the number of time points, and the number of subjects. The power depends on the standard error of the effect estimate, which is a function of these parameters.

Generalized Estimating Equations

GEE is a different approach that does not require specifying the full distribution of the data. Instead, GEE estimates the regression parameters and uses a working correlation matrix to adjust the standard errors. The working correlation matrix can be specified as independent, exchangeable, autoregressive, or unstructured, depending on the expected pattern of correlation over time.

The power analysis for GEE requires specifying the working correlation matrix and the variance of the outcome. The power depends on the standard error of the regression coefficient, which is calculated using the sandwich estimator that accounts for the within-subject correlation.

GEE is often used when the outcome is not normally distributed, such as binary or count outcomes. The power analysis for GEE with binary outcomes requires specifying the prevalence of the outcome and the correlation between repeated measurements.

Sample Size Formulas for Longitudinal Studies

The sample size formula for a longitudinal study depends on the specific design and the intended analysis. For a simple two-group comparison with repeated measurements and an exchangeable correlation structure, the sample size per group can be calculated using a formula that accounts for the number of time points and the ICC.

For a continuous outcome with a mixed model, the sample size per group is approximately:

n = (z_α/2 + z_β)² × 2 × σ² × (1 + (T - 1) × ρ) / (T × d²)

where z_α/2 is the critical value for the significance level, z_β is the critical value for the desired power, σ² is the total variance, T is the number of time points, ρ is the ICC, and d is the effect size.

This formula shows that the required sample size decreases as the number of time points increases, but the decrease is limited by the ICC. When the ICC is high, the term (1 + (T - 1) × ρ) / T approaches ρ, so the sample size is close to the cross-sectional sample size. When the ICC is low, the term approaches 1/T, so the sample size is divided by the number of time points.

For a two-group comparison with a time-by-group interaction, the formula is more complex because the effect of interest is the difference in the change over time between the two groups. The sample size depends on the variance of the change scores, which is a function of the within-subject variance and the correlation between the time points.

For GEE with a binary outcome, the sample size formula uses the variance of the binomial distribution and the working correlation matrix. The formula is:

N = (z_α/2 + z_β)² × [p1(1-p1) + p2(1-p2)] × (1 + (T - 1) × ρ) / (T × (p1 - p2)²)

where p1 and p2 are the proportions in the two groups, and the other terms are as defined above.

These formulas provide a starting point for power analysis, but they make simplifying assumptions about the correlation structure and the variance. In practice, the actual power may differ from the calculated power if the assumptions are not met.

Software for Power Analysis

Several software packages can perform power analysis for longitudinal studies. The choice of software depends on the complexity of the design and the familiarity of the researcher with the tool.

R Packages

R is a free and open-source statistical computing environment that is widely used in bioinformatics and life sciences. Several R packages provide functions for power analysis in longitudinal studies.

The lme4 package fits mixed-effects models and can be used for simulation-based power analysis. The researcher simulates data under the assumed model, fits the model to the simulated data, and calculates the proportion of simulations in which the effect is statistically significant. This approach is flexible and can handle complex designs, but it requires programming skills and can be computationally intensive.

The simr package provides tools for power analysis using simulation with mixed models. The researcher specifies the model, the effect size, and the sample size, and the package simulates the data and calculates the power. The package can also be used to determine the sample size needed to achieve a target power.

The longpower package provides power calculations for longitudinal studies with continuous outcomes. The package implements the formulas for the sample size based on the mixed model and the GEE approach.

The pwr package provides power analysis for basic statistical tests, but it does not handle the correlation structure of longitudinal data. It can be used for the cross-sectional components of the analysis, but not for the full longitudinal design.

G*Power

G*Power is a free software for power analysis that is widely used in the social and biological sciences. It provides power analysis for a range of statistical tests, including t-tests, ANOVA, and regression. However, G*Power does not have a specific module for longitudinal data with repeated measurements and within-subject correlation.

G*Power can be used for power analysis of a repeated-measures ANOVA, which is a special case of a mixed model with a specific correlation structure. The repeated-measures ANOVA module requires the correlation between the repeated measurements and the effect size. The power calculation assumes a specific pattern of correlation, such as sphericity, which may not hold in practice.

For more complex longitudinal designs, G*Power is not sufficient, and the researcher should use R or another specialized software.

Other Software

Other software packages that can perform power analysis for longitudinal studies include SAS, Stata, and PASS. These packages provide procedures for power analysis with mixed models and GEE. The choice of software depends on the availability of the software and the researcher's familiarity with the procedures.

Simulation-Based Power Analysis

Simulation-based power analysis is the most flexible approach for longitudinal studies. The researcher specifies the data-generating model, simulates the data many times, fits the model to each simulated dataset, and calculates the proportion of the simulations in which the effect is statistically significant. This proportion is the estimated power.

The simulation approach has several advantages over the formula-based approach. It can handle any design, any correlation structure, any distribution of the outcome, and any missing data pattern. It can also account for the uncertainty in the parameter estimates by simulating the parameters from a distribution.

The simulation approach also has disadvantages. It requires programming skills and can be computationally intensive, especially for large sample sizes and many simulations. The results depend on the assumptions of the simulation model, and the researcher must be careful to specify the model correctly.

The steps for a simulation-based power analysis are:

  1. Specify the data model, including the fixed effects, the random effects, the variance components, and the correlation structure.
  2. Specify the sample size, the number of time points, and the dropout rate.
  3. Generate the data from the model.
  4. Fit the analysis model to the simulated data.
  5. Test the effect of interest and record whether the test is statistically significant.
  6. Repeat steps 3-5 many times, typically 1000 or more.
  7. Calculate the proportion of significant tests, which is the estimated power.
  8. Repeat the process for different sample sizes to find the sample size that achieves the target power.

The simulation approach is particularly useful for designs with missing data, because the missing data mechanism can be simulated as part of the data generation. The researcher can specify the dropout rate and the pattern of dropout, such as dropout that is random or dropout that is related to the outcome.

Accounting for Dropout

Dropout is a common problem in longitudinal studies. Subjects may leave the study for various reasons, such as moving away, losing interest, or experiencing adverse effects. The dropout reduces the effective sample size at the later time points, which reduces the power of the study.

The power analysis must account for the expected dropout rate. The simplest approach is to inflate the sample size by the expected dropout rate. If the expected dropout rate is 20%, the sample size is increased by 25% (1 / (1 - 0.20) to achieve the target sample size at the end of the study.

However, the dropout rate may not be constant across the time points. The dropout may be higher in the early time points and lower in the later time points, or the dropout may be higher in one group than in the other. The power analysis should account for the pattern of dropout.

The dropout may also be related to the outcome. For example, subjects who are doing poorly may be more likely to drop out. This is called informative dropout, and it can bias the results if not handled properly. The power analysis should account for the possibility of informative dropout, but this is difficult because the dropout mechanism is not known in advance.

The simulation approach can handle dropout by simulating the dropout process. The researcher specifies the dropout rate and the pattern of dropout, and the simulation generates the missing data. The analysis model must be able to handle the missing data, which is typically done using maximum likelihood or multiple imputation.

Effect Size and Variability

The effect size is the magnitude of the difference or association that the study aims to detect. The effect size is a key input to the power analysis, and the power is very sensitive to the effect size. A small effect size requires a large sample size, while a large effect size requires a smaller sample size.

The effect size for a longitudinal study can be defined in several ways. For a two-group comparison, the effect size can be the difference in the mean change between the two groups. For a time-by-group interaction, the effect size can be the difference in the slope between the two groups.

The effect size should be based on the clinically or biologically meaningful difference, not on the difference that is statistically detectable. The researcher should specify the effect size in advance, based on the prior studies, the pilot data, or the clinical judgment.

The variability of the data is also a key input to the power analysis. The variability includes the between-subject variability and the within-subject variability. The between-subject variability is the variation in the baseline levels between the subjects. The within-subject variability is the variation in the measurements within a subject over time.

The variability can be estimated from prior studies, pilot data, or the literature. The researcher should use the best available estimate of the variability, and should consider the uncertainty in the estimate. The power analysis can be repeated with a range of variability estimates to assess the sensitivity of the power to the variability.

Correlation Structure

The correlation structure is the pattern of the correlation between the repeated measurements. The correlation structure can be exchangeable, autoregressive, or unstructured.

The exchangeable correlation structure assumes that the correlation between any two time points is the same. This is also called the compound symmetry structure. The exchangeable correlation is often used when the time points are equally spaced and the correlation does not decay over time.

The autoregressive correlation structure assumes that the correlation between two time points decreases as the time between them increases. The correlation is higher for the adjacent time points and lower for the time points that are further apart. The autoregressive structure is often used when the time points are not equally spaced or when the correlation is expected to decay over time.

The unstructured correlation structure allows the correlation to be different for each pair of time points. This is the most flexible structure, but it requires the estimation of many parameters, which can be difficult with a small sample size.

The choice of the correlation structure affects the power analysis. The power is higher when the correlation is lower, because the repeated measurements provide more independent information. The power is lower when the correlation is higher, because the repeated measurements provide less independent information.

The researcher should choose the correlation structure based on the expected pattern of the correlation. The correlation structure can be estimated from the pilot data or the prior studies. If the correlation structure is unknown, the researcher should consider the range of the correlation structures and the power for each.

Practical Workflow for Power Analysis

The following workflow provides a practical approach to power analysis for a longitudinal study.

Step 1: Define the Research Question

The first step is to define the research question and the primary outcome. The research question should be specific and the primary outcome should be clearly defined. The primary outcome is the outcome that is used for the power analysis.

Step 2: Specify the Design

The second step is to specify the design of the study. This includes the number of groups, the number of time points, the spacing of the time points, and the randomization scheme. The design should be specified in advance, and the power analysis should be based on the design.

Step 3: Estimate the Parameters

The third step is to estimate the parameters for the power analysis. This includes the effect size, the variability, the correlation, and the dropout rate. The parameters should be estimated from the prior studies, the pilot data, or the literature.

Step 4: Choose the Analysis Method

The fourth step is to choose the analysis method. The analysis method should be the method that will be used for the primary analysis of the study. The power analysis should be based on the same method.

Step 5: Calculate the Sample Size

The fifth step is to calculate the sample size. The sample size can be calculated using the formulas, the software, or the simulation. The sample size should be calculated for the target power, which is typically 80% or 90%.

Step 6: Assess the Sensitivity

The sixth step is to assess the sensitivity of the power to the assumptions. The power analysis should be repeated with different values of the parameters to see how the power changes. The sensitivity analysis should be reported in the study protocol.

Step 7: Document the Analysis

The seventh step is to document the power analysis. The documentation should include the parameters, the software, the formulas, and the results. The documentation should be included in the study protocol and the final report.

Records and Measurements

The power analysis should be documented in the study protocol. The protocol should include the following information:

  • The research question and the primary outcome
  • The design of the study, including the number of groups, the number of time points, and the spacing
  • The parameters for the power analysis, including the effect size, the variability, the correlation, and the dropout rate
  • The software and the version used for the power analysis
  • The sample size and the power for the sample size
  • The sensitivity analysis, including the range of the parameters and the power for each

The protocol should be reviewed by the research team and the institutional review board. The protocol should be updated if the parameters change during the study.

The power analysis should be reported in the final report. The report should include the parameters, the software, and the results. The report should also include the sensitivity analysis and the limitations of the power analysis.

Common Failure Patterns

There are several common failure patterns in power analysis for longitudinal studies.

Ignoring the Correlation

The most common failure is to ignore the within-subject correlation and treat the repeated measurements as independent. This leads to an overestimate of the effective sample size and an underpowered study. The researcher should always account for the correlation in the power analysis.

Using the Wrong Effect Size

Another common failure is to use the wrong effect size. The effect size should be the clinically or biologically meaningful difference, not the difference that is statistically detectable. The researcher should specify the effect size in advance and justify it.

Ignoring the Dropout

Another common failure is to ignore the dropout. The dropout reduces the effective sample size at the later time points, which reduces the power. The researcher should account for the dropout in the power analysis.

Using the Wrong Correlation Structure

Another common failure is to use the wrong correlation structure. The correlation structure affects the power, and the wrong structure can lead to an underpowered or overpowered study. The researcher should choose the correlation structure based on the expected pattern of the correlation.

Not Performing the Sensitivity Analysis

Another common failure is to not perform the sensitivity analysis. The power analysis is based on the assumptions, and the assumptions may be wrong. The sensitivity analysis should be performed to assess the robustness of the power to the assumptions.

Limitations of Power Analysis

Power analysis has several limitations. The power analysis is based on the assumptions about the parameters, which may be wrong. The power analysis is also based on the assumptions about the analysis method, which may be wrong. The power analysis is a planning tool, not a guarantee of the results.

The power analysis is also limited by the uncertainty in the parameters. The parameters are estimated from the prior studies or the pilot data, which may not be representative of the study population. The power analysis should be repeated with the range of the parameters to assess the sensitivity.

The power analysis is also limited by the complexity of the longitudinal data. The longitudinal data has the correlation structure, the missing data, and the dropout, which are difficult to model. The power analysis may not capture all the complexity of the data.

The power analysis is also limited by the software. The software may not be able to handle the complex designs, and the researcher may need to use the simulation approach. The simulation approach is flexible, but it requires the programming skills and the computational resources.

Reporting and Transparency

The power analysis should be reported in the study protocol and the final report. The reporting should be transparent and complete, including the parameters, the software, and the results. The reporting should follow the reporting guidelines for the study design.

The EQUATOR Network provides the reporting guidelines for the health research, including the guidelines for the reporting of the sample size and the power. The researcher should consult the EQUATOR Network for the reporting guidelines that apply to the study.

The Committee on Publication Ethics provides the core practices for the publication ethics, including the authorship, the peer review, and the data. The researcher should follow the core practices in the reporting of the study.

The National Institutes of Health provides the grant policy and the application for the research. The researcher should follow the NIH policy for the grant application, including the reporting of the power analysis.

The ORCID for Researchers provides the researcher identity and the record maintenance. The researcher should use the ORCID to the record of the research.

The Data Management and Sharing Policy provides the data management and the sharing expectations for the NIH-funded research. The researcher should follow the policy for the data management and the sharing.

Reporting Guidelines

The reporting of the power analysis should follow the reporting guidelines for the study. The EQUATOR Network provides the reporting guidelines for the health research, including the guidelines for the reporting of the sample size and the power. The guidelines include the CONSORT statement for the randomized trials, the STROBE statement for the observational studies, and the PRISMA statement for the systematic reviews.

The CONSORT statement includes the item for the sample size, which requires the reporting of the sample size calculation, the parameters, and the software. The STROBE statement includes the item for the sample size, which requires the reporting of the sample size and the power.

The reporting of the power analysis should include the following information:

  • The primary outcome and the effect size
  • The parameters for the power analysis, including the variability, the correlation, and the dropout rate
  • The software and the version used for the power analysis
  • The sample size and the power
  • The sensitivity analysis, including the results of the sensitivity analysis

The reporting should be transparent and the limitations of the power analysis should be acknowledged.

Data Management and Sharing

The data management and sharing are important for the reproducibility of the research. The NIH Data Management and Sharing Policy requires the NIH-funded researchers to submit the data management and sharing plan with the grant application. The plan should describe the data that will be collected, the data that will be shared, and the data that will be preserved.

The data management and sharing plan should include the following information:

  • The type of the data and the data format
  • The data standards and the metadata
  • The data sharing and the data access
  • The data preservation and the data repository

The data management and sharing plan should be updated as the study progresses. The plan should be reviewed by the NIH and the data should be shared in the timely manner.

The data sharing is important for the reproducibility of the research. The data should be shared in the repository that is accessible to the research community. The data should be documented and the metadata should be provided.

Professional Escalation Criteria

The power analysis should be performed by the researcher with the statistical expertise. If the researcher does not have the statistical expertise, the researcher should consult the biostatistician. The biostatistician can provide the guidance on the power analysis and the analysis method.

The power analysis should be reviewed by the ethics committee and the funding agency. The ethics committee should review the power analysis to ensure that the study is not underpowered and that the subjects are not exposed to the unnecessary risk. The funding agency should review the power analysis to ensure that the study is adequately powered and that the resources are used efficiently.

The power analysis should be updated if the parameters change during the study. The parameters may change if the pilot data is collected or if the prior studies are published. The power analysis should be updated to reflect the new parameters.

The power analysis should be reported in the final report. The report should include the power analysis and the sensitivity analysis. The report should be transparent and the limitations should be acknowledged.

Frequently Asked Questions

What is the difference between a mixed model and GEE for power analysis?

Mixed models and GEE are both used for the analysis of longitudinal data, but they make different assumptions. Mixed models assume the data follows a normal distribution and the random effects are normally distributed. GEE does not require the full distribution of the data, and it uses the working correlation matrix to adjust the standard errors. The power analysis should be based on the analysis method that will be used for the primary analysis.

How do I estimate the correlation for the power analysis?

The correlation can be estimated from the prior studies, the pilot data, or the literature. The correlation is the ICC, which is the ratio of the between-subject variance to the total variance. The correlation can also be estimated from the data of the similar study. If the correlation is uncertain, the power analysis should be repeated with the range of the correlation values.

What is the effect of the dropout on the power?

The dropout reduces the effective sample size at the later time points, which reduces the power. The sample size should be inflated by the expected dropout rate. The dropout rate should be estimated from the prior studies or the pilot data. The dropout may be higher in the early time points or in the one group, and the power analysis should account for the pattern of the dropout.

How many time points do I need for the longitudinal study?

The number of time points depends on the research question and the expected pattern of the change. The time points should be spaced to capture the change of the outcome. The power increases with the number of time points, but the increase is limited by the correlation. The time points should be chosen based on the research question, not the power analysis.

What is the target power for the longitudinal study?

The target power is typically 80% or 90%. The target power should be chosen based on the research question and the resources. The higher power requires the larger sample size, which increases the cost of the study. The target power should be justified in the study protocol.

Can I use G*Power for the longitudinal study?

G*Power can be used for the power analysis of the repeated-measures ANOVA, which is a special case of the mixed model. G*Power does not support the power analysis for the general mixed model or the GEE. For the complex longitudinal design, the researcher should use R or the other specialized software.

What is the sensitivity analysis for the power analysis?

The sensitivity analysis is the repetition of the power analysis with the different values of the parameters. The sensitivity analysis assesses the robustness of the power to the assumptions. The sensitivity analysis should be performed for the parameters that are uncertain, such as the correlation, the dropout rate, and the effect size.

How do I report the power analysis in the study protocol?

The power analysis should be reported in the study protocol with the parameters, the software, and the results. The protocol should include the primary outcome, the effect size, the variability, the correlation, the dropout rate, and the sample size. The protocol should also include the sensitivity analysis and the limitations of the power analysis.

Using the Evidence

SourceBest use in this topicImportant limitation
Research Methods Resourcesofficial guidanceCheck the linked page for current local requirements
EQUATOR Networkofficial guidanceCheck the linked page for current local requirements
Core Practicesofficial guidanceCheck the linked page for current local requirements

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.