Power Analysis for Pilot Studies

By Dr. Zubair Khalid, DVM, MS, PhD ·

Power Analysis for Pilot Studies

Key Takeaways

  • Pilot studies are designed for feasibility assessment and parameter estimation (e.g., standard deviation, preliminary effect size), not for definitive hypothesis testing or demonstrating statistical significance.
  • Pilot-derived effect sizes are subject to substantial sampling error due to small sample sizes, resulting in wide confidence intervals that must be used for planning, not treated as precise estimates.
  • Power analysis for a confirmatory study should utilize a sensitivity analysis across a range of plausible effect sizes, informed by the confidence interval of the pilot estimate and potentially external benchmarks, to determine a robust sample size.
  • The most conservative approach for sample size planning involves using the lower bound of the effect size confidence interval and the upper bound of the standard deviation confidence interval derived from pilot data.
  • Treating a pilot study's statistically significant finding as proof of efficacy is a common failure pattern that leads to underpowered confirmatory studies and potentially erroneous conclusions.
  • When pilot data quality is compromised (e.g., <10 observations/group, inconsistent protocol, high attrition), external benchmarks from literature or prior experiments should be prioritized for effect size estimation in power calculations.

Quick Answer

  • Pilot data can inform effect size estimates for a confirmatory study, but confidence intervals around those estimates are typically wide and should be used to plan a range of sample sizes.
  • Use pilot results to calculate effect sizes with confidence intervals, then run sensitivity analyses across plausible effect sizes to determine the sample size for the main study.
  • Pilot studies are not designed to detect statistically significant differences, and treating them as such will overestimate power and undermine the confirmatory study.

Purpose and Scope of Pilot Studies in Biological Research

Pilot studies occupy a specific position in the research pipeline. They are small-scale investigations conducted before a confirmatory study to test feasibility, refine protocols, estimate variability, and gather preliminary data. In biological research, pilot studies often involve laboratory assays, animal models, or cell culture systems where the cost per observation is high and the biological variability is poorly characterized.

The primary intent of a pilot study is not to test a scientific hypothesis with inferential statistics. Instead, the pilot study answers operational questions: Can the protocol be executed as written? Are the measurements stable across repeated runs? What is the natural variability of the biological system under study? These questions must be answered before committing resources to a larger confirmatory investigation.

Researchers frequently make a conceptual error when they use pilot data to compute a power analysis for the main study. The error is treating the pilot effect size as a precise estimate of the true effect. Because pilot studies have small sample sizes, the observed effect size is subject to substantial sampling error. The confidence interval around a pilot-derived effect size is typically wide, and the true effect may be considerably smaller or larger than the observed value.

The National Library of Medicine Research Methods Resources gateway provides access to authoritative biomedical texts that describe the role of pilot studies in the research process. These resources emphasize that pilot studies are for feasibility and parameter estimation, not for hypothesis testing.

A pilot study should have explicit objectives that are distinct from the confirmatory study. These objectives include assessing recruitment feasibility, testing measurement instruments, estimating standard deviations, and identifying procedural problems. When the pilot study is framed this way, the statistical analysis of pilot data becomes a parameter estimation exercise instead of a hypothesis test.

The distinction between pilot and confirmatory studies matters for the interpretation of results. A pilot study that shows a promising trend but no statistically significant difference has not failed. It has provided information about variability and effect direction that can inform the design of the main study. Conversely, a pilot study that shows a statistically significant difference has not proven the hypothesis, because the small sample size makes the estimate unstable.

Statistical Foundations of Power Analysis

Power analysis is a calculation that determines the probability of detecting a true effect of a specified size, given a particular sample size, significance level, and variability. The four components of a power calculation are the effect size, the sample size, the significance level (alpha), and the power (1 minus beta, where beta is the probability of a type II error).

The relationship among these components is deterministic. If any three are specified, the fourth can be calculated. In practice, researchers specify the effect size they want to detect, choose an acceptable alpha (commonly 0.05), select a target power (commonly 0.80), and calculate the required sample size.

The effect size is the most difficult component to specify because it depends on the biological system and the measurement method. For continuous outcomes, the standardized effect size is often expressed as Cohen's d, which is the difference between group means divided by the pooled standard deviation. For categorical outcomes, the effect size may be expressed as a risk ratio, odds ratio, or difference in proportions.

The standard deviation is a critical input because it appears in the denominator of the effect size calculation. If the standard deviation is underestimated, the effect size is overestimated, and the required sample size is too small. Pilot data can provide an estimate of the standard deviation, but this estimate has its own uncertainty.

The National Institutes of Health Grants and Funding pages describe how grant reviewers evaluate the statistical justification for proposed sample sizes. Reviewers expect a clear explanation of the effect size used, the source of that effect size, and the assumptions underlying the power calculation. A power analysis that relies on a single pilot-derived effect size without acknowledging uncertainty is a common weakness in grant applications.

Why Pilot Data Overestimates Power

The overestimation of power from pilot data is a well-recognized problem in biostatistics. The mechanism is straightforward. A pilot study with a small sample size produces an effect size estimate that is subject to large sampling error. When this estimate is used in a power calculation, the resulting sample size is too small to detect the true effect with the desired probability.

The problem is compounded by the tendency to select the pilot effect size that is most favorable. If a researcher runs a pilot study and observes a large effect, they may be encouraged to proceed with a confirmatory study. If the pilot effect is small, the researcher may abandon the line of investigation or modify the protocol. This selection process means that the effect sizes that proceed to confirmatory studies are systematically larger than the true effect.

The confidence interval around a pilot effect size provides a way to quantify the uncertainty. A 95 percent confidence interval for the effect size from a pilot study with a small sample will be wide. The lower bound of the confidence interval represents a plausible but less favorable effect size. Using the lower bound in a power calculation produces a larger sample size that is more likely to achieve the target power.

The EQUATOR Network provides reporting guidelines that emphasize transparent reporting of how sample sizes were determined. These guidelines encourage researchers to report the effect size used in the power calculation, the source of that effect size, and the assumptions made. Transparent reporting allows readers to assess whether the power analysis is credible.

Estimating Effect Sizes with Confidence Intervals

The correct approach to using pilot data for power analysis begins with estimating the effect size and its confidence interval. For a continuous outcome, the standardized effect size is calculated as the difference between group means divided by the pooled standard deviation. The confidence interval for this effect size depends on the sample size and the variability of the data.

The confidence interval for the effect size can be calculated using the noncentral t distribution or using bootstrap methods. The noncentral t distribution is appropriate when the data are approximately normally distributed. Bootstrap methods are more flexible and can be used when the data are skewed or when the effect size is a more complex statistic.

The width of the confidence interval is inversely related to the sample size. A pilot study with 10 animals per group will produce a much wider confidence interval than a pilot study with 30 animals per group. The researcher should report the confidence interval alongside the point estimate of the effect size.

The practical implication is that the power analysis for the confirmatory study should be conducted across the range of plausible effect sizes defined by the confidence interval. The researcher should calculate the required sample size for the lower bound, the point estimate, and the upper bound of the effect size confidence interval. The sample size for the confirmatory study should be based on the lower bound of the confidence interval, because this is the most conservative plausible effect size.

The National Library of Medicine Research Methods Resources gateway provides access to biostatistics texts that describe the calculation of confidence intervals for effect sizes and the use of these intervals in sample size planning.

Planning the Confirmatory Study

The confirmatory study is the investigation that will provide the definitive test of the scientific hypothesis. It must be designed with sufficient power to detect the effect size of interest with a high probability. The sample size for the confirmatory study is determined by the power analysis, which requires specification of the effect size, alpha, and target power.

The effect size for the confirmatory study should be chosen based on scientific relevance, not on the pilot point estimate alone. The researcher should consider what effect size is biologically meaningful and worth detecting. This is a scientific judgment that should be made before the pilot data are analyzed to avoid bias.

The pilot data provide an estimate of the standard deviation, which is a key input to the power calculation. The standard deviation estimate from the pilot study should be used with its own confidence interval. The upper bound of the standard deviation confidence interval is the conservative choice for sample size planning, because a larger standard deviation requires a larger sample size.

The power analysis should be conducted using a sensitivity analysis approach. The researcher should calculate the required sample size for a range of effect sizes and standard deviations. This produces a table or graph that shows how the sample size changes as the assumptions change. The final sample size should be chosen to provide adequate power across the plausible range of assumptions.

The EQUATOR Network reporting guidelines for clinical trials and observational studies require that the sample size calculation be described in the methods section. The description should include the effect size, the standard deviation, the alpha, the power, and the software used for the calculation.

Practical Workflow for Pilot Data Analysis

The following workflow provides a structured approach to using pilot data for power analysis. This workflow is applicable to biological research settings including laboratory experiments, animal studies, and field trials.

Step 1: Define the Primary Outcome and Effect Size Metric

The primary outcome is the measurement that will be used to test the main hypothesis. It should be defined before the pilot study is conducted. The effect size metric depends on the type of outcome. For continuous outcomes, the effect size is typically the standardized mean difference. For binary outcomes, the effect size is the difference in proportions or the odds ratio.

The choice of effect size metric affects the power analysis. The standardized mean difference is used for continuous outcomes and is calculated as the difference between group means divided by the pooled standard deviation. The difference in proportions is used for binary outcomes and is calculated as the difference between the event rates in the two groups.

Step 2: Collect Pilot Data

The pilot study should be conducted according to the same protocol that will be used in the confirmatory study. The sample size for the pilot study should be sufficient to provide a reasonable estimate of the standard deviation and the effect size. There is no universal rule for the pilot sample size, but a common recommendation is to use at least 10 to 20 observations per group for a pilot study.

The pilot data should be recorded with the same rigor as the confirmatory study. The data should include the primary outcome, any covariates that will be used in the analysis, and the identifiers for each sample. The data should be stored in a format that can be imported into statistical software.

Step 3: Calculate the Effect Size and Confidence Interval

The effect size is calculated from the pilot data using the appropriate formula for the outcome type. The confidence interval for the effect size is calculated using the appropriate statistical method. For a continuous outcome, the confidence interval for the standardized mean difference is calculated using the noncentral t distribution. For a binary outcome, the confidence interval for the difference in proportions is calculated using the normal approximation or the exact method.

The confidence interval should be reported alongside the effect size. The width of the confidence interval is a measure of the uncertainty in the effect size estimate. A wide confidence interval indicates that the pilot study provides little information about the true effect size.

Step 4: Conduct Sensitivity Analysis

The sensitivity analysis is the core of the power analysis. The sample size for the confirmatory study is calculated for a range of effect sizes and standard deviations. The range of effect sizes should include the point estimate from the pilot and the lower and upper bounds of the confidence interval. The range of standard deviations should include the point estimate and the upper bound of the confidence interval.

The sensitivity analysis produces a table of sample sizes. The researcher should examine the table to understand how the sample size changes as the assumptions change. The final sample size should be chosen to provide adequate power across the plausible range of assumptions.

Step 5: Select the Sample Size for the Confirmatory Study

The sample size for the confirmatory study is selected based on the sensitivity analysis. The most conservative approach is to use the sample size that provides the target power for the lower bound of the effect size confidence interval and the upper bound of the standard deviation confidence interval. This approach ensures that the confirmatory study has adequate power even if the true effect size is smaller than the pilot estimate.

The selected sample size should be feasible within the resources available. If the conservative sample size is too large, the researcher may need to reconsider the study design, the primary outcome, or the target effect size. The tradeoff between power and feasibility should be documented.

Step 6: Document the Power Analysis

The power analysis should be documented in the study protocol and in the grant application. The documentation should include the effect size, the standard deviation, the alpha, the target power, the software used, and the assumptions made. The documentation should also include the sensitivity analysis table.

The NIH Grants and Funding page describes the expectations for the statistical design section of a grant application. The power analysis should be described in enough detail that a reviewer can reproduce the calculation.

At a Glance

Decision PointPilot Study ApproachConfirmatory Study ApproachKey Limitation
Effect size estimationCalculate point estimate from pilot dataUse lower bound of confidence interval for conservative planningPilot effect size has wide confidence interval
Standard deviationEstimate from pilot dataUse upper bound of confidence intervalPilot standard deviation may be underestimated
Power calculationNot appropriate for hypothesis testingCalculate sample size for target powerPower depends on assumptions that are uncertain
Sample sizeSmall, for feasibilityLarger, for adequate powerResource constraints may limit feasibility

Common Failure Patterns in Pilot Power Analysis

Researchers encounter several recurring problems when using pilot data for power analysis. Recognizing these patterns can help avoid the most common errors.

Failure Pattern 1: Treating the Pilot Effect Size as the True Effect Size

The most common error is to use the pilot effect size as the effect size in the power analysis without accounting for uncertainty. This produces a sample size that is too small to detect the true effect. The consequence is a confirmatory study that is underpowered and may fail to detect a real effect.

The solution is to use the confidence interval around the effect size. The power analysis should be conducted across the range of plausible effect sizes, and the sample size should be chosen to provide adequate power for the lower bound of the confidence interval.

Failure Pattern 2: Using the Pilot Study to Test the Hypothesis

Some researchers conduct a pilot study and then perform a hypothesis test on the pilot data. If the pilot study does not show a statistically significant difference, the researcher concludes that the effect does not exist and abandons the line of study. This is a misuse of the pilot study.

The pilot study is not designed to test the hypothesis. The sample size is too small to detect a meaningful effect. A non-significant result in a pilot study does not mean that the effect is absent. It means that the pilot study was not powered to detect the effect.

Failure Pattern 3: Underestimating the Standard Deviation

The standard deviation is a critical parameter in the power analysis. If the standard deviation is underestimated, the effect size is overestimated, and the required sample size is too small. The pilot study provides an estimate of the standard deviation, but this estimate is subject to sampling error.

The solution is to use the upper bound of the confidence interval for the standard deviation in the power analysis. This provides a conservative sample size that is more likely to achieve the target power.

Failure Pattern 4: Ignoring the Confidence Interval

The confidence interval around the effect size is a measure of the uncertainty in the estimate. Ignoring the confidence interval and using only the point estimate is a common error. The confidence interval should be reported and used in the power analysis.

The width of the confidence interval is determined by the pilot sample size. A small pilot study produces a wide confidence interval, which means that the effect size is not well estimated. The power analysis should reflect this uncertainty.

Failure Pattern 5: Using a Single Power Calculation

A single power calculation based on a single effect size and a single standard deviation is not sufficient. The power analysis should be a sensitivity analysis that examines the sample size across a range of plausible assumptions. This provides a more complete picture of the relationship between the assumptions and the required sample size.

The sensitivity analysis should be documented in the study protocol. The final sample size should be chosen to be robust across the plausible range of assumptions.

Records and Measurements for Pilot Studies

The quality of the power analysis depends on the quality of the pilot data. The following records and measurements should be collected during the pilot study.

Primary Outcome Measurements

The primary outcome should be measured with the same protocol that will be used in the confirmatory study. The measurement should be recorded for each sample in the pilot study. The measurement method should be described in the study protocol.

Variability Measurements

The variability of the primary outcome is the key parameter for the power analysis. The standard deviation of the primary outcome should be calculated from the pilot data. The confidence interval for the standard deviation should also be calculated.

Protocol Compliance Records

The pilot study should record the number of samples that were successfully processed, the number of samples that were lost or excluded, and the reasons for exclusion. This information is used to assess the feasibility of the protocol and to estimate the attrition rate for the confirmatory study.

Cost and Time Records

The pilot study should record the cost per sample and the time required to process each sample. This information is used to estimate the total cost and time for the confirmatory study. The cost and time estimates are used to assess the feasibility of the sample size selected in the power analysis.

Data Quality Records

The pilot study should record the quality of the data, including the number of missing values, the number of outliers, and the number of technical failures. This information is used to assess the reliability of the measurement method and to plan for data quality issues in the confirmatory study.

Quality Controls and Reproducibility

The pilot study should be conducted with the same quality controls that will be used in the confirmatory study. This ensures that the variability estimated in the pilot study is representative of the variability in the confirmatory study.

Standard Operating Procedures

The pilot study should follow a written standard operating procedure for each step of the protocol. The standard operating procedure should describe the equipment, the reagents, the timing, and the recording of the measurements. The standard operating procedure should be followed exactly during the pilot study.

Calibration and Quality Controls

The equipment used in the pilot study should be calibrated according to the manufacturer's specifications. The quality controls should be run at the beginning and the end of each batch of samples. The quality control results should be recorded and reviewed.

Replication

The pilot study should include technical replicates to assess the variability of the measurement method. The technical replicates are repeated measurements of the same sample. The variability of the technical replicates is a component of the total variability of the primary outcome.

Data Recording

The data should be recorded in a laboratory notebook or an electronic data capture system. The data should be recorded in a way that is traceable to the sample and the measurement run. The data should be backed up and stored in a secure location.

Reproducibility

The pilot study should be reproducible. This means that another researcher should be able to repeat the pilot study using the standard operating procedure and obtain similar results. The standard operating procedure should be detailed enough to allow this.

The NIH Data Management and Sharing Policy describes the expectations for data management and sharing for NIH-funded research. The policy requires that data be managed and shared in a way that is consistent with the scientific integrity and the reproducibility of the research.

Reporting the Power Analysis

The power analysis should be reported in the study protocol, the grant application, and the final manuscript. The reporting should be transparent and complete, so that a reviewer can assess the credibility of the sample size.

Reporting the Effect Size

The effect size used in the power analysis should be reported with its confidence interval. The source of the effect size should be described. If the effect size is from the pilot study, the pilot sample size and the confidence interval should be reported.

Reporting the Standard Deviation

The standard deviation used in the power analysis should be reported with its confidence interval. The source of the standard deviation should be described. If the standard deviation is from the pilot study, the pilot sample size and the confidence interval should be reported.

Reporting the Sensitivity Analysis

The sensitivity analysis should be reported as a table or a graph. The table should show the sample size for each combination of effect size and standard deviation. The graph should show the power as a function of the sample size for each effect size.

Reporting the Software

The software used for the power analysis should be reported. The software and the version should be identified. The specific procedure or function used for the calculation should be described.

Reporting the Assumptions

The assumptions underlying the power analysis should be reported. This includes the alpha level, the target power, the type of test, and the direction of the test. The assumptions should be justified in the context of the research question.

The EQUATOR Network provides reporting guidelines that describe the expectations for reporting the sample size calculation in the manuscript. The guidelines require that the sample size calculation be described in the methods section with the effect size, the standard deviation, the alpha, the power, and the software.

Limitations of Pilot Data for Power Analysis

The use of pilot data for power analysis has inherent limitations that should be acknowledged in the research plan.

Small Sample Size

The pilot study has a small sample size, which means that the effect size and the standard deviation are estimated with low precision. The confidence intervals are wide, and the point estimates may be far from the true values.

Selection Bias

The pilot study may be subject to selection bias. The samples in the pilot study may not be representative of the population that will be studied in the confirmatory study. This can lead to an effect size estimate that is not generalizable.

Protocol Differences

The protocol used in the pilot study may differ from the protocol used in the confirmatory study. The differences in the protocol can change the variability of the primary outcome. The standard deviation estimated in the pilot study may not be applicable to the confirmatory study.

Measurement Error

The measurement error in the pilot study may be different from the measurement error in the confirmatory study. The measurement error is a component of the total variability. If the measurement error is different, the standard deviation estimate is not applicable.

The Pilot Study Is Not a Miniature Confirmatory Study

The pilot study is not a miniature version of the confirmatory study. The pilot study is designed to answer feasibility questions. The confirmatory study is designed to test the hypothesis. The two studies have different objectives and different statistical designs.

The Committee on Publication Ethics Core Practices describes the expectations for the ethical conduct of research, including the reporting of the research methods. The power analysis should be reported honestly and transparently, without overstating the precision of the estimates.

Professional Escalation Criteria

The following criteria indicate that the power analysis should be reviewed by a biostatistician or a senior researcher.

When the Confidence Interval Is Very Wide

If the confidence interval for the effect size is so wide that the lower bound is near zero or the upper bound is implausibly large, the pilot study has not provided a useful estimate of the effect size. A biostatistician should be consulted to determine whether the pilot study should be repeated with a larger sample size or whether the study should proceed with a different approach.

When the Required Sample Size Is Not Feasible

If the sample size required for the confirmatory study is not feasible within the available resources, a biostatistician should be consulted. The biostatistician can help to explore alternative study designs, such as a paired design, a crossover design, or a sequential design, that may reduce the required sample size.

When the Assumptions Are Uncertain

If the assumptions underlying the power analysis are uncertain, such as the distribution of the outcome or the correlation between the repeated measurements, a biostatistician should be consulted. The biostatistician can help to assess the sensitivity of the power analysis to the assumptions.

When the Pilot Study Has a High Attrition Rate

If the pilot study has a high attrition rate, the confirmatory study may have a similar attrition rate. The attrition rate should be accounted for in the sample size calculation. A biostatistician should be consulted to determine the appropriate adjustment for the attrition.

When the Data Are Not Normally Distributed

If the pilot data are not normally distributed, the power analysis may need to use a different method. A biostatistician should be consulted to determine the appropriate power analysis for the non-normal data.

Decision Framework for Choosing Between Pilot Effect Sizes and External Benchmarks

Researchers face a practical decision when planning a confirmatory study: whether to anchor the power analysis on pilot-derived effect sizes or on external benchmarks from published literature, prior laboratory records, or established biological standards. The choice matters because pilot data and external evidence carry different strengths and weaknesses, and the optimal decision depends on the maturity of the research line, the quality of the pilot data, and the availability of comparable studies.

Decision Rule 1: Assess Pilot Data Quality Before Use

The first decision is whether the pilot data are suitable for effect size estimation at all. Apply three quality checks before considering the pilot effect size as an input to the power analysis.

Check the pilot sample size per group. If the pilot has fewer than 10 observations per group, the confidence interval for the effect size will be so wide that the lower bound may approach zero or reverse direction. In this case, the pilot effect size carries little information beyond the standard deviation estimate.

Check the protocol consistency. The pilot must have been run under the same conditions planned for the confirmatory study. If the pilot used different equipment, different operators, different animal strains, or different reagent lots, the variability estimate may not transfer to the confirmatory setting.

Check the attrition and missing data pattern. If more than 20 percent of pilot samples were lost or excluded, the remaining data may be a biased subset. The National Library of Medicine Research Methods Resources gateway describes how missing data and attrition affect the generalizability of parameter estimates from preliminary studies.

If any of these three checks fail, treat the pilot effect size as a secondary input and prioritize external anchors for the power analysis.

Decision Rule 2: Compare Pilot Effect Size with External Benchmarks

When the pilot passes the quality checks, compare the pilot effect size with external benchmarks from published literature, prior experiments in the same laboratory, or established biological expectations. The comparison serves as a plausibility check.

Calculate the pilot effect size with its confidence interval. Then identify external benchmarks from studies that used a similar outcome measure, similar population, and similar intervention. Record the effect size and the sample size from each external source.

If the pilot point estimate falls inside the range of external benchmarks, the pilot estimate is plausible and can be used as one input in the sensitivity analysis. If the pilot point estimate falls outside the range of external benchmarks, investigate the cause before proceeding. The discrepancy may indicate a protocol difference, a measurement error, or a genuinely different biological context.

The EQUATOR Network reporting guidelines emphasize that the source of the effect size should be transparent. When the pilot effect differs from external benchmarks, the researcher should document the discrepancy and justify the choice of effect size for the power analysis.

Decision Rule 3: Use the More Conservative Input for Sample Size Planning

When the pilot effect size and the external benchmark disagree, the sample size for the confirmatory study should be based on the smaller effect size, which is the more conservative choice. A smaller effect size requires a larger sample size to achieve the same power. This approach protects against overestimating power.

The same logic applies to the standard deviation. If the pilot standard deviation is smaller than the external benchmark, use the larger standard deviation in the power calculation. A larger standard deviation requires a larger sample size.

The NIH Grants and Funding pages describe how reviewers evaluate the justification for the effect size and standard deviation in grant applications. Reviewers expect the researcher to explain why the chosen values are appropriate and to acknowledge the uncertainty in the estimates.

Decision Rule 4: Weight the Inputs by Their Reliability

When both pilot data and external benchmarks are available, the power analysis should not rely on a single input. The sensitivity analysis should include the pilot point estimate, the pilot confidence interval bounds, the external benchmark, and the range of external benchmarks.

The final sample size should be selected to provide adequate power across the full range of plausible effect sizes. If the pilot and external benchmarks are consistent, the range is narrow and the sample size is well determined. If the pilot and external benchmarks disagree, the range is wide and the sample size may be large.

The Committee on Publication Ethics Core Practices describe the expectation that researchers report the basis for their sample size decisions honestly. The power analysis should not be presented as a single deterministic calculation but as a range of plausible scenarios.

Decision Rule 5: Document the Decision Process

The decision framework should be documented in the study protocol and the grant application. The documentation should include the pilot quality checks, the external benchmarks, the comparison between the pilot and external estimates, and the rationale for the chosen effect size and standard deviation.

The documentation should also include the sensitivity analysis table that shows the sample size for each combination of effect size and standard deviation. This table allows the reviewer to see how the sample size changes as the assumptions change.

The NIH Data Management and Sharing Policy describes the expectations for documenting the research process, including the statistical design. The decision framework for the power analysis should be part of the study documentation.

Practical Implementation Steps

The following steps provide a structured approach to applying the decision framework.

Step 1: Record the pilot quality metrics. Record the pilot sample size per group, the protocol version, the attrition rate, and the missing data pattern. Apply the three quality checks.

Step 2: Identify external benchmarks. Search the literature for studies with a similar outcome, population, and treatment. Record the effect size and the standard deviation from each benchmark. If no external benchmarks exist, document this and rely on the pilot data with a wider confidence interval.

Step 3: Compare the pilot and external estimates. Plot the pilot effect size with its confidence interval against the external benchmarks. Record whether the pilot estimate falls within the external range.

Step 4: Select the conservative inputs. Use the smaller effect size and the larger standard deviation from the available sources for the primary power calculation.

Step 5: Run the sensitivity analysis. Calculate the sample size for the pilot point estimate, the pilot confidence interval bounds, the external benchmark, and the external range. Record the sample size for each scenario.

Step 6: Select the final sample size. Choose the sample size that provides the target power across the plausible range of inputs. Document the decision and the rationale.

Common Failure Patterns in the Decision Framework

Failure Pattern 1: Ignoring External Benchmarks. Some researchers rely exclusively on the pilot effect size and ignore the published literature. This can lead to an overestimated effect size and an underpowered confirmatory study. The external benchmarks provide a check on the plausibility of the pilot estimate.

Failure Pattern 2: Cherry-Picking the Favorable Benchmark. When external benchmarks vary, some researchers select the largest effect size to justify a smaller sample size. This is a form of selection bias. The power analysis should use the range of external benchmarks and the conservative end of the range.

Failure Pattern 3: Treating the Pilot as the Only Source of Variability. The pilot standard deviation is an estimate with its own uncertainty. The external benchmarks may provide a more stable estimate of the variability. The power analysis should consider both sources.

Failure Pattern 4: Failing to Document the Decision. The decision framework is only useful if it is documented. The study protocol and the grant application should include the decision tree, the external benchmarks, and the sensitivity analysis.

When to Escalate to a Biostatistician

The decision framework should be escalated to a biostatistician when the pilot and external benchmarks disagree substantially, when the external benchmarks are not available, or when the sensitivity analysis produces a sample size that is not feasible. A biostatistician can help to interpret the discrepancy, to identify additional external sources, or to explore alternative study designs that reduce the required sample size.

The ORCID for Researchers page describes how researchers can maintain a record of their work, including the methods and the data. The decision framework and the sensitivity analysis should be recorded in the research record to support the transparency of the power analysis.

Frequently Asked Questions

What is the main purpose of a pilot study?

The main purpose of a pilot study is to test the feasibility of the protocol, estimate the variability of the primary outcome, and refine the procedures for the confirmatory study. The pilot study is not designed to test the scientific hypothesis or to provide a definitive estimate of the effect size.

How large should a pilot study be?

There is no universal rule for the pilot sample size. The pilot study should be large enough to provide a reasonable estimate of the standard deviation and the variability of the primary outcome. A common recommendation is to use at least 10 to 20 samples per group, but the appropriate size depends on the variability of the outcome and the cost of the samples.

Can I use the pilot effect size directly in the power analysis?

You can use the pilot effect size in the power analysis, but you should not use the point estimate alone. The confidence interval around the effect size should be used to conduct a sensitivity analysis. The sample size for the confirmatory study should be based on the lower bound of the confidence interval to avoid overestimating the power.

What is the difference between a pilot study and a confirmatory study?

A pilot study is a small-scale investigation designed to test the feasibility of the protocol and to estimate the variability of the outcome. A confirmatory study is a large-scale investigation designed to test the scientific hypothesis with adequate power. The two studies have different objectives and different statistical designs.

What should I do if the pilot study does not show a statistically significant difference?

A non-significant result in a pilot study does not mean that the effect does not exist. The pilot study is not powered to detect a meaningful effect. You should use the pilot data to estimate the effect size and the variability, and then use the power analysis to determine the sample size for the confirmatory study.

How do I account for the uncertainty in the pilot effect size?

The uncertainty in the pilot effect size is quantified by the confidence interval. The power analysis should be conducted across the range of effect sizes defined by the confidence interval. The sample size for the confirmatory study should be chosen to provide adequate power for the lower bound of the confidence interval.

What is a sensitivity analysis in the context of power analysis?

A sensitivity analysis is a set of power calculations that examine how the required sample size changes as the assumptions change. The assumptions include the effect size, the standard deviation, the alpha, and the power. The sensitivity analysis produces a table or a graph that shows the sample size for each combination of assumptions.

When should I consult a biostatistician?

You should consult a biostatistician when the confidence interval for the effect size is very wide, when the required sample size is not feasible, when the assumptions are uncertain, when the pilot study has a high attrition rate, or when the pilot data are not normally distributed. A biostatistician can help you design the confirmatory study and interpret the pilot data.

Using the Evidence

SourceBest use in this topicImportant limitation
Research Methods Resourcesofficial guidanceCheck the linked page for current local requirements
EQUATOR Networkofficial guidanceCheck the linked page for current local requirements
Core Practicesofficial guidanceCheck the linked page for current local requirements

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.