Interpreting P-Values and Confidence Intervals in Veterinary Research

By Dr. Zubair Khalid, DVM, MS, PhD ·

Interpreting P-Values and Confidence Intervals in Veterinary Research

Key Takeaways

  • The p-value quantifies the probability of observing data at least as extreme as the study's results, assuming the null hypothesis is true; it is not the probability that the null hypothesis is correct. A statistically significant result (conventionally p < 0.05) indicates that the observed data are unlikely under the null hypothesis, not that the intervention is clinically important.
  • Confidence intervals provide a range of plausible values for the true effect size, conveying both the magnitude of the effect and its precision. A narrow interval around a clinically meaningful effect supports a conclusion of real importance, whereas a wide interval indicates substantial uncertainty, regardless of the p-value.
  • Statistical significance does not equate to clinical significance; a large sample size can yield a statistically significant p-value for a trivial effect, while an underpowered study may fail to detect a clinically important effect. Clinical relevance must be assessed by comparing the effect size and its confidence interval to established thresholds for meaningful outcomes (e.g., economic thresholds in production animals, survival extension in oncology).
  • Non-significant results in underpowered studies should not be interpreted as evidence of "no effect." Instead, they indicate insufficient evidence to reject the null hypothesis, particularly when the confidence interval is wide and encompasses clinically relevant values.
  • Reporting guidelines such as ARRIVE and CONSORT mandate the reporting of effect sizes with confidence intervals, not p-values alone, to facilitate accurate interpretation of study findings and to avoid common misinterpretations like dichotomizing results into "significant" or "non-significant."
  • When appraising research, consider the study's power and sample size justification, the presence of bias and confounding, and the context of multiple comparisons. A non-significant p-value from a well-powered study with a narrow confidence interval around the null provides meaningful evidence against a clinically important effect, whereas a non-significant result from a small, variable study is largely uninformative.

This article explains how to read and interpret p-values and confidence intervals in veterinary studies, with emphasis on the distinction between statistical and clinical significance. It serves veterinary researchers, graduate students, and clinicians who appraise the primary literature. The content addresses a recurring failure mode in veterinary and related biomedical research: treating a p-value threshold as a binary verdict on whether an intervention works, when the p-value is only one element of a broader inferential picture.

The article covers the formal definition of the p-value, the logic of null hypothesis significance testing, the construction and interpretation of confidence intervals, common reporting errors, and practical guidance for applying these concepts to study appraisal. Bayesian alternatives are excluded from scope. The guidance applies across species and study designs, from controlled trials in companion animals to observational studies in production medicine and in vitro experiments.

At a Glance

ParameterCorrect interpretationCommon error
P-valueProbability of observing data at least as extreme as obtained, assuming the null hypothesis is trueTreated as the probability that the null hypothesis is true
Statistical significanceA decision rule based on a pre-specified alpha, conventionally 0.05Equated with clinical importance
Confidence intervalA range of plausible values for the true effect, generated by a procedure that captures the true value in 95% of repetitionsRead as a 95% probability that the true effect lies within the observed interval
Confidence interval widthReflects precision, narrower intervals indicate more precise estimatesIgnored when the interval excludes the null
Non-significant resultInsufficient evidence to reject the null, does not prove no effectReported as "no effect"
Effect sizeThe magnitude of the observed difference or associationOmitted entirely from reporting
Sample size and powerDetermine whether a study can detect a meaningful effectNot considered when interpreting a non-significant p-value

The Definition and Logic of the P-Value

The p-value is the probability of obtaining a test statistic at least as extreme as the one observed, conditional on the null hypothesis being true. It is not the probability that the null hypothesis is correct, nor the probability that the observed result arose by chance alone. The distinction matters because the p-value conditions on the null, not on the data actually observed. A small p-value indicates that the data are unusual under the null model, which shifts evidential weight toward the alternative, but it does not quantify the probability of either hypothesis.

Null hypothesis significance testing combines this p-value with a pre-specified alpha level, conventionally 0.05. When the p-value falls below alpha, the result is declared statistically significant and the null hypothesis is rejected. When the p-value exceeds alpha, the null is not rejected. This decision procedure has well-documented limitations. The threshold is arbitrary, the binary outcome discards information about the strength of evidence, and the procedure is frequently misreported in the literature. A review of animal cognition research found that non-significant results were commonly described as "no effect" in titles and abstracts, a phrasing that overstates what a failure to reject the null can support, particularly in studies with low statistical power Farrar et al., reporting and interpreting non-significant results in animal cognition research.

The same inferential logic applies across veterinary research settings. In an in vitro fermentation experiment evaluating methane mitigation strategies, for example, the p-value for a treatment effect must be interpreted in light of the technical variability inherent to the system, including inoculum source, donor animal diet, and incubation conditions, all of which influence the precision of the estimate Yáñez-Ruiz et al., design and interpretation of in vitro batch culture experiments. A statistically significant difference in gas production between treatments does not guarantee that the difference is biologically meaningful at the rumen level, and a non-significant result does not establish that two treatments are equivalent.

Confidence Intervals as Estimates of Precision

A confidence interval provides a range of values for the population parameter of interest, constructed so that the procedure, repeated across many hypothetical samples, captures the true parameter in a specified proportion of repetitions. A 95% confidence interval does not mean there is a 95% probability that the true effect lies within the observed interval. The true parameter is fixed, the interval is random. The confidence level describes the long-run performance of the method, not the posterior probability of any single interval.

Confidence intervals convey two pieces of information that a p-value cannot. The location of the interval indicates the most plausible magnitude of the effect, and the width indicates precision. A narrow interval around a clinically meaningful effect supports a conclusion that the effect is real and worth acting on. A wide interval that includes both the null and clinically important values indicates genuine uncertainty, regardless of whether the p-value crosses the 0.05 threshold.

Statistical Significance versus Clinical Significance

Statistical significance answers a narrow question: is the observed effect unlikely under the null hypothesis? Clinical significance asks whether the effect matters for patient outcomes, production parameters, or population health. A study with a large sample size can produce a statistically significant p-value for a trivial effect, because the standard error shrinks as sample size grows. Conversely, a small study can fail to reach statistical significance for an effect that would be clinically important if real.

The distinction is particularly relevant in veterinary medicine, where effect sizes that are statistically detectable may be too small to justify the cost, risk, or effort of an intervention. The converse also occurs: a non-significant result in an underpowered study should not be read as evidence that a treatment lacks value. The confidence interval is the most direct tool for assessing clinical significance, because it shows the range of effect sizes consistent with the data. If the entire interval lies above the threshold for a clinically meaningful effect, the evidence supports clinical significance even if the p-value is marginal. If the interval includes both trivial and meaningful effects, the study is simply uninformative on that question.

The Problem of Underpowered Studies

Statistical power is the probability of rejecting the null hypothesis when a specified alternative is true. Studies with low power frequently produce non-significant p-values even when a real effect exists. The p-value distribution in animal cognition research is consistent with widespread low power, meaning that many published non-significant results may be false negatives Farrar et al., reporting and interpreting non-significant results in animal cognition research. Veterinary research faces similar constraints, particularly in fields where animal numbers are limited by cost, ethics, or availability.

When appraising a non-significant result, the reader should ask whether the study had adequate power to detect the effect of interest. This requires knowing the effect size the study was designed to detect, the variability of the outcome, and the sample size. Reporting guidelines such as the ARRIVE guidelines for animal research require specification of sample size estimation and the primary outcome, which supports this appraisal ARRIVE guidelines for reporting animal research. A non-significant result from a study with adequate power and a narrow confidence interval around the null provides meaningful evidence against a clinically important effect. A non-significant result from a small, variable study provides almost no information.

Reporting Standards and the Role of Effect Sizes

The EQUATOR Network maintains a library of reporting guidelines that cover the major study designs used in veterinary research, including CONSORT for randomised trials, PRISMA for systematic reviews, STROBE for observational studies, and REFLECT for livestock trials EQUATOR Network reporting guidelines. These guidelines converge on a common requirement: report effect sizes with confidence intervals, not p-values alone. A p-value without an accompanying effect size and confidence interval leaves the reader unable to judge clinical significance.

The ARRIVE guidelines similarly require that results be reported with measures of precision, and that the interpretation of results address both the size of the effect and its uncertainty ARRIVE guidelines for reporting animal research. Adherence to these standards is variable, but the reader can use them as a checklist when evaluating whether a study has reported the information needed for sound interpretation.

Implications for Meta-Analysis and Systematic Review

The interpretation of individual study results feeds directly into meta-analysis, where the unit of analysis is the effect estimate and its confidence interval, not the binary significant or non-significant classification. A meta-analysis that pools only studies with statistically significant results introduces selection bias and overestimates the true effect. The same logic applies to narrative reviews that cite only significant findings. The confidence interval from each study contributes to the pooled estimate, and studies with wide intervals receive less weight in a properly conducted meta-analysis.

Practical Interpretation in Study Design and Review

Reading a P-Value in Context: The Four-Step Assessment

When you encounter a p-value in a veterinary study, work through a fixed sequence before deciding what the result means for your patients or your next experiment.

First, identify the pre-specified hypothesis and the primary outcome. A p-value attached to a secondary or exploratory endpoint carries less evidential weight, particularly when many endpoints were tested. The risk of false positives accumulates across multiple comparisons, and a single significant p-value among dozens of tests may reflect chance alone.

Second, examine the effect size and its confidence interval. A p-value below 0.05 with a narrow confidence interval that excludes the null value supports a real effect. A p-value below 0.05 with a wide confidence interval that barely excludes the null value supports a statistically detectable effect of uncertain magnitude. The confidence interval tells you the range of plausible effect sizes consistent with the data, and that range is what you should carry into clinical reasoning.

Third, assess the study's power and sample size justification. Studies in veterinary science are frequently underpowered to detect clinically meaningful effects, and non-significant results from such studies are often misinterpreted as evidence of no effect. This pattern is well documented in animal cognition research, where low statistical power produces non-significant p-values even when the null hypothesis is false, and where authors frequently report such results as "no effect" in titles and abstracts. The same logic applies across veterinary disciplines. A non-significant p-value from an underpowered study is uninformative, not exculpatory.

Fourth, evaluate the design for bias and confounding. Randomisation, blinding, and allocation concealment protect the validity of the p-value. If these safeguards are absent or poorly described, the p-value is conditional on assumptions that may not hold. The ARRIVE guidelines specify the minimum information required for transparent and reproducible animal research, and their checklist provides a structured way to audit whether a study's design supports its statistical claims.

Confidence Intervals in Clinical Decision-Making

A confidence interval is the range of values within which the true population parameter is likely to fall, given the observed data and the chosen confidence level. For clinical work, the interval matters more than the p-value because it expresses the magnitude of the effect in the original units of measurement.

Consider a study evaluating a new analgesic protocol in dogs. The mean reduction in pain score is 2.1 points with a 95% confidence interval from 0.4 to 3.8 points. The p-value is 0.02. The interval tells you that the true mean reduction could be as small as 0.4 points, which may be clinically trivial, or as large as 3.8 points, which is substantial. The p-value alone cannot convey this uncertainty.

When you review a confidence interval, ask three questions. Does the interval exclude the null value? If it does, the result is statistically significant at the corresponding alpha level. Does the interval exclude the minimum clinically important difference? If the lower bound sits above that threshold, you can be reasonably confident the effect is clinically meaningful. Does the interval width reflect adequate precision? A wide interval indicates imprecision, often from a small sample or high variability, and should temper your confidence in the point estimate.

Species and production system alter how these questions are answered. In food animal medicine, the minimum clinically important difference may be defined by economic thresholds, such as average daily gain or feed conversion ratio, instead of by biological endpoints alone. In companion animal oncology, the threshold may be a survival extension measured in months. The same statistical output can support different clinical conclusions depending on the context in which it is applied.

The Checklist for Avoiding Common Misinterpretations

Use this checklist when reading or reviewing a veterinary manuscript. Each item addresses a specific failure mode.

MisinterpretationDetection MethodCorrect Approach
"Non-significant means no effect"Check the confidence interval width and the study's power calculationReport the effect size and its confidence interval, state that the study was underpowered to detect a meaningful effect if applicable
"Significant means clinically important"Compare the effect size to the minimum clinically important differenceInterpret the magnitude in clinical units, not in p-value units
"P = 0.049 is meaningfully different from P = 0.051"Examine the confidence intervals around both estimatesTreat the p-value as a continuous measure of evidence, not a binary threshold
"The p-value is the probability the null hypothesis is true"Re-read the study's statistical methods sectionRecognize that the p-value is calculated under the assumption that the null hypothesis is true
"Multiple significant results confirm the hypothesis"Count the total number of tests performedApply a correction for multiple comparisons or interpret results as exploratory
"A wide confidence interval is acceptable if the p-value is significant"Inspect the interval bounds against clinically relevant thresholdsReport the interval and acknowledge the imprecision

Documenting Statistical Findings in Your Own Work

When you write up your own research, the reporting standards matter as much as the analysis. The EQUATOR Network maintains a comprehensive library of reporting guidelines, including CONSORT for randomised trials, PRISMA for systematic reviews, STROBE for observational studies, and REFLECT for livestock trials. These guidelines specify which statistical details must be reported, and following them reduces the risk that your readers will misinterpret your results.

Report the exact p-value instead of a threshold statement such as "p < 0.05" wherever possible. Report confidence intervals for all primary and secondary effect estimates. State the statistical software and version used, the analysis approach, and any assumptions that were tested. Describe how missing data were handled and whether any sensitivity analyzes were performed.

For studies involving animal models, the choice of species and the generalizability of findings to other species must be addressed explicitly. Comparative reviews of animal models, such as those developed for listeriosis, demonstrate that species-specific differences can substantially alter the interpretation of experimental results and that uncertainty about the best model choice often remains. Your statistical reporting should acknowledge these limitations instead of obscure them.

When the Evidence Base Is Limited

Some clinical questions in veterinary medicine have no adequate statistical evidence. In vitro systems, for example, are widely used to evaluate feed additives and methane mitigation strategies in ruminants, but technical factors such as donor animal species, diet, inoculum preparation, and incubation conditions can influence results substantially. The p-values from such studies describe the in vitro system, not the whole animal, and extrapolation to production settings requires additional validation.

Similarly, high-throughput technologies such as RNA sequencing and multi-omics analyzes generate vast datasets that challenge traditional hypothesis-testing frameworks. Data-driven approaches have emerged as alternatives to hypothesis-driven studies, and the statistical interpretation of findings from these approaches requires different tools, including false discovery rate control and replication across independent cohorts. A p-value from a genome-wide association study in cattle, for instance, must be interpreted against the number of variants tested and the prior probability of association.

In these settings, state explicitly what the statistical evidence can and cannot support. Acknowledge genuine uncertainty and identify the specific additional studies needed to resolve it. This is not a weakness in your analysis. It is the accurate representation of the evidence.

Recognized Complications and Failure Modes

The most consequential failure in interpreting p-values and confidence intervals is the dichotomisation of results into "significant" and "non-significant" categories. This practice converts a continuous measure of evidence into a binary decision, discarding information about the magnitude and direction of effects. In animal cognition research, a systematic evaluation found that 84% of titles and 64% of abstracts described non-significant results as "no effect", a phrasing that overstates the evidence for the null hypothesis Farrar et al., reporting on non-significant results in animal cognition research. The correct reading of a non-significant p-value is that the data are compatible with the null hypothesis, not that the null hypothesis is true.

A second failure mode is the misinterpretation of a wide confidence interval as evidence of no effect. A confidence interval spanning zero includes the null value, but it also includes clinically meaningful effect sizes. The interval communicates the range of effects compatible with the data, and a wide interval signals imprecision, not absence of effect. This distinction matters in veterinary studies where sample sizes are often constrained by animal availability, cost, and ethical limits.

A third complication arises when multiple comparisons are made without adjustment. Each additional test increases the probability of at least one false positive. This problem is amplified in high-throughput settings such as transcriptomic analyzes, where hundreds of thousands of associations may be tested simultaneously. The CattleGTEx project, for example, reported hundreds of thousands of genetic associations with gene expression across 23 tissues, a scale that demands explicit multiple-testing control Liu et al., cattle regulatory variant atlas.

ObservationLikely CauseDiscriminating Check
P-value between 0.01 and 0.05 in a small studyFragile result driven by few observationsExamine the confidence interval width and the number of events per group
Confidence interval excludes zero but p-value is non-significantMismatch between test and interval calculationVerify that both used the same statistical model and error structure
Non-significant result reported as "no effect"Conflation of absence of evidence with evidence of absenceCheck whether the confidence interval excludes clinically meaningful effects
Significant result in a study with many endpointsMultiplicity inflationReview whether adjustment for multiple comparisons was applied

Common Errors and Corrective Actions

Less experienced analysts frequently mistake the p-value for the probability that the null hypothesis is true. The p-value is calculated under the assumption that the null hypothesis holds, so it cannot assign a probability to that hypothesis. The correct framing is that the p-value describes the compatibility of the observed data with the null model.

A related error is the interpretation of a p-value as a measure of effect size. A very small p-value can arise from a trivial effect if the sample is large, while a clinically important effect can yield a p-value above 0.05 if the sample is small. The confidence interval, not the p-value, conveys the range of plausible effect magnitudes.

Students also err by treating p = 0.049 and p = 0.051 as fundamentally different outcomes. The threshold of 0.05 is a convention, not a natural boundary. Two studies with nearly identical p-values provide nearly identical evidence, regardless of which side of the threshold they fall. Reporting the exact p-value and the confidence interval allows readers to judge the evidence for themselves.

The corrective action in each case is the same: report effect sizes with confidence intervals, state the exact p-value, and describe the statistical model fully. Reporting guidelines such as those catalogued by the EQUATOR Network specify the minimum information required for transparent statistical reporting.

Limitations of the Current Evidence

The veterinary literature contains many studies with small sample sizes, particularly in species where large cohorts are impractical. Underpowered studies produce unstable estimates and non-reproducible results. The p-value distribution in animal cognition research is consistent with studies conducted at low statistical power, meaning that many non-significant results may be false negatives Farrar et al., reporting on non-significant results in animal cognition research.

Expert opinion still differs on the appropriate response to underpowered data. Some statisticians advocate for a shift away from null hypothesis significance testing toward estimation and uncertainty quantification. Others argue that p-values retain value when used correctly and in conjunction with effect sizes. This disagreement is not resolved, and veterinary researchers should be aware that statistical practice is evolving.

The evidence base is also limited by publication bias. Studies with significant results are more likely to be published than those without, which distorts the literature available for systematic review and meta-analysis. A meta-analysis of lupeol in preclinical cancer models found asymmetrical funnel plots and negative Egger values, indicating publication bias among the included studies Fatma et al., systematic review and meta-analysis of lupeol. Readers should interpret pooled estimates from such analyzes with caution.

When to Seek Specialist Input

Referral to a statistician is warranted when the study design involves clustered data, repeated measures, survival analysis, or hierarchical models. These designs require specialised methods that general-purpose statistical software does not handle correctly by default. A statistician should also be consulted before data collection begins, not after, because sample size calculations and randomisation schemes must be specified in advance.

Laboratory involvement is indicated when data are generated by high-throughput platforms such as RNA sequencing or mass spectrometry. The bioinformatics pipelines for these data require computational expertise beyond standard statistical analysis Uesaka et al., bioinformatics in bioscience and bioengineering. The ARRIVE guidelines specify the minimum information required for transparent reporting of animal research, and compliance should be checked before submission.

Regulatory reporting obligations arise when study findings affect animal health policy, trade, or public health. The WOAH terrestrial animal health standards define international requirements for disease surveillance and reporting. Researchers working with notifiable diseases or production animals destined for trade should confirm their obligations with the relevant national authority before publishing.

Frequently Asked Questions

How should I interpret a non-significant p-value when my study was underpowered?

A non-significant p-value in an underpowered study does not demonstrate that the treatment lacks effect. It indicates only that the observed data are compatible with the null hypothesis, and they are often equally compatible with a meaningful effect that the study could not detect. Researchers in animal cognition frequently report such results as "no effect," yet this interpretation is frequently unjustified when statistical power is low. Examine the confidence interval around the effect estimate. If the interval is wide and includes clinically meaningful values, the result is uninformative instead of negative. Report the effect size and its uncertainty explicitly, and describe the result as inconclusive instead of as evidence of absence. See reporting and interpreting non-significant results in animal cognition research for a detailed analysis of this problem.

What can I do when sample size is constrained by ethical or practical limits?

When animal numbers are limited by ethical review, cost, or availability, you cannot simply increase power. Prioritize precision over hypothesis testing. Use continuous outcomes instead of binary ones, reduce measurement error through standardized protocols, and consider repeated measures designs that use each animal as its own control. Report confidence intervals instead of relying on p-values alone, and be explicit about the width of those intervals when discussing clinical relevance. For in vitro work, standardization of donor animals, inoculum preparation, and incubation conditions can reduce unexplained variability substantially, improving the precision you can achieve with a fixed number of replicates. The technical recommendations for in vitro fermentation experiments provide practical guidance on controlling variability in that specific setting.

Does the interpretation of p-values and confidence intervals differ between species or production systems?

The statistical logic does not change across species, but the clinical context does. A confidence interval that excludes zero in a laboratory mouse study may correspond to an effect too small to matter in a production herd, where cost per animal, handling time, and population-level outcomes dominate decisions. Conversely, in companion animal oncology, a modest survival benefit may be clinically valuable even when the confidence interval is wide. The threshold for clinical significance must be defined from species-specific knowledge of disease biology, treatment costs, and owner expectations. Cross-species extrapolation of effect sizes is particularly hazardous in infectious disease research, where species-specific differences in pathogenesis can alter dose-response relationships substantially, as demonstrated in comparative reviews of listeriosis animal models.

How do I explain statistical uncertainty to a client or referring veterinarian?

Use concrete language tied to the individual patient. Explain that the study result is an estimate, not a certainty, and that the confidence interval describes the range of plausible effects. For example, state that the treatment is associated with a median survival improvement of four months, with the true benefit likely falling between one and seven months. Avoid the word "significant" entirely, as clients reasonably interpret it as "important." Instead, describe whether the effect size is large enough to matter for their animal's condition and whether the precision of the estimate supports a confident recommendation. Acknowledge when the evidence is too imprecise to guide a firm decision, and outline what additional information would help. The MSD Veterinary Manual offers species-specific guidance on communicating prognosis and treatment options in clinical practice.

What should I record in my laboratory notebook or study file regarding statistical decisions?

Document the pre-specified hypothesis, the primary outcome, and the planned analysis before data collection begins. Record the chosen alpha level, whether tests are one- or two-sided, and any planned adjustments for multiple comparisons. Note the target effect size and the power calculation, including the assumptions used. After analysis, record the exact test used, the test statistic, degrees of freedom, p-value, and the confidence interval for the primary effect estimate. If you deviate from the pre-specified plan, document the reason and the date. This audit trail supports transparency and reproducibility, and it aligns with the expectations of ARRIVE guidelines 2.0 for reporting animal research, which require clear description of statistical methods and outcome definitions.

When should I consult a statistician during a veterinary research project?

Consult a statistician before you finalise the study design, not after data collection. Early input is most valuable for sample size estimation, randomisation schemes, and choice of primary outcome. A second consultation is warranted when you encounter unexpected data structures, such as missing observations, clustering within litters or herds, or non-normal distributions. If you plan to use high-dimensional data, such as transcriptomics or genome-wide association studies, statistical input is essential from the outset because the analytic pipeline determines the validity of all downstream inference. The bioinformatics methods review emphasizes that data-driven approaches require computational and statistical expertise that differs substantially from traditional hypothesis testing. Seek help early enough that the statistician can influence design decisions instead of merely diagnose problems after the data are collected.

Related Clinical & Scientific Guides

References and Further Reading

Related Articles

This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.