Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

Validity of a Study: Understanding Internal and External Validity

When you read a research paper or plan your own study, the first question is whether the results mean what the authors claim. Validity is the degree to which a study supports its conclusions. Internal validity asks whether the observed effect is truly caused by the intervention instead of by something else. External validity asks whether the findings apply beyond the study setting, to other people, places, and times. For students, researchers, and life-science professionals, understanding these two concepts is the difference between applying a finding with confidence and repeating a mistake at scale. This article explains both types of validity, the specific threats that undermine them, and practical strategies for assessing and improving study design and measurement.

At a Glance: Validity Concepts and Practical Focus

Concept What It Answers Common Threat Practical Check
Internal validity Did the intervention cause the outcome? History, maturation, testing effects, instrumentation change Was there a control group and randomization? Were measurements identical across groups and time points?
External validity Do the results apply to other settings and populations? Narrow sample, artificial conditions, single site Does the sample match the population you want to apply results to? Were conditions similar to real practice?
Reliability Would the measurement produce the same result if repeated? Poor instrument calibration, inconsistent administration Did the authors report test-retest or inter-rater reliability? Was the measurement tool validated?
Statistical conclusion validity Is the statistical analysis appropriate and powerful enough? Small sample size, violated assumptions, multiple comparisons Was the sample size justified? Were effect sizes and confidence intervals reported?

Internal and external validity are not either-or categories. A study can have strong internal validity and weak external validity, or the reverse. A tightly controlled laboratory experiment may show a clear causal effect but tell you little about how the intervention works on a commercial farm. A large observational survey may reflect real-world conditions but cannot rule out alternative explanations for the associations it finds. The practical task is to identify which threats are plausible in a given study and to design or interpret the study accordingly.

Defining Internal Validity

Internal validity refers to the degree of control exerted over potential confounding variables to reduce alternative explanations for the effects of various treatments. When a study has internal validity, extraneous variables have been controlled, and the researcher can reasonably conclude that the intervention, not some other factor, produced the observed outcome. This definition appears consistently across the research literature, including reviews of exercise science research and clinical trial methodology.

The core logic is causal. You expose a group to an intervention, you measure an outcome, and you want to attribute any change to the intervention. That attribution fails if something else changed at the same time. For example, if you test a new feed additive during a season when ambient temperature also shifted, you cannot tell whether the additive or the temperature caused the change in weight gain. The study lacks internal validity because the temperature is a plausible alternative explanation.

Internal validity is a prerequisite for any causal claim. Without it, you cannot say that the treatment worked, only that the treatment was followed by the outcome. This distinction matters in applied settings. A farmer who changes feeding practices and sees improved milk yield may be observing a real effect, or may be observing the effect of better weather, different forage quality, or the simple act of paying more attention to the animals. The same logic applies to clinical trials, educational interventions, and policy evaluations.

Defining External Validity

External validity is the extent to which study findings can be generalized to other populations, settings, and times. A study has external validity when its conclusions hold beyond the specific participants, materials, and procedures used in the research. For in vitro toxicity studies, external validity has been conceptualized as consisting of two components: applicability and generalisability. Applicability asks whether the study model is relevant to the question being asked. Generalisability asks whether the results extend from the model to the target system, such as from a cell line to a whole organism.

External validity is often in tension with internal validity. The more you control a study to eliminate confounding variables, the more artificial the conditions become. A randomized controlled trial with strict inclusion criteria, standardized protocols, and blinded outcome assessment may produce internally valid results that do not apply to patients with multiple health conditions, variable adherence, or different baseline characteristics. Conversely, a study conducted in real-world conditions may have strong external validity but weak internal validity because the researcher cannot control all the variables.

The practical question for a reader is whether the study population and setting resemble the population and setting where you intend to apply the results. If a clinical trial enrolled only healthy young adults, the results may not apply to older adults with chronic disease. If a feeding trial used a single breed under confined housing, the results may not apply to pasture-based systems with different breeds.

Threats to Internal Validity

Threats to internal validity are specific mechanisms by which alternative explanations can arise. The classic framework, developed by Cook and Campbell and applied across fields from worksite health promotion to nursing research, identifies a set of recurring threats. Understanding these threats allows you to examine a study design and identify which ones are plausible.

History

History refers to events that occur during the study period, outside the intervention, that can affect the outcome. These events are not part of the treatment but coincide with it. In a worksite health promotion program, a company-wide wellness initiative or a change in cafeteria offerings could influence employee health outcomes independently of the program being studied. In an animal trial, a disease outbreak, a change in feed supplier, or a weather event could affect all groups and obscure the treatment effect.

The key feature of a history threat is that it affects the outcome at the same time as the intervention. A control group helps address this threat because both groups experience the same historical events. If the treatment group improves but the control group does not, history is less plausible as an explanation. However, if the historical event affects only the treatment group, such as a power outage in one barn, the control group does not protect against it.

Maturation

Maturation refers to changes that occur naturally over time within the participants, independent of the intervention. Animals grow, heal, age, and adapt. People learn, fatigue, and develop. If you measure weight gain in growing animals, some gain is expected even without any intervention. If you measure learning in students, some improvement occurs simply from practice and development.

Maturation is a particular problem in studies without a control group. A before-and-after design that shows improvement after an intervention cannot distinguish the intervention effect from the natural trajectory of the participants. A control group that experiences the same passage of time provides the counterfactual: what would have happened without the intervention.

Testing Effects

Testing effects occur when the act of measurement itself influences the outcome. Repeated testing can produce practice effects, where participants improve because they become familiar with the test. It can also produce fatigue effects, where performance declines because of the burden of repeated measurement. In animal studies, repeated handling, blood sampling, or behavioral testing can alter the animals' stress levels and physiology, confounding the treatment effect.

The repeated testing effect is a critical internal validity threat in physical therapy research and other fields where outcomes are measured multiple times. The Halili Physical Therapy Statistical Analysis Tool was designed specifically to control for repeated testing effects, study sample uniformity, and increases in type I or type II error. The existence of such tools underscores how seriously researchers treat this threat.

Instrumentation

Instrumentation refers to changes in the measurement instrument or the observers over the course of the study. A scale that drifts out of calibration, a laboratory assay that uses different reagent lots, or an observer who becomes more skilled or more fatigued can all produce changes in measurements that have nothing to do with the intervention.

Instrumentation threats are especially relevant in studies with long durations or multiple measurement points. If baseline measurements are taken by one technician and follow-up measurements by another, differences in technique can create apparent changes. If an automated sensor is recalibrated mid-study, the data before and after recalibration may not be comparable.

Statistical Regression

Statistical regression, also called regression to the mean, occurs when participants are selected because of extreme scores. If you select animals with very low weight gain for a feeding trial, their subsequent weight gain is likely to improve even without an effective intervention, simply because extreme values tend to move toward the mean on repeated measurement. The same logic applies to selecting students with very low test scores, patients with very high blood pressure, or farms with very low productivity.

Regression is a threat when selection is based on an extreme score and the outcome is measured again. A control group selected on the same extreme criterion helps address this threat because both groups would be expected to regress similarly. Without a control group, improvement after intervention may be entirely due to regression.

Selection and Selection Bias

Selection refers to systematic differences between groups that exist before the intervention begins. If the treatment group and control group differ in baseline characteristics, any post-intervention difference could be due to those pre-existing differences instead of the intervention. Random assignment is the primary defense against selection bias because it makes the groups probabilistically similar on all characteristics, measured and unmeasured.

Selection can also interact with other threats. Selection-maturation occurs when groups mature at different rates. Selection-history occurs when groups experience different historical events. These interactions are particularly problematic in quasi-experimental designs where random assignment is not possible.

Mortality or Attrition

Mortality, also called attrition, refers to participants dropping out of the study before it is completed. If the participants who drop out differ systematically from those who remain, the results can be biased. In animal studies, mortality can occur when sick or unthrifty animals are removed from the trial. In human studies, participants who find the intervention burdensome or ineffective may withdraw.

Attrition threatens internal validity when it is differential across groups. If more participants drop out of the control group than the treatment group, the remaining control participants may be unusually motivated or healthy, making the treatment effect appear larger or smaller than it truly is. Researchers should report attrition rates and compare the characteristics of completers and dropouts.

Diffusion or Imitation of Treatments

Diffusion occurs when participants in the control group receive the intervention, or when participants in the treatment group share the intervention with controls. In worksite health promotion programs, employees in the control group may see their colleagues exercising and adopt the behavior themselves. In agricultural trials, animals in different pens may share feed or water, or farm staff may inadvertently treat all animals similarly.

Diffusion dilutes the contrast between groups and makes it harder to detect a true treatment effect. It can also produce a null result that is misinterpreted as evidence that the intervention does not work. Researchers should design studies to minimize cross-contamination, such as physical separation of groups or blinding of staff.

Compensatory Equalization and Rivalry

Compensatory equalization occurs when those responsible for delivering services provide extra benefits to the control group to make up for the fact that they are not receiving the intervention. In a school study, teachers may give extra attention to students in the control group because they feel sorry for them. In an animal trial, staff may provide better care to control animals because they perceive them as disadvantaged.

Compensatory rivalry, also called the John Henry effect, occurs when control group participants work harder because they know they are in the control group. Resentful demoralization is the opposite: control participants give up because they feel deprived. All three threats reduce the apparent effect of the intervention by improving control group performance or worsening treatment group performance.

Ambiguity About the Direction of Causal Influence

In cross-sectional or correlational designs, it can be unclear whether the cause precedes the effect. Does poor housing cause poor animal health, or do farms with poor animal health tend to have poor housing for other reasons? Does stress cause disease, or does disease cause stress? Without temporal ordering, causal claims are ambiguous.

Longitudinal designs, where exposure is measured before outcome, help establish temporal order. Experimental designs, where the researcher controls the timing of the intervention, provide the strongest evidence for causal direction.

Threats to External Validity

External validity threats are conditions that limit the generalizability of study findings. These threats are often less formalized than internal validity threats but are equally important for applied decision-making.

Sample Characteristics

The most common external validity threat is a sample that does not represent the target population. Clinical trials that exclude older adults, pregnant women, or people with multiple health conditions produce results that may not apply to those groups. Animal studies that use a single breed, age class, or production system may not apply to other breeds or systems.

The question is not whether the sample is representative of the general population but whether it is representative of the population to which you want to apply the results. A study of dairy cows in a temperate climate may or may not apply to dairy cows in a tropical climate. The reader must judge the similarity between the study sample and the target population.

Setting and Context

Findings from one setting may not transfer to another. A study conducted in a university research barn with controlled temperature, ventilation, and feeding may not apply to a commercial farm with variable conditions. A study conducted in a high-income country may not apply to a low-income country with different infrastructure, resources, and disease patterns.

The artificiality of research settings is a particular concern for studies that prioritize internal validity. The more controlled the setting, the less it resembles real-world conditions. Researchers should describe the setting in enough detail for readers to judge applicability.

Time and Historical Context

Findings may not generalize across time. Agricultural practices, animal genetics, disease patterns, and market conditions change over time. A feeding trial conducted in 2010 may not apply to current production systems with different genetics and management. A clinical trial conducted before a new standard of care was introduced may not apply to current practice.

The temporal generalizability of a study depends on how stable the underlying relationships are. Basic physiological relationships may be stable over decades. Management-dependent relationships may change quickly as practices evolve.

Treatment and Outcome Definitions

The specific form of the intervention and the way outcomes are measured affect generalizability. A study that tests a specific dose, duration, or delivery method of an intervention may not apply to other doses, durations, or delivery methods. A study that measures a surrogate outcome, such as blood pressure, may not tell you about the clinical outcome of interest, such as heart attacks.

Readers should examine whether the intervention tested is the same as the intervention they intend to use and whether the outcomes measured are relevant to their decisions.

Reliability and Its Relationship to Validity

Reliability refers to the consistency and reproducibility of measurements. A reliable measurement produces the same result when repeated under the same conditions. Reliability is a necessary but not sufficient condition for validity. A measurement can be reliable without being valid, such as a scale that consistently reads five kilograms too heavy. But a measurement cannot be valid if it is unreliable, because the measurement error obscures the true value.

Reliability takes several forms. Test-retest reliability assesses stability over time. Inter-rater reliability assesses consistency across observers. Internal consistency assesses whether items on a scale measure the same construct. The Measure of Indigenous Racism Experiences scale validation study in Guatemala reported Cronbach's alpha for the three dimensions of racism, which is a measure of internal consistency. The study also used confirmatory factor analysis to test construct validity, demonstrating that reliability and validity are assessed together in scale development.

In animal research, reliability concerns include the precision of laboratory assays, the consistency of behavioral observations, and the repeatability of physiological measurements. A study that uses an unreliable measurement tool may fail to detect a true treatment effect or may detect a spurious effect. The systematic review of chemotherapy-induced peripheral neuropathy trials identified lack of a valid and reliable measurement as one of the design flaws that threatened internal validity.

Statistical Conclusion Validity

Statistical conclusion validity refers to the appropriateness of the statistical analysis and the correctness of the conclusions drawn from it. A study can have good internal and external validity but still draw incorrect conclusions if the statistical analysis is flawed. Common statistical conclusion validity threats include low statistical power, violated assumptions, and inappropriate multiple comparisons.

Low statistical power increases the risk of type II error, failing to detect a true effect. The systematic review of chemotherapy-induced peripheral neuropathy trials found that internal validity threats may have resulted in type II error and subsequent dismissal of potentially effective interventions. This finding illustrates a broader principle: a null result is only meaningful if the study had adequate power to detect a clinically important effect.

Type I error, detecting an effect that does not exist, is increased by multiple comparisons and data dredging. If a researcher tests many outcomes or many subgroups, some will appear significant by chance alone. Randomization in single-case experimental designs allows for various randomization statistical tests, which increases data-evaluation and statistical-conclusion validity.

Assessing the Validity of a Study: A Practical Checklist

When you read a research paper, work through the following checklist to assess internal and external validity. This checklist is adapted from the threats described above and from reporting guidelines promoted by the EQUATOR Network, which maintains a collection of reporting guidelines for health research.

Step 1: Examine the Design

Determine whether the study used randomization, a control group, and blinding. Randomized controlled trials provide the strongest internal validity because randomization addresses selection bias and control groups address history, maturation, and testing effects. Blinding addresses performance and detection bias.

If the study did not use randomization, identify which threats are plausible. Quasi-experimental designs, before-and-after designs, and observational studies have specific vulnerability patterns. The multiple baseline design literature shows that even within a single design family, variations differ in their ability to address threats. Nonconcurrent multiple baseline designs are considered less rigorous than concurrent designs because of their presumed limited ability to address the threat of coincidental events, though this skepticism has been challenged.

Step 2: Evaluate the Measurement

Ask whether the outcome measures are valid and reliable. Was the measurement tool validated for the population and setting? Were the same measurement procedures used for all groups and time points? Were observers blinded to treatment assignment? Was the instrument calibrated and maintained?

The validation study of the paediatric inguinal herniotomy simulator assessed face validity, content validity, and construct validity using expert and novice participants. This example shows how validity assessment applies to measurement tools as well as to studies. A measurement tool must be valid for its intended purpose before the study results can be trusted.

Step 3: Assess the Sample

Examine the inclusion and exclusion criteria, the sample size, and the attrition rate. Does the sample represent the population to which you want to apply the results? Was the sample size adequate to detect the expected effect? Did participants drop out, and were the dropouts similar to completers?

The external validation of a machine-learning model for acute kidney injury prediction used 38,287 real ICU patients from a multi-center dataset, developing the model on one hospital and externally validating it on a second independent hospital. This two-stage approach, development followed by external validation, is a strong model for assessing generalizability.

Step 4: Scrutinize the Analysis

Check whether the statistical analysis matches the design. Were the assumptions of the statistical tests met? Were effect sizes and confidence intervals reported in addition to p-values? Were multiple comparisons handled appropriately? Was the analysis pre-specified or post hoc?

The conformal prediction framework for acute kidney injury demonstrated that class-conditional conformal prediction is exactly valid under a change in disease prevalence, so that any residual under-coverage on a new site isolates within-class distribution shift instead of the prevalence difference. This level of analytical rigor is the standard to which studies should aspire.

Step 5: Consider the Setting

Ask whether the study setting resembles the setting where you intend to apply the results. Were the conditions artificial or natural? Were the participants volunteers or representative samples? Were the researchers or practitioners delivering the intervention different from those who would deliver it in practice?

The external validity for in vitro toxicity studies paper argues that external validity assessment is relevant at two stages of evidence synthesis, during study screening and during evidence integration. This two-stage approach acknowledges that external validity is not a single judgment but a process that continues as evidence accumulates.

Strategies to Improve Internal Validity

Researchers can take specific actions to reduce threats to internal validity. These strategies should be planned at the design stage, not applied after data collection.

Randomization

Randomization is the most powerful tool for addressing selection bias. Random assignment makes groups probabilistically similar on all characteristics, measured and unmeasured. In single-case experimental designs, various forms of randomization can be applied, including phase-order randomization, between-intervention case randomization, within-intervention case randomization, and intervention start-point randomization. These randomization strategies control for internal validity concerns that traditional replication alone may not address.

Randomization also enables statistical tests that assume random assignment. The review of randomization in single-case design experiments notes that design randomization allows for various randomization statistical tests to be conducted, increasing data-evaluation and statistical-conclusion validity.

Control Groups

A control group provides the counterfactual: what would have happened without the intervention. The control group experiences the same historical events, the same passage of time, and the same measurement procedures as the treatment group. Differences between groups can then be attributed to the intervention.

The choice of control condition matters. A no-treatment control group answers a different question than an active control group. A placebo control group addresses expectancy effects. Researchers should choose the control condition that best answers their research question while controlling the most plausible threats.

Blinding

Blinding prevents knowledge of treatment assignment from influencing behavior or measurement. Single blinding hides treatment assignment from participants. Double blinding hides it from both participants and researchers. Blinding addresses performance bias, where participants or providers behave differently based on treatment assignment, and detection bias, where outcome assessors measure differently based on their expectations.

Blinding is not always possible. In surgical trials, the surgeon cannot be blinded to the procedure. In animal feeding trials, the farm staff may be able to tell which animals receive which feed. When blinding is impossible, researchers should use objective outcome measures and blinded outcome assessors.

Standardization

Standardizing the intervention, the measurement procedures, and the environment reduces variability and controls instrumentation threats. A written protocol for the intervention ensures that all participants receive the same treatment. A written protocol for measurement ensures that all assessments are conducted the same way. Standardization also improves replicability, allowing other researchers to reproduce the study.

The Halili Physical Therapy Statistical Analysis Tool was designed to address the common lack of interventional standardization in physical therapy research. The tool isolates the average rate of change instead of average change after treatment, allowing for the simultaneous analysis of hundreds of treatment combinations while controlling for repeated testing effects, study sample uniformity, and increases in type I and type II error.

Statistical Control

When randomization is not possible, statistical control can address some threats. Analysis of covariance can adjust for baseline differences. Propensity score methods can balance groups on observed characteristics. Sensitivity analyses can test whether results are robust to plausible confounding.

Statistical control has limitations. It can only adjust for measured variables. Unmeasured confounders remain a threat. The directed acyclic graph approach to discovering internal validity threats in single-case experimental designs provides a formal method for identifying which variables need to be measured and controlled.

Strategies to Improve External Validity

External validity is improved by designing studies that reflect the populations, settings, and conditions where the results will be applied.

Broad Inclusion Criteria

Broad inclusion criteria produce samples that represent the target population. Instead of excluding participants with comorbidities, researchers can include them and analyze them as subgroups. Instead of using a single breed or production system, researchers can recruit multiple breeds or systems.

The tradeoff is that broad inclusion criteria increase variability and may reduce the ability to detect treatment effects. Researchers must balance internal and external validity, recognizing that a study that is perfectly controlled may not be applicable and a study that is perfectly applicable may not be controlled.

Multi-Site Studies

Multi-site studies recruit participants from multiple locations, increasing the diversity of settings, populations, and conditions. The external validation of the acute kidney injury prediction model used two independent hospitals, demonstrating that the model performed well across sites. Multi-site studies also allow researchers to examine whether treatment effects vary across sites.

Multi-site studies are more complex and expensive than single-site studies. They require standardized protocols across sites, coordination of data collection, and analysis that accounts for site effects. The payoff is greater confidence in the generalizability of the findings.

Real-World Conditions

Studies conducted in real-world conditions have greater external validity than studies conducted in artificial settings. Pragmatic trials test interventions under routine conditions, with flexible protocols and broad inclusion criteria. Implementation research examines how interventions work when delivered by regular staff in regular settings.

The tension between internal and external validity is real. Researchers should be explicit about which type of validity they are prioritizing and why. A phase I clinical trial prioritizes internal validity to establish safety and efficacy. A phase IV effectiveness study prioritizes external validity to establish real-world performance.

Replication

Replication is the ultimate test of external validity. If a finding is reproduced in different populations, settings, and times, confidence in its generalizability increases. The NC3Rs Experimental Design Assistant supports researchers in designing robust experiments, including considerations of replication. The EQUATOR Network promotes complete and transparent reporting, which enables replication.

Replication can take several forms. Direct replication repeats the study with the same methods. Conceptual replication tests the same hypothesis with different methods. Systematic replication varies one aspect of the study while holding others constant. Each form contributes to external validity in different ways.

Common Failure Patterns in Validity Assessment

Certain failure patterns recur when researchers and readers assess validity. Recognizing these patterns helps you avoid them.

Confusing Internal and External Validity

A common error is treating a study with strong internal validity as if it also had strong external validity. A randomized controlled trial with strict inclusion criteria may produce internally valid results that do not apply to the broader population. Conversely, a large observational study may have strong external validity but weak internal validity. Readers should assess each type of validity separately.

Overweighting Statistical Significance

A statistically significant result is not necessarily an important result, and a non-significant result is not necessarily a null result. The systematic review of chemotherapy-induced peripheral neuropathy trials found that internal validity threats may have resulted in type II error and dismissal of potentially effective interventions. Readers should examine effect sizes, confidence intervals, and study power, beyond p-values.

Ignoring Measurement Validity

A study can have a perfect design and still produce invalid results if the measurements are unreliable or invalid. The systematic review identified lack of a valid and reliable measurement as a design flaw in clinical trials. Readers should ask whether the outcome measures were validated for the population and setting.

Assuming Generalizability From a Single Study

A single study, no matter how well designed, provides limited evidence for generalizability. The external validity for in vitro toxicity studies paper argues that external validity assessment is relevant at two stages of evidence synthesis, during study screening and during evidence integration. Readers should consider the body of evidence, beyond individual studies.

Failing to Consider Plausible Threats

Researchers and readers often focus on the threats they know about and ignore the ones they do not. The review of overlooked confounding variables in exercise science identified variables that often do not receive adequate attention, including instructions on how to perform the test, volume and frequency of verbal encouragement, knowledge of exercise endpoint, number and gender of observers in the room, influence of music played before and during testing, and the effects of mental fatigue on performance. These overlooked variables can threaten internal validity even in well-designed studies.

Records and Measurements for Validity Assessment

When you document a study for validity assessment, keep the following records. These records support both your own assessment and the assessment of others who read your work.

Design Records

Document the study design, including randomization procedures, control conditions, blinding, and sample size justification. The NC3Rs Experimental Design Assistant provides a tool for planning experiments and documenting design decisions. The EQUATOR Network provides reporting guidelines that specify what should be reported for different study types.

Measurement Records

Document the measurement instruments, their calibration, and their validation status. Record who performed the measurements, when they were performed, and under what conditions. Record any changes to measurement procedures during the study.

Participant Records

Document the recruitment procedures, inclusion and exclusion criteria, and the characteristics of the sample. Record the attrition rate and the reasons for dropout. Compare the characteristics of completers and dropouts.

Analysis Records

Document the statistical analysis plan, including pre-specified analyses, assumptions, and adjustments. Record the software and version used for analysis. Document any deviations from the analysis plan and the reasons for them.

Limitations and Professional Escalation Criteria

Every study has limitations. The task is to identify which limitations matter for the conclusions and to escalate concerns when they threaten the validity of the findings.

When to Question Internal Validity

Question the internal validity of a study when you identify plausible threats that the design does not address. If a study lacks a control group, maturation and history are plausible threats. If a study uses a nonconcurrent multiple baseline design, coincidental events are a plausible threat. If a study has high attrition, selection bias is a plausible threat.

The multiple baseline design literature shows that the consensus in recent textbooks and methodological papers is that nonconcurrent designs are less rigorous than concurrent designs because of their presumed limited ability to address the threat of coincidental events. However, this skepticism has been challenged, with arguments that the primary reliance on across-tier comparisons and the resulting deprecation of nonconcurrent designs are not well-justified. This debate illustrates that validity assessment involves professional judgment, beyond rule application.

When to Question External Validity

Question the external validity of a study when the sample, setting, or time period differs substantially from the population, setting, or time period where you intend to apply the results. If a study enrolled only young healthy adults, question its applicability to older adults with chronic disease. If a study was conducted in a single hospital, question its applicability to other hospitals with different resources and patient populations.

The external validation of the acute kidney injury prediction model demonstrated that marginal conformal prediction delivered 92.1% external coverage but under-covered AKI-positive patients at 31.3%. This finding shows that overall performance can mask important subgroup differences. External validity assessment should examine subgroups, beyond overall performance.

When to Escalate

Escalate your concerns when the validity threats are severe enough to change the conclusions. If a study has multiple design flaws, the results may not be trustworthy. The systematic review of chemotherapy-induced peripheral neuropathy trials found that 22 of 24 phase III clinical trials had two or more design flaws, including sample heterogeneity, malapropos mechanism of action, malapropos intervention dose, malapropos timing of the outcome measurement, confounding variables, lack of a valid and reliable measurement, and suboptimal statistical validity.

Escalation criteria include the following. First, escalate when the study conclusions depend on an assumption that is not met. Second, escalate when the effect size is small and the confidence interval includes the null. Third, escalate when the study has high attrition or differential attrition across groups. Fourth, escalate when the measurement tools have not been validated for the population and setting. Fifth, escalate when the study has not been reported according to relevant reporting guidelines.

Welfare and Safety Context

Validity assessment has direct implications for animal welfare and human safety. A study with poor internal validity can lead to the adoption of ineffective interventions or the rejection of effective ones. The systematic review of chemotherapy-induced peripheral neuropathy trials found that numerous interventions have been declared ineffective based on the results of phase III trials, but internal validity threats may have resulted in type II error and subsequent dismissal of a potentially effective intervention. Patients may benefit from rigorous retesting of several agents to expand and validate the evidence.

In animal agriculture, the stakes are similar. A feeding trial with poor internal validity can lead farmers to adopt an ineffective additive or reject an effective one. A management study with poor external validity can lead to recommendations that do not work in different production systems. The cost of invalid research is paid by animals and by the people who care for them.

The polyvagal theory perspective on safety highlights that feelings of safety emerge from internal physiological states regulated by the autonomic nervous system. When animals feel safe, their nervous systems support the homeostatic functions of health, growth, and restoration. Research that fails to account for the physiological state of animals may produce invalid results. A study that measures stress outcomes without controlling for handling procedures, social environment, or other stressors may attribute to the intervention effects that are actually due to the animals' physiological state.

Frequently Asked Questions

What is the difference between internal and external validity?

Internal validity asks whether the observed effect is truly caused by the intervention instead of by confounding variables. External validity asks whether the findings apply to other populations, settings, and times. A study can have strong internal validity and weak external validity, or the reverse. Internal validity is a prerequisite for causal claims. External validity determines whether those causal claims are useful in other contexts.

How does reliability differ from validity?

Reliability refers to the consistency and reproducibility of measurements. A reliable measurement produces the same result when repeated under the same conditions. Validity refers to whether the measurement measures what it claims to measure and whether the study supports its conclusions. Reliability is necessary but not sufficient for validity. A measurement can be reliable without being valid, but a measurement cannot be valid if it is unreliable.

What are the most common threats to internal validity?

The most common threats are history, maturation, testing effects, instrumentation, statistical regression, selection, mortality or attrition, diffusion of treatments, compensatory equalization, compensatory rivalry, and resentful demoralization. History refers to external events that coincide with the intervention. Maturation refers to natural changes over time. Testing effects occur when measurement influences the outcome. Instrumentation refers to changes in the measurement instrument or observers. Each threat provides an alternative explanation for the observed effect.

How can I improve the internal validity of my study?

Use randomization to address selection bias. Use a control group to address history, maturation, and testing effects. Use blinding to address performance and detection bias. Standardize the intervention and measurement procedures to address instrumentation threats. Use statistical control when randomization is not possible. Plan these strategies at the design stage, not after data collection.

How can I improve the external validity of my study?

Use broad inclusion criteria to represent the target population. Conduct multi-site studies to increase the diversity of settings and conditions. Test interventions under real-world conditions. Replicate the study in different populations, settings, and times. Report the sample and setting in enough detail for readers to judge applicability.

What is statistical conclusion validity?

Statistical conclusion validity refers to the appropriateness of the statistical analysis and the correctness of the conclusions drawn from it. Threats include low statistical power, violated assumptions, and inappropriate multiple comparisons. Low power increases the risk of type II error, failing to detect a true effect. Multiple comparisons increase the risk of type I error, detecting an effect that does not exist.

How do I assess the validity of a study I am reading?

Work through a checklist. Examine the design for randomization, control groups, and blinding. Evaluate the measurement tools for validity and reliability. Assess the sample for representativeness and adequacy. Scrutinize the analysis for appropriateness and power. Consider the setting for similarity to your context. Identify which threats are plausible and whether the design addresses them.

When should I reject a study because of validity concerns?

Reject a study when the validity threats are severe enough to change the conclusions. Escalate when the study conclusions depend on an unmet assumption, when the effect size is small and the confidence interval includes the null, when attrition is high or differential, when measurement tools have not been validated, or when the study has not been reported according to relevant guidelines. A single flaw does not necessarily invalidate a study, but multiple flaws or a severe flaw should raise serious concerns.

Related Articles

References and Further Reading

This article is educational and does not replace institutional policy, professional advice, or applicable safety and regulatory requirements.