Case-Control Studies in Veterinary Epidemiology: Selection and Analysis
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Case-control studies in veterinary epidemiology are retrospective designs that sample based on outcome status (cases with disease, controls without) to efficiently investigate rare diseases, long-latency outcomes, or adverse effects of interventions, such as investigating the association between specific feed additives and the incidence of subacute ruminal acidosis in dairy cattle.
- The odds ratio is the primary measure of association, approximating the relative risk when the outcome is rare (e.g., <10% incidence), and is calculated from a 2x2 table comparing odds of exposure in cases versus controls, with adjusted estimates derived from multivariable logistic regression to control for confounders like age, sex, or herd.
- Rigorous control selection is paramount; controls must represent the exposure distribution of the source population from which cases arose, necessitating clear definition of this population (e.g., all animals presenting to a referral hospital, or all animals in a defined geographic region) and employing methods like incidence density sampling or risk-set sampling.
- Matching (individual or frequency) on potential confounders such as age, breed, herd, or time is crucial to control for their influence on the exposure-outcome relationship, but overmatching on intermediate variables (e.g., matching on body condition score when studying diet and metabolic disease) or variables strongly associated with exposure can induce bias and attenuate the true effect.
- Common biases in veterinary case-control studies include selection bias (e.g., referral bias where animals with certain exposures are more or less likely to be referred), recall bias (owners of affected animals remembering exposures more thoroughly), and information bias (differential misclassification of exposure or outcome), which can be mitigated through careful protocol design and objective data collection methods like using medical records or validated diagnostic tests (e.g., ELISA for antibody titers, PCR for pathogen detection).
- Reporting standards like STROBE-Vet are essential for transparency, requiring detailed documentation of case definitions (e.g., histopathological confirmation of neoplasia, positive RT-PCR for influenza), control selection methods, matching strategies, and bias assessment, enabling readers to critically evaluate the study's validity and generalizability across species and production systems.
This article provides a procedural reference for veterinary researchers designing, conducting, and interpreting case-control studies in animal populations. It addresses the logic of the design, the principles of control selection, matching strategies, and the analytical framework centered on the odds ratio. The intended reader is a veterinarian or graduate student with working knowledge of epidemiologic concepts who requires practical guidance on study architecture and common pitfalls. The scope is cross-species, with attention to differences between companion animal, livestock, and wildlife settings where they affect design decisions. Cohort methods are excluded, the focus rests entirely on the retrospective, outcome-based sampling that defines the case-control approach.
At a Glance
| Parameter | Decision or Fact |
|---|---|
| Core design logic | Sample on outcome status, measure exposure retrospectively |
| Primary measure of association | Odds ratio, with 95% confidence interval |
| Control selection goal | Controls must represent the population that produced the cases |
| Matching | Individual or frequency matching on confounders such as age, sex, herd, or time |
| Overmatching risk | Matching on an intermediate variable or exposure proxy attenuates the exposure effect |
| Key biases | Selection bias, recall bias, information bias, and confounding |
| Minimum reporting standard | STROBE-Vet statement for observational veterinary studies |
| Validity assessment | Risk of bias tools specific to case-control designs |
Conceptual Foundation of the Case-Control Design
A case-control study identifies individuals with the outcome of interest, the cases, and a comparison group without that outcome, the controls, then compares their prior exposure histories. The design is inherently retrospective in exposure ascertainment, although the underlying exposure may have been recorded prospectively before outcome occurrence, as in a nested case-control study conducted within an established cohort. The defining feature is sampling on outcome status, which makes the design efficient for rare diseases, long-latency outcomes, and conditions where follow-up would be prohibitively expensive or prolonged.
The case-control design is particularly valuable in veterinary medicine for investigating disease outbreaks, uncommon neoplasms, and adverse effects of interventions. In production animal settings, the design supports investigation of herd-level exposures associated with disease clusters. In companion animal practice, it enables study of conditions such as lymphoma or renal failure where incidence is low and prospective enrollment would require impractically large populations. The efficiency of the design comes at a cost: the investigator must exercise rigorous control over selection and measurement to preserve validity. Methodological quality assessment tools for observational studies, including those specific to case-control designs, should be applied during protocol development, not after data collection is complete, as described in the review of risk of bias assessment instruments by Ma and colleagues.
The Odds Ratio as the Measure of Association
The odds ratio approximates the relative risk when the disease is rare in the source population, a condition known as the rare disease assumption. In a case-control study, the odds of exposure among cases is compared with the odds of exposure among controls. The resulting ratio estimates the odds of disease given exposure relative to the odds of disease given no exposure. When disease incidence is low, typically below 10 percent, the odds ratio closely approximates the risk ratio. When the outcome is common, the odds ratio overestimates the relative risk, and the investigator should interpret the magnitude accordingly.
The odds ratio is calculated from the familiar two-by-two table crossing exposure status with case or control status. Unadjusted estimates are obtained by direct calculation. Adjusted estimates require multivariable logistic regression, which permits simultaneous control of multiple confounders. The logistic model treats case status as the dependent variable and exposure plus covariates as predictors. The exponentiated coefficient for the exposure term yields the adjusted odds ratio. Confidence intervals derived from the model reflect the precision of the estimate and should always accompany the point estimate.
Selection of Cases
Case definition is the first and most consequential decision. The definition must be explicit, reproducible, and based on objective criteria. Histopathologic confirmation, culture results, imaging findings, or a validated clinical scoring system are preferable to subjective clinical judgment alone. The investigator should state the diagnostic criteria in the protocol and apply them identically throughout the study. For conditions with a spectrum of severity, the case definition should specify whether mild, subclinical, or atypical presentations are included. Including a heterogeneous case group dilutes the exposure effect if different disease subtypes have different etiologies. Restricting the case definition to a homogeneous subtype increases power but reduces generalizability.
Prevalent cases, those identified at a single point in time, differ from incident cases, those identified at the time of first diagnosis. Prevalent cases are easier to recruit but introduce survivor bias. Animals that died rapidly from the disease are excluded, and long-term survivors are overrepresented. If survival is related to exposure, the odds ratio will be distorted. Incident case ascertainment through active surveillance or prospective identification within a defined population is preferred whenever feasible. The source population from which cases arise must be clearly defined, because the control group must be drawn from that same population.
Control Selection in Veterinary Case-Control Studies
Control selection determines the validity of a case-control study more than any other design decision. The goal is to sample individuals from the population that gave rise to the cases, such that controls represent the exposure distribution in that source population. In veterinary practice, the source population is often implicit instead of enumerated, which creates the central challenge.
Defining the Source Population
The source population is the set of animals that would have been identified as cases had they developed the disease during the study period. For a referral hospital study, the source population includes all animals in the hospital's catchment that would have been referred for the condition of interest. For a production animal study, the source population may be all herds enrolled in a health scheme or all animals in a defined geographic region.
Three practical approaches exist for defining the source population in veterinary studies:
| Approach | Definition | Suitable When | Common Failure Mode |
|---|---|---|---|
| Hospital-based | All animals presenting to the same institution(s) during the study period | Referral populations, uncommon diseases, limited resources | Referral bias distorts exposure prevalence |
| Population-based | All animals in a defined geographic or administrative area | Notifiable diseases, production animal populations, well-characterized catchments | Incomplete case ascertainment |
| Nested within cohort | Cases arising from an existing prospective cohort | Biobanked samples exist, longitudinal data already collected | Cohort attrition reduces available controls |
The nested approach deserves particular attention. When a case-control study is nested within an existing cohort, controls are sampled from cohort members who remained at risk at the time the case was diagnosed. This design preserves the temporal relationship between exposure and outcome and reduces recall bias because exposure data are collected prospectively. The CDC principles of epidemiology in public health practice describe the nested design as a hybrid that retains many strengths of the cohort approach while maintaining the efficiency of case-control sampling.
Control-to-Case Ratio and Sampling Methods
A 1:1 ratio provides the minimum statistical efficiency, but increasing the ratio improves power. The marginal gain diminishes beyond 4:1, and ratios above this are rarely justified unless controls are inexpensive to ascertain and the disease is rare. For veterinary studies where diagnostic confirmation of controls may require testing, the cost of additional controls must be weighed against the modest power gain.
Controls should be sampled using incidence density sampling, also called risk-set sampling. Under this scheme, for each case, controls are selected from animals that were at risk of developing the disease at the time the case was diagnosed. This approach allows the odds ratio to estimate the incidence rate ratio directly, provided the disease is rare. An alternative, cumulative sampling, selects controls from animals that remained disease-free at the end of the study period. Cumulative sampling can produce biased estimates when exposure prevalence changes over time, and it is generally inferior for veterinary applications where disease incidence varies seasonally or by production cycle.
Time-based matching is often necessary. Matching controls to cases on calendar time, such as year of diagnosis or season, controls for temporal changes in exposure prevalence, diagnostic practices, and referral patterns. In production animal medicine, matching on production cycle or cohort is frequently essential because management changes over time.
Matching Strategies and Their Consequences
Matching selects controls that are similar to cases on specified variables. The purpose is to control confounding by the matching variables and to improve efficiency in the analysis. Matching does not eliminate confounding, it requires that the analysis account for the matching design.
Individual Matching
Individual matching assigns one or more controls to each case based on exact or approximate values of the matching variables. Common matching variables in veterinary studies include age, breed, sex, herd or flock, and geographic location. Age matching is frequently necessary because many diseases have strong age associations and because age may influence exposure accumulation. Breed matching controls for genetic susceptibility and breed-related management practices.
The matching ratio may be fixed, such as 2 controls per case, or variable. Variable ratios occur when some cases have more available controls than others. The analysis must accommodate variable ratios, typically through conditional logistic regression.
Frequency Matching
Frequency matching, also called category matching, selects controls so that the distribution of matching variables matches that of the cases. For example, if 40% of cases are dairy cattle and 60% are beef cattle, controls are selected to reflect this same proportion. Frequency matching is simpler to implement than individual matching and is often preferred when the matching variables are categorical.
Risks of Overmatching
Matching on a variable that is not a confounder reduces precision without improving validity. More seriously, matching on an intermediate variable, one that lies on the causal pathway between exposure and outcome, induces bias. For example, matching on body condition score in a study of dietary risk factors for ketosis would obscure the effect of diet because body condition is partly determined by diet.
A further risk is that matching on a variable strongly associated with exposure can create selection bias. This occurs because the control group's exposure distribution is forced to resemble the case group's distribution. The problem is most severe when the matching variable is associated with exposure but not independently with disease. Methodological quality assessment tools for primary and secondary medical studies emphasize that the choice of matching variables should be justified by prior knowledge of the causal structure, not by convenience.
Analysis Implications
When matching is used, the analysis must reflect the design. Unmatched analysis of matched data produces odds ratios biased toward the null. Conditional logistic regression is the standard approach for individually matched studies. For frequency-matched designs, the matching variables should be included as covariates in an unconditional logistic regression model.
Sources of Bias in Control Selection
Selection bias arises when the probability of being selected as a control is related to exposure status. In veterinary studies, the most common sources are referral patterns, owner compliance, and diagnostic intensity.
Referral bias occurs when animals with certain exposures are more or less likely to be referred to a specialist center. A study of risk factors for gastric dilatation-volvulus conducted at a referral hospital may overrepresent large-breed dogs if smaller breeds with the condition are managed in primary care. The control group drawn from the same referral population may not reflect the exposure distribution of the underlying source population.
Survivorship bias affects case-control studies when prevalent instead of incident cases are enrolled. Prevalent cases have survived with the disease, and their exposure history may differ systematically from that of incident cases. Enrolling incident cases, those newly diagnosed within a defined period, reduces this problem.
Recall bias is less prominent in veterinary studies because exposure information is often obtained from medical records or owners instead of from the animals themselves. However, owners of affected animals may search their memory more thoroughly for potential causes, and they may have altered management in response to early signs of disease. The estrogen replacement therapy and risk of Alzheimer disease study illustrates how exposure measurement can be affected by disease status, a concern that applies equally to veterinary studies relying on owner-reported exposures.
Odds Ratio Calculation with Veterinary Examples
The odds ratio is calculated from the standard 2 by 2 table:
| Cases | Controls | |
|---|---|---|
| Exposed | a | b |
| Unexposed | c | d |
The odds ratio is (a × d) / (b × c). The interpretation depends on the sampling design. In an unmatched case-control study with incident cases and population-based controls, the odds ratio estimates the incidence rate ratio when the disease is rare. When the disease is not rare, the odds ratio overestimates the rate ratio, and the magnitude of overestimation increases with disease frequency.
Consider a study of risk factors for bovine respiratory disease in feedlot cattle. Cases are 120 animals diagnosed with respiratory disease within 30 days of arrival. Controls are 240 animals from the same feedlots that remained healthy during the same period, matched on arrival date and source. Exposure is defined as transport duration exceeding 12 hours.
| Cases | Controls | |
|---|---|---|
| Transport > 12 h | 72 | 96 |
| Transport ≤ 12 h | 48 | 144 |
The odds ratio is (72 × 144) / (96 × 48) = 10368 / 4608 = 2.25. Animals transported for more than 12 hours have 2.25 times the odds of developing respiratory disease compared with animals transported for shorter durations, assuming the matching was accounted for in the analysis.
For matched designs, the analysis uses discordant pairs. In a 1:1 matched study, only pairs where the case and control differ in exposure status contribute information. The matched odds ratio is the ratio of discordant pairs where the case is exposed to pairs where the control is exposed. This is the McNemar odds ratio, and it is calculated as b / c, where b is the number of pairs with an exposed case and unexposed control, and c is the number of pairs with an unexposed case and exposed control.
The plasma sphingomyelin level study demonstrates the use of multivariate logistic regression to adjust for multiple risk factors simultaneously. In veterinary studies, logistic regression allows adjustment for continuous variables such as age or body weight that cannot be matched exactly, and it permits the inclusion of interaction terms when effect modification is suspected.
Practical Documentation and Reporting
The reporting of a case-control study should allow readers to assess the validity of control selection. The following elements should be documented:
- The source population and how it was defined
- The sampling frame for controls and the method of selection
- The matching variables and the rationale for each
- The control-to-case ratio and whether it was fixed or variable
- The time period of case ascertainment and control sampling
- The response rates for cases and controls, including refusals and losses
- The method used to verify that controls were at risk during the study period
The WOAH animal health surveillance standards emphasize the importance of documenting case definitions and population denominators in surveillance activities. The same principle applies to case-control studies: the case definition must be applied identically to cases and controls, and the control selection procedure must be described with sufficient detail to permit replication.
Species-specific considerations affect control selection. In companion animal studies, controls are often selected from the same hospital population, and the choice of control group, such as animals with a different disease versus healthy animals, influences the interpretation. In production animal studies, herd-level clustering must be addressed, and controls are frequently selected from the same herds as cases to control for herd-level factors. In wildlife studies, sampling logistics and the inability to enumerate the source population often force pragmatic compromises that should be acknowledged as limitations.
Recognized Failure Modes and Early Detection
Case-control studies in veterinary settings fail in predictable patterns. The most consequential failure is selection bias arising from control recruitment that does not reflect the exposure distribution of the source population. This occurs when controls are drawn from a referral hospital while cases originate from primary care practices, or when controls are recruited from owners who respond to advertisements while cases are identified through active surveillance. Detection requires comparing the demographic and geographic profile of controls against the catchment population described in the practice or laboratory database. A control group that is younger, more urban, or skewed toward one breed than the case group signals a recruitment problem.
Recall bias is the second major failure mode. Owners of affected animals search their memory for exposures with greater intensity than owners of unaffected animals, particularly when the exposure is plausibly linked to the disease. This asymmetry inflates the odds ratio. Early detection is possible through the inclusion of a sham exposure question, one that has no plausible biological link to the outcome. If the sham exposure shows an elevated odds ratio, recall bias is operating.
Misclassification of disease status is a third failure mode. Cases confirmed by histopathology or culture are reliable, but cases diagnosed on clinical signs alone may include phenocopies that dilute the association. Detection requires a diagnostic confirmation protocol applied uniformly to cases and controls, with a stated proportion of cases that received definitive testing. If that proportion falls below 80%, the case definition should be revised.
Common Errors and Corrective Actions
Less experienced investigators frequently confuse the case-control design with a cohort study and attempt to calculate incidence or risk. The case-control design yields only the odds ratio, which approximates the relative risk only when the disease is rare in the source population. When the disease is common, the odds ratio overstates the association. The corrective action is to report the odds ratio with its confidence interval and to state explicitly that incidence cannot be estimated from this design.
A second common error is matching on variables that are also exposure pathways. Matching on a variable that lies on the causal pathway between exposure and disease removes the very association under study. For example, matching on body condition score when investigating dietary risk factors for obesity-related neoplasia will attenuate the odds ratio toward the null. The corrective action is to construct a directed acyclic graph before matching decisions are made, identifying which variables are confounders, which are mediators, and which are colliders.
A third error is the use of hospital controls without considering why those animals were hospitalized. Animals admitted for trauma may have different exposure histories than the general population, and animals admitted for vaccination may have systematically different preventive care. The corrective action is to document the admission diagnoses of controls and to test whether the odds ratio changes when controls with diagnoses plausibly related to the exposure are excluded.
Limitations of the Evidence Base
The veterinary literature contains relatively few case-control studies with adequate sample sizes, and the quality of exposure measurement varies widely. Many studies rely on owner-reported exposures collected years after the disease event, which introduces both recall bias and measurement error. Exposure validation studies, in which owner reports are compared against medical records or environmental sampling, are uncommon in veterinary settings. The evidence base is therefore strongest for diseases with objective exposure measures, such as serological status or tissue residues, and weakest for behavioral or management exposures.
Expert opinion differs on the acceptability of hospital-based controls. Some authorities argue that hospital controls are acceptable when the catchment populations of cases and controls are demonstrably similar, while others require population-based controls drawn from census or registration data. The World Organization for Animal Health surveillance standards emphasize that control selection must be documented in sufficient detail to permit external evaluation of the study's validity WOAH animal health surveillance standards. The CDC principles of epidemiology similarly stress that the source population must be defined before controls are selected CDC principles of epidemiology in public health practice.
Troubleshooting Table
| Observation | Likely Cause | Discriminating Check |
|---|---|---|
| Odds ratio above 10 with wide confidence interval | Selection bias or recall bias | Compare control demographics with source population, check sham exposure result |
| Odds ratio near 1.0 despite strong prior evidence | Overmatching on causal pathway variable | Review matching variables against directed acyclic graph |
| Controls younger or healthier than cases | Differential recruitment | Compare age distribution and comorbidity counts between groups |
| Exposure prevalence implausibly high in cases | Recall bias or leading questionnaire design | Blinded interviewers, objective exposure records |
| Odds ratio changes when hospital controls excluded | Hospital control selection bias | Stratified analysis by control source |
Referral, Consultation, and Reporting
When a case-control study informs regulatory decisions, such as the declaration of a notifiable disease or the withdrawal of a product from the market, the investigator should consult the relevant veterinary authority before publication. The WOAH terrestrial animal health code specifies reporting obligations for listed diseases, and these obligations apply regardless of study design WOAH terrestrial animal health code. Statistical consultation is warranted when the analysis plan includes conditional logistic regression for matched data, when exposure variables are correlated, or when the number of cases is small relative to the number of covariates. Laboratory involvement is required when exposure measurement depends on assays that have not been validated for the species under study. A veterinary epidemiologist with experience in case-control methodology should be consulted before the study begins, not after data collection is complete.
Frequently Asked Questions
How many controls per case are needed when the budget is tight?
A 1:1 ratio is the minimum acceptable design and remains statistically efficient for detecting large effects. Increasing to 2 or 3 controls per case improves statistical power, but the marginal gain diminishes beyond 4 controls per case. When resources are constrained, prioritize control quality over quantity. A single well-defined control from the same source population is more valuable than several poorly characterized controls. If the disease is rare and cases accumulate slowly, consider frequency matching instead of individual matching to reduce administrative cost. Document the achieved ratio and its effect on precision in the final report. The CDC principles of epidemiology provide practical guidance on matching and sample size decisions.
Can a case-control study be done when diagnostic records are incomplete?
Yes, but the study must be redesigned around what can be verified. Restrict case definition to animals with laboratory-confirmed diagnosis or histopathology, even if this reduces case numbers. Use a single pathologist or laboratory to confirm diagnoses where possible. For exposure data, rely on records that were created before diagnosis, such as herd health logs, feed delivery receipts, or vaccination records, to reduce recall bias. If exposure information is missing for more than 20% of eligible cases, assess whether missingness is related to exposure status. Consider a pilot review of 20 to 30 records to estimate data completeness before full enrollment. The risk of bias assessment tools described by Ma and colleagues can help identify which design elements are most vulnerable to poor data quality.
How do I select controls when the source population includes multiple herds or premises?
Controls should be sampled from the same herds or premises that produced the cases, in proportion to their contribution to the case series. If cases cluster in a few herds, use frequency matching on herd to prevent confounding by herd-level factors such as management style or regional disease pressure. When a case arises from a herd, select one or more controls from that same herd within a defined time window. This approach controls for stable environmental and management factors without requiring measurement. For diseases with long latent periods, such as bovine spongiform encephalopathy or chronic neoplasia, consider whether controls should be matched on age cohort to ensure equal opportunity for exposure. The WOAH terrestrial animal health standards describe surveillance frameworks that can inform control selection in production animal settings.
What should I do when the ideal laboratory assay is unavailable for exposure measurement?
Use the best available assay and state its limitations explicitly. A less specific assay will bias the odds ratio toward the null, while a less sensitive assay may bias it away from the null if misclassification differs between cases and controls. If the assay is expensive, consider testing all cases but a random subset of controls, then adjust the analysis for the sampling fraction. Store additional samples for future testing if a better assay becomes available. Blinding the laboratory to case-control status prevents differential misclassification. When exposure is measured from stored samples, verify that storage conditions and duration are comparable between cases and controls. The nested case-control study of organochlorine residues and breast cancer demonstrates how stored specimens can be analyzed after case identification while maintaining blinding.
How do I explain the study design to a referring veterinarian or herd owner?
Frame the study in terms of the question it answers: whether a specific exposure is more common in affected animals than in unaffected animals from the same population. Explain that controls are not healthy animals in an absolute sense, but animals that would have been included as cases if they had developed the disease. Emphasize that participation requires no change in clinical care and that individual animal results will not be reported back in identifiable form. Describe the time commitment honestly, including any sample collection or record review. For herd owners, clarify that the study cannot prove causation and that results will inform future prevention strategies instead of immediate treatment decisions. The AVMA practice resources offer communication guidance for discussing research participation with clients.
When is a case-control study the wrong design for my research question?
A case-control study is unsuitable when the exposure is rare in the source population, because controls will have very few exposed individuals and the odds ratio will be imprecise. It is also inappropriate when the disease has a long and variable latent period and historical exposure data are unreliable. If you need to estimate incidence or attributable risk in the population, a cohort or cross-sectional design is required. Case-control studies cannot establish temporal sequence when exposure and outcome are measured simultaneously, so they are weak for rapidly progressive diseases where subclinical exposure may have altered behavior or management. For emerging diseases with no validated case definition, defer the study until diagnostic criteria are established. The MSD Veterinary Manual provides disease-specific guidance on diagnostic confirmation that can inform case definitions.
Related Clinical & Scientific Guides
- Evaluating Veterinary Surveillance System Attributes
- Network Analysis for Infectious Disease Spread in Animal Populations
- Randomized Controlled Trials in Veterinary Field Settings
References and Further Reading
- Methodological quality (risk of bias) assessment tools for primary and secondary medical studies: what are they and which is better?. 2020.
- Estrogen replacement therapy and risk of Alzheimer disease.. 1996.
- Blood levels of organochlorine residues and risk of breast cancer.. 1993.
- Plasma sphingomyelin level as a risk factor for coronary artery disease.. 2000.
- Estrogen plus progestin and risk of venous thrombosis.. 2004.
- The causes of cancer: quantitative estimates of avoidable risks of cancer in the United States today.. 1981.
- WOAH Animal Health Surveillance Standards. WOAH.
- CDC Principles of Epidemiology in Public Health Practice. CDC.
- MSD Veterinary Manual, Professional Edition. MSD Veterinary Manual.
Related Articles
- Understanding Ecological Studies in Veterinary Epidemiology
- Cohort Studies in Veterinary Medicine: Design and Analysis
- Confounding in Veterinary Studies: Identification and Control
- Understanding Bias in Veterinary Epidemiological Studies
- Designing Cross-Sectional Studies in Veterinary Populations
This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.