Diagnostic Test Accuracy Studies in Veterinary Medicine: Design and Reporting
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Diagnostic test accuracy studies critically evaluate how well a test distinguishes diseased from non-diseased animals, informing decisions from herd culling to individual treatment. Primary accuracy measures are sensitivity and specificity, reported with confidence intervals, and the choice of threshold significantly impacts this balance.
- The reference standard, the benchmark for true disease status, must be defined a priori and its limitations, particularly the absence of a perfect ante-mortem standard for diseases like paratuberculosis or bovine tuberculosis, must be explicitly handled and reported.
- Cross-sectional cohort designs are preferred for accuracy studies as they preserve the disease spectrum, unlike case-control designs which tend to inflate accuracy estimates by excluding atypical cases.
- Bias sources such as verification bias (reference testing only positive index tests), spectrum bias (unrepresentative population), and incorporation bias (index test part of reference standard) can distort accuracy estimates and require careful mitigation and reporting.
- Reporting standards like STARD (for human medicine) and veterinary-specific adaptations (e.g., STRADAS-paraTB) are crucial for ensuring transparency, interpretability, and reproducibility of diagnostic accuracy study results, detailing population characteristics, reference standards, and blinding procedures.
- Test performance is not a fixed property but varies by species, production system, disease prevalence, and the chosen threshold, necessitating species-specific validation and careful consideration of the intended clinical use when interpreting results.
Diagnostic test accuracy studies answer a deceptively simple question: how well does a test distinguish animals with a target condition from those without it? The answer determines whether a test can be used for screening, confirmation, or monitoring, and it shapes clinical decisions from herd-level culling to individual patient treatment. This article provides veterinary researchers with a structured approach to designing, conducting, and reporting diagnostic test accuracy studies across species. It covers the conceptual foundations of accuracy measurement, the selection of reference standards, the handling of bias and variation, and the reporting standards that make study results interpretable and reproducible.
The reader is assumed to be familiar with clinical epidemiology and statistical inference. The article does not cover therapeutic trials, and it does not provide species-specific test protocols. Instead, it offers a transferable framework for evaluating any diagnostic modality, from point-of-care devices to laboratory immunoassays, in any domestic or production animal population. The guidance draws on established reporting standards and on systematic reviews of veterinary diagnostic tests that illustrate both best practice and common pitfalls.
At a Glance
| Parameter | Decision or Fact |
|---|---|
| Primary accuracy measures | Sensitivity and specificity, reported with confidence intervals |
| Reference standard | Must be defined a priori, imperfect reference standards require explicit handling |
| Study design | Cross-sectional cohort preferred, case-control designs inflate accuracy estimates |
| Population | Spectrum of disease severity and stage must reflect target use population |
| Reporting standard | STARD for human medicine, STRADAS-paraTB for paratuberculosis, EQUATOR network for general guidance |
| Bias sources | Verification bias, incorporation bias, spectrum bias, and partial verification |
| Threshold selection | ROC analysis or Youden index, thresholds must be pre-specified or clearly justified |
| Meta-analysis | Hierarchical summary ROC (HSROC) methods recommended for pooling accuracy data |
| Species variation | Test performance differs by species, production system, and disease prevalence |
The Logic of Diagnostic Accuracy
A diagnostic test accuracy study measures the agreement between an index test and a reference standard in a defined population. The index test is the test under evaluation. The reference standard is the best available method for determining true disease status, which may be a combination of clinical examination, laboratory testing, culture, histopathology, or post-mortem findings. The fundamental challenge is that the reference standard is rarely perfect. In veterinary medicine, ante-mortem reference standards with high sensitivity and specificity are often unavailable, particularly for chronic infections such as paratuberculosis or bovine tuberculosis. This limitation must be acknowledged in the study design and in the interpretation of results, as it directly affects the validity of accuracy estimates.
Accuracy is expressed through sensitivity and specificity. Sensitivity is the proportion of diseased animals correctly identified by the index test. Specificity is the proportion of non-diseased animals correctly identified. These two measures are inversely related, and the relationship between them determines the false-positive and false-negative proportions of a test. The choice of threshold, where the index test result is dichotomised into positive and negative, shifts this balance. A low threshold increases sensitivity at the cost of specificity, while a high threshold does the reverse. The optimal threshold depends on the clinical consequence of each error type, which differs by context. In a screening program for a notifiable disease, sensitivity may be prioritized to avoid missing infected animals, whereas in a confirmatory setting, specificity may be more important to avoid culling healthy animals.
Reference Standards and Their Limitations
The reference standard defines the ground truth against which the index test is compared. Its selection is the single most consequential design decision in a diagnostic accuracy study. A reference standard with imperfect sensitivity will misclassify diseased animals as non-diseased, causing the index test to appear less sensitive than it truly is. An imperfectly specific reference standard will misclassify non-diseased animals as diseased, causing the index test to appear less specific. These effects are predictable and can be modelled, but they cannot be eliminated without a perfect reference standard.
In livestock medicine, the absence of a highly sensitive and specific ante-mortem reference standard is a recurring problem. For paratuberculosis, fecal culture and PCR are commonly used, but both have limited sensitivity in subclinically infected animals. The consensus-based reporting standards developed for paratuberculosis, termed STRADAS-paraTB, explicitly address this issue by requiring authors to describe the reference standard, its performance characteriztics, and the rationale for its selection. The standards also require authors to state whether the reference standard was applied to all animals or only to a subset, and to describe how discrepancies between the index test and reference standard were resolved.
For bovine tuberculosis, the tuberculin skin test and the gamma-interferon assay are the primary ante-mortem diagnostic tools, but neither achieves perfect accuracy. The relationship between sensitivity and specificity in these tests is well documented, and the choice of test or combination of tests depends on the purpose of testing, the prevalence of disease in the population, and the consequences of false results. A review of ante-mortem tuberculosis diagnosis in cattle emphasizes that no currently available test allows a perfectly accurate determination of infection status, and that test performance is influenced by factors such as the stage of infection, the immune status of the animal, and the strain of Mycobacterium bovis involved.
Study Designs for Accuracy Evaluation
The cross-sectional cohort design is the preferred approach for evaluating diagnostic test accuracy. In this design, a consecutive or random sample of animals from the target population undergoes both the index test and the reference standard within a defined time window. This design preserves the natural spectrum of disease, including the full range of disease severity, duration, and stage, which is essential for estimating accuracy that generalizes to clinical practice.
The case-control design, in which known diseased and known non-diseased animals are selected separately, is generally discouraged for accuracy studies because it inflates sensitivity and specificity estimates. The inflation occurs because the design excludes animals with mild or atypical disease and animals with conditions that mimic the target disease, both of which are common sources of misclassification in practice. Case-control designs may be acceptable for preliminary evaluations of a new test, but their results should not be used to guide clinical decisions without confirmation in a cross-sectional study.
Experimental challenge studies are sometimes used in veterinary medicine to generate animals with known infection status. These studies offer the advantage of a defined infection time and a confirmed reference standard, but they do not represent the natural spectrum of disease. Challenge models typically produce acute, high-dose infections that differ immunologically and pathologically from natural exposure. The STRADAS-paraTB guidelines acknowledge the potential use of experimental challenge studies while requiring authors to describe the challenge strain, dose, route, and timing, and to discuss the generalizability of the results to naturally infected populations.
Sampling Strategies and Population Selection
The population from which animals are drawn determines the generalizability of accuracy estimates. A test evaluated in a high-prevalence referral population will show different predictive values than the same test applied to a low-prevalence screening population, even when sensitivity and specificity remain constant. For veterinary accuracy studies, the target population should be defined by species, production class, age, herd health status, and the clinical question the test is intended to answer.
Convenience sampling of animals already presenting for clinical investigation tends to over-represent advanced disease. This spectrum bias inflates sensitivity because severely affected animals produce stronger test signals. Conversely, sampling only healthy herdmates for the non-diseased group can underestimate false positives that would occur in animals with concurrent inflammatory conditions. The consensus-based reporting standards for paratuberculosis diagnostic accuracy studies explicitly address this problem, recommending that authors describe how the study population was assembled and whether it reflects the target population for the intended test use.
For herd-level tests, the sampling unit requires particular attention. A test validated on individual animals may perform differently when applied to pooled samples or when herd-level interpretation is applied. The prevalence of infection within the herd, the within-herd clustering of disease, and the sampling fraction all influence the diagnostic accuracy of herd tests. Studies should report whether sampling was random, systematic, or convenience-based, and should state the number of herds and animals within herds.
Threshold Selection and Its Consequences
Every quantitative test requires a threshold that separates positive from negative results. The choice of threshold is a trade-off between sensitivity and specificity, and the optimal threshold depends on the consequences of misclassification. For a screening test where false negatives allow disease spread, a lower threshold that favours sensitivity may be appropriate. For a confirmatory test where false positives trigger culling or trade restrictions, a higher threshold that favours specificity is usually preferred.
The review of ante mortem tuberculosis diagnosis in cattle illustrates this principle clearly. The single intradermal tuberculin test and the comparative cervical test use different interpretation criteria, and the choice between them depends on whether the testing program prioritizes detection of infected animals or avoidance of false positives in herds vaccinated or exposed to environmental mycobacteria. The same test, read at different thresholds, serves different purposes within the same disease control program.
Thresholds should be prespecified where possible. Post hoc threshold selection, where the threshold is chosen to maximize accuracy in the study sample, produces optimiztic estimates that will not replicate in new populations. If multiple thresholds are evaluated, the analysis should report accuracy at each threshold with confidence intervals, allowing readers to select the threshold appropriate to their context.
Worked Example: Sensitivity, Specificity, and Likelihood Ratios
Consider a study evaluating a point-of-care beta-hydroxybutyrate test for hyperketonemia in dairy cows, using serum BHB concentration as the reference standard. Suppose the test is applied to 200 cows in early lactation. The reference standard classifies 40 cows as hyperketonemic and 160 as non-hyperketonemic. The index test identifies 32 of the 40 affected cows as positive and 152 of the 160 unaffected cows as negative.
Sensitivity is 32 divided by 40, or 80 percent. Specificity is 152 divided by 160, or 95 percent. The positive likelihood ratio is sensitivity divided by (1 minus specificity), which is 0.80 divided by 0.05, or 16. A cow with a positive test result is 16 times more likely to be hyperketonemic than non-hyperketonemic. The negative likelihood ratio is (1 minus sensitivity) divided by specificity, which is 0.20 divided by 0.95, or 0.21. A cow with a negative test result is approximately one-fifth as likely to be hyperketonemic as a non-hyperketonemic cow.
Likelihood ratios are particularly useful because they can be combined with pre-test probability to estimate post-test probability. If the expected prevalence of hyperketonemia in a herd is 20 percent, the pre-test odds are 0.25. Multiplying by the positive likelihood ratio of 16 gives post-test odds of 4.0, corresponding to a post-test probability of 80 percent. This calculation allows clinicians to interpret test results in the context of their own herd prevalence instead of relying on predictive values from the study population.
The systematic review and meta-analysis of point-of-care hyperketonemia tests demonstrates how such estimates are synthesised across studies. The authors used hierarchical summary receiver operating characteriztic methods to compare the accuracy of blood, urine, and milk tests, and identified the optimal threshold for each method. This approach recognizes that sensitivity and specificity are not fixed properties of a test but vary with threshold and population.
A Reporting Checklist for Veterinary Accuracy Studies
The following checklist adapts the STARD framework to veterinary contexts. It is intended for authors preparing manuscripts and for reviewers assessing submissions.
| Item | Reporting requirement | Veterinary-specific considerations |
|---|---|---|
| Study identification | Title and abstract identify the study as diagnostic accuracy | State species and production system |
| Study rationale | Clinical question and intended test use | Screening, confirmatory, or monitoring |
| Study population | Inclusion and exclusion criteria, sampling method | Herd versus individual sampling, disease spectrum |
| Reference standard | Description of reference test and its own accuracy | Note absence of a perfect reference standard where applicable |
| Index test | Full protocol, including threshold and interpretation | State whether threshold was prespecified |
| Blinding | Whether index test and reference standard were read independently | Describe how blinding was maintained |
| Sample size | Justification for number of animals and herds | Account for clustering and expected prevalence |
| Flow diagram | Number of animals at each stage, exclusions, missing results | Include animals with invalid or inconclusive test results |
| Accuracy estimates | Sensitivity, specificity, likelihood ratios with confidence intervals | Report at the threshold used, also at the optimal threshold |
| Harms | Adverse events or practical limitations of the test | Include cost, time, and equipment requirements |
The reporting guidelines available through the EQUATOR Network provide a broader catalogue of standards for related study designs, including systematic reviews and observational studies. Veterinary authors should consult these resources alongside species-specific extensions such as the paratuberculosis standards.
Species and Context Modifications
The correct study design and reporting approach varies with the production system and the purpose of testing. In companion animal practice, accuracy studies typically enrol patients presenting with clinical signs, and the reference standard may be histopathology or long-term follow-up. In production animal medicine, testing is often applied at herd level for surveillance or trade purposes, and the reference standard may be a combination of culture, serology, and post-mortem examination.
The World Organization for Animal Health terrestrial animal health standards specify validation requirements for tests used in international trade. These standards require documentation of diagnostic sensitivity and specificity, repeatability, and reproducibility across laboratories. Tests intended for trade certification must meet higher evidence standards than tests used for within-herd management decisions.
Resource constraints also change the acceptable trade-off between sensitivity and specificity. A low-cost test with moderate accuracy may be preferable to an expensive high-accuracy test if it allows more animals to be tested within a fixed budget. Studies should report the practical costs and logistical requirements of each test so that readers can judge whether the accuracy estimates are achievable in their own setting.
Recognized Failure Modes in Accuracy Studies
Diagnostic accuracy studies fail in predictable patterns. The most consequential failure is verification bias, where only animals with a positive or suspicious index test result undergo reference standard testing. This inflates sensitivity and deflates specificity because false negatives remain undiscovered. Detection requires auditing the study flow diagram for the proportion of enrolled animals that completed reference testing, stratified by index test result.
Spectrum bias arises when the study population does not reflect the target population's disease severity, stage, or comorbidity profile. A test evaluated only in severely affected hospital patients will perform differently in ambulatory screening. The discriminating check is comparison of the study's inclusion criteria and case mix against the intended clinical application. The consensus-based reporting standards developed for paratuberculosis test accuracy studies explicitly require authors to describe the population spectrum and sampling design so readers can judge generalizability Gardner et al., STRADAS-paraTB reporting standards.
Imperfect reference standard bias occurs when the comparator itself misclassifies disease status. In bovine tuberculosis, no ante-mortem test achieves perfect accuracy, and the relationship between sensitivity and specificity determines the false-positive and false-negative proportions for every available assay de la Rua-Domenech et al., review of ante mortem tuberculosis diagnosis. When the reference standard is imperfect, apparent test accuracy is biased toward the reference standard's own error structure. Latent class analysis or discrepant resolution should be considered, though both introduce their own assumptions.
Threshold effects and publication bias distort the literature. Studies reporting multiple cut-offs without pre-specification capitalise on chance. Meta-analyzes that include only published studies risk overestimating accuracy because negative or null results are less likely to reach print. The systematic review of point-of-care hyperketonemia tests identified substantial between-study heterogeneity and recommended that future work standardize thresholds and reference definitions before pooling Tatone et al., systematic review of hyperketonemia point-of-care tests.
| Observation | Likely cause | Discriminating check |
|---|---|---|
| Sensitivity far higher than published values | Verification bias or spectrum bias | Audit reference testing completion by index result, compare case mix |
| Specificity collapses in a second population | Threshold overfitting to the derivation sample | Re-analyze at pre-specified cut-offs, validate externally |
| Confidence intervals implausibly narrow | Ignored clustering by herd or repeated measures | Check whether herd-level clustering was modelled |
| Results vary across similar studies | Different reference standards or thresholds | Compare case definitions and cut-off values item by item |
| Accuracy improves with each additional cut-off analyzed | Post hoc threshold selection | Confirm thresholds were pre-registered or justified a priori |
Common Errors and Corrective Actions
Students and early-career researchers frequently conflate diagnostic accuracy with clinical utility. A test with high sensitivity and specificity may still fail to improve outcomes if the disease is rare, the test is costly, or treatment is ineffective. Accuracy parameters describe test performance, not patient benefit. The corrective action is to frame the research question around a clinical consequence, not a statistical property.
A second recurring error is treating the reference standard as infallible. When no gold standard exists, authors should state this explicitly and discuss how imperfect reference classification affects their estimates. The paratuberculosis reporting guidelines were developed precisely because livestock diseases often lack an ante-mortem reference standard with high sensitivity and specificity, and they require authors to describe the reference standard's own limitations Gardner et al., STRADAS-paraTB reporting standards.
Sample size miscalculation is common. Accuracy studies need sufficient numbers of both diseased and non-diseased animals, also a large total sample. A study with 500 healthy animals and 20 diseased animals will produce a precise specificity estimate and an unstable sensitivity estimate. Power calculations should target the width of the confidence interval for each accuracy parameter separately.
Inappropriate handling of correlated data is another frequent fault. Animals from the same herd or repeated measurements from the same animal are not independent observations. Ignoring this clustering produces artificially narrow confidence intervals. Mixed-effects models or cluster-robust variance estimators are required.
Limitations of the Current Evidence
The veterinary diagnostic accuracy literature is thinner than its human counterpart. Many published studies are small, single-center, and use convenience sampling. The systematic review of milk protein biomarkers for mastitis identified only 33 manuscripts meeting inclusion criteria across all dairy ruminant species, and the authors noted substantial heterogeneity in reference definitions and immunoassay formats Giagu et al., systematic review of milk protein mastitis markers. This limits the strength of any pooled estimate.
Expert opinion still differs on several points. Whether experimental challenge studies provide valid accuracy estimates for naturally occurring disease remains contested. Challenge studies offer controlled infection timing and a known disease status, but they may not reproduce the spectrum of lesions, bacterial loads, and host responses seen in field conditions. The paratuberculosis guidelines accommodate both experimental and observational designs but require authors to justify their choice Gardner et al., STRADAS-paraTB reporting standards.
The choice between population-level and individual-level accuracy targets also divides opinion. Herd-level test interpretation, common in paratuberculosis and tuberculosis control programs, requires different statistical handling than individual diagnosis. The tuberculosis literature distinguishes clearly between tests used for herd screening and those used for individual animal certification, and the optimal sensitivity-specificity trade-off differs by purpose de la Rua-Domenech et al., review of ante mortem tuberculosis diagnosis.
Referral, Consultation, and Regulatory Reporting
Referral to a specialist biostatistician or epidemiologist is warranted when the study involves latent class analysis, complex sampling designs, or hierarchical modeling of clustered data. These methods are not reliably implemented without formal statistical training, and errors in model specification are difficult for reviewers to detect.
Laboratory involvement is appropriate when reference standard testing requires specialised facilities, such as mycobacterial culture, histopathology, or molecular typing. Early consultation with the diagnostic laboratory is essential to confirm sample handling requirements, transport conditions, and turnaround times. Laboratories should be blinded to index test results to prevent information bias.
Regulatory reporting obligations arise when a study involves notifiable diseases, animals intended for the food chain, or products subject to licensing oversight. Tuberculosis is a notifiable disease in many jurisdictions, and positive results on screening tests trigger statutory reporting regardless of the study context de la Rua-Domenech et al., review of ante mortem tuberculosis diagnosis. Researchers should consult their institutional animal care committee, national veterinary authorities, and the World Organization for Animal Health terrestrial standards before commencing work on regulated pathogens WOAH terrestrial animal health standards. Requirements differ between countries and production systems, and local regulations take precedence over general guidance.
Frequently Asked Questions
How Do I Choose a Reference Standard When No Perfect Test Exists?
When no gold standard is available, use a composite reference standard that combines multiple independent sources of information, such as culture, histopathology, and clinical history. For paratuberculosis, consensus-based reporting standards explicitly acknowledge the widespread lack of an ante-mortem reference standard with high sensitivity and specificity, and they recommend transparent description of how infection status was defined. Alternatively, latent class analysis can estimate accuracy without a perfect reference, but this requires careful modeling assumptions. Whichever approach you select, state the reference standard's known limitations and justify why it was considered adequate for your study question. The consensus-based reporting standards for paratuberculosis diagnostic accuracy studies provide item-by-item guidance on describing reference standard selection.
What Can I Do When Budget Constraints Limit Sample Size or Laboratory Methods?
Reduce scope deliberately instead of compromising measurement quality. Test fewer animals but maintain the sampling frame's representativeness, or restrict the study to one herd and one season while acknowledging limited generalizability. For hyperketonemia in dairy cows, a systematic review and meta-analysis identified gaps in the literature partly attributable to small primary studies with heterogeneous thresholds, so design your study to report confidence intervals honestly. Consider biobanked samples if available, or collaborate across institutions to share specimens. If a quantitative laboratory method is unaffordable, a validated point-of-care test may serve as the index test, but you must then evaluate it against the best feasible reference standard and report the discrepancy. The systematic review of point-of-care hyperketonemia tests demonstrates how meta-analysis can pool small studies to generate clinically useful accuracy estimates.
How Should I Handle Accuracy Testing When the Target Condition Differs Across Species?
Accuracy parameters do not transfer automatically between species because pathogenesis, immune responses, and infection prevalence differ. For bovine tuberculosis, the review of ante-mortem diagnostic techniques shows that skin test sensitivity and specificity vary with disease stage, prior exposure to environmental mycobacteria, and the population being tested. Re-estimate sensitivity and specificity in each target species and production system instead of citing values from another species. When designing a multi-species study, stratify analyzes by species and report accuracy separately for each stratum. Also consider species-specific reference standards, since post-mortem examination protocols and culture methods may differ. If sample sizes preclude species-specific estimates, state this limitation explicitly and avoid pooled accuracy claims.
What Records Must I Keep During an Accuracy Study?
Maintain a complete audit trail from enrollment through data analysis. Record the sampling frame, inclusion and exclusion criteria applied at each stage, the number of animals approached versus enrolled, and reasons for non-participation. Document the timing of index test and reference standard collection, including any interval between them, since disease progression can cause misclassification. Preserve raw instrument outputs, laboratory reports, and the identities of personnel performing each test. Keep a log of protocol deviations and how they were resolved. The EQUATOR Network reporting guidelines library lists resources that specify the minimum information required for transparent research reporting, and following these standards during data collection makes later reporting straightforward. Store records securely for at least the period required by your institution or funding body.
How Do I Explain Diagnostic Accuracy Results to a Referring Veterinarian or Herd Owner?
Translate sensitivity and specificity into predictive values at the prevalence relevant to their population. Explain that a positive result on a screening test does not confirm disease, particularly in low-prevalence herds where false positives outnumber true positives. Use the bovine tuberculosis example, where the relationship between sensitivity and specificity determines the false-positive and false-negative proportions, to illustrate why no test is perfect. Present likelihood ratios as a practical tool: a likelihood ratio above 10 substantially increases post-test probability, while one below 0.1 substantially decreases it. Recommend confirmatory testing for positive screening results when a second test exists. Provide written summaries that include the test's accuracy estimates, the population in which they were derived, and the uncertainty around those estimates.
When Should I Use a Reporting Guideline instead of a General Style Guide?
Use a reporting guideline whenever you submit a diagnostic accuracy study for publication, regardless of the journal's stated requirements. The STARD statement was developed because incomplete reporting was associated with overly optimiztic estimates of test performance, and veterinary adaptations such as STRADAS-paraTB address species-specific considerations like herd testing and experimental challenge studies. General style guides do not cover study design elements such as blinding, sampling strategy, or threshold selection. Consult the EQUATOR Network to identify the appropriate guideline for your study type before writing. Journals increasingly require completed checklists at submission, and reviewers use them to assess methodological rigour.
Related Clinical & Scientific Guides
- Conducting Systematic Reviews of Veterinary Diagnostic Test Accuracy
- Bias in Veterinary Research: Types, Sources, and Mitigation
- Cluster Randomized Trials in Veterinary Research: Design and Analysis
References and Further Reading
- Consensus-based reporting standards for diagnostic test accuracy studies for paratuberculosis in ruminants.. 2011.
- A systematic review and meta-analysis of the diagnostic accuracy of point-of-care tests for the detection of hyperketonemia in dairy cows.. 2016.
- Ante mortem diagnosis of tuberculosis in cattle: a review of the tuberculin tests, gamma-interferon assay and other ancillary diagnostic techniques.. 2006.
- Enhancing the reporting and transparency of rheumatology research: a guide to reporting guidelines.. 2013.
- Milk proteins as mastitis markers in dairy ruminants - a systematic review.. 2022.
- Termination of Resuscitation Rules and Survival Among Patients With Out-of-Hospital Cardiac Arrest: A Systematic Review and Meta-Analysis.. 2024.
- ARRIVE Guidelines 2.0 for Reporting Animal Research. PLOS Biology, 2020.
- EQUATOR Network Reporting Guidelines. EQUATOR Network.
- MSD Veterinary Manual, Professional Edition. MSD Veterinary Manual.
Related Articles
- Appraising Diagnostic Accuracy Studies in Veterinary Medicine
- Meta-Analysis of Veterinary Diagnostic Test Accuracy
- Conducting Systematic Reviews of Veterinary Diagnostic Test Accuracy
- Conducting Pharmacovigilance Studies in Veterinary Medicine
- Designing Questionnaire Studies for Veterinary Research
This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.