Conducting Systematic Reviews of Veterinary Diagnostic Test Accuracy
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Veterinary diagnostic test accuracy (DTA) systematic reviews address the performance of an index test against a reference standard for a specific target condition in a defined population, crucial for surveillance, trade testing, and individual diagnosis.
- The review question must be precisely formulated using a PICO variant (Population, Index test, Comparator, Target condition, Reference standard), specifying details like species, breed, production class, and the technical characteristics of the index test (e.g., specific immunoassay generation, PCR target gene).
- QUADAS-2 is the standard risk of bias tool, requiring adaptation of signaling questions for veterinary contexts, particularly regarding patient selection (e.g., avoiding case-control designs), index test thresholds, and reference standard independence.
- Data extraction must capture detailed information for a two-by-two table, including index test thresholds, reference standard execution, animal flow, and species-specific characteristics (breed, age, production class, prevalence), as these significantly influence test performance and generalizability.
- Heterogeneity is anticipated and must be explored through meta-regression or narrative synthesis, considering sources such as disease spectrum, reference standard quality, index test thresholds, and study setting, to understand variations in test performance across studies.
- Reporting should adhere to PRISMA-DTA guidelines for the review and STARD for primary studies, ensuring transparency and reproducibility, with particular attention to potential biases like spectrum bias, verification bias, and incorporation bias, which are common in veterinary DTA literature.
Systematic reviews of diagnostic test accuracy (DTA) answer a specific clinical question: how well does an index test identify the target condition in a defined population, relative to an accepted reference standard? In veterinary medicine, these reviews inform test selection for herd health surveillance, regulatory trade testing, and individual patient diagnosis. This article provides a step-by-step procedural guide for veterinary researchers planning, conducting, and reporting such reviews. It covers protocol development, search construction, study selection, risk of bias assessment with QUADAS-2, data extraction, and synthesis planning. Therapeutic systematic reviews are outside the scope of this article.
The reader is assumed to be familiar with clinical epidemiology and diagnostic reasoning but may be new to systematic review methodology. The procedures described apply across species and production systems, from companion animal point-of-care assays to food animal surveillance tests. Where species-specific or regional considerations affect conduct, these are noted explicitly.
At a Glance
| Parameter | Decision or fact |
|---|---|
| Review question format | PICO variant: Population, Index test, Comparator test, Target condition, Reference standard |
| Protocol registration | Register prospectively where a registry accepts veterinary DTA reviews, follow PRISMA reporting when publishing |
| Primary databases | PubMed/MEDLINE, Embase, CAB Abstracts, Web of Science, add regional databases for non-English literature |
| Risk of bias tool | QUADAS-2, with signaling questions tailored to veterinary context |
| Reference standard requirement | Must be defined independently of the index test, imperfect reference standards require explicit handling |
| Heterogeneity assessment | Anticipate sources before data extraction, record test platform, population spectrum, and study setting |
| Reporting guideline | PRISMA for systematic reviews, STARD for included primary studies, ARRIVE for animal research reporting |
| Minimum search reporting | Full search strings for at least one database, with date limits and filters stated |
The Logic of Diagnostic Accuracy Reviews
A diagnostic accuracy review differs fundamentally from a therapeutic review. In a therapeutic review, the intervention is assigned and outcomes are compared between groups. In a DTA review, each primary study cross-tabulates index test results against reference standard results within the same subjects, producing sensitivity and specificity estimates. The unit of analysis is the test result, not the patient or animal, and the measure of interest is the accuracy of that result under specified conditions.
The target condition must be defined with care. In veterinary medicine, the target condition may be an infectious agent detected by PCR, a pathological state confirmed by histopathology, or a production outcome such as mastitis confirmed by culture. The reference standard is the best available method for establishing the true presence or absence of that condition. When no perfect reference standard exists, as with many chronic or subclinical infections, the review must address this limitation explicitly. The EQUATOR Network's reporting guideline library includes STARD, which specifies the minimum information primary diagnostic accuracy studies should report, and reviewers should use STARD when appraising the completeness of included primary studies.
Formulating the Review Question
The review question determines every subsequent step. Use the Population, Index test, Comparator, Target condition, Reference standard framework. Specify the population precisely, including species, breed or production class, age range, disease prevalence context, and clinical setting. An index test evaluated in a referral hospital population will perform differently from the same test applied in a field screening program, and the review question must reflect this.
Define the index test with sufficient technical detail. For example, specify whether the question concerns a particular immunoassay generation, a specific PCR target gene, or a defined cytological scoring system. The MSD Veterinary Manual provides species-specific descriptions of many diagnostic tests and their clinical applications, which can help reviewers clarify the technical scope of their question. Comparator tests are optional but useful when the question concerns whether a new test outperforms an existing one.
Search Strategy Construction
The search must balance sensitivity against precision. A DTA search that is too narrow will miss relevant studies, one that is too broad will return an unmanageable volume of irrelevant records. Begin with the index test terms and the target condition terms, combined with a diagnostic accuracy filter. Published filters exist for human DTA searches, but veterinary applications require modification because animal study terminology differs.
Search at least three databases. PubMed and Embase provide broad biomedical coverage, while CAB Abstracts includes substantial veterinary and agricultural literature. Web of Science adds citation indexing. For food animal and wildlife topics, regional databases may be necessary. The WOAH terrestrial animal health standards define internationally recognized reference tests for many notifiable diseases, and reviewers should consult these standards when identifying appropriate reference standards for regulatory test questions.
Do not restrict by language without justification. Non-English studies may contain relevant accuracy data, particularly for diseases with regional importance. When resources permit, include non-English databases and arrange translation of potentially eligible abstracts. Record the full search strategy for each database, including date limits, filters, and any vocabulary variations. This transparency is required for reproducibility and is a core expectation of the ARRIVE guidelines for reporting animal research, which emphasize complete methodological description.
Study Selection and Eligibility Criteria
Define eligibility criteria before screening begins. These criteria operationalise the review question and must specify the study design, population, index test, reference standard, and reported outcomes. Include only primary studies that report sufficient data to construct a two-by-two table, or that provide sensitivity and specificity with confidence intervals. Case reports, narrative reviews, and expert opinion are excluded.
Selection proceeds in two stages. First, screen titles and abstracts against the eligibility criteria, removing obviously irrelevant records. Second, retrieve full texts of remaining records and assess them against the same criteria. Two reviewers should perform each stage independently, with disagreements resolved by discussion or a third reviewer. Record the number of records excluded at each stage and the reasons for exclusion, as required by PRISMA reporting standards. The EQUATOR Network's reporting guideline library provides the current PRISMA checklist, which specifies the items to report in the final manuscript.
Risk of Bias Assessment with QUADAS-2
QUADAS-2 is the standard tool for assessing risk of bias in diagnostic accuracy studies. It evaluates four domains: patient selection, index test, reference standard, and flow and timing. Each domain is assessed for risk of bias, and the first three domains are also assessed for applicability concerns. Signaling questions guide the assessment, and reviewer judgments are recorded with supporting text.
In veterinary applications, several signaling questions require adaptation. Patient selection should consider whether the study used a case-control design, which can inflate accuracy estimates, or a consecutive series, which better reflects clinical populations. The index test domain must consider whether thresholds were prespecified or derived from the data, a particular concern for quantitative assays where cut-offs may be optimized post hoc. The reference standard domain must verify that the reference standard was applied independently of index test results and that it is appropriate for the target condition. Flow and timing addresses whether all animals received the reference standard and whether the time interval between index test and reference standard was clinically acceptable.
The QUADAS-2 tool as applied in published systematic reviews demonstrates how signaling questions are operationalised across different clinical contexts. Reviewers should pilot the tool on several included studies before full application, and disagreements between reviewers should be resolved by consensus. The results of the risk of bias assessment should be presented graphically and used to inform sensitivity analyzes in the synthesis phase.
Data Extraction and Verification
Extraction forms for diagnostic test accuracy reviews differ from those used in therapeutic reviews. Capture the following domains for each included study: index test description including threshold, comparator or reference standard including its execution and interpretation, patient or animal flow through the study, and the two-by-two table of test results against the reference standard. Record whether the index test and reference standard were interpreted independently and whether the reference standard was applied to all subjects regardless of index test result.
Two reviewers should extract data independently into a piloted form. Discrepancies are resolved by discussion or a third reviewer. When a study reports multiple thresholds or multiple index tests, extract each as a separate data row with explicit annotation of the clinical context. For veterinary studies, record species, breed, age, production class, and disease prevalence in the study population, as these modify test performance and generalizability.
Verify numerical data against the original publication. Common errors include transposing sensitivity and specificity, using the wrong denominator for prevalence, and copying confidence intervals from the wrong row of a table. When a study reports insufficient data to construct the two-by-two table, contact the corresponding author. If no response is received, document the missing data and consider whether the study can be included in a meta-analysis or only in narrative synthesis.
Data Synthesis and Heterogeneity
Synthesis begins with a descriptive summary of study characteriztics and test performance before any pooling is considered. Forest plots of sensitivity and specificity with confidence intervals display the range of results across studies. When studies are sufficiently homogeneous in population, index test, and reference standard, meta-analysis using bivariate random effects models or hierarchical summary receiver operating characteriztic (HSROC) models is appropriate. These methods preserve the correlation between sensitivity and specificity across studies and produce a summary ROC curve instead of a single pair of summary estimates.
Heterogeneity is expected in diagnostic accuracy reviews and must be explored instead of ignored. Sources include differences in disease spectrum, reference standard quality, index test threshold, and study design. Meta-regression can examine whether test performance varies by species, sample type, or disease stage. When heterogeneity is substantial, pooling may be misleading and narrative synthesis with tabulated results is preferable. The decision to pool should be made on clinical and methodological grounds before statistical criteria are applied.
Publication bias assessment in diagnostic accuracy reviews is less developed than in therapeutic reviews. Funnel plot asymmetry may reflect small-study effects, but the interpretation is complicated by the bivariate nature of sensitivity and specificity. Report what was done and acknowledge the limitations of these methods in the veterinary context.
The Protocol Template
A protocol should be registered before the review begins. The template below follows the structure expected by systematic review registries and journal reviewers.
| Protocol element | Content required |
|---|---|
| Title | Population, index test, reference standard, and review type |
| Background | Clinical rationale, current test use, and known uncertainty |
| Objectives | Primary and secondary questions in PICO or PIRT format |
| Eligibility criteria | Population, index test, reference standard, study design, language, publication date |
| Search strategy | Databases, date ranges, search terms, and filters for each database |
| Study selection | Screening process, number of reviewers, disagreement resolution |
| Risk of bias assessment | QUADAS-2 domains and signaling questions, tailored to the review |
| Data extraction | Items to be extracted, piloting process, verification steps |
| Synthesis | Planned descriptive and quantitative methods, heterogeneity exploration |
| Sensitivity analyzes | Planned analyzes to test robustness of findings |
| Reporting | PRISMA flow diagram, PRISMA-DTA checklist, and dissemination plan |
The protocol should state how studies with missing data will be handled, how multiple reports of the same study will be identified, and how studies in languages other than English will be managed. Veterinary reviews may need to include grey literature such as conference abstracts and institutional reports, and the protocol should specify the approach to these sources.
A Tailored QUADAS-2 Checklist for Veterinary Diagnostic Studies
QUADAS-2 assesses four domains: patient selection, index test, reference standard, and flow and timing. The standard tool requires tailoring to veterinary contexts. The signaling questions below address species-specific and production-system-specific concerns.
Domain 1: Patient selection
- Was a consecutive or random sample of animals enrolled?
- Was a case-control design avoided? If used, is the justification explicit and does it risk spectrum bias?
- Did the study avoid inappropriate exclusions, such as animals with equivocal results or animals with comorbidities?
- Is the study population described with sufficient detail on species, breed, age, sex, and production class?
- Does the spectrum of disease severity and stage reflect the population in which the test will be used?
Domain 2: Index test
- Was the index test described with sufficient detail to allow replication, including sample type, collection method, and storage conditions?
- Was the threshold prespecified or determined from the data? Data-derived thresholds overestimate accuracy.
- Were index test results interpreted without knowledge of the reference standard?
- Were the personnel performing and interpreting the test blinded to clinical status and other test results?
- Was the test performed and interpreted by personnel with appropriate training and experience?
Domain 3: Reference standard
- Is the reference standard acceptable for the target condition in the species studied?
- Were reference standard results interpreted without knowledge of the index test?
- Was the reference standard applied consistently across all animals?
- For latent or chronic infections, was the reference standard capable of detecting the full spectrum of disease?
- For production animals, does the reference standard reflect the relevant production or trade outcome?
Domain 4: Flow and timing
- Was the time interval between index test and reference standard appropriate for the disease course?
- Did all animals receive the same reference standard regardless of index test result?
- Were all enrolled animals included in the analysis, or were withdrawals and missing results reported and explained?
- Was the disease status stable between index test and reference standard?
Concerns about applicability should be assessed separately for each domain. A study of experimentally infected animals may have low risk of bias but poor applicability to field conditions. A study using a research laboratory assay may not reflect the performance of the same test in a commercial diagnostic laboratory. These distinctions matter for clinical recommendations and should be reported explicitly.
Reporting and Interpretation
The review report should follow the PRISMA-DTA reporting guideline, which extends the general PRISMA statement to diagnostic accuracy reviews. Reporting guidelines improve completeness and transparency, and their use is associated with better description of methods and results across the medical literature. The EQUATOR Network maintains a searchable library of reporting guidelines including PRISMA-DTA and the ARRIVE guidelines for animal research. The ARRIVE guidelines specify the minimum information required for transparent reporting of animal studies and are relevant when primary studies are assessed for inclusion.
Interpretation of findings should address the clinical context in which the test will be used. A test with high sensitivity may be appropriate for screening in a low-prevalence population, while a test with high specificity may be preferred for confirmation in a high-prevalence population. The review should state the implications for veterinary practice, including whether the evidence supports adoption, continued use with caution, or further evaluation. Where evidence is insufficient, the review should identify the specific gaps that future primary studies should address.
The final section of the report should discuss limitations of the review itself, including the possibility of missed studies, the effect of excluding non-English publications, and the influence of poor reporting in primary studies on the confidence of conclusions. These limitations should be stated plainly without undermining the value of the synthesis.
Recognized Complications and Failure Modes
Systematic reviews of diagnostic test accuracy fail in characteriztic ways. The most consequential is spectrum bias introduced during study selection. When primary studies enrol animals with advanced disease and clearly healthy controls, sensitivity and specificity are overestimated relative to the target population. Detect this early by tabulating disease severity, stage, and comorbidity burden from each included study during data extraction. If severity data are missing, contact the corresponding author before proceeding to synthesis.
Verification bias arises when only a subset of animals receives the reference standard. This occurs commonly in veterinary studies where necropsy or histopathology is the reference standard and owners decline postmortem examination. Record the proportion of index-test-positive and index-test-negative animals that actually received the reference standard. A large discrepancy signals verification bias and should trigger sensitivity analysis or exclusion.
Incorporation bias occurs when the index test forms part of the reference standard, inflating agreement. This is a particular risk in clinicopathological studies where the same laboratory values inform both the test under evaluation and the diagnostic criteria. Examine the reference standard definition in each primary study and document any overlap with index test components.
Publication bias remains difficult to detect in diagnostic accuracy reviews because small negative studies are rarely registered prospectively. Funnel plot asymmetry testing has limited power with fewer than ten studies. Instead, search grey literature sources including conference proceedings and institutional repositories, and compare effect estimates between published and unpublished sources when available.
Common Errors and Corrective Actions
Less experienced reviewers frequently conflate diagnostic accuracy with clinical utility. Accuracy measures describe test performance under defined conditions, not whether testing improves patient outcomes. Keep these questions separate throughout the review.
Another recurring error is applying human-derived quality assessment criteria without adaptation. The QUADAS-2 tool requires tailoring to veterinary contexts, particularly for the reference standard domain where ante-mortem and post-mortem standards differ substantially across species. Use the EQUATOR Network reporting guideline library to identify species-appropriate reporting standards such as REFLECT for livestock studies.
Reviewers often fail to distinguish between analytical and clinical sensitivity. An immunoassay may detect analyte at low concentration yet perform poorly in clinical populations due to interfering substances. Record the analytical platform and manufacturer generation for each study, as these details explain heterogeneity that would otherwise appear inexplicable. This approach mirrors the finding that successive generations of TSH receptor antibody immunoassays showed measurable improvements in diagnostic accuracy for Graves' disease.
A further error is double-counting data from multiple publications of the same cohort. Track author lists, institution, recruitment dates, and animal identifiers across included studies. Where overlap is suspected, contact authors for clarification or include only the most complete report.
Limitations of Current Evidence
The veterinary diagnostic accuracy literature remains sparse relative to human medicine. Many index tests have been evaluated in a single center with a single breed or production system, limiting generalizability. Reference standards vary widely, and histopathology, culture, and composite clinical diagnoses each carry their own imperfections.
Expert opinion differs on several points. Whether to pool results across species when the biological basis of the target condition is similar remains contested. Some reviewers argue that cross-species pooling is justified for conditions with shared pathophysiology, while others insist on species-specific analyzes. Similarly, there is no consensus on the minimum number of studies required before meta-analysis is appropriate, with some authorities recommending at least four and others accepting fewer when heterogeneity is low.
Threshold effects pose particular difficulty. Different studies may use different cut-offs for the same index test, and the optimal threshold may vary by species, age, or production setting. Hierarchical summary ROC methods accommodate threshold variation but require sufficient data. When data are insufficient, a narrative synthesis with explicit threshold reporting is preferable to an inappropriate meta-analysis.
Escalation and Consultation
Referral to a specialist systematic review methodologist or veterinary epidemiologist is warranted when the review question involves complex latent class analysis, when multiple index tests are compared indirectly, or when the evidence base contains substantial verification bias that requires statistical correction.
Laboratory involvement becomes necessary when interpreting assay methodology across generations or manufacturers. Clinical pathologists can clarify whether differences in reported accuracy reflect true test performance or methodological variation in assay calibration. This is particularly relevant for immunoassays where MSD Veterinary Manual professional resources describe species-specific assay limitations.
Regulatory reporting obligations arise when the review identifies a commercially available test with accuracy so poor that continued use poses an animal welfare or public health risk. In such cases, consult the relevant national veterinary authority and the WOAH terrestrial animal health standards for reportable disease surveillance requirements. Professional obligations may also include publishing a correction or letter to the editor when the review contradicts widely cited accuracy figures.
| Observation | Likely cause | Discriminating check |
|---|---|---|
| Sensitivity far higher than in clinical experience | Spectrum bias from severe-case enrollment | Compare disease severity distributions between included studies and target population |
| Specificity varies widely across studies | Threshold differences or reference standard variation | Tabulate cut-offs and reference standard definitions by study |
| Two publications report identical data | Duplicate publication of same cohort | Compare author lists, recruitment dates, and animal identifiers |
| Meta-analysis shows extreme heterogeneity | Mixing assay generations or species | Stratify by assay generation and species before pooling |
| Accuracy improves over time | Assay technology advancement | Record manufacturer generation and compare by subgroup |
Frequently Asked Questions
How do I conduct a diagnostic accuracy systematic review when the budget only covers one database?
A single-database search is feasible but must be acknowledged as a limitation. MEDLINE via PubMed is the most common starting point, as demonstrated in systematic reviews of cone-beam computed tomography and TSH receptor antibody assays, both of which relied on PubMed as a primary or sole source De Grauwe et al., 2019, Tozzoli et al., 2012. Supplement the database search with citation tracking of included studies and contact authors of relevant conference abstracts. Report the single-database restriction explicitly in the methods and discuss the risk of missed studies in the limitations. If resources allow only one database, prioritize PubMed for veterinary diagnostic questions, then hand-search reference lists of all included full texts.
What should I do when the reference standard in primary studies is imperfect or unavailable?
An imperfect reference standard is common in veterinary diagnostics, particularly for diseases confirmed only at necropsy or by long-term follow-up. Do not exclude such studies outright. Instead, document the reference standard used in each study and consider a sensitivity analysis restricted to studies using the strongest reference standard. The QUADAS-2 tool requires you to judge the reference standard's risk of bias and applicability concerns separately from the index test. When no acceptable reference standard exists, consider whether a diagnostic accuracy review is appropriate at all, or whether a broader review of test utility with narrative synthesis would better serve the question.
How do the search and appraisal steps differ when the target species is exotic, wildlife, or production animals?
The core methodology is identical across species, but practical adaptations matter. Production animal questions may benefit from searching agricultural databases such as CAB Abstracts alongside biomedical databases. Wildlife and exotic species often have sparse primary literature, so broaden eligibility criteria to include related domestic species and clearly label extrapolation as indirect evidence. The ARRIVE guidelines emphasize transparent reporting of animal characteriztics, which is directly relevant when appraising veterinary diagnostic studies ARRIVE guidelines 2.0. For production animal tests with trade implications, consult the WOAH terrestrial animal health standards for guidance on test validation expectations and reporting conventions WOAH terrestrial animal health code.
What records must I keep to make the review reproducible and auditable?
Maintain a complete audit trail from search inception to final synthesis. Save the full search strategy for each database, including the exact date of execution, and export all results to a reference manager. Record screening decisions at title, abstract, and full-text stages with reasons for exclusion at the full-text stage. Keep a data extraction form that maps each item to the corresponding QUADAS-2 signaling question. Document protocol deviations with dates and justification. Reporting guidelines such as PRISMA specify the minimum information required for transparent systematic reviews, and the EQUATOR Network maintains a searchable library of these standards EQUATOR Network reporting guidelines. Store all files in a version-controlled repository accessible to co-authors.
How do I explain the limitations of a diagnostic accuracy review to a referring clinician or practice owner?
Frame the explanation around what the review can and cannot tell them. A systematic review summarizes the available evidence on test performance, but it cannot guarantee performance in their specific population or setting. Explain that sensitivity and specificity estimates from pooled analyzes may not apply to their caseload if spectrum of disease differs. Use the analogy of a test validated in a referral hospital population performing differently in first-opinion practice. Direct them to species-specific clinical resources for practical guidance on test selection MSD Veterinary Manual. Emphasize that the review identifies gaps in evidence, which is valuable for deciding whether to trust a test result or seek confirmatory testing.
When is it appropriate to stop the review and conclude that meta-analysis is not possible?
Stop quantitative synthesis when fewer than three studies report the same index test, reference standard, and outcome metric, or when study populations differ so markedly that pooling would be misleading. The systematic review of cerebral blood flow thresholds in stroke illustrates this decision: the authors could not perform meta-analysis because reported thresholds varied widely across studies, and they concluded that clinical use of specific thresholds could not be recommended without further evaluation Bandera et al., 2006. If heterogeneity is extreme, present a narrative synthesis with forest plots showing individual study estimates without a pooled summary. State explicitly that the absence of meta-analysis reflects data limitations, not a failure of the review process.
Related Clinical & Scientific Guides
- Bias in Veterinary Research: Types, Sources, and Mitigation
- Cluster Randomized Trials in Veterinary Research: Design and Analysis
- Performing Economic Evaluations of Veterinary Interventions
References and Further Reading
- CBCT in orthodontics: a systematic review on justification of CBCT in a pediatric population prior to orthodontic treatment.. 2019.
- Ultrasonography for confirmation of endotracheal tube placement: a systematic review and meta-analysis.. 2015.
- Does the medical literature remain inadequately described despite having reporting guidelines for 21 years? - A systematic review of reviews: an update.. 2018.
- The Impact of Mild Cognitive Impairment on Gait and Balance: A Systematic Review and Meta-Analysis of Studies Using Instrumented Assessment.. 2017.
- Cerebral blood flow threshold of ischemic penumbra and infarct core in acute ischemic stroke: a systematic review.. 2006.
- TSH receptor autoantibody immunoassay in patients with Graves' disease: improvement of diagnostic accuracy over different generations of methods. Systematic review and meta-analysis.. 2012.
- ARRIVE Guidelines 2.0 for Reporting Animal Research. PLOS Biology, 2020.
- EQUATOR Network Reporting Guidelines. EQUATOR Network.
- MSD Veterinary Manual, Professional Edition. MSD Veterinary Manual.
Related Articles
- Conducting Systematic Reviews of Veterinary Therapeutic Interventions
- Meta-Analysis of Veterinary Diagnostic Test Accuracy
- Diagnostic Test Accuracy Studies in Veterinary Medicine: Design and Reporting
- Appraising Diagnostic Accuracy Studies in Veterinary Medicine
- Narrative Reviews in Veterinary Medicine: When and How to Write Them
This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.