# Stratified Sampling for Disease Prevalence Estimation


## Key Takeaways

- Stratified sampling is essential for accurate disease prevalence estimation in heterogeneous animal populations by partitioning the population into strata based on risk factors (e.g., species, age, production type, geography) to ensure representation and reduce overall variance.
- Allocation methods, such as proportional allocation (matching sample size to stratum population size) and optimal allocation (incorporating stratum variance and sampling cost), are critical design decisions that influence the efficiency and precision of the prevalence estimate.
- Weighted prevalence estimation is mandatory when sampling fractions differ across strata; the stratum-specific prevalence is weighted by its proportion of the total population to derive an unbiased population-level estimate.
- Variance estimation in stratified designs requires specific formulas that account for within-stratum variances and sampling fractions, with the finite population correction being important when sampling a significant proportion of a stratum.
- Reporting requirements for stratified prevalence studies necessitate clear documentation of stratum definitions, allocation methods, weighting schemes, and stratum-specific prevalences to ensure interpretability and defensibility of the findings.
- Recognized failure modes include incomplete sampling frames, stratum misclassification, differential non-response, and inappropriate unit of analysis, all of which can introduce bias and necessitate careful auditing and potential post-stratification adjustments.

---

Prevalence estimation in animal populations is rarely a simple matter of counting affected individuals. When the target population is heterogeneous, with subpopulations that differ in disease risk, management, or geographic exposure, an unweighted sample will misrepresent the true burden of disease. Stratified sampling addresses this by partitioning the population into defined strata, sampling within each stratum, and combining stratum-specific estimates into a population-level prevalence with a known variance. This article explains how to design, execute, and analyze such schemes in veterinary research, with emphasis on allocation methods, weighting, and variance estimation. It is written for veterinary researchers and graduate students who need a procedural reference for field studies, surveillance programs, and prevalence surveys across species and production systems.

The central question this article answers is practical: given a finite budget and a heterogeneous animal population, how many animals should you sample from each subpopulation, and how should you combine the results to obtain an unbiased prevalence estimate with a defensible confidence interval? The methods described apply to cross-sectional surveys of livestock, companion animals, and wildlife, and they extend naturally to active surveillance programs where national or regional prevalence estimates are required. Cluster sampling and non-probability sampling are outside the scope of this article, the focus is exclusively on designs where every stratum is sampled and every sampling unit within a stratum has a known, non-zero selection probability.

## At a Glance

| Parameter | Decision or Fact |
| --- | --- |
| Stratum definition | Partition the population by variables associated with disease risk, such as species, age class, production type, or geographic region |
| Allocation method | Proportional allocation matches sample size to stratum population size, optimal allocation also incorporates stratum variance and sampling cost |
| Weighting | Each sampled animal represents a known fraction of its stratum, the weighted mean of stratum prevalences estimates population prevalence |
| Variance estimation | Use the stratified estimator of variance, which combines within-stratum variances weighted by the square of the stratum weight |
| Sample size calculation | Requires stratum population sizes, expected prevalence per stratum, desired precision, and confidence level, consult a standard epidemiologic reference for the formula |
| Post-stratification | If stratum membership is known only after sampling, adjust weights retrospectively to correct for unequal selection probabilities |
| Reporting | Report stratum-specific prevalences alongside the overall estimate, and state the allocation method and weighting scheme used |

## Why Stratification Improves Prevalence Estimates

A simple random sample from a heterogeneous population can, by chance, over-represent low-risk or high-risk segments. The resulting estimate carries unnecessary variance, and the confidence interval may be wide enough to obscure meaningful differences between subpopulations. Stratification converts this uncontrolled variation into a design feature. By guaranteeing that every stratum contributes observations, the researcher ensures that rare but important subpopulations are represented, and the variance of the overall estimate is reduced when the strata are internally homogeneous with respect to disease status.

The logic depends on the relationship between the stratification variable and the outcome. A useful stratum is one within which disease prevalence is relatively uniform, but between which prevalence differs. Age, sex, breed, production system, and geographic region are common stratification variables in veterinary surveys. For example, a national survey of equine piroplasmosis in Spain allocated sampling across autonomous communities, recognizing that tick exposure and management practices vary substantially by region, and the resulting study detected clear geographic differences in seroprevalence [Camino et al., institutional publication](https://pubmed.ncbi.nlm.nih.gov/32918303/). Similarly, a helminth prevalence survey in horses in Brandenburg, Germany used randomised stratified sampling of farms and identified age, breed, and sex as risk factors for high strongyle egg shedding, findings that would have been obscured by an unstratified design [Hinney et al., institutional publication](https://pubmed.ncbi.nlm.nih.gov/21472400/).

The precision gain is not automatic. If the stratification variable is unrelated to disease status, stratification adds complexity without reducing variance, and the design may even be slightly less efficient than simple random sampling because degrees of freedom are lost to stratum estimation. The researcher should therefore choose stratification variables based on prior knowledge of disease biology, pilot data, or published risk factor studies, not on convenience alone.

## Defining Strata and Sampling Frames

The sampling frame is the complete list of units from which the sample is drawn. For a stratified design, the frame must allow every unit to be assigned to exactly one stratum, and the frame must be complete for each stratum. In livestock populations, the frame is often a national or regional register of holdings, with stratum membership determined by variables such as herd size, production type, or administrative region. In companion animal studies, the frame may be a list of veterinary practices or a census of households, with strata defined by postal area or urban versus rural location.

The number of strata requires a balance between homogeneity and practicality. More strata allow finer control of variance but increase the complexity of sampling and analysis, and each stratum must have enough sampled units to support a stable prevalence estimate. A common rule is to aim for at least 30 to 50 observations per stratum when the prevalence estimate within that stratum is of interest in its own right. When strata are only needed to improve the overall estimate, smaller stratum samples may suffice, but the variance contribution of each stratum must still be estimable.

## Allocation Methods

Proportional allocation assigns each stratum a sample size proportional to its share of the population. If a stratum contains 40% of the animals, it receives 40% of the sample. This method is simple, self-weighting, and efficient when the prevalence is similar across strata. Its weakness is that small strata receive few observations, and a stratum with a very different prevalence may be represented by only a handful of animals, producing an unstable stratum estimate.

Optimal allocation, sometimes called Neyman allocation, assigns sample sizes in proportion to the product of stratum population size and stratum standard deviation. Strata with higher expected variability in disease status receive more observations. This minimizes the variance of the overall prevalence estimate for a fixed total sample size, but it requires prior estimates of stratum-specific prevalence or variance, which may not be available. In practice, researchers often use pilot data, published literature, or conservative assumptions about prevalence to inform the allocation. When sampling costs differ between strata, the allocation can be further adjusted to minimize cost for a target variance, a method known as optimal allocation with cost.

The choice between proportional and optimal allocation should be made explicit in the study protocol and reported in the final publication. The [CDC principles of epidemiology in public health practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html) describe the rationale for these allocation methods and the conditions under which each is preferred.

## Weighted Prevalence and Variance Estimation

Once sampling is complete, the raw proportion of test-positive animals in the combined sample does not, in general, estimate the population prevalence. That raw proportion is biased whenever the sampling fraction differs across strata and the true stratum-specific prevalences differ. The correct estimator weights each stratum's observed prevalence by that stratum's share of the target population.

For a population divided into \(H\) strata, with \(N_h\) animals in stratum \(h\) and \(n_h\) sampled, the weighted prevalence is:

\[
\hat{p}_{w} = \sum_{h=1}^{H} \left( \frac{N_h}{N} \right) \hat{p}_h
\]

where \(N = \sum N_h\) and \(\hat{p}_h\) is the observed prevalence in stratum \(h\). The weight \(N_h/N\) is the stratum's population proportion, not its sample proportion. This is the standard estimator described in [CDC principles of epidemiology in public health practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html).

The variance of the weighted prevalence under proportional allocation simplifies because each stratum's sampling fraction is constant. Under optimal or arbitrary allocation, the variance must account for the differing sampling fractions:

\[
\widehat{Var}(\hat{p}_{w}) = \sum_{h=1}^{H} \left( \frac{N_h}{N} \right)^2 \left( 1 - \frac{n_h}{N_h} \right) \frac{\hat{p}_h (1 - \hat{p}_h)}{n_h - 1}
\]

The term \((1 - n_h/N_h)\) is the finite population correction. It matters when the sampling fraction exceeds roughly 5% of the stratum. In large livestock populations where \(n_h/N_h\) is small, the correction is negligible and can be omitted. In small herds or when a large proportion of a stratum is sampled, omitting it overstates the variance and produces unnecessarily wide confidence intervals.

A 95% confidence interval is constructed as:

\[
\hat{p}_w \pm 1.96 \times \sqrt{\widehat{Var}(\hat{p}_w)}
\]

This normal approximation performs adequately when each stratum contributes at least 5 to 10 positive and 5 to 10 negative observations. When stratum-specific counts are smaller, exact binomial methods applied to the weighted estimate are preferable, though they require computational tools.

## Worked Example: Proportional and Optimal Allocation

Consider a regional survey of bovine tuberculosis seroprevalence in a population of 10,000 cattle distributed across three production strata: 5,000 dairy cattle, 3,000 beef cattle, and 2,000 youngstock. A total sample of 500 animals is planned.

### Proportional Allocation

Proportional allocation assigns each stratum a sample size proportional to its population share:

\[
n_h = n \times \frac{N_h}{N}
\]

| Stratum | \(N_h\) | \(n_h\) | Sampled positive | \(\hat{p}_h\) |
|---------|---------|---------|------------------|---------------|
| Dairy | 5,000 | 250 | 25 | 0.100 |
| Beef | 3,000 | 150 | 9 | 0.060 |
| Youngstock | 2,000 | 100 | 2 | 0.020 |
| **Total** | **10,000** | **500** | **36** | |

The weighted prevalence is:

\[
\hat{p}_w = 0.5(0.100) + 0.3(0.060) + 0.2(0.020) = 0.072
\]

The unweighted raw proportion is \(36/500 = 0.072\) in this case because proportional allocation makes the sampling fraction identical across strata. This equality is the defining property of proportional allocation: the raw proportion is unbiased, and the weighted and unweighted estimates coincide.

The variance is:

\[
\widehat{Var}(\hat{p}_w) = 0.5^2 \left(1 - \frac{250}{5000}\right)\frac{0.1(0.9)}{249} + 0.3^2 \left(1 - \frac{150}{3000}\right)\frac{0.06(0.94)}{149} + 0.2^2 \left(1 - \frac{100}{2000}\right)\frac{0.02(0.98)}{99}
\]

\[
\widehat{Var}(\hat{p}_w) = 0.25(0.95)(0.000361) + 0.09(0.95)(0.000379) + 0.04(0.95)(0.000198) = 0.0000857 + 0.0000324 + 0.0000075 = 0.0001256
\]

The standard error is \(\sqrt{0.0001256} = 0.0112\). The 95% confidence interval is \(0.072 \pm 1.96(0.0112)\), giving \(0.050\) to \(0.094\).

### Optimal Allocation

Optimal allocation, also called Neyman allocation, minimizes the variance of the weighted estimator for a fixed total sample size. It assigns sample sizes proportional to the product of stratum size and stratum standard deviation:

\[
n_h = n \times \frac{N_h \sqrt{\hat{p}_h(1 - \hat{p}_h)}}{\sum_{h=1}^{H} N_h \sqrt{\hat{p}_h(1 - \hat{p}_h)}}
\]

The complication is that optimal allocation requires prior knowledge of the stratum-specific prevalences. In practice, these are taken from pilot data, published estimates, or the proportional allocation sample itself. A two-phase approach is common: collect a small proportional sample, estimate stratum variances, then allocate the remaining sample optimally.

Using the proportional sample results above as priors:

| Stratum | \(N_h\) | \(\sqrt{\hat{p}_h(1 - \hat{p}_h)}\) | \(N_h \times\) SD | \(n_h\) |
|---------|---------|--------------------------------------|--------------------|--------|
| Dairy | 5,000 | 0.300 | 1,500 | 268 |
| Beef | 3,000 | 0.237 | 711 | 127 |
| Youngstock | 2,000 | 0.140 | 280 | 50 |
| **Total** | | | **2,491** | **445** |

The remaining 55 samples would be distributed proportionally to the same products, yielding final stratum samples of approximately 268, 127, and 50 after rounding. Note that the total in the table sums to 445 because the allocation formula distributes the planned 500, the discrepancy arises from rounding and the need to verify the arithmetic against the full formula.

Optimal allocation places more samples in the dairy stratum because its prevalence is highest and its variance contribution is largest. This produces a narrower confidence interval than proportional allocation for the same total sample size. The variance under optimal allocation is:

\[
\widehat{Var}(\hat{p}_w) = \frac{\left(\sum N_h \sqrt{\hat{p}_h(1 - \hat{p}_h)}\right)^2}{N^2 n} - \frac{\sum N_h \hat{p}_h(1 - \hat{p}_h)}{N^2}
\]

Substituting the values gives approximately \(0.000112\), corresponding to a standard error of \(0.0106\) and a 95% confidence interval of \(0.051\) to \(0.093\). The gain over proportional allocation is modest here because the stratum prevalences do not differ dramatically. When stratum prevalences differ by a factor of 5 or more, optimal allocation can reduce the required sample size by 20% to 30%.

## Post-Stratification

Post-stratification adjusts the weights after data collection when the sample does not match the population distribution across known strata. This situation arises when the sampling frame is incomplete, when non-response or sample loss is differential across strata, or when the frame is updated after sampling. The estimator is identical to the weighted estimator above, but the weights are applied retrospectively.

Post-stratification cannot repair a fundamentally broken sampling design. If a stratum was undersampled because the frame omitted part of the population, weighting adjusts only for the known strata, not for the omitted segment. The [WOAH terrestrial animal health code](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/) requires that surveillance designs be documented before sampling begins, and post-stratification should be reported as a deviation from the planned design.

Post-stratification is most useful when administrative data provide reliable stratum totals but the sampling process could not enforce quotas. For example, a farm-level survey of equine helminths that samples horses presented at clinics cannot control the breed or age distribution of the sample. Applying post-stratification weights based on the regional horse population partially corrects the resulting selection bias, as illustrated in a stratified survey of horse farms in Brandenburg, Germany, where farm-level prevalence estimates were adjusted for the distribution of farm types [prevalence of helminths in horses in Brandenburg](https://pubmed.ncbi.nlm.nih.gov/21472400/).

The variance of a post-stratified estimator is larger than that of a proportionally allocated estimator with the same sample size, because the sample sizes within strata are random instead of fixed. The variance formula must include a term for the variability of the stratum sample sizes. Most statistical software implements this correction automatically when the post-stratification weights are declared.

## Software and Computation

Spreadsheet software is adequate for the calculations shown above, but dedicated statistical packages are preferable for surveys with more than a few strata. The survey package in R, the svy commands in Stata, and the Complex Samples module in SPSS all implement stratified estimators with appropriate variance formulas. These tools also handle the finite population correction automatically when population stratum sizes are supplied.

The choice of software matters less than the correctness of the variance specification. A common error is to analyze stratified data as though they came from a simple random sample. This produces standard errors that are too large when proportional allocation was used, because stratification removes the between-stratum component of variance. The direction of the bias reverses under optimal allocation, where the analysis must reflect the unequal sampling fractions.

## Reporting Requirements

A stratified prevalence estimate is interpretable only when the report includes the stratum definitions, the population and sample sizes per stratum, the allocation method, the observed stratum-specific prevalences, and the variance formula used. The [WOAH animal health surveillance standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/) require this level of detail for internationally notifiable diseases. For research publications, the same information is expected under standard reporting guidelines for observational studies.

The confidence interval should be reported with the point estimate, and the design effect should be stated. The design effect is the ratio of the variance under the actual design to the variance that would have been obtained from a simple random sample of the same size. A design effect below 1 indicates that stratification improved precision, a value above 1 indicates that the allocation was inefficient relative to simple random sampling. Reporting the design effect allows readers to compare studies that used different designs and to plan future surveys using the published variance estimates.

## Recognized Failure Modes

Stratified sampling fails in predictable ways, and most failures originate before data collection begins. The most consequential error is constructing strata from a sampling frame that does not match the target population. If the frame omits a subpopulation, such as small herds in a national livestock registry, the prevalence estimate will be biased even with perfect allocation and laboratory performance. Detect this early by comparing the frame against independent sources, such as movement records, veterinary practice lists, or producer organizations, and quantify the proportion of the target population missing.

A second failure mode is stratum misclassification. Animals or herds assigned to the wrong stratum distort both the allocation and the weighted estimate. This occurs when stratum definitions rely on outdated information, such as herd size recorded at the time of a previous census. Verify stratum assignments on a subsample during fieldwork and document the date and source of the classification variable.

Non-response and refusal create missing data that is rarely random. If large herds decline participation more often than small herds, the realised sample becomes disproportionate and the design-based weights no longer reflect the population. Track response rates by stratum from the outset. When response is below 80% in any stratum, compare responders with non-responders on available covariates and consider whether post-stratification adjustment can recover representativeness.

A third failure mode is the use of an inappropriate unit of analysis. In the Brandenburg helminth study, the farm was the experimental unit even though individual horses were sampled, because transmission and management operate at farm level. Analyzing individual animals without accounting for the farm unit inflates precision and produces confidence intervals that are too narrow. Specify the primary sampling unit in the analysis plan and use variance estimation that respects it.

| Observation | Likely cause | Discriminating check |
|---|---|---|
| Prevalence estimate far from expected range | Sampling frame omits a subpopulation | Compare frame coverage against independent registries |
| Weighted estimate changes markedly after post-stratification | Original stratum weights were inaccurate | Recalculate weights using updated population counts |
| Confidence intervals unexpectedly narrow | Unit of analysis ignores clustering within primary units | Re-estimate variance with cluster-robust methods |
| Response rate differs sharply across strata | Non-response correlated with stratum characteriztics | Compare responders and non-responders on available covariates |
| Several strata contribute zero positive results | Allocation too small in low-prevalence strata | Review stratum-specific sample sizes against expected prevalence |

## Common Errors and Corrective Action

Less experienced investigators often allocate proportionally without considering cost or variance differences across strata. Proportional allocation is simple and self-weighting, but it can leave rare subpopulations with too few samples to produce stable estimates. The corrective action is to compute optimal allocation and compare the variance reduction against the added cost before committing to a design.

A related error is treating the overall sample size calculation as a single formula applied to the pooled population. Stratified designs require stratum-specific sample sizes, and the total is the sum of those sizes. When prevalence is expected to vary widely across strata, as it did across species in the Rift Valley fever study in Uganda where cattle seroprevalence reached 20.5% while goats were 3.6%, a pooled calculation will misallocate effort.

Students frequently confuse stratification with post-stratification. Stratification requires a sampling frame with stratum identifiers available before selection. Post-stratification is a corrective adjustment applied after sampling when stratum membership is known only for the sampled units. Using post-stratification when a proper frame exists wastes the variance reduction that design-based stratification would have provided.

Another common error is ignoring the finite population correction when sampling fraction is high. In small populations, such as a regional herd of a few hundred animals, the correction materially reduces variance. Omitting it overstates uncertainty and may lead to unnecessarily large sample sizes.

## Limitations of Current Evidence

The evidence base for stratified sampling in veterinary prevalence studies is built largely from field applications instead of methodological trials. Several studies cited in this article used stratified designs successfully, including the national equine piroplasmosis survey in Spain and the Q fever seroprevalence study in Irish cattle. These reports demonstrate feasibility but do not quantify how much precision stratification gained over simple random sampling in their specific settings.

Expert opinion differs on the minimum number of strata and the point at which stratification ceases to add value. Some authorities recommend keeping strata broad enough to ensure at least 30 sampled units per stratum for variance estimation. Others accept smaller strata when the primary goal is ensuring representation instead of stratum-specific estimation. The choice depends on whether stratum-level prevalence will be reported, and if so, the required precision for those estimates.

There is also unresolved debate about optimal allocation when prevalence is unknown before sampling. Optimal allocation requires prior estimates of stratum variances, which are often unavailable. Using guesses from the literature or pilot data can produce allocation that is worse than proportional if the priors are badly wrong. A pragmatic compromise is proportional allocation with a minimum sample size per stratum to protect rare subpopulations.

## Escalation and Reporting

Referral to a specialist epidemiologist or biostatistician is warranted when the sampling frame is incomplete, when the population structure is complex enough that stratum boundaries are uncertain, or when the study results will inform regulatory decisions. The same applies when prevalence estimates will be compared across regions or time periods, because design differences can confound such comparisons.

Laboratory involvement should occur before sampling, not after. Confirm that the diagnostic test has known sensitivity and specificity for the target species and that sample types and transport conditions match the laboratory's validated protocols. Test performance affects prevalence estimates directly, and the [CDC principles of epidemiology](https://www.cdc.gov/csels/dsepd/ss1978/index.html) emphasize that diagnostic accuracy must be incorporated into the interpretation of any prevalence measure.

Regulatory reporting obligations depend on the disease and jurisdiction. When a notifiable disease is detected during a prevalence survey, the investigator must follow the reporting pathway of the relevant animal health authority. The [World Organization for Animal Health surveillance standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/) and the [WOAH terrestrial animal health code](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/) describe the international framework for notification, but national requirements take precedence and vary by country and species. Confirm the local reporting list before beginning fieldwork so that positive findings do not create avoidable delays in notification.

## Frequently Asked Questions

### How Do I Choose Between Proportional and Optimal Allocation When My Budget Is Fixed?

Proportional allocation is simpler and requires no prior knowledge of within-stratum variability, making it the default when prevalence estimates are the sole objective. Optimal allocation reduces variance for a given total sample size, but it demands estimates of stratum-specific standard deviations, which are rarely available before data collection. When budget constraints bind, use proportional allocation and accept the variance penalty. If one stratum is known to be highly variable or epidemiologically critical, oversample that stratum deliberately and apply weighted analysis afterward. The gain in precision is usually modest unless stratum variances differ by several fold, so reserve optimal allocation for settings where prior data, such as pilot surveys, justify the added complexity.

### What Sample Size Do I Need per Stratum When Strata Are Very Unequal in Size?

There is no universal minimum, but very small strata create unstable variance estimates and fragile prevalence figures. A practical rule is to sample at least 30 units per stratum when stratum-level estimates are planned, and fewer when strata exist only to improve the overall estimate. When a stratum contains very few eligible units, consider sampling all of them. If a stratum is so small that its contribution to the national estimate is negligible, you may combine it with a neighbouring stratum before analysis, provided the combination remains biologically meaningful. Document any such merging in the methods section, because reviewers will expect transparency about how the sampling frame was finalised.

### Can I Use Stratified Sampling for Wildlife or Free-Ranging Populations?

Yes, but the sampling frame differs from production animal settings. Wildlife populations rarely have complete registries, so strata are often defined by geographic zones, habitat types, or management units instead of by individual animal identification. Capture methods, such as netting, trapping, or darting, introduce accessibility bias that can distort prevalence estimates even when the stratification plan is sound. The [CDC principles of epidemiology](https://www.cdc.gov/csels/dsepd/ss1978/index.html) describe how non-probability elements enter field studies and why they must be reported. When capture probability varies by stratum, consider adjusting for it in analysis or acknowledging that the estimate applies to the accessible portion of the population instead of the entire population.

### What Should I Do When the Sampling Frame Is Incomplete or Outdated?

An incomplete frame undermines the probabilistic basis of stratified sampling. If the registry omits recently established herds or includes dissolved ones, your prevalence estimate will be biased even with perfect laboratory work. Begin by auditing the frame against independent sources, such as movement records, veterinary practice lists, or local authority registers. Where discrepancies are substantial, update the frame before sampling. If updating is impossible, document the coverage gaps and consider post-stratification to adjust for known imbalances. The [WOAH animal health surveillance standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/) require that surveillance systems describe their target population and sampling approach, and incomplete frames should be disclosed in any formal report.

### How Do I Explain Stratified Sampling to a Farm Owner or Practice Manager?

Focus on the practical consequence: stratified sampling gives a more accurate picture of disease status without requiring more samples overall. Explain that the population is divided into groups, such as age cohorts or housing systems, and that each group contributes samples in proportion to its size or importance. Use a concrete example from their operation, such as sampling more young stock if respiratory disease is suspected to concentrate there. Emphasize that the goal is not to test every animal but to obtain a result that can be trusted for decisions. Avoid statistical terminology unless the client requests it, and be prepared to show how the result will inform treatment or biosecurity choices.

### What Records Must I Keep to Defend the Sampling Design in a Peer-Reviewed Report?

Retain the sampling frame as it existed at the time of selection, the randomisation method and seed, the allocation formula, and the final list of selected units with replacement records. Document every deviation from the plan, including refusals, missing samples, and laboratory failures, because these affect the validity of the weighted estimate. The [MSD Veterinary Manual](https://www.msdvetmanual.com/) advises that diagnostic test results are only interpretable when the sampling method and population are clearly described. Keep stratum-specific counts of eligible, sampled, and successfully tested units. This allows reviewers to verify the weighting calculations and to assess whether non-response differed by stratum, which is the most common hidden threat to stratified survey validity.

## Related Clinical & Scientific Guides

* [Evaluating Veterinary Surveillance System Attributes](/knowledge/veterinary-medicine/veterinary-epidemiology/evaluating-veterinary-surveillance-system-attributes)
* [Network Analysis for Infectious Disease Spread in Animal Populations](/knowledge/veterinary-medicine/veterinary-epidemiology/network-analysis-infectious-disease-spread-animal-populations)
* [Randomized Controlled Trials in Veterinary Field Settings](/knowledge/veterinary-medicine/veterinary-epidemiology/randomized-controlled-trials-veterinary-field-settings)


## References and Further Reading

- [Sero-molecular survey and risk factors of equine piroplasmosis in horses in Spain.](https://pubmed.ncbi.nlm.nih.gov/32918303/). 2021.
- [Prevalence of helminths in horses in the state of Brandenburg, Germany.](https://pubmed.ncbi.nlm.nih.gov/21472400/). 2011.
- [Public views on the donation and use of human biological samples in biomedical research: a mixed methods study.](https://pubmed.ncbi.nlm.nih.gov/23929915/). 2013.
- [Prevalence of Coxiella burnetii (Q fever) antibodies in bovine serum and bulk-milk samples.](https://pubmed.ncbi.nlm.nih.gov/21073765/). 2011.
- [Prevalence and risk factors for echinococcal infection in a rural area of northern Chile: a household-based cross-sectional study.](https://pubmed.ncbi.nlm.nih.gov/25167140/). 2014.
- [Rift Valley fever seroprevalence and abortion frequency among livestock of Kisoro district, South Western Uganda (2016): a prerequisite for zoonotic infection.](https://pubmed.ncbi.nlm.nih.gov/30176865/). 2018.
- [WOAH Animal Health Surveillance Standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/). WOAH.
- [CDC Principles of Epidemiology in Public Health Practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html). CDC.
- [MSD Veterinary Manual, Professional Edition](https://www.msdvetmanual.com/). MSD Veterinary Manual.

## Related Articles

- [Measuring Disease Frequency: Incidence and Prevalence in Animal Populations](/knowledge/veterinary-medicine/veterinary-epidemiology/measuring-disease-frequency-incidence-prevalence-animal-populations)
- [Using Capture-Recapture Methods to Estimate Animal Disease Prevalence](/knowledge/veterinary-medicine/veterinary-epidemiology/using-capture-recapture-methods-estimate-animal-disease-prevalence)
- [Cluster Sampling in Veterinary Field Studies](/knowledge/veterinary-medicine/veterinary-epidemiology/cluster-sampling-veterinary-field-studies)
- [Compartmental Models in Veterinary Disease Dynamics](/knowledge/veterinary-medicine/veterinary-epidemiology/compartmental-models-veterinary-disease-dynamics)
- [Purposive Sampling in Veterinary Outbreak Investigations](/knowledge/veterinary-medicine/veterinary-epidemiology/purposive-sampling-veterinary-outbreak-investigations)

> This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.