# Spatial Cluster Detection Methods in Veterinary Epidemiology


## Key Takeaways

- The spatial scan statistic (Kulldorff) is the primary method for detecting disease clusters in veterinary epidemiology, comparing observed to expected cases within a moving circular window to identify statistically significant aggregations beyond chance.
- Local Indicators of Spatial Association (LISA), such as local Moran's I, are crucial for identifying individual geographic units (e.g., herds, counties) that contribute disproportionately to overall spatial autocorrelation, distinguishing hot spots from cold spots.
- Accurate population-at-risk data are essential for population-adjusted models (e.g., Poisson, Bernoulli) used in scan statistics, with data sources including census, herd registers, or slaughter surveillance, while case-only methods are useful when denominators are unavailable, particularly in wildlife surveillance.
- The choice of spatial scale (e.g., county, herd, point-level) and the precision of geographic coordinates significantly impact cluster detection power and the size and detectability of identified clusters.
- Reporting of cluster detection results must include the cluster location, radius, relative risk, p-value, and secondary clusters, with careful interpretation of overlapping secondary clusters as potential artifacts of the primary signal.
- Common failure modes include edge effects in local methods, multiple testing without correction, and coordinate imprecision, necessitating checks like false discovery rate correction and sensitivity analysis for window shapes.

---

Spatial cluster detection identifies geographic areas where disease occurrence exceeds expectation, supporting outbreak source tracing, targeted surveillance, and hypothesis generation in animal populations. This article provides a procedural reference for veterinary researchers selecting and applying cluster detection methods, with emphasis on scan statistics and local indicators of spatial association. It assumes familiarity with epidemiologic study design and basic geographic information system operations.

The methods described here address a recurring question in veterinary field investigations: given case locations and a background population, where does disease aggregate beyond chance? Answering that question requires decisions about the case definition, the spatial scale of analysis, the null hypothesis, and the statistical test most suited to the data structure. The article covers the conceptual basis of cluster detection, the principal test families, practical implementation considerations, and interpretation of results in livestock, wildlife, and companion animal contexts.

## At a Glance

| Parameter | Consideration |
|---|---|
| Case definition | Must be consistent across space and time, laboratory confirmation preferred where feasible |
| Population at risk | Required for Poisson and Bernoulli models, use census, herd registers, or slaughter data |
| Spatial scale | County, herd, or point-level data change cluster size and power |
| Primary method | Spatial scan statistic (Kulldorff) for most outbreak investigations |
| Alternative methods | Cuzick-Edwards k-NN, Moran's I, LISA, K functions for exploratory analysis |
| Statistical inference | Monte Carlo simulation for scan statistics, permutation for LISA |
| Reporting | Report cluster location, radius, relative risk, p-value, and secondary clusters |
| Software | SaTScan, R packages (spatstat, spdep), commercial GIS modules |

## Conceptual Foundations of Spatial Cluster Detection

Spatial clustering in veterinary epidemiology refers to the non-random geographic concentration of disease events. Clusters arise from shared exposure sources, local transmission dynamics, or heterogeneous surveillance intensity. The analytic goal is to distinguish true clustering from the background variation expected under a stated null hypothesis, typically complete spatial randomness or a spatially uniform risk surface.

Cluster detection methods differ from spatial regression and mapping techniques. Mapping describes where disease occurs, cluster detection tests whether the observed pattern deviates from expectation and locates the deviation. The distinction matters for study design. A researcher mapping bovine tuberculosis prevalence by county may observe visual hot spots, but only a formal cluster test can assign statistical confidence to those hot spots.

The choice of null hypothesis shapes every subsequent decision. The most common null model assumes cases occur independently with probability proportional to the population at risk. This model requires accurate denominator data. When denominators are unavailable, some methods operate on case locations alone, testing for spatial randomness without population adjustment. These case-only methods are less powerful but remain useful when population data are incomplete, as often occurs in wildlife disease surveillance.

## The Spatial Scan Statistic

The spatial scan statistic, developed by Kulldorff, is the most widely used method for veterinary cluster detection. The procedure places a circular window over the study region, moves it across all possible centroids, and expands the radius continuously. For each window position and size, the method compares the observed number of cases inside the window with the expected number under the null hypothesis. The window with the maximum likelihood ratio is declared the most likely cluster.

The scan statistic accommodates several probability models. The Bernoulli model applies when data consist of cases and non-cases, such as infected and uninfected herds. The Poisson model applies when case counts are available with a population denominator, such as slaughterhouse lesion counts per county. The purely spatial version analyzes a single time period, the space-time version adds a temporal dimension and can detect emerging clusters before they become geographically widespread.

Monte Carlo simulation generates the statistical significance of detected clusters. The method creates thousands of random datasets under the null hypothesis, recalculates the maximum likelihood ratio for each, and compares the observed statistic to this simulated distribution. The resulting p-value is valid even though the scan statistic examines an enormous number of candidate windows. Reported output includes the cluster location, radius, number of observed and expected cases, relative risk, and p-value.

Secondary clusters are reported in order of likelihood. Researchers should interpret secondary clusters cautiously when they overlap the primary cluster, as they may represent artifacts of the primary signal instead of independent aggregations. The scan statistic's circular window performs poorly for elongated clusters, such as those following river valleys or transportation corridors. Elliptical and irregularly shaped window variants exist but require additional parameters and computational time.

## Local Indicators of Spatial Association

Local indicators of spatial association, or LISA statistics, decompose global spatial autocorrelation into per-location contributions. The local Moran's I and local Geary's C identify individual units, such as herds or counties, that contribute disproportionately to overall clustering. These methods require a spatial weights matrix defining which locations are neighbors, typically based on distance thresholds or contiguity.

LISA methods classify each location into one of four categories: high-high clusters, low-low clusters, high-low outliers, and low-high outliers. This classification supports the identification of both disease hot spots and cold spots. The methods are computationally efficient and integrate naturally with geographic information system workflows.

Moran's I, the global counterpart to local Moran's I, tests for overall spatial autocorrelation across the study region. A veterinary investigation of bovine tuberculosis in Argentina used Moran's I and found no global clustering by county, yet the same study detected significant local clustering using nearest-neighbor and scan statistics. This example illustrates a central principle: global tests can miss localized clustering, and negative global results do not rule out meaningful local aggregations.

## Distance-Based and Second-Order Methods

Cuzick and Edwards' k-nearest-neighbor test evaluates whether cases are closer to other cases than expected under random labeling. The test statistic counts the number of cases among each case's k nearest neighbors. The choice of k is arbitrary and influences power, small k values detect tight clusters, while larger values detect diffuse clustering. The method requires no population denominator, making it suitable for case-only data.

Ripley's K function and its derivatives characterize clustering across multiple distance scales simultaneously. The K function compares the observed number of pairs within a given distance to the expected number under spatial randomness. Plotting the K function against distance reveals the scale at which clustering occurs. The related cross-K function compares the spatial relationship between two event types, such as infected and non-infected herds, and has been applied to assess spatial correlation between human taeniasis and porcine cysticercosis in the Democratic Republic of Congo.

These second-order methods describe pattern structure instead of identifying specific clusters. They serve best as exploratory tools preceding or complementing scan statistics. A complete analysis might use the K function to establish the characteriztic cluster scale, then apply the scan statistic with that scale in mind.

## Analytical Workflow for Cluster Detection

The practical application of cluster detection methods follows a structured sequence. The workflow begins with data assembly, proceeds through exploratory analysis, and culminates in formal cluster testing. Each stage has distinct decision points that materially affect the results.

### Data Preparation and Case Definition

The case definition determines what constitutes an event. For livestock diseases, the unit of analysis may be an individual animal, a herd, a premises, or an administrative unit such as a county or parish. The choice of unit changes the interpretation of results. Herd-level analysis identifies clusters of affected operations, whereas animal-level analysis identifies clusters of affected individuals, which may be more sensitive when within-herd prevalence is low.

Spatial data quality requires scrutiny before any analysis begins. Coordinate accuracy, missing georeferences, and positional error all influence cluster detection performance. In the Argentine bovine tuberculosis study, records from slaughter surveillance were aggregated to the county level because premises-level coordinates were not available, and this aggregation constrained the analysis to county-scale inference [CDC principles of epidemiology in public health practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html). Investigators should document the positional accuracy of each record and assess whether the spatial resolution supports the intended inference.

Population-at-risk data are equally important. Cluster detection methods compare observed case counts against expected counts derived from the underlying population. For livestock, the population denominator may be the number of animals, the number of herds, or the number of premises in each spatial unit. When denominator data are incomplete, the analysis may identify clusters that reflect data availability instead of disease biology. The Q-fever investigation in the Netherlands used farm size as a denominator and restricted the analysis to farms with more than 40 animals, a threshold chosen to stabilize rate estimates [GIS-based source detection for the urban Q-fever outbreak](https://pubmed.ncbi.nlm.nih.gov/20230650/).

### Exploratory Analysis Before Formal Testing

Global clustering tests assess whether cases are more spatially concentrated than expected under a null hypothesis of random distribution. These tests do not locate clusters, they establish whether cluster detection is warranted. Moran's I statistic is the most common global measure for continuous or binary data. In the Argentine bovine tuberculosis analysis, Moran's I returned a value of 0.009 with P = 0.089, indicating no global spatial autocorrelation of county-level prevalence [spatial statistics and monitoring data for bovine tuberculosis clustering in Argentina](https://pubmed.ncbi.nlm.nih.gov/12419600/). The absence of global clustering did not preclude local clustering, and subsequent local tests identified significant clusters in dairy districts.

The distinction between global and local tests is critical. Global tests average spatial association across the entire study region, and localized clustering may be diluted when the study area is large relative to the cluster extent. A negative global test does not justify abandoning cluster detection when disease ecology suggests focal transmission. Conversely, a positive global test does not identify where the cluster is located.

### Selecting the Cluster Detection Method

The choice of method depends on the data type, the hypothesised cluster shape, and the study objective. The table below summarizes the principal methods and their selection criteria.

| Method | Data type | Cluster shape | Primary output | Best suited to |
|---|---|---|---|---|
| Spatial scan statistic (Kulldorff) | Counts, Bernoulli, Poisson, survival | Circular or elliptic, variable size | Most likely cluster with relative risk and P value | Outbreak detection, surveillance, hypothesis generation |
| Local Moran's I (LISA) | Continuous or binary at areal units | Irregular, defined by adjacency | Local autocorrelation with cluster type (high-high, low-low) | Identifying hot spots and cold spots in endemic disease |
| Getis-Ord Gi* | Continuous or count at areal units | Irregular, defined by distance band | Z scores for hot and cold spots | Ranking spatial units by intensity of clustering |
| Cuzick-Edwards k-NN | Case-control binary | Irregular, defined by nearest neighbours | Test statistic for case clustering | Case-control data without population denominators |
| K-function and cross-K function | Point locations | Distance-based, no fixed shape | Spatial dependence across distance bands | Assessing clustering at multiple spatial scales |

The spatial scan statistic is the most widely used method in veterinary epidemiology because it accommodates variable cluster size, adjusts for the underlying population, and provides a likelihood-based significance test. The method was applied to identify local clusters of Taenia solium taeniasis and porcine cysticercosis in the Democratic Republic of Congo, where the scan statistic located villages with elevated infection risk within a rural health zone [geospatial and age-related patterns of Taenia solium taeniasis in the rural health zone of Kimpese](https://pubmed.ncbi.nlm.nih.gov/26996821/). The scan statistic requires the user to specify the maximum cluster size, typically as a percentage of the total population at risk. A maximum of 50% is the default in SaTScan, but smaller values, such as 10% to 25%, are often more appropriate for detecting focal clusters in livestock populations.

Local indicators of spatial association are preferable when the analysis unit is an administrative area and the objective is to classify each unit as a cluster member, a cluster boundary, or a spatial outlier. These methods are computationally efficient and produce interpretable maps, but they require a defined spatial weights matrix and are sensitive to the choice of adjacency or distance criteria.

### Running the Spatial Scan Statistic

SaTScan is the reference software for the spatial scan statistic. The analysis requires four input files: a case file, a population file, a coordinate file, and a configuration file. The case file contains the count of cases in each spatial unit, the population file contains the corresponding population at risk, and the coordinate file provides the geographic location of each unit, either as latitude and longitude or as projected coordinates.

The probability model must match the data structure. The Bernoulli model applies when each individual is classified as case or non-case, and the Poisson model applies when cases are rare relative to the population. For livestock data, the Poisson model is appropriate when the number of affected animals is small relative to the number at risk, as in the bovine tuberculosis slaughter surveillance data [spatial statistics and monitoring data for bovine tuberculosis clustering in Argentina](https://pubmed.ncbi.nlm.nih.gov/12419600/). The Bernoulli model is appropriate when the data are structured as case-control or when the population denominator is not available.

The number of Monte Carlo replications determines the precision of the P value. A minimum of 999 replications is standard, and 9999 replications provide more stable estimates for cluster boundaries. The random number seed should be recorded to allow replication of the analysis.

### Interpreting Cluster Output

The scan statistic produces a most likely cluster and a set of secondary clusters. The most likely cluster is the window with the maximum likelihood ratio. Secondary clusters are ranked by their likelihood ratio, but overlapping clusters are often reported as separate entries. The user must decide whether overlapping secondary clusters represent independent signals or artefacts of the scanning window.

The relative risk within the cluster is the ratio of the observed to expected cases, adjusted for the population distribution. A relative risk of 5.0 with a P value below 0.05 indicates a five-fold elevation in disease risk within the cluster window. Confidence intervals for the relative risk are not routinely provided by SaTScan, and the user should compute them separately if required.

The Q-fever investigation illustrates the interpretive value of the scan statistic when combined with attack rate analysis. The study calculated relative risks for concentric zones around suspect farms and identified a relative risk of 31.1 for persons living within 2 kilometres of an affected dairy goat farm compared with those living more than 5 kilometres away [GIS-based source detection for the urban Q-fever outbreak](https://pubmed.ncbi.nlm.nih.gov/20230650/). This zone-based approach complements the scan statistic by providing a dose-response gradient that supports causal inference.

### Documentation and Reporting

Cluster detection results should be reported with sufficient detail for replication. The report should include the case definition, the population denominator, the spatial unit of analysis, the coordinate system, the probability model, the maximum cluster size, the number of Monte Carlo replications, and the random number seed. The coordinates and radius of each reported cluster should be given, along with the observed and expected case counts, the relative risk, and the P value.

The spatial scan statistic identifies clusters but does not explain their cause. A cluster may reflect a common source, such as a contaminated water supply or a shared livestock market, or it may reflect surveillance artefact, such as differential reporting intensity. The Cryptosporidium study in Egypt found that cluster analysis revealed differences in the distribution of infections between animals and humans, suggesting different transmission dynamics that would not have been apparent from prevalence estimates alone [frequencies and spatial distributions of Cryptosporidium in livestock animals and children in Ismailia province](https://pubmed.ncbi.nlm.nih.gov/25084317/). Follow-up investigations should combine cluster results with risk factor data and field visits to determine whether the cluster represents a true outbreak, an endemic hot spot, or a reporting bias.

### Species and Production System Considerations

The correct choice of cluster detection method and its parameters depends on the production system. In intensive pig and poultry operations, the premises is the natural unit of analysis, and the population denominator is the number of animals on each premises. In pastoral and free-range systems, animal movements complicate the definition of the population at risk, and the analysis may need to account for seasonal movement patterns.

Disease ecology also affects the choice of method. For zoonotic diseases with environmental reservoirs, such as anthrax, the spatial distribution of cases may reflect soil characteriztics instead of animal-to-animal transmission. The anthrax analysis in Azerbaijan used a combination of spatial analysis and cluster detection to document a westward-to-eastward shift in case distribution over time, a pattern that would not have been detected by a single cross-sectional cluster analysis [changing patterns of human anthrax in Azerbaijan](https://pubmed.ncbi.nlm.nih.gov/25032701/). For such diseases, the scan statistic should be applied separately to distinct time periods to detect changes in cluster location.

Surveillance data quality varies by species and region. Slaughter surveillance data, as used in the Argentine bovine tuberculosis study, are subject to selection bias because only a fraction of the national herd is slaughtered under federal inspection [spatial statistics and monitoring data for bovine tuberculosis clustering in Argentina](https://pubmed.ncbi.nlm.nih.gov/12419600/). Passive surveillance data under-report subclinical disease, and cluster detection on such data may identify clusters of diagnostic effort instead of clusters of infection. The WOAH terrestrial animal health standards require member countries to document the sensitivity of their surveillance systems, and this documentation should inform the interpretation of cluster detection results [WOAH terrestrial animal health standards](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/).

## Recognized Complications and Failure Modes

Spatial cluster detection methods fail in characteriztic ways, and recognizing the failure mode early prevents wasted effort and misleading conclusions. The most common complication is the edge effect, where points near the study boundary have fewer neighbours available and therefore appear less clustered than they truly are. The spatial scan statistic handles this internally by allowing candidate windows to extend beyond the study region, but local indicators such as Moran's I and LISA do not. For LISA-based analyzes, apply an explicit edge correction or interpret boundary-adjacent results with caution.

A second frequent complication is multiple testing. The spatial scan statistic evaluates thousands of candidate windows, and the reported p-value already accounts for this through Monte Carlo simulation. Local indicators do not share this property. When many local statistics are computed simultaneously, the expected number of false positives rises. Apply a false discovery rate correction, such as the Benjamini-Hochberg procedure, to LISA outputs before declaring any location significant.

Precision of coordinates is a third failure mode. Data collected from memory, from coarse administrative units, or from handheld GPS units with poor satellite visibility can place cases hundreds of metres from their true location. Cluster methods are sensitive to this noise, particularly when the expected cluster radius is small. Record the coordinate accuracy at the time of data collection and stratify or exclude records with unacceptable positional error.

## Common Errors and Corrective Actions

Less experienced analysts often conflate the absence of global clustering with the absence of local clusters. Global tests such as Moran's I average spatial autocorrelation across the entire study area, and a non-significant global result does not rule out a strong local cluster. Conversely, a significant global result does not identify where the cluster is. Run both global and local procedures and report them as complementary, not interchangeable.

A second recurring error is the misuse of the Bernoulli model in the spatial scan statistic. The Bernoulli model requires cases and non-cases from a defined population at risk. Using it with case locations only, or with an arbitrarily selected set of control points, produces biased cluster locations and inflated significance. Use the Poisson model when population denominators are available at the census tract or county level, and reserve the Bernoulli model for case-control designs where the control selection is defensible.

A third error is the interpretation of cluster boundaries as biologically meaningful. The most likely cluster reported by the scan statistic is the single window that maximizes the likelihood ratio, and its boundary is an artefact of the window shape, usually circular or elliptical. The true disease focus may extend beyond or fall entirely within that boundary. Report the cluster center and radius, and treat the boundary as a starting point for field investigation instead of a precise delineation.

## Limitations of the Current Evidence

The evidence base for spatial cluster detection in veterinary medicine is uneven. Most published applications come from livestock diseases with strong reporting systems, such as bovine tuberculosis in Argentina, where slaughterhouse surveillance data supported county-level cluster detection [Perez et al., spatial statistics and monitoring data for bovine tuberculosis clustering](https://pubmed.ncbi.nlm.nih.gov/12419600/). In contrast, diseases of wildlife, companion animals, and low-resource production systems are underrepresented, and the performance of cluster methods under sparse reporting is not well characterized.

Expert opinion still differs on the choice of cluster window shape. The circular window used by the default spatial scan statistic is poorly suited to elongated features such as river valleys or transport corridors. Elliptical windows address this but require the analyst to specify a shape parameter, and the results can be sensitive to that choice. Some authors advocate for flexible windows that adapt to the data, while others prefer the simplicity and reproducibility of fixed shapes. No consensus exists, and sensitivity analysis across window shapes is the current best practice.

The integration of cluster detection with routine surveillance also remains contested. The [WOAH animal health surveillance standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/) describe reporting obligations but do not prescribe specific cluster detection methods. Whether cluster analysis should run continuously on incoming surveillance data, or only during outbreak investigations, depends on the disease, the surveillance system, and the resources available. Published examples such as the anthrax analysis in Azerbaijan demonstrate the value of retrospective cluster detection over multi-year datasets, but prospective applications are less common [Kracalik et al., changing patterns of human anthrax in Azerbaijan](https://pubmed.ncbi.nlm.nih.gov/25032701/).

## Escalation and Referral Criteria

Cluster detection results should trigger escalation when they meet any of three conditions: the cluster involves a notifiable disease, the cluster suggests an ongoing common source, or the cluster is large enough to affect trade or movement decisions. Notifiable disease reporting follows the requirements of the relevant national authority, and the [WOAH terrestrial animal health code](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/) provides the international framework for notification and trade-related measures.

Laboratory involvement is warranted when cluster detection identifies a spatial concentration of cases but the diagnostic basis is uncertain. Confirmatory testing of a sample of cases within the cluster distinguishes a true disease focus from a diagnostic artefact, such as a change in testing practices or a contaminated laboratory batch. The [CDC principles of epidemiology in public health practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html) describe the role of laboratory confirmation in outbreak verification.

Specialist consultation is appropriate when the cluster analysis is part of a larger investigation involving multiple species or human health. The Q-fever outbreak in the Netherlands, where a dairy goat farm was identified as the source of human cases through GIS-based attack rate analysis, illustrates the value of integrating veterinary and public health expertise [Schimmer et al., GIS identification of a dairy goat farm as the source of an urban Q-fever outbreak](https://pubmed.ncbi.nlm.nih.gov/20230650/). A veterinary epidemiologist with spatial analysis training should review the analytical choices before results are used for regulatory action.

## Troubleshooting Table

| Observation | Likely Cause | Discriminating Check |
|---|---|---|
| All clusters appear at study boundaries | Edge effects in LISA or distance-based methods | Re-run with edge correction, compare with scan statistic results |
| Cluster p-values are extremely small across the entire map | Multiple testing without correction | Apply false discovery rate correction, check for global autocorrelation |
| Bernoulli model produces clusters with no population denominator | Case-only data used with Bernoulli model | Verify that non-case locations exist, switch to Poisson model if denominators are available |
| Cluster boundaries shift dramatically with window shape | Genuine cluster shape differs from assumed window | Run sensitivity analysis with circular, elliptical, and flexible windows |
| Significant global clustering but no local clusters | Global test detects broad trend, local tests lack power | Examine smoothed maps, consider increasing the maximum cluster size |
| Cluster appears where no cases were reported | Coordinate error or geocoding mismatch | Audit coordinate accuracy, verify case addresses against original records |

## Frequently Asked Questions

### How Much Do Spatial Cluster Detection Methods Cost to Implement?

The cost varies widely depending on the method chosen. Free and open-source software, including R packages and SaTScan, can perform the spatial scan statistic and local indicators of spatial association at no financial cost. The main expenses are personnel time for data preparation and interpretation, plus any costs associated with collecting accurate geographic coordinates. Handheld GPS units are inexpensive, but geocoding existing address data may require commercial software or services. For a single outbreak investigation, a basic analysis can be completed in days using free tools. Ongoing surveillance programs that require repeated analyzes will need dedicated computing time and trained staff, but the marginal cost per analysis remains low once the data pipeline is established.

### What Should I Do When I Lack Precise Geographic Coordinates?

When exact point locations are unavailable, aggregate data at the smallest administrative unit you have, such as village, county, or postal code. The spatial scan statistic accommodates this structure through Poisson models that use population-at-risk denominators. Be aware that results depend on the size and shape of these units, a problem known as the modifiable areal unit problem. Coarse units can mask clusters that fall within a single unit or create artificial boundaries. If you have only partial coordinates, consider whether the missing data are spatially biased before proceeding. The [CDC principles of epidemiology in public health practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html) provide guidance on handling incomplete surveillance data. Document the coordinate precision and its limitations in your report so readers can judge the strength of the evidence.

### How Do Cluster Detection Results Differ Between Wildlife and Domestic Livestock?

Wildlife populations present distinct challenges. Home ranges, migration corridors, and seasonal movements mean that a static point location poorly represents an animal's exposure history. Detection methods that assume fixed locations will misattribute cases near range boundaries. Livestock are usually tied to farm premises, making point-based analysis more defensible, though animals moved between markets or pastures complicate this. Population-at-risk denominators are also harder to define for wildlife, so prevalence-based cluster statistics may be unreliable. For wildlife, consider using kernel density estimates or movement-informed exposure surfaces before applying scan statistics. The [WOAH animal health surveillance standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/) describe surveillance expectations that apply across species and production systems.

### What Records Should I Keep to Support Future Cluster Analyzes?

Maintain the raw case dataset with unique animal or herd identifiers, species, case definition used, and the date of diagnosis or sample collection. Record the geographic coordinate, the method used to obtain it, and its estimated accuracy. Keep the population-at-risk data, including the source and the date it was collected, because cluster results are sensitive to denominator quality. Document every analytical decision, including the software version, the scan statistic settings, the maximum cluster size, and the number of Monte Carlo replications. Preserve the original data files and the analysis scripts or log files. This level of documentation allows another analyst to reproduce your results and supports regulatory review. The [WOAH terrestrial animal health code](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/) outlines reporting expectations that may apply to notifiable diseases.

### How Do I Explain Cluster Findings to a Farm Owner or Production Manager?

Focus on what the cluster means for their operation, not on the statistical mechanics. Describe whether the cluster suggests a common source, such as contaminated water or a shared livestock market, and what practical steps they can take while confirmatory testing proceeds. Be explicit about uncertainty. A detected cluster is a statistical signal, not proof of a specific cause. Explain that the analysis identifies where cases are concentrated, and that follow-up investigation is needed to identify why. Avoid language that implies blame or certainty about transmission pathways. The [MSD Veterinary Manual professional edition](https://www.msdvetmanual.com/) offers guidance on communicating clinical and epidemiologic findings to clients. Provide written recommendations that include biosecurity measures and a timeline for re-evaluation.

### When Should I Seek Help From a Specialist Epidemiologist?

Seek specialist input when the analysis will inform regulatory action, when the disease is notifiable, or when the stakes of a false positive or false negative are high. Also consult a specialist if your data have complex features, such as spatial heterogeneity in population density, temporal components, or repeated sampling of the same animals. The spatial scan statistic and related methods have assumptions that are easy to violate, and the [AVMA practice resources](https://www.avma.org/resources-tools) can help you identify qualified consultants. If your preliminary analysis produces a cluster that would trigger a control program, have the results independently reviewed before acting. Specialists can also advise on advanced methods, such as adjusting for covariates or combining multiple data sources, that are beyond the scope of routine software use.

## Related Clinical & Scientific Guides

* [Evaluating Veterinary Surveillance System Attributes](/knowledge/veterinary-medicine/veterinary-epidemiology/evaluating-veterinary-surveillance-system-attributes)
* [Network Analysis for Infectious Disease Spread in Animal Populations](/knowledge/veterinary-medicine/veterinary-epidemiology/network-analysis-infectious-disease-spread-animal-populations)
* [Randomized Controlled Trials in Veterinary Field Settings](/knowledge/veterinary-medicine/veterinary-epidemiology/randomized-controlled-trials-veterinary-field-settings)


## References and Further Reading

- [Geospatial and age-related patterns of Taenia solium taeniasis in the rural health zone of Kimpese, Democratic Republic of Congo.](https://pubmed.ncbi.nlm.nih.gov/26996821/). 2017.
- [Frequencies and spatial distributions of Cryptosporidium in livestock animals and children in the Ismailia province of Egypt.](https://pubmed.ncbi.nlm.nih.gov/25084317/). 2015.
- [Changing patterns of human anthrax in Azerbaijan during the post-Soviet and preemptive livestock vaccination eras.](https://pubmed.ncbi.nlm.nih.gov/25032701/). 2014.
- [Use of spatial statistics and monitoring data to identify clustering of bovine tuberculosis in Argentina.](https://pubmed.ncbi.nlm.nih.gov/12419600/). 2002.
- [Detecting spatial regimes in ecosystems.](https://pubmed.ncbi.nlm.nih.gov/28000431/). 2017.
- [The use of a geographic information system to identify a dairy goat farm as the most likely source of an urban Q-fever outbreak.](https://pubmed.ncbi.nlm.nih.gov/20230650/). 2010.
- [WOAH Animal Health Surveillance Standards](https://www.woah.org/en/what-we-do/animal-health-and-welfare/disease-data-collection/). WOAH.
- [CDC Principles of Epidemiology in Public Health Practice](https://www.cdc.gov/csels/dsepd/ss1978/index.html). CDC.
- [MSD Veterinary Manual, Professional Edition](https://www.msdvetmanual.com/). MSD Veterinary Manual.

## Related Articles

- [Spatial Epidemiology in Veterinary Science: Mapping Disease Clusters](/knowledge/veterinary-medicine/veterinary-epidemiology/spatial-epidemiology-veterinary-science-mapping-disease-clusters)
- [Participatory Epidemiology Methods for Livestock Disease Surveillance](/knowledge/veterinary-medicine/veterinary-epidemiology/participatory-epidemiology-methods-livestock-disease-surveillance)
- [Using Simulation Models in Veterinary Epidemiology](/knowledge/veterinary-medicine/veterinary-epidemiology/using-simulation-models-veterinary-epidemiology)
- [Cluster Sampling in Veterinary Field Studies](/knowledge/veterinary-medicine/veterinary-epidemiology/cluster-sampling-veterinary-field-studies)
- [Basic Reproductive Ratio (R0) in Veterinary Epidemiology](/knowledge/veterinary-medicine/veterinary-epidemiology/basic-reproductive-ratio-r0-veterinary-epidemiology)

> This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.