Using Simulation Models in Veterinary Epidemiology
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Stochastic and agent-based simulation models are crucial for understanding disease spread by explicitly representing individual variability, chance events, and contact structures, differing from deterministic compartmental models.
- Agent-based models represent individual animals or herds as discrete agents with specific rules, enabling the capture of emergent phenomena and the simulation of heterogeneous populations and targeted interventions.
- Contact network data, derived from livestock movement records or proximity logger data, are essential for defining transmission pathways and evaluating the impact of interventions that reconfigure network topology.
- Model validation requires comparing simulation output to historical outbreak data, employing sensitivity analysis to assess parameter uncertainty, and reporting distributions of outcomes rather than single point estimates.
- Transparent reporting of model assumptions, parameter sources, and code availability is critical for reproducibility and external review, with adherence to WOAH standards for surveillance and reporting being paramount for policy support.
- Simulation models are complementary to field trials, valuable for comparing candidate control strategies, exploring ethically or practically infeasible scenarios, and extrapolating findings across different production systems.
Simulation models have become indispensable tools in veterinary epidemiology for understanding how pathogens move through animal populations and for comparing the likely outcomes of alternative control interventions. These models translate biological assumptions about transmission, host behavior, and intervention effects into computational experiments that can be run faster and more safely than field trials. This article provides a technical reference for veterinary researchers who design, interpret, or critically appraise simulation studies in animal disease epidemiology. It covers the conceptual foundations of stochastic and agent-based approaches, the structure of movement and contact networks, practical considerations for model construction and validation, and the interpretation of simulation output for policy support.
The reader is assumed to be familiar with basic epidemiological measures such as incidence, prevalence, and the basic reproductive ratio, and to have some exposure to quantitative methods. The focus here is on simulation approaches that represent variability and individual-level processes, which differ fundamentally from deterministic compartmental models. Those deterministic frameworks are covered separately in this reference series and are excluded from the present discussion.
Simulation modeling in veterinary epidemiology serves two primary purposes. The first is explanatory: models help identify which mechanisms most plausibly generate observed patterns of disease spread. The second is predictive or comparative: models allow researchers to project the consequences of different control strategies under controlled, repeatable conditions. Both purposes rely on the same underlying logic, namely that a model is an explicit, simplified representation of a system, and that its value depends on the transparency of its assumptions and the rigour of its evaluation.
At a Glance
| Parameter or Decision | What the Reader Needs to Know |
|---|---|
| Model type selection | Choose stochastic or agent-based models when individual variability, chance events, or contact structure materially affect outcomes, deterministic compartmental models are covered elsewhere |
| Core components of a simulation model | Population structure, contact processes, transmission rules, disease progression, intervention modules, and time step |
| Stochasticity | Random variation in transmission, duration of infectious periods, and intervention effectiveness, requires multiple replicate runs |
| Contact network data | Livestock movement records, proximity logger data, and census data define who can transmit to whom |
| Validation approach | Compare model output to historical outbreak data, use sensitivity analysis, and report uncertainty bounds |
| Control strategy evaluation | Compare scenarios under identical model conditions, report distributions of outcomes, not single point estimates |
| Reporting standards | Document assumptions, parameter sources, and code availability to support reproducibility and external review |
Conceptual Foundations of Simulation in Epidemiology
Mathematical modeling has a long tradition in infectious disease epidemiology, and the nonlinear dynamics of pathogen transmission create genuine challenges for identifying key determinants and designing effective mitigation strategies. Simulation approaches address these challenges by explicitly representing variability, interconnectedness, and complexity within a system. The expansion of computing power and the emergence of novel data sources, including proximity loggers and global positioning systems, have considerably broadened the modeling toolbox available to veterinary epidemiologists. These advances are described in a review of complex system modeling for veterinary epidemiology, which discusses the principles and challenges of agent-based, network, and related approaches.
Stochastic Simulation
Stochastic simulation models incorporate random variation into the processes that drive disease spread. Where a deterministic model produces a single predicted trajectory from a given set of starting conditions, a stochastic model produces a distribution of possible trajectories. This distinction matters in veterinary epidemiology because real outbreaks are shaped by chance events: an infected animal may or may not contact a susceptible neighbour within a given time window, and the duration of infectiousness varies between individuals. When outbreaks are small, when pathogen introduction events are rare, or when control interventions are applied to small populations, stochastic effects can determine whether an epidemic takes off at all.
The practical consequence is that stochastic models must be run many times, typically hundreds or thousands of replicates, and the output summarized as distributions. Reporting the median, interquartile range, and tail probabilities of outcomes such as epidemic size, duration, or total culls is more informative than reporting a single number. The transparent and flexible multiscale stochastic modeling framework EMULSION was developed specifically to address the challenges of readability, reproducibility, and development time in stochastic mechanistic models, and it illustrates the broader movement toward standardized model description in this field.
Agent-Based Models
Agent-based models represent individual animals, herds, or farms as discrete entities called agents, each governed by rules that determine how it interacts with other agents and with the environment. This approach is particularly suited to veterinary problems because livestock populations are organized into production units with heterogeneous sizes, management practices, and contact patterns. An agent-based model can represent a smallholder herd differently from a large commercial operation, assign different movement frequencies to different production types, and allow intervention effects to act on specific agent categories.
The power of agent-based modeling lies in its ability to capture emergent phenomena, patterns that arise from the local interactions of individuals instead of from a central equation. For example, the spatial clustering of outbreaks can emerge from local contact rules even when the model contains no explicit spatial term. The trade-off is computational cost and parameter uncertainty. Agent-based models require detailed data on individual behavior, and their outputs can be sensitive to assumptions that are difficult to verify empirically.
Model Selection and Design Decisions
The choice among stochastic simulation, agent-based, and hybrid approaches depends on the research question, data availability, and the biological system under investigation. Agent-based modeling is a powerful simulation technique that considers individual behaviors by defining rules that govern how agents within given populations interact with one another and the environment. This makes agent-based models particularly suitable when individual heterogeneity, spatial structure, or behavioral responses drive transmission dynamics.
Stochastic compartmental models remain computationally efficient and are well suited to questions about population-level outcomes when the population can be reasonably treated as homogeneous within compartments. They require fewer parameters and are easier to communicate to non-modellers. However, they cannot represent individual-level variation in contact patterns, movement histories, or intervention compliance without substantial structural extensions.
Agent-based models excel when the system features identifiable individuals whose attributes and behaviors materially affect transmission. Examples include farm-level models where herd size, production type, and movement history determine risk, or wildlife models where territorial behavior and social structure shape contact networks. The cost is computational burden, parameter uncertainty, and the difficulty of validating emergent behavior against field data.
Hybrid models, such as those supported by the EMULSION framework, combine multiple scales within a single simulation. A model might represent within-herd transmission stochastically while simulating between-herd spread through an agent-based movement network. These approaches address the three major challenges of predictive epidemiology: readability of model code, reproducibility, and development time relative to outbreak time scale.
| Model Type | Data Requirements | Computational Cost | Best Suited For | Key Limitations |
|---|---|---|---|---|
| Stochastic compartmental | Aggregate incidence, population sizes, contact rates | Low to moderate | Population-level intervention comparisons, R0 estimation | No individual heterogeneity, poor spatial resolution |
| Agent-based | Individual-level attributes, movement or contact data, spatial coordinates | High | Contact network effects, targeted interventions, behavioral responses | Parameter-rich, difficult to validate, computationally intensive |
| Hybrid multiscale | Data at multiple organizational levels | Very high | Within-host to between-host coupling, metapopulation systems | Complex to implement, requires multidisciplinary expertise |
Species and production system change the correct choice. For intensively housed swine or poultry populations with frequent batch movements, network-based agent-based models capture the shipment patterns that dominate transmission. For extensive beef systems where pasture contact and wildlife interfaces matter, spatial agent-based models with environmental transmission pathways are more appropriate. Companion animal populations, with their heterogeneous owner behavior and clinic-based contact structures, may be better represented by stochastic models that incorporate owner compliance as a probability distribution.
Building a Simple Stochastic Model
A practical stochastic simulation project proceeds through six stages. The sequence applies across species and production systems, though the specific decisions at each stage vary.
Stage 1: Define the question and outcome. Specify the target population, the pathogen, and the decision the model will inform. A model designed to compare vaccination strategies for foot-and-mouth disease requires different structure than one estimating time to detection under passive surveillance. The outcome measures, such as epidemic duration, total cases, or probability of extinction, must be defined before model construction.
Stage 2: Choose the model architecture. Select the state variables and transition events. For a stochastic compartmental model, define the compartments relevant to the natural history of the infection. For an agent-based model, define agent attributes, behavioral rules, and interaction processes. The WOAH terrestrial animal health code provides standardized case definitions and surveillance requirements that should inform how infection states are defined and detected.
Stage 3: Parameterise from data and literature. Estimate transition rates, contact frequencies, and initial conditions from surveillance data, movement records, or published studies. Where parameters are uncertain, define plausible ranges instead of point estimates. The CDC principles of epidemiology provide the standard measures of disease frequency and association that underpin parameter estimation from field data.
Stage 4: Implement the simulation. Write the code or use a modeling platform that separates model specification from simulation engine. The readability of the model specification matters for reproducibility and for allowing epidemiologists, biologists, and economists to validate assumptions at any stage of development.
Stage 5: Run and analyze. Execute multiple stochastic realisations to capture the distribution of outcomes. A single run of a stochastic model is a single sample from the outcome distribution. Report medians, credible intervals, and extinction probabilities across runs, also the mean trajectory.
Stage 6: Validate and report. Compare model output against independent data not used in parameterisation. Report the model structure, parameter values, and assumptions transparently so that others can reproduce the work.
Network-Based Simulation Approaches
Livestock movement data have become a central input for simulation models because livestock movements are important in spreading infectious diseases and many countries require farmers to report movements to authorities. Network analysis characterizes the relationships among farms and other livestock operations, providing information on their role in acquiring and spreading infection that traditional movement studies cannot supply.
Network metrics guide model construction and intervention design. Node-level measures such as in-degree and out-degree identify farms that receive or send many shipments. Betweenness centrality identifies premises that bridge otherwise disconnected parts of the network. These measures predict which farms, if infected or if targeted for control, would have disproportionate effects on epidemic spread.
Simulation studies have demonstrated that disease dynamics can be altered by placing targeted restrictions on contact formation to reconfigure network topology. Preventing farms with high in-degree from selling to farms with high out-degree, a connection type predicted to have a disproportionately strong role in spreading disease, produced significantly lower endemic prevalences in simulation compared to baseline networks. This finding has direct application to movement restriction policies during outbreaks.
Network-based simulation requires movement data of sufficient quality and temporal coverage. Where movement reporting is incomplete, models must account for unobserved contacts through stochastic contact generation or sensitivity analysis. Production system differences matter: auction market networks have different structural properties than direct farm-to-farm sale networks, and models must represent the correct contact structure for the system under study.
Spatial Simulation and Control Strategy Evaluation
Spatial analysis examines the distributions of events in space and can be applied to investigate the spread of animal disease. Epidemiological simulation modeling can study the hypothetical spread of foot-and-mouth disease and evaluate control strategies that might decrease outbreak impact or eradicate the virus from an area. Spatial simulation models incorporate farm locations, distance-dependent transmission kernels, and geographically targeted interventions such as ring culling or vaccination zones.
The choice of spatial scale affects model behavior. Farm-level spatial models treat each premise as a node with a location. Within-herd models represent individual animals and may couple to between-herd spread through movement or local transmission. The appropriate scale depends on the pathogen, the data available, and the intervention being evaluated.
Control strategies evaluated through spatial simulation include depopulation rings, vaccination rings, movement restrictions, and tracing-based culling. Each strategy has different resource requirements, welfare implications, and logistical constraints. Simulation allows these strategies to be compared under identical outbreak scenarios, providing decision support that field observation cannot offer.
Documentation and Reporting Standards
Transparent reporting of simulation models is essential for credibility and reproducibility. The model specification should state all assumptions about transmission processes, parameter values and their sources, initial conditions, and the random number generation approach. The WOAH animal health surveillance standards provide a framework for describing surveillance systems that models represent, including case definitions, detection probabilities, and reporting delays.
Documentation should distinguish between parameters estimated from data, parameters taken from published literature, and parameters set by expert opinion. Each category carries different uncertainty, and this distinction should be preserved through sensitivity analysis. Model code should be version-controlled and archived so that published results can be reproduced exactly.
Species-specific reporting considerations apply. For production animals, the production type, housing system, and movement patterns must be described. For wildlife, the population estimation method and the spatial extent of the study area require documentation. For companion animals, the population denominator and the contact structure between households and veterinary clinics affect interpretation.
Recognized Complications and Failure Modes
Simulation models fail in characteriztic ways. The most consequential failure is structural misspecification, where the model omits a transmission pathway that operates in the field. A model built exclusively around farm-to-farm livestock movements will misattribute spread driven by fomites, wildlife bridges, or airborne plumes. Detection requires external validation against outbreak data that were not used in calibration. If the model cannot reproduce the spatial and temporal pattern of a known epidemic, structural revision is indicated instead of parameter tuning.
Parameter identifiability problems arise when multiple parameter sets produce nearly identical output. This non-identifiability is common in stochastic models with many correlated parameters. Early detection uses profile likelihood analysis or Markov chain Monte Carlo diagnostics. When two parameters trade off against each other, such as transmission rate and infectious period, the model may fit well yet make poor predictions under control interventions that affect only one parameter. Report parameter uncertainty intervals alongside point estimates.
Stochastic models also suffer from inadequate replication. A single realisation of a stochastic process can mislead. Variance estimates from too few runs produce false confidence. The discriminating check is to plot the convergence of the outcome distribution as replicate count increases. When the mean and variance stabilize, replication is adequate. When they drift, the model requires more runs or variance reduction techniques.
Numerical instability, particularly in agent-based models with many interacting entities, can produce artefactual oscillations or extinction events. Detection involves running the model with smaller time steps and comparing outputs. If results change materially, the time step is too coarse for the process dynamics.
| Observation | Likely Cause | Discriminating Check |
|---|---|---|
| Model predicts no outbreak where field data show spread | Missing transmission pathway | Compare model structure against outbreak investigation findings |
| Wide credible intervals on intervention effect | Non-identifiable parameters | Profile likelihood analysis, examine parameter correlations |
| Outcome variance shrinks then grows with more runs | Insufficient replication or unstable dynamics | Plot convergence of mean and variance against replicate count |
| Oscillatory output at fixed time step | Numerical instability | Halve time step, compare results |
| Model fits calibration data but fails on new data | Overfitting or structural error | Hold out independent data for validation |
Common Errors in Model Development
Less experienced modellers frequently conflate model complexity with model fidelity. Adding agents, spatial detail, or behavioral rules does not guarantee better prediction. Each added component introduces parameters that must be estimated, and each estimation carries uncertainty. The corrective action is to begin with the simplest stochastic structure that captures the question, then add complexity only when sensitivity analysis demonstrates that the added component changes conclusions.
A second error is treating the model output as a point forecast instead of a distribution. Stochastic models produce a range of possible outcomes. Presenting the mean without the spread misrepresents the uncertainty that decision-makers need. Report quantiles, particularly the 5th and 95th percentiles, and describe the full distribution for critical outcomes such as epidemic size or duration.
A third error involves circular reasoning in control strategy evaluation. If the model assumes a control intervention works in a particular way, the simulation will naturally show benefit. For example, a model that reduces transmission rate by a fixed proportion when vaccination is applied cannot reveal whether vaccination would achieve that reduction in practice. The corrective action is to model the mechanism of the intervention, not its assumed effect, and to subject intervention assumptions to sensitivity analysis.
A fourth error is ignoring the contact network structure that underlies transmission. Livestock movements create networks with characteriztic features that drive disease dynamics. Models that assume homogeneous mixing will misestimate spread in systems where a small number of high-degree premises dominate transmission. Network-based approaches that preserve degree distributions and demographic characteriztics of movements provide more realistic transmission pathways.
Limitations of Current Evidence
The evidence base for veterinary simulation models is uneven. Foot-and-mouth disease has received extensive modeling attention, with reviews covering simulation approaches and spatial analysis for control strategy evaluation. Other diseases, particularly those with wildlife reservoirs or vector transmission, have thinner modeling literatures. Expert opinion still differs on how to represent wildlife-livestock interfaces, on the appropriate spatial scale for transmission events, and on how to parameterise behavioral responses to control measures.
Model validation against real outbreaks remains the exception instead of the rule. Many published models demonstrate internal consistency but lack external validation because suitable outbreak data are unavailable or inaccessible. This limitation should be stated explicitly in model reports. The reproducibility crisis in computational science also affects veterinary epidemiology. Models described only in prose cannot be reimplemented reliably. Structured modeling frameworks that make model components explicit as readable text files improve transparency and allow domain experts to review assumptions.
Escalation and Referral
Simulation modeling is a specialist activity. When the modeling question informs regulatory policy, trade decisions, or national disease control, involve veterinary epidemiologists with formal modeling training. The World Organization for Animal Health provides international standards for surveillance and disease reporting that should frame any modeling exercise intended to inform official control programs. Models used to support trade-related decisions must align with the terrestrial animal health standards, including provisions for surveillance, notification, and safe movement of animals and products.
Laboratory involvement is warranted when parameter estimation requires pathogen-specific data, such as transmission rates, infectious periods, or diagnostic test performance. These parameters often come from experimental infection studies or field investigations. When model outputs are sensitive to these parameters, the uncertainty should be communicated to the laboratory scientists who can design studies to reduce it.
Regulatory reporting obligations arise when a model identifies a plausible risk of introduction or spread of a notifiable disease. In such circumstances, the modeling team should alert the relevant veterinary authority, even when the model is exploratory. The authority can then decide whether the modelled scenario warrants enhanced surveillance or contingency planning. Modellers should document the assumptions and limitations of their analysis in the report to the authority, making clear that the model is a decision-support tool, not a prediction of certain events.
Frequently Asked Questions
How much time and computational expertise do I need to build a useful simulation model?
A practical stochastic or agent-based model can be developed in weeks instead of months if you restrict scope to a single transmission pathway and a clearly defined control question. The main time cost is not coding but parameter estimation and validation against field data. Modern frameworks such as EMULSION use domain-specific languages that separate model design from implementation, allowing epidemiologists to specify structure and processes as structured text files that a generic simulation engine executes automatically. This reduces programming burden and improves reproducibility. For most veterinary applications, a desktop workstation suffices, large national-scale networks may require cluster computing, but farm-level models rarely do.
What should I do when movement or contact data are incomplete or unavailable?
Start with a simpler network representation instead of abandoning simulation. If farm-to-farm movement records are missing, use distance-based contact kernels or published contact rates from comparable production systems. Network analysis methods for livestock movements can characterize connectivity even with partial data, and sensitivity analysis will show whether conclusions depend on the missing information. When data gaps are substantial, present results as relative comparisons between control strategies instead of absolute predictions. Clearly state data limitations in the reporting, and consider collecting targeted movement data from a small sample of farms to anchor the model.
How do simulation approaches differ between intensively housed livestock and free-ranging wildlife?
The contrast is primarily in contact structure and parameter identifiability. Intensively housed livestock have well-defined premises boundaries, recorded movements, and known population sizes, which supports network-based simulation using movement data. Free-ranging wildlife requires agent-based approaches that model individual movement, home ranges, and environmental contact points, with parameters estimated from telemetry or proximity logger data. Agent-based modeling principles in complex systems apply to both, but wildlife models carry greater uncertainty in contact rates and population density. For wildlife, focus on relative strategy comparisons and report confidence intervals from multiple stochastic realisations.
What are the minimum documentation standards for a simulation study to be publishable?
Document the model purpose, all assumptions, parameter sources with justification, and the stochastic realisation protocol. Specify the number of iterations, random seed handling, and warm-up periods. Provide a complete model description in a format readable by non-modellers, as recommended for transparent and flexible multiscale stochastic models. Report sensitivity analyzes for uncertain parameters and state which outputs were primary versus exploratory. Follow the WOAH standards for animal health surveillance reporting when the model informs surveillance design. Journals increasingly require model code and data availability statements.
How should I explain simulation model results to farm owners or veterinary supervisors?
Frame results as comparisons between control options, not as point predictions of case numbers. State the range of outcomes across stochastic realisations, for example "in 80 percent of simulated outbreaks, movement restriction reduced total infected premises by half or more." Explain that the model encodes current understanding of transmission and that uncertainty reflects real variability. Avoid presenting a single epidemic curve as the expected outcome. For regulatory discussions, reference international standards such as the WOAH terrestrial animal health code when model outputs inform trade-related control decisions.
Can simulation models replace field trials for evaluating disease control strategies?
No. Simulation models are complementary to field studies, not substitutes. They are most valuable for comparing candidate strategies before committing resources to field trials, for exploring scenarios that cannot be tested ethically or practically, and for extrapolating results across production systems. Simulation modeling for foot-and-mouth disease control has informed policy where field experiments would be impossible. However, model outputs require external validation against outbreak data whenever available. Use models to prioritize which strategies merit field evaluation, then validate model predictions against the trial results and revise the model accordingly.
Related Clinical & Scientific Guides
- Evaluating Veterinary Surveillance System Attributes
- Network Analysis for Infectious Disease Spread in Animal Populations
- Randomized Controlled Trials in Veterinary Field Settings
References and Further Reading
- EMULSION: Transparent and flexible multiscale stochastic models in human, animal and plant epidemiology.. 2019.
- Complex system modeling for veterinary epidemiology.. 2015.
- Epidemiological simulation modeling and spatial analysis for foot-and-mouth disease control strategies: a comprehensive review.. 2011.
- A review of network analysis terminology and its application to foot-and-mouth disease modeling and policy development.. 2009.
- Controlling infectious disease through the targeted manipulation of contact network structure.. 2015.
- A standardized framework to identify optimal animal models for efficacy assessment in drug development.. 2019.
- WOAH Animal Health Surveillance Standards. WOAH.
- CDC Principles of Epidemiology in Public Health Practice. CDC.
- MSD Veterinary Manual, Professional Edition. MSD Veterinary Manual.
Related Articles
- Regression Analysis in Veterinary Epidemiology: Logistic and Poisson Models
- Compartmental Models in Veterinary Disease Dynamics
- Sensitivity Analysis in Veterinary Disease Models
- Basic Reproductive Ratio (R0) in Veterinary Epidemiology
- Effect Modification and Interaction in Veterinary Epidemiology
This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.