Genomic Epidemiology: Integrating Pathogen Genomics into Outbreak Investigations
Genomic epidemiology is the practice of using pathogen genome sequence data alongside traditional epidemiological information to trace transmission pathways, identify outbreak sources, and guide control measures. For public health professionals, researchers, and analysts, this means moving beyond case counts and exposure histories to include the genetic relatedness of the organisms causing illness. This article provides a practical framework for incorporating genomic data into outbreak investigations, from specimen selection through sequencing, bioinformatics analysis, interpretation, and reporting. The focus is on operational decisions that determine whether genomic data will be actionable during an active investigation.
The Role of Pathogen Genomics in Modern Outbreak Detection
Whole-genome sequencing has transformed public health microbiology by providing a resolution that traditional subtyping methods cannot match. PulseNet, a network of public health and regulatory laboratories, changed the landscape of foodborne illness surveillance through molecular subtyping, and next-generation sequencing technologies have made whole-genome sequencing of foodborne bacterial pathogens a realistic and superior alternative to those earlier methods [6]. This shift changes what questions investigators can ask and answer during an outbreak.
Genomic data can distinguish between cases that are part of the same transmission chain and those that are epidemiologically linked but genetically distinct. This distinction matters for resource allocation. When a cluster of infections appears in a hospital, investigators need to know whether they are seeing one ongoing transmission event or multiple independent introductions. A genomic investigation of a Bacillus cereus bacteraemia outbreak in Italy during the summer of 2023 used both short- and long-read sequencing combined with bioinformatics analysis to identify the putative source and understand transmission dynamics [7]. The investigation revealed a complex, polyclonal contamination pattern traced to contaminated hospital laundry, with multiple sequence types found in both clinical and environmental samples [7]. SNP-based phylogenetic analysis provided strong evidence linking human and environmental isolates, with close genetic relatedness observed between isolates from patients and those from laundered scrubs, transport trucks, and bed linens [7]. Improved laundry procedures resolved the outbreak [7].
This example illustrates a core principle of genomic epidemiology. The method works best when it is integrated into the investigation from the start, applied alongside traditional methods instead of after they have stalled. Real-time genomic surveillance can detect plasmid transfer events that traditional infection prevention and control methods may miss. In a multispecies outbreak of NDM-5-producing Enterobacterales at an acute care hospital, initial detection occurred via traditional infection prevention and control methods and was supplemented by real-time whole-genome sequencing surveillance performed weekly [10]. Long-read sequencing and hybrid assemblies resolved the NDM-encoding plasmids, and the investigation characterized 10 horizontal plasmid transfer events and 6 bacterial transmission events between patients in varying hospital units [10]. The combination of traditional and prospective genomic methods was essential for identifying and containing a plasmid-associated outbreak [10].
Core Principles of Genomic Epidemiology
Genetic Relatedness as a Transmission Proxy
The central assumption of genomic epidemiology is that organisms with highly similar genomes are more likely to be linked by recent transmission than organisms with divergent genomes. The resolution of this approach depends on the pathogen's mutation rate, the time frame of the outbreak, and the amount of genetic diversity present in the background population. For slowly mutating pathogens, even unrelated cases may have identical or nearly identical genomes, limiting the ability to rule out transmission. For rapidly mutating pathogens such as RNA viruses, the accumulation of mutations over the course of an outbreak can provide a molecular clock that helps date transmission events.
The application of this principle requires careful interpretation. In a Streptococcus suis outbreak in Guangxi, China, in June 2016, investigators determined the genetic characteristics of six clinically isolated strains by serotyping, PCR, and whole-genome sequencing, and performed genome epidemiology analysis on these and 961 publicly available S. suis genomes [11]. Sporadic and outbreak cases were distinguished by whole-genome sequencing and phylogenomics [11]. The approach helped prevent and control S. suis epidemics in the region [11]. The key decision was not simply whether the genomes matched, but whether the genetic clusters corresponded to the epidemiological case definitions.
Phylogenetic Context and Lineage Assignment
Phylogenetic analysis places outbreak sequences in the context of global or regional diversity. This context serves several purposes. It can identify the likely geographic origin of an outbreak strain, reveal whether multiple lineages are circulating simultaneously, and show whether the outbreak strain has been seen before in the region. During the 2015-2016 Zika virus outbreak in Cape Verde, analysis of complete genomes from three isolates indicated the strain was of the Asian lineage, and the sequences formed a distinct monophylogenetic group with unique amino acid changes in the envelope protein [9]. Phylogeographic and serologic evidence supported earlier introduction of the lineage into Cape Verde, possibly from northeast Brazil, suggesting cryptic circulation before the initial wave of cases was detected in October 2015 [9]. The genomic data extended the known timeline of the outbreak and pointed to a source region.
Lineage assignment has become a standard first step in many investigations. During the 2022 mpox outbreak in New York City, investigators sequenced 1,138 specimens from 758 individuals and performed phylogenomic analyses alongside 2,967 global mpox sequences [8]. Nextclade lineage assignment revealed a NYC-specific B.1.12 lineage, with phylogenetic analysis showing unique clusters in NYC and North America [8]. The investigation also identified that 6.4% of individuals with multiple specimens had distinct genomic profiles, with 4.2% likely due to co-infections with distinct mpox strains [8]. This finding has direct implications for surveillance strategies, because co-infections can complicate the interpretation of transmission chains.
Integration with Epidemiological Data
Genomic data are most powerful when analyzed jointly with epidemiological data. A phylogenetic tree alone cannot prove transmission. It must be interpreted alongside case interviews, exposure histories, temporal data, and geographic information. The combined visualization of genomic and epidemiological data for outbreaks is an active area of method development [22]. Investigators should expect to iterate between the genomic analysis and the epidemiological investigation, using each to refine the other.
In the dengue virus outbreak in Thiès, Senegal, in late 2018, investigators reported complete viral genomes from 17 previously undetected dengue cases and identified 19 cases in a cohort of 198 individuals with fever [12]. Three co-circulating serotypes were detected, with DENV 3 the most frequent [12]. Sequences were most similar to recent sequences from West Africa, suggesting ongoing local circulation of viral populations, though detailed inference was limited by the scarcity of available genomic data [12]. The investigators did not find clear associations with reported clinical signs or symptoms, highlighting the importance of testing for diagnosing febrile diseases [12]. This example shows both the value and the limits of genomic data when epidemiological context is incomplete.
At a Glance
| Decision Point | What to Consider | Practical Consequence |
|---|---|---|
| Specimen selection | Prioritize specimens from cases with clear epidemiological links, early in the outbreak timeline, and from diverse exposure groups | Poor selection produces a biased sample that cannot support transmission inference |
| Sequencing platform | Short-read platforms offer high accuracy for SNP detection, long-read platforms resolve plasmids and repetitive regions | Platform choice affects the types of genetic events you can detect, including horizontal gene transfer |
| Analysis approach | SNP-based phylogenetics for bacterial outbreaks, lineage assignment and mutation analysis for viral outbreaks | The analytical method must match the pathogen's biology and the outbreak question |
| Data sharing | Use national platforms and public repositories to compare sequences across jurisdictions | Isolated analysis misses cross-border links and delays cluster detection |
| Reporting timeline | Generate preliminary results within days and refine as more data become available | Delayed reporting reduces the opportunity for intervention during an active outbreak |
Building a Genomic Epidemiology Workflow
Specimen Collection and Selection
The quality of genomic epidemiology depends on the quality and representativeness of the specimens sequenced. A common error is to sequence convenience samples instead of specimens selected according to the outbreak investigation's needs. Investigators should define the sampling strategy before sequencing begins. The strategy should address the following questions. Which cases should be sequenced? How many specimens per case? Should environmental or animal specimens be included? What is the expected turnaround time from specimen collection to sequence data?
For bacterial outbreaks, culture is typically required before sequencing. This step introduces a delay but also provides an isolate that can be stored and re-analyzed. For viral outbreaks, direct sequencing from clinical specimens is often possible, which shortens the time from collection to data. The choice affects the workflow design and the interpretation of results. A positive culture confirms the presence of viable organisms, while direct sequencing may detect nucleic acid from non-viable organisms.
The number of specimens sequenced should be guided by the outbreak's size and the question being asked. Sequencing every case in a large outbreak may be unnecessary if the goal is to confirm a single source. Sequencing only a few cases may be insufficient to distinguish between a point source and ongoing transmission. Investigators should document the sampling fraction and the criteria for inclusion so that the analysis can account for potential biases.
Sequencing Platforms and Library Preparation
The choice of sequencing platform affects the resolution, cost, and turnaround time of the investigation. Short-read platforms such as Illumina are the most widely used and provide high accuracy for single nucleotide polymorphism detection. Long-read platforms such as Oxford Nanopore Technologies provide longer reads that can resolve plasmids, repetitive regions, and structural variation. A multi-country survey of next-generation sequencing and bioinformatics capacity found that Illumina and Oxford Nanopore Technologies were the most predominant platforms, reported by 89.6% and 68.8% of respondents respectively [18]. The median annual throughput reported was 1,940 samples in high-income countries, 850 in upper-middle-income countries, 1,205 in lower-middle-income countries, and 950 in low-income countries [18].
The choice between platforms is not always either-or. The Bacillus cereus outbreak investigation used both short- and long-read sequencing technologies combined with bioinformatics analysis [7]. The NDM-5 outbreak investigation used long-read sequencing and hybrid assemblies to resolve plasmids [10]. Hybrid approaches combine the accuracy of short reads with the contiguity of long reads, which is particularly valuable when mobile genetic elements are involved.
Library preparation methods can introduce bias. A fully automated workflow that integrates enzymatic lysis, extraction, and library preparation in a single-operator, single-cartridge system was designed to reduce hands-on time from 8 to 10 hours to less than 45 minutes [13]. Using mixed microbial communities, the workflow produced high-quality sequencing libraries with an average yield of 80.5 ng/µL and an average quality score of 33.6, and demonstrated a 2.46-fold improvement in Gram-positive representation relative to a standard lysis protocol [13]. This example shows that sample preparation choices can systematically affect which organisms are detected and how well they are represented in the data.
Bioinformatics Pipelines and Quality Control
Bioinformatics analysis is the step where raw sequence data become interpretable results. The analysis typically includes quality assessment, read trimming, assembly or reference-based mapping, variant calling, and phylogenetic inference. Each step has parameters that can affect the final result, and these parameters should be documented and made reproducible.
Quality control is not optional. Raw sequencing data contain adapter sequences, low-quality reads, and potential contamination. These artifacts can produce false variants that distort phylogenetic inference. Investigators should establish minimum quality thresholds for read depth, genome coverage, and variant quality before interpreting results. The thresholds should be defined in advance and applied consistently across all samples in the investigation.
A review of strategies for integrating whole-genome sequencing into antimicrobial resistance surveillance introduced key bioinformatics tools across five domains: pathogen identification, molecular epidemiology, resistance gene detection, virulence profiling, and mobile genetic element analysis [14]. The review proposed structured workflows tailored for both web-based and locally installed environments [14]. The choice between web-based and local tools depends on data sensitivity, internet access, and computational capacity. Web-based tools are easier to use and maintain but require uploading sequence data to external servers. Local tools provide more control over data but require bioinformatics expertise and computing infrastructure.
Phylogenetic Analysis and Interpretation
Phylogenetic analysis reconstructs the evolutionary relationships among the sequenced organisms. The output is typically a tree or network that shows which sequences are most closely related. The interpretation of this tree requires understanding what it can and cannot show. A phylogenetic tree shows genetic relatedness, not transmission events. Two cases with identical sequences may or may not be linked by direct transmission. They could share a common source. Two cases with different sequences are unlikely to be linked by direct transmission, unless the pathogen is evolving rapidly within hosts.
The porcine deltacoronavirus study implemented an approach to analyze the epidemiology of the virus following its emergence in the pig population, performing an integrated analysis of full genome sequence data from 21 newly sequenced viruses along with comprehensive epidemiological surveillance data collected globally over 15 years [5]. The study found four distinct phylogenetic lineages that differ in their geographic circulation patterns, and identified more frequent intra- and interlineage recombination and higher virus genetic diversity in the Chinese lineages compared with the USA lineage [5]. Most recombination breakpoints were located in the ORF1ab gene instead of in genes encoding structural proteins [5]. The study also identified five amino acids under positive selection in the spike protein, with three positively selected sites located in the N-terminal domain of the S1 subunit and two located in or near the fusion peptide of the S2 subunit [5]. Phylogeographic investigations highlighted notable South-North transmission as well as frequent long-distance dispersal events in China that could implicate human-mediated transmission [5].
This study demonstrates the analytical depth that is possible when genomic data are combined with epidemiological surveillance data. It also shows that the interpretation of phylogenetic patterns requires domain knowledge about the pathogen's biology, including recombination, selection, and geographic structure.
Data Management and Sharing
Repositories and Platforms
Genomic epidemiology depends on the ability to compare sequences across investigations, jurisdictions, and time periods. Public repositories such as the National Center for Biotechnology Information provide access to sequence data, genome assemblies, and associated metadata [2]. The European Bioinformatics Institute offers training and resources for working with these data [1]. Investigators should deposit sequence data in public repositories in a timely manner, following the data sharing policies of their funding agencies and institutions.
The National Institutes of Health Genomic Data Sharing Policy sets expectations for the sharing of genomic data generated with NIH funding [3]. Investigators should be familiar with the policy requirements before starting their work, because the policy affects how data can be shared, where they can be deposited, and what access controls are required. The policy is designed to promote data sharing while protecting participant privacy and confidentiality.
National platforms can facilitate data sharing across jurisdictions. The AusTrakka platform in Australia was established and deployed nationally to address barriers to genomic data sharing across jurisdictions, enhance interoperability and usability, and improve governance of public health genomic data [15]. An evaluation of the platform using the US Centers for Disease Control Updated Guidelines for Evaluating Public Health Surveillance Systems found that users reported a very high degree of usefulness as a centralized platform to enable sharing sequence data across jurisdictions, facilitate multijurisdictional outbreak investigations, and clarify transmission chains [15]. The evaluation used a mixed-methods approach with quantitative analysis of utilization data and qualitative interviews with 63 key informants representing all jurisdictions across Australia and New Zealand [15].
FAIR Principles and Metadata Standards
The FAIR Guiding Principles describe the characteristics that data resources should have to be findable, accessible, interoperable, and reusable [4]. These principles apply to genomic epidemiology data as much as to any other scientific data. Sequence data without metadata are of limited value. Investigators should record the specimen type, collection date, geographic location, and relevant clinical or epidemiological information for every sequenced sample.
Metadata standards vary by pathogen and by platform. Investigators should use the standards adopted by their reference laboratory or national surveillance system. The absence of standardized metadata is a common barrier to data integration. A readiness assessment for genomics-enabled antimicrobial resistance surveillance across the Philippines, Thailand, and Malaysia found that common barriers included limited system interoperability and data integration, long turnaround times, and limited governance frameworks [16]. The assessment used structured workshops with participants from sentinel sites, national reference laboratories, and research institutes [16]. All countries had strong foundations, with processes for isolate referral and critical AMR alerts at sentinel sites and sequencing platforms at national reference laboratories [16].
Data Security and Governance
Genomic data from outbreak investigations may include information that could identify individuals, particularly when combined with epidemiological metadata. Investigators must follow the data governance requirements of their institution and jurisdiction. This includes secure storage, access controls, and de-identification where required. The governance framework should be established before the investigation begins, not after data have been collected.
The survey of next-generation sequencing and bioinformatics capacity found that only 57.7% of respondents stored data in multiple locations, and 32.5% lacked any data backup [18]. This finding has direct implications for outbreak investigations. The loss of sequence data during an active investigation can delay or prevent the identification of transmission chains. Investigators should establish backup procedures and test them regularly.
Practical Implementation Steps
Step 1: Define the Outbreak Question
Before collecting specimens or generating sequences, the investigation team should write down the specific questions that genomic data are expected to answer. Examples include the following. Is this a point source outbreak or ongoing transmission? Are the cases linked to a common exposure? Is there evidence of multiple introductions? Are the outbreak strains related to strains seen in other regions or time periods? The questions determine the sampling strategy, the analytical methods, and the reporting format.
Step 2: Establish the Sampling Strategy
The sampling strategy should specify which cases will be sequenced, how many specimens per case, and whether environmental or animal specimens will be included. The strategy should be documented and shared with the laboratory and bioinformatics teams. The sampling fraction should be recorded so that the analysis can account for potential biases.
Step 3: Coordinate with the Laboratory
The laboratory needs to know the expected number of specimens, the turnaround time required, and the sequencing platform to be used. The laboratory should also be informed about any special requirements, such as the need for long-read sequencing to resolve plasmids or the need for direct sequencing from clinical specimens.
Step 4: Run the Bioinformatics Pipeline
The bioinformatics analysis should follow a documented pipeline with defined quality thresholds. The pipeline should include quality assessment, read trimming, assembly or mapping, variant calling, and phylogenetic inference. The parameters and software versions should be recorded for reproducibility.
Step 5: Interpret Results Jointly with Epidemiological Data
The genomic results should be interpreted in the context of the epidemiological investigation. This step requires close collaboration between the bioinformatics team and the epidemiological team. The interpretation should be documented, including any uncertainties or alternative explanations.
Step 6: Report Findings and Recommend Actions
The findings should be reported to the outbreak investigation team and relevant decision makers. The report should include the genomic evidence, the limitations of the analysis, and the recommended actions. The report should be updated as new data become available.
Records and Measurements
What to Record
Investigators should maintain a record of the following items for each outbreak investigation. The case definitions and inclusion criteria. The specimens collected and the criteria for sequencing. The sequencing platform and library preparation method. The bioinformatics pipeline, including software versions and parameters. The quality metrics for each sequenced sample. The phylogenetic analysis and its interpretation. The actions taken based on the genomic findings.
Turnaround Time
The turnaround time from specimen collection to reportable results is a critical operational metric. Delays at any step in the workflow reduce the opportunity for intervention. The multi-country survey found that long turnaround times were a common barrier to genomics-enabled surveillance [16]. Investigators should track the time from collection to sequencing, from sequencing to analysis, and from analysis to reporting. These metrics should be reviewed after each investigation to identify bottlenecks.
Cost Tracking
The cost of genomic epidemiology includes specimen collection, laboratory processing, sequencing, bioinformatics analysis, and data management. The genomics costing tool was initially developed to estimate the costs of SARS-CoV-2 sequencing and associated bioinformatics, and has been expanded to support a wider range of pathogens and laboratory settings [18]. Investigators should track costs to inform future budgeting and to demonstrate the value of genomic surveillance to decision makers.
Common Failure Patterns
Sampling Bias
The most common failure in genomic epidemiology is sampling bias. If the sequenced specimens are not representative of the outbreak cases, the phylogenetic analysis will produce misleading results. For example, sequencing only the most severe cases may miss mild or asymptomatic cases that are part of the same transmission chain. Sequencing only cases from one geographic area may miss the source of the outbreak.
Inadequate Quality Control
Poor quality sequence data can produce false variants that distort the phylogenetic analysis. Investigators should apply quality thresholds consistently and should not interpret results from samples that fail quality checks. The temptation to include low-quality samples because they are the only available data should be resisted.
Overinterpretation of Phylogenetic Trees
A phylogenetic tree is a hypothesis about evolutionary relationships, not a definitive record of transmission events. Investigators should avoid overinterpreting the tree topology, particularly when the number of sequenced samples is small or the genetic diversity is low. The tree should be interpreted alongside epidemiological data, not in isolation.
Delayed Data Sharing
Sequence data that are not shared promptly cannot contribute to multi-jurisdictional investigations. Investigators should deposit data in public repositories as soon as quality checks are complete, following the data sharing policies of their institution and funding agency. Delayed sharing can allow outbreaks to spread across jurisdictions before they are detected.
Lack of Reproducibility
Bioinformatics analyses that cannot be reproduced are of limited value. Investigators should document the software versions, parameters, and reference data used in the analysis. The documentation should be sufficient for another analyst to reproduce the results.
Limitations and Interpretation Boundaries
Pathogen Biology
The resolution of genomic epidemiology depends on the pathogen's mutation rate and population structure. For pathogens with low genetic diversity, such as Mycobacterium tuberculosis, even unrelated cases may have identical genomes. For pathogens with high recombination rates, such as some RNA viruses, the phylogenetic signal may be obscured by recombination events. Investigators should understand the biology of the pathogen before interpreting genomic data.
Sampling Completeness
Genomic epidemiology can only describe the relationships among the organisms that were sequenced. If the sampling is incomplete, the analysis may miss transmission links or suggest links that do not exist. The sampling fraction should be reported alongside the results.
Temporal Resolution
The mutation rate of the pathogen determines the temporal resolution of the analysis. For slowly mutating pathogens, the genomic data may not be able to distinguish between cases that occurred weeks apart. For rapidly mutating pathogens, the data may be able to estimate the timing of transmission events. The temporal resolution should be stated in the report.
Bioinformatics Uncertainty
Variant calling and phylogenetic inference involve statistical uncertainty. The uncertainty should be quantified and reported. Confidence intervals on branch lengths and support values for tree nodes should be included in the report. Investigators should not present the phylogenetic tree as a definitive result without acknowledging the uncertainty.
Safety and Regulatory Context
Specimen Handling
Specimens collected for genomic epidemiology may contain infectious agents. Laboratory personnel should follow biosafety protocols appropriate for the pathogen and specimen type. The protocols should be established before the investigation begins and should be reviewed by the institutional biosafety committee if required.
Data Privacy
Genomic data from human clinical specimens may be subject to privacy regulations. Investigators should follow the data governance requirements of their institution and jurisdiction. The requirements may include de-identification, secure storage, and restrictions on data sharing.
Regulatory Reporting
Some pathogens are subject to mandatory reporting requirements. Investigators should be aware of the reporting requirements for the pathogen and jurisdiction involved. The genomic findings may trigger additional reporting obligations, such as notification of public health authorities about novel resistance mechanisms or unusual transmission patterns.
Professional Escalation Criteria
Investigators should escalate to senior public health authorities or specialized reference laboratories when the following conditions are present. The outbreak is large or rapidly growing. The pathogen has unusual resistance or virulence characteristics. The genomic analysis suggests transmission across jurisdictional boundaries. The investigation requires specialized analytical methods that are not available locally. The findings have implications for public health policy or clinical practice.
The readiness assessment for genomics-enabled antimicrobial resistance surveillance in the ASEAN region identified potential impacts of genomics-enabled surveillance including faster detection of outbreaks, timely interventions, and a better understanding of national trends and inter-hospital transmission [16]. These impacts are realized when the genomic findings are acted upon by public health authorities. Investigators should ensure that their findings reach the decision makers who can implement control measures.
Training and Capacity Building
The integration of genomic epidemiology into outbreak investigations requires trained personnel. A scoping review of existing training programmes in genomics and bioinformatics for pathogen surveillance in Africa found that training programmes were predominantly short-term, with 60.9% covering both genomics and bioinformatics [17]. Key outcomes included enhanced technical skills and career development [17]. A structured survey of pathogen genomics training initiatives identified 81 courses from 17 countries, with over half targeting academic or research audiences and 46% targeting public health professionals [19]. Beginner-level courses accounted for 58% of offerings, while only 6% were classified as advanced [19]. Bioinformatics or genomic data analysis was widely represented, while specialized areas such as biostatistics and systems administration were less frequently included [19].
The training landscape shows that most courses are short and beginner-level, which may not be sufficient for the advanced analytical skills needed for complex outbreak investigations. Investigators should seek ongoing training and mentorship, and institutions should invest in workforce development. The European Bioinformatics Institute offers training resources for bioinformatics [1]. The National Center for Biotechnology Information provides data resources and documentation [2].
Reporting Genomic Epidemiology Findings
Structure of the Report
The report of a genomic epidemiology investigation should include the following sections. The outbreak background and the questions addressed. The sampling strategy and the specimens sequenced. The laboratory methods and sequencing platforms. The bioinformatics pipeline and quality metrics. The phylogenetic analysis and its interpretation. The limitations of the analysis. The recommended actions and the rationale.
Visualizing Genomic and Epidemiological Data
The combined visualization of genomic and epidemiological data is important for communicating findings to diverse audiences [22]. A phylogenetic tree with epidemiological metadata overlaid can show the relationship between genetic clusters and case characteristics such as location, time, and exposure. The visualization should be designed to be interpretable by non-specialists, including public health officials and clinicians.
Timeliness of Reporting
The report should be produced as quickly as possible while maintaining quality. Preliminary results can be shared with the investigation team before the full analysis is complete, with clear caveats about the provisional nature of the findings. The report should be updated as new data become available. The on-demand, hospital-based genomic epidemiology approach for SARS-CoV-2 nosocomial outbreak investigations demonstrates the value of rapid reporting during an active outbreak [20]. A systematic review of local genomic epidemiology investigations of SARS-CoV-2 during the early pandemic response provides additional context on the range of approaches used [21].
Frequently Asked Questions
What is the difference between genomic epidemiology and traditional molecular epidemiology?
Traditional molecular epidemiology uses methods such as pulsed-field gel electrophoresis, multilocus sequence typing, or PCR-based typing to compare organisms. Genomic epidemiology uses whole-genome sequence data, which provides much higher resolution. Whole-genome sequencing can distinguish between organisms that appear identical by traditional methods, and it can identify the specific genetic changes that distinguish outbreak strains from background strains [6]. The higher resolution allows investigators to identify transmission chains that would be missed by traditional methods.
How many specimens should be sequenced during an outbreak investigation?
The number depends on the outbreak size, the pathogen, and the questions being asked. There is no universal threshold. The sampling strategy should be defined before sequencing begins and should be guided by the outbreak investigation's needs. Sequencing every case may be unnecessary for a point source outbreak with a clear exposure. Sequencing only a few cases may be insufficient to distinguish between a point source and ongoing transmission. The sampling fraction should be documented and reported.
What is the role of long-read sequencing in outbreak investigations?
Long-read sequencing provides longer reads that can resolve plasmids, repetitive regions, and structural variation. This is particularly valuable when mobile genetic elements are involved, such as in outbreaks of plasmid-mediated antimicrobial resistance. In the NDM-5 outbreak investigation, long-read sequencing and hybrid assemblies were used to resolve the NDM-encoding plasmids [10]. Short-read platforms provide high accuracy for SNP detection but may not fully resolve repetitive regions.
How should phylogenetic trees be interpreted during an outbreak?
A phylogenetic tree shows genetic relatedness, not transmission events. Two cases with identical sequences may or may not be linked by direct transmission. The tree should be interpreted alongside epidemiological data, including case interviews, exposure histories, temporal data, and geographic information. The interpretation should acknowledge the uncertainty in the analysis, including confidence intervals on branch lengths and support values for tree nodes.
What are the main barriers to implementing genomic epidemiology in low-resource settings?
Common barriers include limited system interoperability and data integration, long turnaround times, and limited governance frameworks [16]. The survey of next-generation sequencing and bioinformatics capacity found that only 57.7% of respondents stored data in multiple locations, and 32.5% lacked any data backup [18]. Training and workforce development are also significant barriers, with most available courses being short and beginner-level [19].
How should genomic data be shared during an outbreak?
Sequence data should be deposited in public repositories such as the National Center for Biotechnology Information [2] in a timely manner, following the data sharing policies of the funding agency and institution. The National Institutes of Health Genomic Data Sharing Policy sets expectations for data sharing [3]. National platforms such as AusTrakka can facilitate data sharing across jurisdictions [15]. The FAIR Guiding Principles describe the characteristics that data resources should have to be findable, accessible, interoperable, and reusable [4].
What quality checks should be applied to sequence data before interpretation?
Quality checks should include assessment of read quality, read depth, genome coverage, and variant quality. Thresholds should be defined in advance and applied consistently across all samples. Samples that fail quality checks should not be interpreted. The bioinformatics pipeline should be documented, including software versions and parameters, so that the analysis can be reproduced.
When should genomic findings be escalated to senior public health authorities?
Escalation is appropriate when the outbreak is large or rapidly growing, when the pathogen has unusual resistance or virulence characteristics, when the genomic analysis suggests transmission across jurisdictional boundaries, when specialized analytical methods are needed, or when the findings have implications for public health policy or clinical practice. The readiness assessment for genomics-enabled surveillance in the ASEAN region identified potential impacts including faster detection of outbreaks and timely interventions [16].
Related Bioinformatics Guides
- Structural and Evolutionary Dynamics of Coronavirus Spike Protein: Integrating Cryo-EM, Molecular Dynamics, and Phylogenetic Surveillance
- Predicting Cross-Species Viral Spillover: Integrating Structural Modeling, Receptor Binding Dynamics, and Genomic Surveillance
- Ethical Considerations in Computational Genomics
- Deep Learning for Functional Genomics
- Genomic Selection in Animal Breeding
References and Further Reading
- EMBL-EBI Training. European Bioinformatics Institute.
- NCBI Data Resources. National Center for Biotechnology Information.
- Genomic Data Sharing Policy. National Institutes of Health.
- The FAIR Guiding Principles. Scientific Data.
- Genomic Epidemiology, Evolution, and Transmission Dynamics of Porcine Deltacoronavirus.. Molecular biology and evolution, 2020.
- Genomic Epidemiology: Whole-Genome-Sequencing-Powered Surveillance and Outbreak Investigation of Foodborne Bacterial Pathogens.. Annual review of food science and technology, 2016.
- Genomic epidemiology of a Bacillus cereus bacteraemia outbreak linked to contaminated hospital laundry.. Microbial genomics, 2025.
- Genomic epidemiology of mpox virus during the 2022 outbreak in New York City.. Nature communications, 2025.
- Genomic Epidemiology of 2015-2016 Zika Virus Outbreak in Cape Verde.. Emerging infectious diseases, 2020.
- Real-time genomic epidemiologic investigation of a multispecies plasmid-associated hospital outbreak of NDM-5-producing Enterobacterales infections.. International journal of infectious diseases : IJID : official publication of the International Society for Infectious Diseases, 2024.
- Genomic epidemiological investigation of a Streptococcus suis outbreak in Guangxi, China, 2016.. Infection, genetics and evolution : journal of molecular epidemiology and evolutionary genetics in infectious diseases, 2019.
- Genomic investigation of a dengue virus outbreak in Thiès, Senegal, in 2018.. Scientific reports, 2021.
- Pathogen to read: a rapid, automated workflow for public health-ready microbial sequencing.. 2026.
- Strategies for integrating whole-genome sequencing into antimicrobial resistance surveillance.. 2026.
- High acceptability of a national platform for public health genomic data sharing and surveillance in Australia: a mixed methods study.. 2026.
- Pathogen genomics to inform control of antimicrobial resistant pathogens: an assessment of readiness and potential impact in the ASEAN region. 2026.
- Harnessing genomic and bioinformatics for surveillance of pathogens in Africa: a scoping review of existing training and gaps in training.. 2026.
- Next-generation sequencing and bioinformatics capacity: findings from a multi-country survey to guide the genomics costing tool 2.0.. 2026.
- Mapping pathogen genomics training provision: a structured analysis within a global consortium network.. 2026.
- On-demand, hospital-based, severe acute respiratory coronavirus virus 2 (SARS-CoV-2) genomic epidemiology to support nosocomial outbreak investigations: A prospective molecular epidemiology study. Antimicrobial Stewardship and Healthcare Epidemiology, 2023.
- Local genomic epidemiology investigations of SARS-CoV-2 during the early pandemic response: A global systematic review. Epidemiology and Infection, 2026.
- Combined visualization of genomic and epidemiological data for outbreaks. Epidemiology and Infection, 2024.
- Implementation and outcomes of a rapid response genomic hospital epidemiology programme at an academic medical centre over 7 years. Lancet Microbe, 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.