# Graph-Based Variant Calling: How to Improve Accuracy in Complex Genomic Regions with Pan-Genome Graphs


## Key Takeaways

- Graph-based variant calling addresses the "reference bias" inherent in linear reference genomes, which systematically underrepresents or misidentifies variants in complex genomic regions like segmental duplications and repetitive sequences. By representing multiple haplotypes and alternative alleles in a pan-genome graph, reads can align to alternative paths, improving accuracy for alleles divergent from a single consensus.

- Pan-genome graphs are constructed by aligning multiple genome assemblies or variant data, collapsing shared sequences into nodes and divergent sequences into edges, thereby encoding a spectrum of genetic variation. This data structure allows for the integrated representation of individual variation, moving beyond the limitations of a single linear reference.

- Tools like `vg` provide a framework for building, aligning to, and calling variants from graph-based references, extending beyond traditional variant calling to improve analyses like peak detection in functional genomics assays (e.g., ChIP-seq). `Paragraph` is specifically designed for accurate structural variant genotyping from short-read data by modeling variants using sequence graphs.

- While graph-based methods offer superior accuracy in complex genomic regions, they generally require more computational resources (memory and runtime) than linear approaches. Algorithmic optimizations, such as cache-friendly de Bruijn graph reassembly, are improving efficiency, but researchers must still consider these tradeoffs for their specific workflows.

- For pathogen surveillance without high-quality reference genomes, reference-free methods like `ska lo` traverse colored De Bruijn graphs to identify within-strain variants, achieving high sensitivity for SNPs and indels without the need for a pre-built graph. This approach is particularly valuable for diverse or poorly characterized organisms.

- Integrating graph-based variant calling into clinical applications requires careful validation and interpretation, as these methods can identify clinically relevant variants missed by linear callers. Tools like StratoMod aim to predict such missed variants, aiding in risk-reward analyses for pipeline design and clinical reporting.

---

Variant calling from short-read sequencing data traditionally relies on aligning reads to a single linear reference genome, but this approach systematically misses or misrepresents variation in complex genomic regions. Graph-based variant calling addresses this limitation by representing multiple sequences simultaneously in a pan-genome graph, allowing reads to align to alternative paths and improving variant detection accuracy. This article explains when and how to use graph-based methods, compares them with linear callers, and provides practical guidance for integrating them into bioinformatics workflows for germline and somatic variant calling, variant filtering, and clinical applications.

## The Reference Bias Problem in Linear Variant Calling

Linear reference genomes represent a single consensus sequence that cannot capture the full spectrum of genetic diversity within a species. When sequencing reads contain alleles that differ substantially from the reference sequence, alignment algorithms struggle to place them correctly, leading to reference bias. This bias manifests as reduced mapping quality, incorrect base placement, and ultimately missed or falsely called variants.

The consequences of reference bias are most severe in complex genomic regions. These include segmental duplications, regions with high sequence homology, structural variant breakpoints, and areas containing novel insertion sequences not present in the reference. In these contexts, reads originating from alternative alleles may fail to map or map with low confidence, causing variant callers to overlook true biological variation.

Graph-based reference genomes offer a solution by encoding multiple haplotypes and alternative alleles within a single data structure. Instead of forcing all reads to align against one linear sequence, graph aligners can match reads against multiple paths simultaneously. This approach reduces reference bias because reads carrying alternative alleles find their correct genomic context within the graph.

The practical implication for researchers is straightforward. If your study focuses on well-characterized regions of the genome with low diversity, linear callers may perform adequately. However, if your work involves structural variants, repetitive elements, immune genes, or populations with high genetic diversity, graph-based methods can recover variants that linear approaches miss.

## What Is a Pan-Genome Graph

A pan-genome graph is a data structure that represents the collective genetic variation of a population or species. Unlike a linear reference that stores one sequence, a graph stores multiple sequences as nodes and edges. Nodes represent conserved sequence segments, while edges connect alternative paths that represent different alleles or haplotypes.

Graph-based representations are considered the future direction for reference genomes because they allow integrated representation of the steadily increasing data on individual variation. Currently available tools allow de novo assembly of graph-based reference genomes, alignment of new read sets to the graph representation, and certain analyses like variant calling and haplotyping [<a href="#ref-1">1</a>].

The construction of a pan-genome graph begins with collecting genome assemblies or variant information from multiple individuals. These sequences are aligned to each other, and the resulting alignments are used to build a graph where shared sequences collapse into common nodes and divergent sequences branch into alternative paths. The resulting graph can represent single-nucleotide polymorphisms, insertions and deletions, and structural variants in a unified framework.

For researchers, the key conceptual shift is understanding that a graph reference is not a single genome but a collection of possible genomes. When you align reads to a graph, you are asking which combination of paths best explains the observed sequencing data. This flexibility is what enables improved variant calling in complex regions.

## Graph-Based Variant Calling Tools and Their Applications

Several tools have been developed to perform variant calling using graph-based references. Each tool approaches the problem differently, and understanding these differences helps researchers select the appropriate method for their specific application.

### vg and General-Purpose Graph Alignment

The vg toolkit provides a comprehensive framework for building, manipulating, and aligning to graph-based reference genomes. It includes tools for graph construction, read alignment, variant calling, and haplotyping. The toolkit has been used to build pan-genome graphs for multiple species and supports both short and long read alignment.

The Graph Peak Caller, built on top of vg, demonstrates the utility of graph-based approaches beyond traditional variant calling. This tool extends the MACS2 peak calling algorithm to work with read data aligned to a graph-based reference genome. Validation using a pan-genome of Arabidopsis thaliana showed that the graph-based approach can trace variants within peaks that are not part of the linear reference genome and find peaks that are generally more motif-enriched than those found by MACS2 [<a href="#ref-1">1</a>].

This example illustrates a broader principle. Graph-based methods are not limited to variant calling but can improve any genomic analysis that depends on accurate read placement. If your workflow involves ChIP-seq, ATAC-seq, or other functional genomics assays, graph-based alignment may improve peak detection in regions where the linear reference is inadequate.

### Paragraph for Structural Variant Genotyping

Paragraph is a graph-based genotyper specifically designed for structural variants. It models structural variants using sequence graphs and variant annotations, allowing accurate genotyping from short-read data. In validation studies using whole-genome sequence data from three samples with long-read structural variant calls as the truth set, Paragraph demonstrated better accuracy than other existing genotypers. The tool was then applied at scale to a cohort of 100 short-read sequenced samples of diverse ancestry, showing that it can be used for population-scale studies [<a href="#ref-2">2</a>].

The practical value of Paragraph lies in its ability to genotype known structural variants accurately. If you have a set of structural variants identified through long-read sequencing or population studies, Paragraph can determine which alleles are present in your short-read samples. This is particularly useful for clinical applications where specific structural variants are associated with disease.

### ska lo for Reference-Free Variant Calling

The ska lo algorithm takes a different approach by performing variant calling without alignment to a reference genome. It traverses a colored De Bruijn graph to identify within-strain variants in pathogen whole-genome sequencing data. The algorithm builds variant groups, which are sets of variant combinations, and achieves high sensitivity in single-nucleotide polymorphism calls while also enabling detection of insertions and deletions, as well as positioning of variants on a reference genome for recombination analyses [<a href="#ref-3">3</a>].

This reference-free approach is particularly valuable for pathogen genomic epidemiology, where traditional reference-based alignment can introduce biases and require significant computational resources. The ability to position variants on a reference genome for recombination analyses adds flexibility, allowing researchers to combine reference-free variant detection with reference-based downstream analyses.

For researchers working with pathogens or other organisms without high-quality reference genomes, ska lo offers a practical alternative. The tool is freely available as part of the SKA program, making it accessible for public health surveillance applications [<a href="#ref-3">3</a>].

## Comparing Graph-Based and Linear Variant Callers

The choice between graph-based and linear variant callers depends on your specific research questions, data characteristics, and computational resources. No single pipeline is optimal across the entire human genome, so researchers need to make tradeoffs when designing pipelines for their application [<a href="#ref-4">4</a>].

### Accuracy in Complex Regions

Graph-based methods excel in regions where linear references fail. These include difficult-to-map regions, homopolymer stretches, and areas with high structural variation. The StratoMod study demonstrated that graph-based references improve recall in hard-to-map regions compared to linear references, identifying specific regions where graph-based methods excelled and quantifying the improvement [<a href="#ref-4">4</a>].

The practical implication is that graph-based calling can recover clinically relevant variants that linear pipelines miss. StratoMod presents a method of predicting clinically relevant variants likely to be missed, which is an improvement over current pipelines that only filter variants likely to be false. This capability supports precise risk-reward analyses when designing variant calling pipelines [<a href="#ref-4">4</a>].

### Computational Cost

Graph-based methods generally require more computational resources than linear approaches. Building and storing graphs consumes memory, and aligning reads to graphs is more complex than linear alignment. However, optimization efforts have improved efficiency. Cache-friendly optimization of de Bruijn graph-based local reassembly in variant calling has produced algorithms up to twice as fast as original implementations on commodity computers, with potentially greater speedups on systems with less complex memory subsystems [<a href="#ref-5">5</a>].

For practical workflows, this means that graph-based calling is feasible on standard computing infrastructure but may require longer runtimes or additional memory compared to linear callers. Researchers should benchmark both approaches on their specific data to understand the computational tradeoffs.

### Workflow Integration

Graph-based tools can be integrated into existing variant calling workflows at multiple points. Some researchers use graph-based alignment followed by linear variant calling, while others use graph-based callers throughout. The choice depends on the specific tools and the nature of the research question.

For germline variant calling, graph-based references can improve single-nucleotide polymorphism and insertion-deletion detection in diverse populations. For somatic variant calling, graph-based methods may help identify variants in repetitive or complex regions that are relevant to cancer genomics. The key is to understand the strengths and limitations of each approach and select the appropriate tool for your specific application.

## At a Glance: Choosing Between Linear and Graph-Based Variant Calling

| Scenario | Linear Reference Calling | Graph-Based Calling | Recommendation |
|----------|-------------------------|---------------------|----------------|
| Well-characterized genome regions with low diversity | High accuracy, low computational cost | Similar accuracy, higher computational cost | Use linear calling for efficiency |
| Diverse populations or underrepresented ancestry | Misses variants due to reference bias | Recovers variants in hard-to-map regions | Use graph-based methods for comprehensive detection |
| Structural variant genotyping | Limited accuracy for complex structural variants | Improved accuracy using sequence graphs | Use Paragraph or similar graph-based genotypers |
| Pathogen surveillance without high-quality reference | Reference bias and computational burden | Reference-free k-mer approaches like ska lo | Use reference-free methods for diverse pathogens |
| Clinical variant detection | May miss clinically relevant variants | Predicts and recovers missed variants | Use graph-based methods with interpretable error prediction |

## Practical Workflow for Graph-Based Variant Calling

Implementing graph-based variant calling requires careful planning and execution. The following workflow outlines the key steps and considerations for researchers who want to incorporate graph-based methods into their analysis pipelines.

### Step 1: Define Your Research Question and Genomic Regions of Interest

Before selecting tools, clarify what types of variants you need to detect and which genomic regions are most important for your study. If your research focuses on structural variants, Paragraph may be the most appropriate tool. If you need comprehensive variant detection across the genome, a general-purpose graph aligner like vg combined with a variant caller may be more suitable.

Consider the diversity of your study population. If you are working with samples from populations that are underrepresented in reference genomes, graph-based methods may provide greater benefits. Similarly, if your regions of interest include known complex loci such as the major histocompatibility complex region, killer-cell immunoglobulin-like receptor genes, or other highly polymorphic areas, graph-based approaches are likely to improve accuracy.

### Step 2: Select or Construct a Pan-Genome Graph

You have two options for obtaining a pan-genome graph. You can use a pre-built graph if one exists for your species and population of interest, or you can construct your own graph from genome assemblies or variant data.

Constructing a graph requires collecting high-quality genome assemblies from multiple individuals representing the diversity of your study population. The vg toolkit provides tools for graph construction, but this process requires substantial computational resources and expertise. Alternatively, you can build graphs from existing variant databases, which may be simpler but may not capture all relevant variation.

For pathogen studies, reference-free approaches like ska lo may be more appropriate, as they do not require a reference genome or graph construction. These methods build graphs directly from the sequencing data, which can be advantageous when working with diverse or poorly characterized organisms [<a href="#ref-3">3</a>].

### Step 3: Align Reads to the Graph

Read alignment to a graph reference differs from linear alignment. Graph aligners must consider multiple possible paths when placing reads, which increases computational complexity. The vg toolkit provides the map command for this purpose, which can handle both short and long reads.

When aligning reads to a graph, pay attention to mapping quality metrics. Reads that align uniquely to a single path have high mapping quality, while reads that align to multiple paths with similar scores have lower mapping quality. These metrics are important for downstream variant calling and filtering.

### Step 4: Call Variants

After alignment, you can call variants using graph-aware variant callers or adapt existing callers to work with graph alignments. Some tools, like Paragraph, perform genotyping directly from graph alignments. Others require additional processing steps.

For structural variant genotyping, Paragraph accepts aligned reads and structural variant annotations as input and produces genotype calls for each variant. The tool models structural variants using sequence graphs, which allows it to accurately determine which alleles are present in each sample [<a href="#ref-2">2</a>].

For single-nucleotide polymorphism and insertion-deletion calling, you may need to convert graph alignments to a format compatible with existing callers or use graph-specific calling tools. The choice depends on your specific workflow and the tools you are comfortable using.

### Step 5: Filter and Annotate Variants

Variant filtering is essential for removing false positives and identifying high-confidence calls. Filtering criteria may include read depth, mapping quality, variant quality scores, and allele balance. The specific thresholds depend on your sequencing platform, coverage, and research question.

For clinically relevant variants, consider using tools like StratoMod that predict which variants are likely to be missed by your pipeline. This information can guide additional validation efforts and help you understand the limitations of your approach [<a href="#ref-4">4</a>].

Annotation of variants provides biological context and helps prioritize candidates for further analysis. Standard annotation tools can be used with graph-based variant calls, as the output formats are generally compatible with existing downstream tools.

## Data Inputs and Quality Considerations

The quality of graph-based variant calling depends heavily on the quality of input data. Understanding the requirements for each data type helps researchers design experiments and troubleshoot problems.

### Sequencing Data Requirements

Short-read sequencing data for graph-based calling should meet the same quality standards as data for linear calling. Coverage depth, read length, and base quality scores all influence variant calling accuracy. Higher coverage provides more evidence for variant calls, while longer reads can span repetitive regions more effectively.

For structural variant genotyping with Paragraph, the tool was validated on whole-genome sequence data and applied to a cohort of 100 samples of diverse ancestry. This demonstrates that the approach works with standard whole-genome sequencing data, but the specific coverage and quality requirements may vary depending on the complexity of the variants being genotyped [<a href="#ref-2">2</a>].

### Reference and Graph Quality

The quality of your pan-genome graph directly affects variant calling accuracy. Graphs built from high-quality genome assemblies will represent variation more accurately than graphs built from lower-quality data. When constructing your own graphs, use assemblies that have been validated and represent the diversity of your study population.

For pathogen studies using reference-free approaches, the quality of the sequencing data is even more critical, as the graph is built directly from the reads. High error rates in sequencing data can introduce spurious branches in the De Bruijn graph, leading to false variant calls [<a href="#ref-3">3</a>].

### Computational Infrastructure

Graph-based variant calling requires more computational resources than linear approaches. Memory requirements for storing graphs can be substantial, particularly for large genomes or graphs representing many individuals. Alignment to graphs is also more computationally intensive than linear alignment.

Cache-friendly optimization of graph-based algorithms has improved efficiency, with some implementations running up to twice as fast as original versions on commodity computers [<a href="#ref-5">5</a>]. However, researchers should still plan for increased runtime and memory usage when using graph-based methods.

## Records and Measurements for Quality Control

Maintaining detailed records of your variant calling workflow is essential for reproducibility and quality control. The following measurements should be documented for each analysis.

### Alignment Statistics

Record the percentage of reads that align to the graph reference, the distribution of mapping qualities, and the number of reads that align to multiple paths. These statistics provide insight into the quality of your graph and the complexity of your data.

For graph alignments, also record the distribution of reads across different paths in the graph. If most reads align to a single dominant path, your graph may not be capturing sufficient diversity. Conversely, if reads are distributed across many paths, your graph may be overfit to specific individuals.

### Variant Calling Metrics

Document the number of variants called, the transition-to-transversion ratio, and the distribution of variant quality scores. These metrics help identify systematic errors in variant calling. For example, an unusually high number of variants in repetitive regions may indicate alignment errors.

For structural variant genotyping, record the genotype concordance rates if you have validation data. Paragraph demonstrated better accuracy than other existing genotypers in validation studies, but your specific results will depend on your data and variants of interest [<a href="#ref-2">2</a>].

### Computational Performance

Track runtime and memory usage for each step of your workflow. This information helps you plan future analyses and identify bottlenecks. If runtime becomes prohibitive, consider optimizing your workflow or using more efficient implementations.

The cache-friendly optimization of de Bruijn graph-based local reassembly demonstrates that algorithmic improvements can significantly reduce runtime without impacting accuracy [<a href="#ref-5">5</a>]. Staying current with tool updates and optimization efforts can help you maintain efficient workflows.

## Common Failure Patterns in Graph-Based Variant Calling

Understanding common failure modes helps researchers troubleshoot problems and interpret results appropriately.

### Graph Construction Errors

Errors in graph construction can propagate through the entire variant calling process. Misassembled sequences, incorrect variant representation, and missing haplotypes all reduce the accuracy of downstream analyses. If your graph contains errors, reads may align incorrectly, leading to false variant calls.

To identify graph construction errors, examine alignment statistics and variant calls in known well-characterized regions. If you observe unexpected variants in regions that are typically conserved, your graph may contain errors.

### Overfitting to Graph Construction Samples

Pan-genome graphs built from a limited set of individuals may overfit to the variation present in those individuals. Samples that carry alleles not represented in the graph may still experience reference bias, reducing the benefits of graph-based calling.

To mitigate this issue, construct graphs from diverse individuals that represent the range of variation in your study population. For human studies, include individuals from multiple ancestral backgrounds to capture global genetic diversity.

### Computational Resource Exhaustion

Graph-based methods can exhaust available memory or exceed runtime limits, particularly for large genomes or complex graphs. If your analysis fails due to resource limitations, consider using more efficient tools, reducing the complexity of your graph, or increasing your computational resources.

The cache-friendly optimization of graph algorithms shows that careful attention to memory access patterns can substantially improve performance [<a href="#ref-5">5</a>]. When selecting tools, consider their computational efficiency and whether they have been optimized for your specific use case.

### Misinterpretation of Graph-Based Results

Graph-based variant calling produces results that may differ from linear approaches, and researchers may misinterpret these differences. Variants that are called only by graph-based methods are not necessarily false positives. They may represent true variation that linear methods miss due to reference bias.

Conversely, variants that are called only by linear methods may be artifacts of misalignment. When comparing results between approaches, validate unexpected calls using orthogonal methods such as PCR, Sanger sequencing, or long-read sequencing.

## Limitations of Graph-Based Variant Calling

While graph-based methods offer significant advantages, they also have limitations that researchers should understand.

### Complexity and Learning Curve

Graph-based variant calling requires understanding new data structures, tools, and concepts. The learning curve can be steep, particularly for researchers trained primarily in linear reference-based methods. Training resources from EMBL-EBI Training can help researchers develop the necessary skills for bioinformatics data analysis [<a href="#ref-6">6</a>].

The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers learn graph-based methods in a practical context [<a href="#ref-7">7</a>]. Similarly, nf-core documentation describes community pipeline standards that can guide workflow implementation [<a href="#ref-8">8</a>]. The Carpentries offers foundational computing, data, shell, Git, and programming training that supports bioinformatics skill development [<a href="#ref-9">9</a>].

### Computational Requirements

The increased computational requirements of graph-based methods may be prohibitive for some research groups. Building and storing graphs requires substantial memory, and alignment to graphs is more computationally intensive than linear alignment. Researchers with limited computational resources may need to prioritize which analyses benefit most from graph-based approaches.

### Reference-Free Methods Have Different Limitations

Reference-free approaches like ska lo avoid some problems of reference-based methods but introduce their own limitations. These methods may have sensitivity limitations in high mutation density regions, and the interpretation of results can be more challenging without a reference context [<a href="#ref-3">3</a>].

The ska lo algorithm was designed to address sensitivity limitations of existing k-mer-based methods, achieving high sensitivity in single-nucleotide polymorphism calls while enabling detection of insertions and deletions [<a href="#ref-3">3</a>]. However, researchers should validate results from reference-free methods using complementary approaches when possible.

## Safety and Regulatory Context for Clinical Applications

For researchers working in clinical or diagnostic settings, understanding the regulatory context of variant calling is essential. Graph-based methods may identify variants that linear methods miss, which has implications for clinical reporting.

### Variant Interpretation Challenges

Clinically relevant variants identified only through graph-based methods may not be present in standard variant databases, complicating interpretation. The StratoMod approach of predicting clinically relevant variants likely to be missed represents an improvement over current pipelines that only filter variants likely to be false [<a href="#ref-4">4</a>].

When reporting variants identified through graph-based methods, clearly document the methods used and the evidence supporting each call. This transparency supports appropriate clinical interpretation and follow-up testing.

### Validation Requirements

Clinical variant calling pipelines require validation to demonstrate accuracy and reliability. Validation should include comparison with orthogonal methods and assessment of performance in challenging genomic regions. The improved accuracy of graph-based methods in complex regions may justify their use in clinical settings, but validation data specific to your workflow is essential.

### Professional Escalation Criteria

When graph-based variant calling identifies variants with potential clinical significance, follow established protocols for confirmation and reporting. If you are uncertain about the interpretation of a variant, consult with colleagues who have expertise in the relevant genomic region or clinical condition.

For research applications, document any unexpected findings and consider whether they warrant further investigation. Variants identified only through graph-based methods may represent genuine biological variation that was previously inaccessible.

## Reproducibility and Workflow Documentation

Reproducibility is a core requirement for bioinformatics analyses, and graph-based variant calling introduces additional complexity that must be documented carefully.

### Version Control and Environment Management

Record the exact versions of all tools used in your workflow, including graph construction tools, aligners, and variant callers. Tool versions can significantly affect results, and version changes may explain differences between analyses.

Use containerization or environment management tools to ensure that your analysis environment is reproducible. The nf-core documentation describes community pipeline standards that emphasize reproducibility and configuration management [<a href="#ref-8">8</a>]. Bioconductor provides documentation for reproducible genomic-analysis workflows that can be adapted for graph-based methods [<a href="#ref-10">10</a>].

### Data Management

Document the source and version of all reference data, including genome assemblies used for graph construction and any variant databases used for annotation. Maintain clear records of file paths, checksums, and processing dates.

For large-scale studies, consider using workflow management systems that track data provenance automatically. These systems can help ensure that analyses are reproducible and that results can be traced back to specific inputs and parameters.

### Reporting Standards

When reporting results from graph-based variant calling, clearly describe the methods used, including the graph construction approach, alignment parameters, and variant calling criteria. This information allows other researchers to evaluate your results and reproduce your analysis.

For publications, consider depositing your analysis code and configuration files in a public repository. This practice supports the broader goal of reproducible research and allows others to build on your work.

## A Decision Framework for Selecting Graph-Based Variant Calling Approaches

Selecting the right graph-based variant calling strategy requires a structured evaluation of your specific research context, data characteristics, and biological questions. The following framework provides a practical method for matching your project requirements to the appropriate graph-based approach, based on evidence from validated implementations.

### Step 1: Assess Your Genomic Complexity Profile

Begin by cataloging the genomic regions that matter most for your study. Create a table listing each region of interest, its known complexity features, and the variant types you need to detect. Complexity features include segmental duplications, homopolymer stretches, structural variant breakpoints, and regions with high sequence homology. The StratoMod study demonstrated that no single pipeline is optimal across the entire human genome, so researchers must make tradeoffs when designing pipelines for their application [<a href="#ref-4">4</a>].

For each region, classify the expected variant types. Single-nucleotide polymorphisms and small insertions or deletions may be adequately handled by linear callers in simple regions, but structural variants and complex rearrangements require graph-based approaches. The Graph Peak Caller study showed that graph-based representations allow integrated representation of the steadily increasing data on individual variation, with tools available for de novo assembly of graph-based reference genomes, alignment of new read sets to the graph representation, and analyses like variant calling and haplotyping [<a href="#ref-1">1</a>].

### Step 2: Evaluate Your Reference Data Options

Determine whether a suitable pan-genome graph already exists for your species and population. Pre-built graphs save substantial time and computational resources compared to constructing your own. If no appropriate graph exists, you must decide between constructing a graph from genome assemblies or using a reference-free approach.

For pathogen studies, reference-free approaches like ska lo may be more appropriate, as they do not require a reference genome or graph construction. The ska lo algorithm traverses a colored De Bruijn graph to identify within-strain variants in pathogen whole-genome sequencing data, building variant groups that represent sets of variant combinations [<a href="#ref-3">3</a>]. This approach is particularly valuable when working with diverse or poorly characterized organisms.

When constructing your own graph, collect high-quality genome assemblies from multiple individuals representing the diversity of your study population. The vg toolkit provides tools for graph construction, but this process requires substantial computational resources and expertise. Alternatively, you can build graphs from existing variant databases, which may be simpler but may not capture all relevant variation.

### Step 3: Match Tools to Your Variant Types and Scale

Different graph-based tools excel at different tasks. For structural variant genotyping, Paragraph models structural variants using sequence graphs and variant annotations, allowing accurate genotyping from short-read data. Validation studies using whole-genome sequence data from three samples with long-read structural variant calls as the truth set demonstrated that Paragraph has better accuracy than other existing genotypers, and the tool was applied at scale to a cohort of 100 short-read sequenced samples of diverse ancestry [<a href="#ref-2">2</a>].

For comprehensive variant detection across the genome, general-purpose graph aligners like vg combined with variant callers may be more suitable. The Graph Peak Caller, built on top of vg, extends the MACS2 peak calling algorithm to work with read data aligned to a graph-based reference genome. Validation using a pan-genome of Arabidopsis thaliana showed that the graph-based approach can trace variants within peaks that are not part of the linear reference genome and find peaks that are generally more motif-enriched than those found by MACS2 [<a href="#ref-1">1</a>].

For pathogen genomic epidemiology, ska lo offers a simple, fast, and effective approach that extends the range of reference-free variant-calling methods. The tool achieves high sensitivity in single-nucleotide polymorphism calls while also enabling detection of insertions and deletions, as well as positioning of variants on a reference genome for recombination analyses [<a href="#ref-3">3</a>].

### Step 4: Calculate Computational Feasibility

Graph-based methods generally require more computational resources than linear approaches. Building and storing graphs consumes memory, and aligning reads to graphs is more complex than linear alignment. However, optimization efforts have improved efficiency. Cache-friendly optimization of de Bruijn graph-based local reassembly in variant calling has produced algorithms up to twice as fast as original implementations on commodity computers, with potentially greater speedups on systems with less complex memory subsystems [<a href="#ref-5">5</a>].

Estimate your memory and runtime requirements before committing to a graph-based approach. Consider the size of your genome, the complexity of your graph, and the number of samples you need to process. For large cohorts, benchmark a small subset of samples first to project total resource requirements.

### Step 5: Plan for Validation and Interpretation

Graph-based variant calling produces results that may differ from linear approaches, and researchers may misinterpret these differences. Variants that are called only by graph-based methods are not necessarily false positives. They may represent true variation that linear methods miss due to reference bias. Conversely, variants that are called only by linear methods may be artifacts of misalignment.

When comparing results between approaches, validate unexpected calls using orthogonal methods such as PCR, Sanger sequencing, or long-read sequencing. For clinically relevant variants, consider using tools like StratoMod that predict which variants are likely to be missed by your pipeline. StratoMod presents a method of predicting clinically relevant variants likely to be missed, which is an improvement over current pipelines that only filter variants likely to be false [<a href="#ref-4">4</a>].

## A Record System for Graph-Based Variant Calling Projects

Maintaining detailed records of your graph-based variant calling workflow is essential for reproducibility, troubleshooting, and quality control. The following record system provides a structured approach to documenting your analyses.

### Project Metadata Records

Create a project-level record that captures the overall context of your analysis. Include the research question, the biological samples being analyzed, the sequencing platform and coverage, and the version of every tool used in the workflow. Record the source and version of all reference data, including genome assemblies used for graph construction and any variant databases used for annotation.

Document the graph construction approach in detail. If you built your own graph, record the assemblies used, the alignment parameters, and the graph building commands. If you used a pre-built graph, record its source, version, and any known limitations. This information is critical for reproducing your analysis and for interpreting results in the context of the graph's composition.

### Alignment Quality Records

For each sample, record the percentage of reads that align to the graph reference, the distribution of mapping qualities, and the number of reads that align to multiple paths. These statistics provide insight into the quality of your graph and the complexity of your data.

For graph alignments, also record the distribution of reads across different paths in the graph. If most reads align to a single dominant path, your graph may not be capturing sufficient diversity. Conversely, if reads are distributed across many paths, your graph may be overfit to specific individuals. These observations can guide decisions about whether to refine your graph or adjust your analysis approach.

### Variant Calling Metrics Records

Document the number of variants called, the transition-to-transversion ratio, and the distribution of variant quality scores. These metrics help identify systematic errors in variant calling. For example, an unusually high number of variants in repetitive regions may indicate alignment errors.

For structural variant genotyping, record the genotype concordance rates if you have validation data. Paragraph demonstrated better accuracy than other existing genotypers in validation studies, but your specific results will depend on your data and variants of interest [<a href="#ref-2">2</a>]. Maintain a comparison log that tracks variants called by graph-based methods versus linear methods, noting which variants are unique to each approach.

### Computational Performance Records

Track runtime and memory usage for each step of your workflow. This information helps you plan future analyses and identify bottlenecks. If runtime becomes prohibitive, consider optimizing your workflow or using more efficient implementations.

The cache-friendly optimization of de Bruijn graph-based local reassembly demonstrates that algorithmic improvements can significantly reduce runtime without impacting accuracy [<a href="#ref-5">5</a>]. Staying current with tool updates and optimization efforts can help you maintain efficient workflows.

## Troubleshooting Common Graph-Based Variant Calling Failures

Understanding common failure modes helps researchers troubleshoot problems and interpret results appropriately. The following troubleshooting guide addresses the most frequently encountered issues.

### Graph Construction Errors

Errors in graph construction can propagate through the entire variant calling process. Misassembled sequences, incorrect variant representation, and missing haplotypes all reduce the accuracy of downstream analyses. If your graph contains errors, reads may align incorrectly, leading to false variant calls.

To identify graph construction errors, examine alignment statistics and variant calls in known well-characterized regions. If you observe unexpected variants in regions that are typically conserved, your graph may contain errors. Compare your graph-based results with linear reference-based results in these regions to identify discrepancies that warrant investigation.

### Overfitting to Graph Construction Samples

Pan-genome graphs built from a limited set of individuals may overfit to the variation present in those individuals. Samples that carry alleles not represented in the graph may still experience reference bias, reducing the benefits of graph-based calling.

To mitigate this issue, construct graphs from diverse individuals that represent the range of variation in your study population. For human studies, include individuals from multiple ancestral backgrounds to capture global genetic diversity. The StratoMod study identified hard-to-map regions where graph-based methods excelled and quantified the improvement, providing guidance on which regions benefit most from diverse graph construction [<a href="#ref-4">4</a>].

### Computational Resource Exhaustion

Graph-based methods can exhaust available memory or exceed runtime limits, particularly for large genomes or complex graphs. If your analysis fails due to resource limitations, consider using more efficient tools, reducing the complexity of your graph, or increasing your computational resources.

The cache-friendly optimization of graph algorithms shows that careful attention to memory access patterns can substantially improve performance [<a href="#ref-5">5</a>]. When selecting tools, consider their computational efficiency and whether they have been optimized for your specific use case. Benchmark tools on a small test dataset before committing to a full-scale analysis.

### Misinterpretation of Graph-Based Results

Graph-based variant calling produces results that may differ from linear approaches, and researchers may misinterpret these differences. Variants that are called only by graph-based methods are not necessarily false positives. They may represent true variation that linear methods miss due to reference bias.

Conversely, variants that are called only by linear methods may be artifacts of misalignment. When comparing results between approaches, validate unexpected calls using orthogonal methods such as PCR, Sanger sequencing, or long-read sequencing. For structural variants, validation may require specialized approaches such as optical mapping or targeted long-read sequencing.

### Sensitivity Limitations in High Mutation Density Regions

Reference-free approaches like ska lo may have sensitivity limitations in high mutation density regions. The ska lo algorithm was designed to address these limitations, achieving high sensitivity in single-nucleotide polymorphism calls while enabling detection of insertions and deletions [<a href="#ref-3">3</a>]. However, researchers should validate results from reference-free methods using complementary approaches when possible.

If you observe unexpectedly low variant counts in regions known to have high diversity, consider whether your approach is adequately capturing variation. Compare results across multiple tools and approaches to identify systematic biases in your workflow.

## Professional Escalation Criteria for Graph-Based Variant Calling

Knowing when to escalate problems to colleagues, collaborators, or specialized services is important for maintaining analysis quality and avoiding costly errors.

### Escalate When Graph Construction Fails Repeatedly

If you cannot construct a valid pan-genome graph after multiple attempts, escalate the issue. Graph construction failures may indicate problems with input assemblies, alignment parameters, or computational resources. Consult with colleagues who have experience building graphs for your species or consider using a pre-built graph from a trusted source.

### Escalate When Results Conflict Between Methods

If graph-based and linear methods produce substantially different results that you cannot explain, escalate the issue. Discrepancies may indicate problems with your graph, your alignment parameters, or your variant calling approach. Consult with bioinformatics experts who can help you interpret the differences and determine which results are more reliable.

### Escalate When Clinically Relevant Variants Are Ambiguous

When graph-based variant calling identifies variants with potential clinical significance, follow established protocols for confirmation and reporting. If you are uncertain about the interpretation of a variant, consult with colleagues who have expertise in the relevant genomic region or clinical condition. The StratoMod approach of predicting clinically relevant variants likely to be missed represents an improvement over current pipelines that only filter variants likely to be false [<a href="#ref-4">4</a>], but clinical interpretation requires specialized expertise.

### Escalate When Computational Requirements Exceed Available Resources

If your graph-based analysis consistently exceeds available computational resources, escalate the issue to your institution's high-performance computing support team or consider cloud-based solutions. The cache-friendly optimization of graph algorithms demonstrates that efficiency improvements are possible [<a href="#ref-5">5</a>], but you may need specialized infrastructure for large-scale analyses.

## Training and Skill Development for Graph-Based Methods

Developing proficiency in graph-based variant calling requires targeted training and practice. Several organizations provide structured learning opportunities that can help researchers build the necessary skills.

### Foundational Bioinformatics Training

The Carpentries offers foundational computing, data, shell, Git, and programming training that supports bioinformatics skill development [<a href="#ref-9">9</a>]. These skills are essential for working with graph-based tools, which often require command-line proficiency and version control practices.

EMBL-EBI Training provides learning pathways and data-resource training for practical analysis education [<a href="#ref-6">6</a>]. These resources can help researchers understand the underlying data structures and algorithms used in graph-based variant calling.

### Workflow-Specific Training

The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers learn graph-based methods in a practical context [<a href="#ref-7">7</a>]. These tutorials often include step-by-step instructions for specific analyses, making them valuable for researchers who prefer hands-on learning.

nf-core documentation describes community pipeline standards that emphasize reproducibility and configuration management [<a href="#ref-8">8</a>]. Understanding these standards can help researchers integrate graph-based tools into larger analysis pipelines and ensure that their workflows are reproducible.

### Reproducible Analysis Practices

Bioconductor provides documentation for reproducible genomic-analysis workflows that can be adapted for graph-based methods [<a href="#ref-10">10</a>]. These resources emphasize the importance of version control, environment management, and data provenance tracking for ensuring that analyses can be reproduced and verified.

For researchers planning to use graph-based methods in clinical or diagnostic settings, additional training in variant interpretation and regulatory requirements may be necessary. The improved accuracy of graph-based methods in complex regions may justify their use in clinical settings, but validation data specific to your workflow is essential.

## Frequently Asked Questions

### What types of genomic regions benefit most from graph-based variant calling?

Graph-based methods provide the greatest benefit in complex genomic regions where linear references fail. These include segmental duplications, regions with high sequence homology, structural variant breakpoints, and areas containing novel insertion sequences not present in the reference. The StratoMod study identified hard-to-map regions where graph-based methods excelled and quantified the improvement in recall [<a href="#ref-4">4</a>]. If your regions of interest include these challenging areas, graph-based calling is likely to recover variants that linear approaches miss.

### How do I choose between linear and graph-based variant callers for my project?

Consider your research question, study population, and computational resources. If you are working with well-characterized regions in populations well-represented by reference genomes, linear callers may be sufficient and more efficient. If you need comprehensive variant detection in diverse populations or complex regions, graph-based methods are worth the additional computational cost. For structural variant genotyping, Paragraph offers improved accuracy compared to other genotypers [<a href="#ref-2">2</a>]. For pathogen surveillance without high-quality references, reference-free approaches like ska lo may be most appropriate [<a href="#ref-3">3</a>].

### What is the difference between graph-based alignment and graph-based variant calling?

Graph-based alignment refers to the process of mapping reads to a graph reference, which allows reads to align to alternative paths representing different alleles. Graph-based variant calling uses these alignments to identify variants, either by comparing aligned reads to the graph or by genotyping known variants. Some tools perform both functions, while others specialize in one aspect. The vg toolkit provides general-purpose graph alignment, while Paragraph specializes in structural variant genotyping using sequence graphs [<a href="#ref-2">2</a>].

### Can graph-based methods improve somatic variant calling in cancer genomics?

Graph-based methods can improve variant detection in repetitive or complex regions that are relevant to cancer genomics. The ability to represent structural variants and complex rearrangements in a graph may help identify somatic variants that linear methods miss. However, somatic variant calling has additional considerations, including tumor purity, copy number alterations, and clonal heterogeneity. Researchers should validate graph-based somatic calls using orthogonal methods and consider the specific challenges of their cancer type and sample type.

### How much computational resources do graph-based methods require?

Graph-based methods generally require more memory and runtime than linear approaches. Building and storing graphs consumes substantial memory, particularly for large genomes or graphs representing many individuals. Alignment to graphs is more computationally intensive than linear alignment. However, optimization efforts have improved efficiency, with cache-friendly implementations running up to twice as fast as original versions on commodity computers [<a href="#ref-5">5</a>]. Researchers should benchmark tools on their specific data to understand resource requirements.

### What training resources are available for learning graph-based variant calling?

Several organizations provide training resources for bioinformatics methods. EMBL-EBI Training offers learning pathways and data-resource training for practical analysis education [<a href="#ref-6">6</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-7">7</a>]. The Carpentries offers foundational computing, data, shell, Git, and programming training that supports bioinformatics skill development [<a href="#ref-9">9</a>]. Bioconductor provides documentation for reproducible genomic-analysis workflows [<a href="#ref-10">10</a>]. These resources can help researchers develop the skills needed for graph-based variant calling.

### How do I validate variants identified only through graph-based methods?

Variants identified only through graph-based methods may represent true variation that linear methods miss, but they should be validated before drawing conclusions. Orthogonal validation methods include PCR amplification followed by Sanger sequencing, long-read sequencing, or independent analysis with different bioinformatics tools. For structural variants, validation may require specialized approaches such as optical mapping or targeted long-read sequencing. The appropriate validation method depends on the variant type and your research question.

### What are the limitations of reference-free variant calling approaches?

Reference-free approaches like ska lo avoid reference bias and reduce computational requirements, but they have limitations. These methods may have sensitivity issues in high mutation density regions, and interpreting results without a reference context can be challenging. The ska lo algorithm addresses some of these limitations by achieving high sensitivity in single-nucleotide polymorphism calls and enabling detection of insertions and deletions [<a href="#ref-3">3</a>]. However, researchers should validate reference-free results using complementary approaches when possible and consider whether reference-based methods provide additional context for their specific application.

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Deep Learning for Annotating Structural Variants in Viral Genomes](/knowledge/bioinformatics/deep-learning-for-annotating-structural-variants-in-viral-genomes)
- [Human Reference Genomes: Navigating and Downloading hg38 and hg19 FASTA Assemblies](/knowledge/bioinformatics/human-reference-genomes-hg38-hg19-assemblies)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Graph Peak Caller: Calling ChIP-seq peaks on graph-based reference genomes.](https://pubmed.ncbi.nlm.nih.gov/30779737). PLoS computational biology, 2019.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Paragraph: a graph-based structural variant genotyper for short-read sequence data.](https://pubmed.ncbi.nlm.nih.gov/31856913). Genome biology, 2019.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).](https://pubmed.ncbi.nlm.nih.gov/40171940). Molecular biology and evolution, 2025.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [StratoMod: predicting sequencing and variant calling errors with interpretable machine learning.](https://pubmed.ncbi.nlm.nih.gov/39397114). Communications biology, 2024.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Cache Friendly Optimisation of de Bruijn Graph Based Local Re-Assembly in Variant Calling.](https://pubmed.ncbi.nlm.nih.gov/30452377). IEEE/ACM transactions on computational biology and bioinformatics, 2020.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.