How to Call Structural Variants from Oxford Nanopore Ultra-Long Reads: A Step-by-Step Tutorial with Sniffles2
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Ultra-long Oxford Nanopore reads are critical for structural variant (SV) detection, enabling single reads to span entire deletions, insertions, duplications, inversions, and translocations (>50 bp) by providing direct breakpoint junction evidence.
- The workflow necessitates rigorous quality control of basecalled data, filtering for minimum read length (e.g., >1 kb) and appropriate quality scores to mitigate false split-read signals from low-quality short reads.
- Minimap2 with the
map-ontpreset is recommended for aligning ultra-long reads, with careful consideration of secondary and supplementary alignment handling to balance SV detection in repetitive regions against false-positive rates. - Sniffles2 parameter tuning for ultra-long reads involves adjusting minimum read support (potentially to 1), tandem-repeat filtering, and minimum SV length to optimize sensitivity and specificity based on coverage and biological context.
- Output VCF interpretation requires filtering based on read support (RE field), quality scores (QUAL), and potentially population frequency, followed by manual validation using genome browsers like IGV or Ribbon to confirm read-level evidence at breakpoint junctions.
Structural variant (SV) calling from Oxford Nanopore ultra-long reads requires a specific workflow: basecall and quality-check your raw signal data, align reads to a reference genome with a long-read-aware aligner, sort and index the alignment, run Sniffles2 with parameters tuned for ultra-long read length and coverage, then filter and interpret the output VCF using read support and population frequency evidence. This tutorial walks through each stage with concrete commands, parameter decisions, quality thresholds, and troubleshooting steps for the errors most commonly encountered when working with ultra-long nanopore data.
The audience for this workflow includes biology students, researchers, laboratory professionals, and life-science practitioners who have generated or plan to generate Oxford Nanopore sequencing data and need to identify deletions, insertions, duplications, inversions, and translocations larger than 50 base pairs. The tutorial assumes basic familiarity with the Linux command line and a working installation of Sniffles2, minimap2, and samtools. If you need to refresh foundational shell and data skills, The Carpentries Lessons provide structured training in command-line computing, and the EMBL-EBI Training portal offers bioinformatics learning pathways that cover sequence analysis fundamentals.
At a Glance
The table below summarizes the key decision points in an ultra-long-read SV calling workflow. Use it as a quick reference before starting an analysis, then consult the detailed sections that follow for parameter rationale and troubleshooting.
| Workflow Stage | Primary Tool | Key Parameter or Decision | Common Failure Mode |
|---|---|---|---|
| Basecalling and quality control | Dorado or Guppy | Quality score cutoff and read length filtering | Retaining low-quality short reads that inflate coverage and create false split-read signals |
| Read alignment | minimap2 | Map-ont preset, secondary alignment suppression | Misalignment at homopolymer and tandem repeat regions producing spurious SV calls |
| Alignment processing | samtools | Sort and index, optional duplicate marking | Unsorted alignments causing Sniffles2 to fail or produce duplicate calls |
| SV calling | Sniffles2 | Tandem-repeat and low-depth filtering flags | Calling SVs in regions with insufficient read support or excessive coverage noise |
| Output filtering | Sniffles2 or bcftools | Minimum support, quality, and allele frequency thresholds | Retaining low-confidence calls that do not replicate across samples or platforms |
| Visualization and validation | IGV or Ribbon | Manual inspection of split reads and breakpoint junctions | Accepting calls without visual confirmation of read-level support |
Understanding Structural Variants and Why Ultra-Long Reads Matter
Structural variants are genomic rearrangements that alter the physical structure of chromosomes. They include deletions, insertions, duplications, inversions, and translocations, and they are typically defined as events larger than 50 base pairs. This size threshold distinguishes SVs from smaller indels and single-nucleotide variants, and it has important consequences for detection strategy. Short-read sequencing platforms generate reads of 150 to 300 base pairs, which makes it difficult to span large rearrangements or resolve repetitive regions where many SVs occur. Long-read platforms such as Oxford Nanopore Technologies produce reads that can exceed 100 kilobases, and ultra-long preparations push into the megabase range for a subset of reads.
The practical advantage of ultra-long reads for SV calling is that a single read can span an entire deletion or insertion, providing direct evidence of the breakpoint junction. When a read aligns across a deletion, the alignment software must split the read into two segments that map to the reference on either side of the deleted region. The presence of these split-read signatures, combined with the alignment gap they create, gives the SV caller high-confidence evidence that a deletion exists. For insertions, the inserted sequence may be visible as an unaligned portion of the read, or it may align to a different genomic location, revealing the insertion source.
Ultra-long reads also improve SV detection in complex genomic regions. The human genome contains large segmental duplications, tandem repeats, and satellite arrays that are difficult to assemble or align with short reads. A study of amylase gene copy number variation in Indigenous Andean populations used ultra-long-read sequencing to demonstrate that recombination-based mechanisms drive the formation of high-copy haplotypes at the structurally complex AMY1 locus (Rapid adaptive increase of amylase gene copy number in Indigenous Andeans). This work illustrates how ultra-long reads can resolve the architecture of repetitive loci that short-read approaches cannot fully characterize. The NCBI Data Resources portal provides access to reference genomes and sequence databases that are useful for understanding the genomic context of the regions you are analyzing.
Preparing Your Input Data
Basecalling and Quality Control
The first decision in any nanopore SV calling workflow is how to convert raw electrical signal data into nucleotide sequences. Oxford Nanopore devices produce FAST5 or POD5 files containing the raw signal, and basecalling software translates that signal into FASTQ or BAM output. The choice of basecaller model affects downstream accuracy, particularly in homopolymer regions and areas with modified bases. For SV calling, the primary concern is that basecalling errors create mismatches and small indels that can confuse the aligner and the SV caller.
After basecalling, run a quality assessment on the resulting FASTQ files. Calculate read length distributions, quality score distributions, and total yield. For ultra-long-read SV calling, you want to know what fraction of your data exceeds 50 kilobases, because those reads provide the longest spanning evidence. You also want to check for adapter contamination and chimeric reads, which can produce false SV signals if they align in unexpected orientations.
Set a minimum read length threshold for your analysis. Reads shorter than 1 kilobase are generally not useful for SV calling because they cannot span most structural variant breakpoints. However, do not discard short reads entirely if you plan to use them for other purposes such as SNP genotyping. A study of target adaptive sampling long-read sequencing demonstrated that off-target reads, which are typically discarded in targeted workflows, can be used to accurately genotype common single-nucleotide polymorphisms across the genome (Assessing the efficacy of target adaptive sampling long-read sequencing through hereditary cancer patient genomes). This finding suggests that retaining a copy of your unfiltered data is prudent even when your primary analysis focuses on SVs.
Read Length Filtering and Coverage Estimation
Coverage depth is a critical parameter for SV calling confidence. Ultra-long-read experiments often produce lower total coverage than short-read experiments because the sequencing yield per flow cell is lower and the read length distribution is skewed toward shorter fragments. For SV calling, you need enough reads spanning each breakpoint to distinguish true events from sequencing artifacts.
A practical approach is to calculate the N50 read length and the total number of bases that exceed your minimum length threshold. The N50 is the length at which half of the total sequenced bases are in reads of that length or longer. If your N50 is below 10 kilobases, you are working with standard long reads instead of ultra-long reads, and you should adjust your expectations for SV detection sensitivity accordingly.
Coverage estimation for SV calling should account for the fact that not all reads will align uniquely. Reads from repetitive regions may map to multiple locations, and the aligner will either place them randomly or mark them as secondary alignments. Sniffles2 can use secondary alignments to identify SVs in repetitive regions, but this increases the risk of false positives. A conservative workflow filters secondary alignments and relies on primary alignments for initial SV discovery, then revisits repetitive regions with specialized parameters if the biological question requires it.
Aligning Ultra-Long Reads to a Reference Genome
Choosing an Aligner and Preset
Minimap2 is the most widely used aligner for Oxford Nanopore reads, and it is the aligner that Sniffles2 documentation assumes for most workflows. The map-ont preset configures minimap2 for Oxford Nanopore error profiles, including the higher indel rates and the specific substitution patterns produced by nanopore basecallers. Using the correct preset is essential because the alignment scoring parameters directly influence where split reads are placed and how gaps are penalized.
For ultra-long reads, consider whether you need to adjust the minimizer window size or the k-mer length. The default parameters for map-ont work well for reads up to several hundred kilobases, but extremely long reads can benefit from a larger k-mer size to reduce spurious alignments in repetitive regions. The tradeoff is that larger k-mers reduce sensitivity for reads with higher error rates. If you are working with a newly released basecaller model, test a small subset of your data with different minimap2 parameters and compare alignment rates and SV call consistency before committing to a full run.
Handling Secondary and Supplementary Alignments
Long reads frequently produce supplementary alignments, which are additional alignment segments for the same read that are chained together to represent a split alignment. Supplementary alignments are essential for SV calling because they provide the breakpoint evidence for deletions and inversions. Sniffles2 uses the CIGAR strings and supplementary alignment records to reconstruct the read structure and identify where it deviates from the reference.
Secondary alignments are different from supplementary alignments. A secondary alignment is an alternative placement for the same read that is not chained to the primary alignment. These occur when a read matches multiple genomic locations, which is common in repetitive regions. For SV calling, secondary alignments can be useful for detecting SVs in segmental duplications, but they also increase the false-positive rate. The standard workflow marks secondary alignments and excludes them from the initial SV call, then optionally includes them for a second pass focused on repetitive regions.
Sorting and Indexing Alignments
Sniffles2 requires sorted and indexed BAM files as input. After alignment, sort the BAM file by coordinate position and create an index. The samtools sort and samtools index commands perform these operations. Sorting is not optional for Sniffles2, the tool assumes that alignments are processed in genomic order, and unsorted input produces incorrect results or crashes.
If you are processing multiple samples, establish a consistent file naming convention that includes the sample identifier, the reference genome version, and the analysis date. This convention supports reproducibility and makes it easier to trace results back to the input data. The nf-core documentation provides community standards for pipeline configuration and reproducibility that are useful models for organizing your own analysis, even if you are not using a formal pipeline framework.
Running Sniffles2 for Structural Variant Calling
Basic Invocation and Output Files
Sniffles2 is invoked with a sorted BAM file and produces a VCF file containing the structural variant calls. The basic command is:
sniffles --input sample.sorted.bam --vcf sample.sniffles.vcf
This command runs Sniffles2 with default parameters, which are appropriate for many standard long-read datasets. For ultra-long reads, you will likely need to adjust parameters related to minimum read support, tandem repeat handling, and coverage depth. The output VCF contains one record per SV call, with genotype information, read support counts, and quality scores.
Sniffles2 also produces a support file that contains detailed information about the reads supporting each SV call. This file is useful for manual validation and for understanding why a particular call was made. The support file can be large for high-coverage datasets, so consider whether you need it for every sample or only for samples that require detailed follow-up.
Parameter Selection for Ultra-Long Reads
The minimum number of reads supporting an SV call is one of the most important parameters. The default value is 2, which means that at least two reads must show evidence for the same variant. For ultra-long-read data with lower coverage, you may need to lower this threshold to 1 to detect SVs that are covered by only a single long read. However, single-read calls have a higher false-positive rate, and you should validate them carefully.
The tandem repeat filtering parameter controls how Sniffles2 handles SVs that fall within tandem repeat regions. Tandem repeats are prone to alignment errors, and the expansion or contraction of a repeat array can be mistaken for a deletion or insertion. Sniffles2 can filter calls in these regions or adjust the breakpoint positions to account for repeat-induced ambiguity. For ultra-long reads, the repeat-spanning capability of individual reads can help resolve tandem repeat SVs, but the filtering parameter remains important for controlling false positives.
The minimum SV length parameter sets the lower bound for reported variants. The default is 50 base pairs, which matches the standard SV definition. If you are interested in smaller events, you can lower this threshold, but be aware that the false-positive rate increases as the size threshold decreases. Conversely, if you are only interested in large SVs, raising the threshold reduces the number of calls you need to filter.
Genotype and Population Frequency Considerations
Sniffles2 assigns a genotype to each SV call based on the read support pattern. A heterozygous call has support from a subset of reads, while a homozygous call has support from the majority of reads. The genotype quality score reflects the confidence in the genotype assignment. For ultra-long-read data, the genotype quality is generally higher than for short-read data because individual reads provide more information about the haplotype phase.
If you are analyzing multiple samples from the same population or study, you can merge the VCF files and calculate population-level allele frequencies. This step is important for filtering out private or low-frequency variants that are more likely to be artifacts. The Galaxy Training Network provides accessible tutorials on variant filtering and population genetics workflows that are useful for this stage of analysis.
Interpreting the Sniffles2 Output VCF
Understanding VCF Fields for Structural Variants
The Sniffles2 VCF output follows the standard VCF format with additional fields specific to structural variants. The INFO column contains the SV type, length, and supporting evidence. The FORMAT column contains genotype information for each sample. The key fields to examine are:
- SVTYPE: The type of structural variant (DEL, INS, DUP, INV, BND)
- SVLEN: The length of the variant in base pairs
- RE: The number of reads supporting the variant
- QUAL: The quality score for the variant call
The RE field is particularly important for assessing call confidence. A call with high read support is more reliable than a call with minimal support, regardless of the quality score. For ultra-long-read data, a single read that spans an entire deletion breakpoint can be more informative than multiple shorter reads that only partially cover the region.
Filtering Low-Confidence Calls
After Sniffles2 produces the initial VCF, apply filtering criteria to remove low-confidence calls. The most common filters are minimum read support, minimum quality score, and maximum allele frequency. The specific thresholds depend on your biological question and the coverage of your dataset.
For a typical whole-genome SV calling experiment, a minimum read support of 3 to 5 reads is a reasonable starting point. This threshold balances sensitivity and specificity for most datasets. If you are working with lower coverage, you may need to accept calls with fewer supporting reads, but you should validate those calls with an independent method.
Population frequency filtering is useful when you have multiple samples. Variants that appear in a single sample at low frequency are more likely to be artifacts or somatic events. Variants that appear across many samples are more likely to be genuine polymorphisms or shared structural variants. The vvv2_display tool, developed for viral genome analysis, demonstrates a useful approach to variant summarization: it consolidates results into a visual genome map and a tab-separated file listing only high-confidence variants with their frequencies and affected genes (vvv2_align_SE, vvv2_align_PE/vvv2_display: Galaxy-Based Workflows and Tool Designed to Perform, Summarize and Visualize Variant Calling and Annotation in Viral Genome Assemblies). This design principle of separating high-confidence calls from the full variant list is directly applicable to SV analysis.
Manual Validation with Read-Level Visualization
Automated filtering reduces the number of false positives, but manual validation remains important for high-impact calls. Load the BAM file and the VCF file into a genome browser such as IGV or Ribbon and examine the read alignments at each candidate SV breakpoint. Look for consistent split-read patterns, discordant read pairs, and coverage changes that support the SV call.
For ultra-long reads, the validation process is often straightforward because a single read can span the entire variant. If you see a read that aligns continuously across the reference region with a large gap, that gap is strong evidence for a deletion. If you see a read with an insertion that does not align to the reference, that unaligned sequence is evidence for an insertion. The visual confirmation step is especially important for calls that were made with minimal read support or in repetitive regions.
Optimizing the Workflow for Different Biological Questions
Whole-Genome SV Discovery
For whole-genome SV discovery, the goal is to identify as many true SVs as possible while keeping the false-positive rate manageable. This application benefits from the default Sniffles2 parameters with modest filtering. Run the initial call with a minimum read support of 2, then filter to a minimum support of 3 or higher for the final call set.
Whole-genome SV discovery also benefits from comparing your calls to known SV databases. The NCBI Data Resources provide access to dbVar, the database of genomic structural variation, which contains curated SV calls from multiple studies and platforms. Comparing your calls to dbVar can help you identify which of your calls are novel and which are known polymorphisms.
Targeted and Adaptive Sampling Applications
Targeted adaptive sampling is an Oxford Nanopore technology that enriches sequencing for specific genomic regions by rejecting reads that do not originate from the target area. This approach reduces the cost of sequencing specific genes or regions of interest. A study of hereditary cancer patient genomes demonstrated that target adaptive sampling long-read sequencing can identify single-nucleotide variants with accuracy comparable to short-read platforms and can elucidate complex structural variations, including SINE-R/VNTR/Alu elements affecting the APC gene (Assessing the efficacy of target adaptive sampling long-read sequencing through hereditary cancer patient genomes).
For targeted SV calling, the workflow is similar to whole-genome analysis, but the parameters may need adjustment because the coverage in the target region is much higher than in a whole-genome experiment. The high coverage can create noise that leads to false-positive calls, so you may need to increase the minimum read support threshold. The off-target reads from adaptive sampling can be used for other purposes, as demonstrated by the SNP genotyping application in the hereditary cancer study, so retain those reads even if they are not part of your primary SV analysis.
Transcriptome and Fusion Detection
Structural variants in transcriptome data include fusion genes, which are chimeric transcripts formed by the joining of two different genes. Fusion detection requires a different approach than genomic SV calling because the reads are derived from RNA instead of DNA, and the splicing patterns complicate alignment. A study of pediatric B-cell acute lymphoblastic leukemia developed FUSILLI, a long-read fusion detection algorithm that demonstrated higher sensitivity than existing methods for detecting fusion oncogenes in nanopore whole-transcriptome sequencing data (Long-Read Whole-Transcriptome Sequencing and Selective Gene Panel Profiling Enable Sensitive Detection of Fusion Oncogenes in Pediatric B-Cell Acute Lymphoblastic Leukemia).
If your biological question involves fusion detection, consider whether Sniffles2 is the appropriate tool or whether a transcriptome-specific caller such as FUSILLI would be more suitable. Sniffles2 can detect genomic rearrangements that produce fusions, but it does not directly model the transcript-level consequences of those rearrangements. The choice of tool depends on whether you are interested in the genomic event or the expressed fusion transcript.
Common Failure Patterns and Troubleshooting
Sniffles2 Crashes or Produces No Output
The most common cause of Sniffles2 failure is unsorted or unindexed BAM input. Verify that your BAM file is sorted by coordinate and that a corresponding BAI index file exists. If the BAM file is large, ensure that you have sufficient memory and disk space for the analysis. Sniffles2 can use substantial memory for high-coverage datasets, and running out of memory causes the process to terminate without producing output.
Another cause of empty output is a mismatch between the reference genome used for alignment and the reference expected by Sniffles2. Sniffles2 does not require a reference FASTA file for basic SV calling, but it uses the reference for some annotations. If you provide a reference file, ensure that it is the same version used for alignment.
Excessive False-Positive Calls
If Sniffles2 produces an unusually high number of SV calls, the most likely causes are low-quality alignments, insufficient read support thresholds, or repetitive regions. Check the alignment statistics for your BAM file, including the percentage of reads that aligned and the mapping quality distribution. Low mapping quality indicates that reads are aligning ambiguously, which produces spurious SV signals.
Increase the minimum read support threshold and apply tandem repeat filtering to reduce false positives. Examine the distribution of SV types and lengths in your output. A healthy SV call set has a mix of deletion and insertion calls with a size distribution that matches the expected biology of your organism. An excess of one SV type or an unusual size distribution suggests a systematic alignment or calling artifact.
Low Sensitivity for True SVs
If you are missing SVs that you expect to find, the most likely causes are insufficient coverage, overly aggressive filtering, or alignment parameters that are too strict. Check the coverage in the regions where you expect SVs. If the coverage is below 10x, consider whether you need additional sequencing or whether you should lower the minimum read support threshold.
Alignment parameters that are too strict can cause reads to be clipped or discarded instead of aligned across breakpoints. Review the alignment statistics for the fraction of reads that are chimeric or supplementary. If this fraction is very low, the aligner may be failing to split reads at true breakpoints. Consider adjusting the minimap2 parameters or trying a different aligner.
Records and Measurements for Reproducible SV Calling
Documenting the Analysis Environment
Reproducibility requires documentation of the software versions, parameters, and reference genome used for each analysis. Record the version of the basecaller, the alignment software, and Sniffles2. Record the exact command lines used for each step, including all parameter values. Store this information in a plain text file or a structured format such as a YAML configuration file.
The nf-core documentation emphasizes the importance of reproducible pipeline configuration and provides standards for documenting software versions and parameters. Even if you are not using nf-core pipelines, adopting similar documentation practices improves the reproducibility of your analysis. The Bioconductor project also provides guidance on reproducible genomic analysis workflows, including version control and package management.
Tracking Quality Metrics Across Samples
Maintain a spreadsheet or table that records key quality metrics for each sample. Include the total number of reads, the N50 read length, the total number of bases, the alignment rate, and the number of SV calls before and after filtering. These metrics allow you to compare samples and identify outliers that may indicate sequencing or analysis problems.
For multi-sample studies, track the number of SVs per sample and the distribution of SV types. A sample with an unusually high or low number of SVs may have a biological explanation, such as a genome with many rearrangements, or a technical explanation, such as poor sequencing quality. The vvv2_display approach of consolidating variant results into a summary file with frequencies and affected genes provides a useful model for tracking variant calls across samples (vvv2_align_SE, vvv2_align_PE/vvv2_display: Galaxy-Based Workflows and Tool Designed to Perform, Summarize and Visualize Variant Calling and Annotation in Viral Genome Assemblies).
Limitations and Interpretation Boundaries
What Sniffles2 Cannot Detect
Sniffles2 detects SVs that produce detectable alignment signatures, but some structural variants are invisible to this approach. Balanced rearrangements such as inversions and translocations can be detected if the breakpoints fall within reads, but very small inversions or those in repetitive regions may be missed. Copy-neutral events that do not change the total amount of DNA, such as inversions, are more difficult to detect than deletions or duplications because they do not produce coverage changes.
Sniffles2 also has limitations in highly repetitive regions. Tandem repeats, segmental duplications, and satellite DNA produce ambiguous alignments that confound SV calling. The tandem repeat filtering parameter helps control false positives in these regions, but it also reduces sensitivity for true SVs that occur within repeats. Ultra-long reads improve the situation because they can span entire repeat arrays, but the fundamental ambiguity of repetitive sequence remains.
Platform and Technology Dependencies
SV calling results depend on the sequencing platform, basecaller version, and analysis software versions. Oxford Nanopore data processed with different basecaller models will produce different SV call sets, even for the same underlying DNA. This platform dependence means that SV calls should be interpreted in the context of the specific technology used to generate them.
Cross-platform validation is recommended for high-impact SV calls. If you identify a deletion that is biologically important, consider validating it with an independent method such as PCR, optical mapping, or short-read sequencing. The long-read sequencing study of hereditary cancer genomes demonstrated that long-read approaches can identify complex structural variations that short-read platforms miss, but it also showed that different platforms have complementary strengths (Assessing the efficacy of target adaptive sampling long-read sequencing through hereditary cancer patient genomes).
Professional Escalation Criteria
When to Seek Additional Expertise
If you encounter persistent problems with SV calling that you cannot resolve through parameter adjustment, consider seeking help from bioinformatics support services or collaborators with long-read sequencing expertise. The EMBL-EBI Training portal and the Galaxy Training Network provide educational resources that can help you build the skills needed to troubleshoot your own analyses, but some problems require specialized knowledge.
Escalate to a bioinformatics specialist if you observe any of the following: Sniffles2 consistently crashes with a specific error message that you cannot resolve, your SV call set contains an implausibly high number of variants, or you need to validate a clinically or agriculturally significant SV call that will inform a consequential decision. For clinical applications, the stakes are higher, and the validation requirements are more stringent.
Data Sharing and Publication Standards
When publishing SV calling results, follow the community standards for data deposition and reporting. Deposit raw sequencing data in a public repository such as the NCBI Sequence Read Archive, and deposit the SV call set in a database such as dbVar. Provide detailed methods that include software versions, parameters, and reference genome versions so that other researchers can reproduce your analysis.
The nf-core documentation and the Bioconductor project both emphasize the importance of reproducible workflows and transparent reporting. Adopting these standards improves the credibility of your results and facilitates comparison with other studies.
Frequently Asked Questions
What is the minimum read length needed for structural variant calling with nanopore data?
There is no single minimum read length that works for all SV calling applications, but reads shorter than 1 kilobase are generally not useful for detecting SVs larger than 50 base pairs. The key consideration is whether a read can span the breakpoints of the variant you want to detect. For a deletion of 10 kilobases, you need reads that are longer than 10 kilobases to span the entire event. Ultra-long reads, defined as reads exceeding 100 kilobases, provide the best spanning evidence for large SVs and for resolving complex rearrangements in repetitive regions.
How much coverage do I need for reliable SV calling with Sniffles2?
The coverage requirement depends on the size and type of SVs you want to detect and the tolerance for false positives. A minimum of 10x to 15x coverage is a reasonable starting point for whole-genome SV discovery with long reads. Higher coverage improves sensitivity for smaller SVs and increases confidence in genotype calls. For ultra-long-read experiments where coverage is lower, you can compensate by using the spanning evidence from individual long reads, but you should validate low-support calls manually.
Should I use Sniffles2 or another SV caller for my nanopore data?
Sniffles2 is the most widely used SV caller for Oxford Nanopore data and is a good default choice for most applications. Other callers such as cuteSV, SVIM, and pbsv have different strengths and may perform better for specific data types or biological questions. If you are analyzing transcriptome data for fusion detection, a transcriptome-specific caller such as FUSILLI may be more appropriate than a genomic SV caller (Long-Read Whole-Transcriptome Sequencing and Selective Gene Panel Profiling Enable Sensitive Detection of Fusion Oncogenes in Pediatric B-Cell Acute Lymphoblastic Leukemia). Consider testing multiple callers on a subset of your data and comparing the results before committing to a single tool.
How do I distinguish real SVs from alignment artifacts?
Real SVs produce consistent read-level evidence across multiple independent reads. Alignment artifacts typically appear as calls supported by a single read or by reads with low mapping quality. Validate candidate SVs by examining the read alignments in a genome browser, checking for consistent split-read patterns and coverage changes. Compare your calls to known SV databases to identify variants that have been observed in other studies. Population frequency information from multiple samples can also help distinguish genuine polymorphisms from artifacts.
Can Sniffles2 detect SVs in repetitive regions of the genome?
Sniffles2 can detect SVs in repetitive regions, but the accuracy is lower than in unique regions because of alignment ambiguity. The tandem repeat filtering parameter controls how the tool handles calls in these regions. Ultra-long reads improve detection in repetitive regions because individual reads can span entire repeat arrays, providing direct evidence for the variant structure. The AMY1 copy number study demonstrated that ultra-long-read sequencing can resolve structurally complex loci that are difficult to characterize with other approaches (Rapid adaptive increase of amylase gene copy number in Indigenous Andeans). However, some repetitive regions remain intractable, and SVs in these regions may require specialized analysis approaches or validation with orthogonal methods.
What is the difference between primary, secondary, and supplementary alignments in SV calling?
Primary alignments are the best alignment for a read and are used for most downstream analysis. Supplementary alignments are additional segments of the same read that are chained to the primary alignment to represent a split alignment, which is essential evidence for SV breakpoints. Secondary alignments are alternative placements for the same read that are not chained to the primary alignment, and they occur when a read matches multiple genomic locations. Sniffles2 uses primary and supplementary alignments for SV calling, and secondary alignments can be included for analysis of repetitive regions.
How do I merge SV calls from multiple samples for population analysis?
Merge the VCF files from individual samples using a tool such as SURVIVOR or bcftools, then calculate allele frequencies across the merged call set. The merging process requires defining criteria for when two calls from different samples represent the same variant, typically based on breakpoint proximity and SV type and length similarity. After merging, filter the call set based on population frequency to remove private or low-frequency variants that are more likely to be artifacts.
What validation methods should I use to confirm important SV calls?
The choice of validation method depends on the type and size of the SV and the resources available. PCR amplification across the breakpoint is a direct validation method for deletions and insertions. Optical mapping provides genome-wide validation for large SVs. Short-read sequencing can validate SVs that produce discordant read pair or split-read signatures, though short reads have limited power for large events. For clinically significant SVs, orthogonal validation with an independent technology is strongly recommended.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Olink Proteomics: A Practical Guide to Panel Selection and Data Interpretation
- Spatial Transcriptomics Data Analysis: A Guide to Preprocessing, Integration, and Interpretation
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- vvv2_align_SE, vvv2_align_PE/vvv2_display: Galaxy-Based Workflows and Tool Designed to Perform, Summarize and Visualize Variant Calling and Annotation in Viral Genome Assemblies.. 2025.
- Rapid adaptive increase of amylase gene copy number in Indigenous Andeans.. 2026.
- Assessing the efficacy of target adaptive sampling long-read sequencing through hereditary cancer patient genomes.. 2024.
- Long-Read Whole-Transcriptome Sequencing and Selective Gene Panel Profiling Enable Sensitive Detection of Fusion Oncogenes in Pediatric B-Cell Acute Lymphoblastic Leukemia.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.