A Decision Guide to Structural Variant Callers for PacBio HiFi: pbsv, Sniffles2, and cuteSV Compared

By Dr. Zubair Khalid, DVM, MS, PhD ·

A Decision Guide to Structural Variant Callers for PacBio HiFi: pbsv, Sniffles2, and cuteSV Compared

Key Takeaways

  • PacBio HiFi structural variant (SV) calling requires careful tool selection, as pbsv, Sniffles2, and cuteSV exhibit distinct strengths impacting variant detection sensitivity, genotyping accuracy, and computational resource utilization.
  • Sniffles2 demonstrates superior speed and accuracy across various coverage levels (5-50x) and is specifically designed for population-level and mosaic SV detection, producing fully genotyped VCF files.
  • pbsv offers PacBio ecosystem integration and baseline performance, making it suitable for standard HiFi SV calling within production pipelines, though it may lack the specialized capabilities of research-driven tools.
  • cuteSV provides extensive parameter control and is often used as a baseline for genotyping improvement studies, particularly when paired with phasing-based tools like SVUPP to enhance genotype likelihoods.
  • General-purpose SV callers have limitations in paralogous regions and segmental duplications, where dedicated haplotype-based tools like Paraphase are necessary to detect a significant fraction of clinically relevant variants missed by standard approaches.
  • For clinical applications, comprehensive validation, including prospective clinical utility studies and careful documentation of coverage, data quality, and specific caller performance metrics, is paramount, especially when dealing with complex alleles or pharmacogenomic applications requiring integrated CNV, SV, and phasing analysis.

Structural variant calling from PacBio HiFi sequencing data requires a deliberate choice among available software tools, and that choice materially affects which genomic alterations you detect, how much compute you consume, and what downstream interpretation you can trust. This article gives you a structured decision framework for selecting among three widely used callers: pbsv, Sniffles2, and cuteSV. You will find performance context from published benchmarks, practical workflow steps, record-keeping guidance, common failure patterns, and criteria for escalating to specialized tools when general-purpose callers fall short.

The decision matters because structural variants (SVs) are technically challenging to identify, and long-read sequencing remains the most accurate approach for resolving complex genomic alterations [<a href="#ref-1">1</a>]. PacBio HiFi reads provide high per-base accuracy, but the choice of caller determines sensitivity for different SV types, genotyping accuracy, runtime, and memory footprint. Your selection should follow from your variant-type priorities, coverage depth, available computational resources, and whether you need population-level genotyping or clinical-grade accuracy in difficult genomic regions.

Scope and Reader Context

This guide serves biology students, researchers, laboratory professionals, and life-science practitioners who have PacBio HiFi data and need to choose an SV caller. You likely have sequencing output in BAM or FASTQ format, a reference genome, and a research question that involves deletions, insertions, duplications, inversions, or translocations. You may also be comparing HiFi against Oxford Nanopore Technologies (ONT) data, or you may be deciding whether to invest in higher coverage to improve SV detection.

The three callers covered here represent distinct design philosophies. pbsv is the PacBio-supported caller that integrates with the PacBio workflow ecosystem. Sniffles2 is a research-driven caller with demonstrated speed and accuracy improvements over earlier methods [<a href="#ref-1">1</a>]. cuteSV is a widely used caller that appears frequently in benchmarking studies and has been used as a baseline for newer genotyping approaches [<a href="#ref-2">2</a>]. Understanding how these tools differ in clustering strategy, filtering logic, and genotyping output will help you match the tool to your biological question.

At a Glance

The table below summarizes the key decision factors for the three callers. Use it as a starting point, then read the detailed sections that follow for workflow guidance and limitations.

Decision FactorpbsvSniffles2cuteSV
Primary design focusPacBio HiFi integration and production pipelinesSpeed, accuracy across coverage levels, population genotypingGeneral-purpose SV calling with flexible parameter control
Reported performance contextBaseline in multiple benchmark studies [<a href="#ref-3">3</a>]11.8x faster and 29% more accurate than prior state-of-the-art across 5-50x coverage and multiple SV types [<a href="#ref-1">1</a>]Used as comparison baseline in genotyping improvement studies [<a href="#ref-2">2</a>]
Genotyping capabilityProduces genotype calls in VCF outputProduces fully genotyped VCF files for family and population level calling [<a href="#ref-1">1</a>]Produces genotype calls, can be paired with phasing tools for improved accuracy [<a href="#ref-2">2</a>]
Mosaic SV detectionNot specifically designed for mosaic callingSupports detection of mosaic SVs in bulk long-read data [<a href="#ref-1">1</a>]Not specifically designed for mosaic calling
Best suited forStandard HiFi SV calling within PacBio workflowsComplex alleles, population studies, mosaic detection, multi-technology dataUsers who want parameter control and integration with phasing approaches
Computational profileModerate resource use typical of production toolsFaster runtime than prior methods [<a href="#ref-1">1</a>]Resource use varies with parameter settings and coverage

Structural Variant Types and Why They Matter for Caller Choice

Structural variants encompass a range of genomic alterations that differ in size, mechanism, and detectability. The most commonly reported SV types are deletions, insertions, duplications, inversions, and translocations. Each caller has different sensitivity across these types, and your study design should account for those differences.

Deletions and insertions in the 50 to 10,000 base pair range are particularly difficult to resolve with short-read sequencing, but long-read callers perform better on the same datasets [<a href="#ref-3">3</a>]. If your research question centers on this size range, HiFi data with an appropriate caller is a sound choice. However, you should verify that your chosen caller has demonstrated performance in this range instead of assuming all callers handle it equally.

Inversions and translocations require read-level evidence that spans the breakpoints. HiFi reads can span these regions because of their length, but the caller must correctly interpret the alignment signatures. Sniffles2 uses a repeat-aware clustering approach that improves accuracy in repetitive regions [<a href="#ref-1">1</a>], which is relevant for inversion breakpoints that often fall in or near repeats.

Copy-number variants (CNVs) and gene conversions add another layer of complexity. In paralogous regions, standard HiFi variant callers detected 95 of 125 known clinically relevant variants, while the remaining 30 were only identified by a dedicated haplotype-based caller called Paraphase [<a href="#ref-4">4</a>]. This finding has direct implications for your caller choice: if your study involves segmental duplications, gene-pseudogene pairs, or other highly homologous regions, a general-purpose SV caller may miss a substantial fraction of variants. You should plan for a specialized follow-up tool instead of expecting any single general caller to capture everything.

Core Principles of Structural Variant Calling from HiFi Data

Read Alignment and Its Influence on SV Detection

Every SV caller depends on read alignment to the reference genome. HiFi reads have high per-base accuracy, which reduces alignment ambiguity compared to noisier long-read technologies. However, alignment ambiguities still occur in repetitive and paralogous regions [<a href="#ref-4">4</a>]. The caller you choose applies its own clustering and filtering logic to the aligned reads, so the same alignment file can produce different SV calls depending on the caller.

You should generate alignments with a mapper that is appropriate for HiFi data and that produces the auxiliary tags callers expect. Most callers document their recommended mapper and alignment parameters. Deviating from those recommendations can reduce sensitivity or increase false positives. If you are comparing callers, use the same alignment file as input to all callers so that differences in output reflect caller behavior instead of alignment differences.

Coverage Depth and Its Effect on Sensitivity and Precision

Coverage depth is one of the strongest determinants of SV calling performance. Sniffles2 was evaluated across 5 to 50x coverage and demonstrated improvements over prior methods across that range [<a href="#ref-1">1</a>]. Lower coverage reduces cost and input DNA requirements, but it also reduces sensitivity. A benchmark study showed that Blackbird, a hybrid approach using synthetic long reads and low-coverage long reads, achieved F1-scores of 0.835 for deletions and 0.808 for insertions at 5x coverage, comparable to pbsv and Sniffles2 using 10x PacBio HiFi coverage [<a href="#ref-3">3</a>].

This finding has a practical implication: if your budget or DNA input is limited, you may be able to reduce coverage and still obtain useful SV calls, but you should expect lower sensitivity than you would get at higher coverage. You should also recognize that the comparison in the Blackbird study used F1-scores, which balance precision and recall. Your tolerance for false positives versus false negatives should inform your coverage decision.

Variant Size Range and Caller Sensitivity

The size range of SVs you care about should influence your caller choice. Most SVs in the 50 to 10,000 base pair range cannot be resolved with short-read sequencing, but long-read callers perform better on the same datasets [<a href="#ref-3">3</a>]. Within that range, individual callers may have different sensitivity profiles. Some callers are more sensitive for small insertions, while others excel at large deletions.

You should establish your variant size range of interest before selecting a caller. If your study spans a broad size range, you may need to combine calls from multiple tools or use a caller with demonstrated performance across the full range. The Sniffles2 evaluation covered multiple SV types and a range of sizes [<a href="#ref-1">1</a>], which supports its use for broad SV surveys.

Practical Workflow for Caller Selection and Execution

Step 1: Define Your Variant Priorities and Study Design

Before running any caller, write down your primary variant types, size range, coverage depth, and downstream analysis needs. This specification will guide every subsequent decision. Include the following in your study design notes:

  • Primary SV types of interest (deletions, insertions, duplications, inversions, translocations)
  • Size range of interest
  • Coverage depth of your HiFi data
  • Whether you need population-level genotyping or single-sample calling
  • Whether your regions of interest include paralogous or repetitive sequences
  • Available computational resources (CPU, memory, wall time)
  • Downstream analysis requirements (phasing, annotation, clinical interpretation)

Step 2: Assess Your Data Quality and Coverage

Run quality control on your HiFi reads before calling SVs. Check read length distribution, per-base accuracy, and coverage uniformity. Low-quality reads or uneven coverage will degrade SV calling regardless of caller choice. If your coverage is below 10x, consider whether your study can tolerate reduced sensitivity or whether you should sequence deeper.

For clinical or diagnostic applications, coverage and data quality requirements may be higher. The evaluation of HiFi sequencing for paralogous gene analysis used 86 individuals with 125 known clinically relevant variants [<a href="#ref-4">4</a>], and the authors suggested that long-read genome sequencing may be ready for first-tier diagnostic use in specific contexts, guided by prospective clinical utility studies [<a href="#ref-4">4</a>]. If you are working toward clinical adoption, document your coverage and quality metrics carefully.

Step 3: Select Your Caller Based on the Decision Tree

Use the following decision logic to select your primary caller:

  • If you need population-level or family-level genotyped VCF files, choose Sniffles2. It was specifically designed to solve family-level to population-level SV calling and produces fully genotyped VCF files [<a href="#ref-1">1</a>].
  • If you suspect mosaic SVs in your sample, choose Sniffles2. It enables detection of mosaic SVs in bulk long-read data [<a href="#ref-1">1</a>].
  • If you are working within the PacBio ecosystem and want the vendor-supported caller, choose pbsv. It has been used as a baseline in benchmark studies [<a href="#ref-3">3</a>] and integrates with PacBio workflows.
  • If you want parameter control and plan to use phasing-based genotyping improvement, choose cuteSV. It can be paired with tools like SVUPP that incorporate read phasing information into genotype likelihoods [<a href="#ref-2">2</a>].
  • If your regions of interest include paralogous genes or segmental duplications, plan to add a specialized caller such as Paraphase regardless of your general-purpose caller choice [<a href="#ref-4">4</a>].

Step 4: Run the Caller with Documented Parameters

Record the exact command, parameters, reference version, and input files for every caller run. This record is essential for reproducibility and for troubleshooting when results differ between callers. Use the parameter recommendations from each caller's documentation as your starting point, and only deviate when you have a documented reason.

For Sniffles2, the published evaluation used coverage from 5 to 50x and included both ONT and HiFi data [<a href="#ref-1">1</a>]. Your parameter settings should match your sequencing technology and coverage. For cuteSV, the SVUPP study demonstrated that genotyping accuracy can be improved by incorporating phasing information [<a href="#ref-2">2</a>], so you may want to plan for a phasing step if genotyping accuracy is critical.

Step 5: Validate Your Calls

Validation is a separate step from calling. You should assess precision and recall on a benchmark set if one is available for your organism and variant types. The Genome in a Bottle (GIAB) benchmark for HG002 was used in the Blackbird evaluation [<a href="#ref-3">3</a>], and similar benchmarks may exist for your species of interest.

If no benchmark exists, consider orthogonal validation approaches such as PCR, targeted sequencing, or manual inspection of read alignments at candidate loci. Document your validation results and include them in your analysis report.

Step 6: Document and Report

Your final report should include the caller version, parameters, reference genome version, coverage statistics, and validation results. This documentation allows others to reproduce your analysis and interpret your findings in context. If you publish your results, include these details in the methods section.

Options and Tradeoffs Among the Three Callers

pbsv: Production Integration and Baseline Performance

pbsv is the PacBio-supported structural variant caller and is the default choice for many HiFi users who want a supported tool within the PacBio ecosystem. It has been used as a baseline in benchmark studies [<a href="#ref-3">3</a>], which means you can find published performance comparisons to contextualize your results.

The main tradeoff with pbsv is that it may not have the same sensitivity as research-driven callers for certain variant types or coverage levels. The Blackbird study compared pbsv and Sniffles2 at 10x HiFi coverage [<a href="#ref-3">3</a>], and both served as baselines for the hybrid approach. If you use pbsv, you should verify its performance on your specific variant types and coverage instead of assuming it matches the best-performing research tool.

pbsv is a reasonable choice when you want vendor support, integration with PacBio analysis tools, and a stable production workflow. It is less suitable when you need population-level genotyping improvements or mosaic SV detection, which are not its primary design focus.

Sniffles2: Speed, Accuracy, and Population Genotyping

Sniffles2 was designed to address the technical challenges of SV calling with long reads and implements a repeat-aware clustering approach with fast consensus sequence generation and coverage-adaptive filtering [<a href="#ref-1">1</a>]. The published evaluation reported that Sniffles2 is 11.8 times faster and 29% more accurate than state-of-the-art SV callers across different coverages (5 to 50x), sequencing technologies (ONT and HiFi), and SV types [<a href="#ref-1">1</a>].

The population-level genotyping capability is a distinct advantage. Sniffles2 produces fully genotyped VCF files and was used to accurately identify causative SVs around MECP2 in 11 probands, including highly complex alleles with three overlapping SVs [<a href="#ref-1">1</a>]. It also detected mosaic SVs in brain tissue from a patient with multiple system atrophy, revealing diversity within the cingulate cortex that impacted genes involved in neuron function and repetitive elements [<a href="#ref-1">1</a>].

The main tradeoff with Sniffles2 is that its speed and accuracy advantages were demonstrated in a specific evaluation context, and your results may vary depending on your data and parameters. You should also note that the 29% accuracy improvement was relative to prior state-of-the-art callers at the time of publication, and the field continues to evolve.

cuteSV: Parameter Control and Phasing Integration

cuteSV is a general-purpose SV caller that gives you substantial parameter control. It has been used as a comparison baseline in studies of genotyping improvement [<a href="#ref-2">2</a>], which means you can find published performance data for your reference.

The SVUPP study demonstrated that cuteSV2 genotyping can be improved by incorporating read phasing information into genotype likelihoods [<a href="#ref-2">2</a>]. SVUPP achieved higher accuracy than cuteSV2, Sniffles2, and kanpig for genotyping SVs without close neighbor SVs, using both ONT and PacBio HiFi data [<a href="#ref-2">2</a>]. This finding suggests that if genotyping accuracy is your priority, you should consider pairing cuteSV with a phasing-based genotyping improvement tool.

The main tradeoff with cuteSV is that you take on more responsibility for parameter optimization. The flexibility that allows you to tune the caller for your data also means you need to understand the parameter space and validate your choices. If you prefer a caller with sensible defaults and less tuning burden, Sniffles2 or pbsv may be more appropriate.

Observations and Measurements for Caller Comparison

Runtime and Resource Consumption

Runtime differences between callers can be substantial. Sniffles2 was reported to be 11.8 times faster than prior state-of-the-art callers [<a href="#ref-1">1</a>], which translates to meaningful wall-clock savings on large datasets. If you are processing many samples or working under time constraints, runtime should factor into your caller choice.

Memory consumption also varies between callers and depends on coverage and parameter settings. You should measure peak memory usage for your specific dataset instead of relying on published values from different data. Record runtime and memory for each caller run so you can make informed decisions for future projects.

Genotyping Accuracy

Genotyping accuracy is distinct from detection accuracy. A caller may detect an SV correctly but assign an incorrect genotype. The SVUPP study focused specifically on genotyping and showed that incorporating phasing information improves genotype likelihoods [<a href="#ref-2">2</a>]. If your downstream analysis depends on accurate genotypes, such as in association studies or clinical interpretation, you should evaluate genotyping accuracy separately from detection accuracy.

For pharmacogenomics applications, accurate copy number calling, structural variation identification, variant calling, and phasing within each pharmacogene copy are all required for star-allele calling [<a href="#ref-5">5</a>]. Aldy 4, a dedicated star-allele caller, demonstrated near-perfect accuracy across different sequencing technologies including PacBio HiFi [<a href="#ref-5">5</a>]. This example illustrates that specialized tools may be necessary when your downstream interpretation requires more than general SV detection.

Sensitivity in Difficult Regions

Standard HiFi variant callers detected 95 of 125 known clinically relevant variants across 11 paralogous loci, while the remaining 30 were only identified by Paraphase, a dedicated haplotype-based caller [<a href="#ref-4">4</a>]. This finding quantifies the limitation of general-purpose callers in paralogous regions. If your study includes such regions, you should expect to miss variants with any general-purpose caller and plan for specialized follow-up.

The same study demonstrated that HiFi sequencing combined with Paraphase detected all known variants, including SNVs, InDels, CNVs, SVs, and gene conversions [<a href="#ref-4">4</a>]. This result supports a workflow that combines general-purpose SV calling with specialized haplotype-based calling for difficult regions.

Records and Measurements for Reproducible SV Calling

Essential Records for Each Caller Run

Maintain a structured record for every SV calling run. Include the following fields:

  • Caller name and version
  • Reference genome build and version
  • Input BAM or FASTQ file paths and checksums
  • Exact command line with all parameters
  • Coverage statistics for the input data
  • Runtime and peak memory usage
  • Output VCF file path and checksum
  • Validation results if applicable

This record allows you to reproduce your analysis exactly and to diagnose discrepancies between callers. It also supports good scientific practice for publication and for collaboration with other researchers.

Quality Metrics to Track

Track the following quality metrics for each run:

  • Number of SVs called by type (deletion, insertion, duplication, inversion, translocation)
  • Size distribution of called SVs
  • Genotype quality metrics from the VCF
  • Transition-transversion ratio if applicable to your organism
  • False positive rate on a benchmark set if available

These metrics help you identify anomalous runs and compare caller performance on your specific data.

Common Failure Patterns and Troubleshooting

Failure Pattern 1: Low Sensitivity in Repetitive Regions

If your caller misses SVs in repetitive regions, the cause may be alignment ambiguity instead of caller deficiency. HiFi reads can span repetitive regions, but the aligner may place reads incorrectly. Sniffles2 implements repeat-aware clustering [<a href="#ref-1">1</a>], which addresses this issue to some degree. If you observe low sensitivity in repeats, check your alignment parameters and consider whether a repeat-aware caller is appropriate.

Failure Pattern 2: High False Positive Rate at Low Coverage

Low coverage increases false positives because the caller has less evidence to distinguish true SVs from alignment artifacts. The Blackbird study showed that reducing coverage significantly impacts the performance of existing SV callers [<a href="#ref-3">3</a>]. If you see a high false positive rate, verify your coverage and consider whether you need deeper sequencing or a caller with coverage-adaptive filtering.

Failure Pattern 3: Missing Variants in Paralogous Regions

If your study includes paralogous genes or segmental duplications, general-purpose callers will miss a substantial fraction of variants. The Paraphase study showed that standard HiFi variant callers detected only 95 of 125 known variants in paralogous loci [<a href="#ref-4">4</a>]. This is not a caller bug but a fundamental limitation of alignment-based approaches in highly homologous regions. You need a dedicated haplotype-based caller for these regions.

Failure Pattern 4: Genotype Errors for SVs with Close Neighbors

The SVUPP study found that genotyping accuracy was higher for SVs without close neighbor SVs [<a href="#ref-2">2</a>]. If your SVs of interest cluster together, you may see genotype errors with standard callers. Consider using a phasing-based genotyping improvement tool such as SVUPP, which incorporates read phasing information into genotype likelihoods [<a href="#ref-2">2</a>].

Failure Pattern 5: Runtime or Memory Exhaustion

If your caller run exceeds available resources, check your parameter settings and coverage. Some callers have parameters that control the clustering or filtering complexity. You may need to adjust these parameters or reduce the genomic region being analyzed. Record the resource usage of successful runs to calibrate expectations for future runs.

Limitations of General-Purpose SV Callers

Paralogs and Segmental Duplications

The most significant limitation of general-purpose SV callers is reduced sensitivity in paralogous regions. The Paraphase evaluation demonstrated that 30 of 125 known clinically relevant variants were missed by standard HiFi variant callers and only detected by a dedicated haplotype-based caller [<a href="#ref-4">4</a>]. If your study includes such regions, you must use specialized tools to achieve comprehensive variant detection.

Mosaic Variants

Mosaic SVs, which are present in only a subset of cells, require specialized detection approaches. Sniffles2 supports mosaic SV detection in bulk long-read data [<a href="#ref-1">1</a>], but not all callers have this capability. If you suspect mosaicism in your sample, verify that your chosen caller can detect it.

Complex Alleles

Highly complex alleles with multiple overlapping SVs present a challenge for any caller. Sniffles2 was able to identify causative SVs around MECP2, including alleles with three overlapping SVs [<a href="#ref-1">1</a>]. This capability is not universal, and you should validate your caller's performance on complex alleles if your study involves them.

Genotyping Accuracy

Detection and genotyping are separate challenges. A caller may detect an SV but assign an incorrect genotype. The SVUPP study showed that genotyping accuracy can be improved by incorporating phasing information [<a href="#ref-2">2</a>]. If your downstream analysis depends on accurate genotypes, you should evaluate genotyping accuracy separately and consider phasing-based improvement tools.

Safety and Regulatory Context for Clinical Applications

If your SV calling results will inform clinical decisions, you must consider regulatory and validation requirements. The Paraphase study suggested that long-read genome sequencing may be ready for wider implementation, possibly as a first-tier diagnostic approach for individuals with suspected variants in paralogous regions [<a href="#ref-4">4</a>]. However, the authors emphasized that clinical adoption should be guided by prospectively designed clinical utility studies, alongside evaluation of sensitivity, specificity, and cost-effectiveness [<a href="#ref-4">4</a>].

For pharmacogenomics applications, accurate star-allele calling requires copy number calling, structural variation identification, variant calling, and phasing within each pharmacogene copy [<a href="#ref-5">5</a>]. Aldy 4 demonstrated near-perfect accuracy across different sequencing technologies [<a href="#ref-5">5</a>], but you should verify that any tool you use for clinical pharmacogenomics meets the required validation standards.

You should also be aware that variant calling results can have implications for genetic counseling and patient management. If your analysis identifies variants of potential clinical significance, you should have a pathway for confirmation and clinical interpretation. This pathway should be established before you begin the analysis, not after you have results.

Professional Escalation Criteria

You should escalate to specialized tools or seek expert consultation in the following situations:

  • Your regions of interest include paralogous genes, segmental duplications, or gene-pseudogene pairs. Standard callers will miss a substantial fraction of variants in these regions [<a href="#ref-4">4</a>], and you need a dedicated haplotype-based caller such as Paraphase.
  • You need accurate star-allele calls for pharmacogenomics. General-purpose SV callers do not provide the integrated copy number, structural variation, variant calling, and phasing required for star-allele calling [<a href="#ref-5">5</a>]. Use a dedicated tool such as Aldy 4.
  • You suspect mosaic SVs in your sample. Not all callers support mosaic detection, and you need a caller with demonstrated mosaic capability such as Sniffles2 [<a href="#ref-1">1</a>].
  • Your genotyping accuracy is critical and you have SVs with close neighbors. Consider phasing-based genotyping improvement tools such as SVUPP [<a href="#ref-2">2</a>].
  • You are working toward clinical adoption. Engage with clinical genomics experts and follow prospective clinical utility study requirements [<a href="#ref-4">4</a>].

Building a Reproducible SV Calling Benchmark for Your Own Data

Published benchmarks provide useful context, but they cannot tell you how a caller will perform on your specific organism, coverage profile, and variant size distribution. The most reliable way to select among pbsv, Sniffles2, and cuteSV is to build a small, targeted benchmark using your own data or a closely related reference sample. This section gives you a practical framework for constructing that benchmark, recording the results, and using the outcomes to make a defensible caller choice.

Why Published Benchmarks Are Not Sufficient for Your Decision

Published performance numbers come from specific datasets, reference builds, and parameter settings that may not match your situation. The Sniffles2 evaluation reported an 11.8 times speed improvement and 29% accuracy gain over prior callers across 5 to 50x coverage, multiple sequencing technologies, and several SV types [<a href="#ref-1">1</a>]. Those numbers are useful for understanding the caller's design strengths, but they do not guarantee identical performance on your data. The Blackbird study compared pbsv and Sniffles2 at 10x PacBio HiFi coverage on the HG002 Genome in a Bottle benchmark [<a href="#ref-3">3</a>], which is a human sample with well-characterized variants. If you work on a non-human species, a plant genome with high heterozygosity, or a cancer sample with aneuploidy, the published numbers may not transfer directly.

A second limitation is that published benchmarks often report aggregate F1-scores that combine precision and recall across all SV types. Your study may prioritize deletions over insertions, or you may care more about avoiding false positives than maximizing sensitivity. An aggregate score hides these distinctions. Building your own benchmark lets you measure performance for the variant types and size ranges that matter to your specific research question.

A third limitation is version drift. Callers update frequently, and the version you install may differ from the version used in a published study. Parameter defaults also change between versions. A benchmark run on your installed versions gives you current information instead of historical context.

Designing a Benchmark That Answers Your Question

Step 1: Select a Benchmark Sample

Choose a sample that resembles your real data as closely as possible. If you work on human samples, the HG002 Genome in a Bottle sample is a reasonable choice because it has extensive validated variant calls and was used in the Blackbird evaluation [<a href="#ref-3">3</a>]. If you work on a non-human species, look for a reference sample with validated SVs from orthogonal methods such as PCR, optical mapping, or assembly comparison. If no validated sample exists, you can create a synthetic benchmark by simulating HiFi reads from a reference genome with known SVs inserted, but you should recognize that simulated data does not capture all real-world complexities such as base modification, coverage dropout, and alignment artifacts.

For clinical or diagnostic applications, the choice of benchmark sample carries additional weight. The Paraphase evaluation used 86 individuals with 125 known clinically relevant variants across 11 paralogous loci [<a href="#ref-4">4</a>], which provided a realistic assessment of caller performance in difficult regions. If your study targets paralogous genes, your benchmark should include such regions instead of relying on a genome-wide benchmark that may not represent those challenges.

Step 2: Define Your Variant Truth Set

Your benchmark needs a set of known SVs to compare against caller output. For human samples, curated truth sets such as those from Genome in a Bottle provide validated variants with confidence annotations. For other species, you may need to construct a truth set from multiple evidence sources. Record the following for each truth variant:

  • Chromosome and position
  • Variant type (deletion, insertion, duplication, inversion, translocation)
  • Size in base pairs
  • Genotype if known
  • Confidence level (high, medium, low)
  • Whether the variant falls in a repetitive or paralogous region

This annotation lets you stratify your benchmark results by variant type, size range, and genomic context. Stratification is essential because a caller may perform well on simple deletions but poorly on insertions in repetitive regions.

Step 3: Establish Your Evaluation Metrics

Define your metrics before running any caller. The most common metrics are precision, recall, and F1-score, but you should also consider genotype accuracy if your downstream analysis depends on genotypes. The SVUPP study demonstrated that genotyping accuracy is a separate challenge from detection accuracy, and incorporating phasing information into genotype likelihoods improved results [<a href="#ref-2">2</a>]. If you need accurate genotypes, measure genotype concordance separately from detection concordance.

For each caller, calculate:

  • Precision: the fraction of called SVs that match a truth variant
  • Recall: the fraction of truth variants that were called
  • F1-score: the harmonic mean of precision and recall
  • Genotype concordance: the fraction of correctly genotyped calls among true positives
  • Runtime and peak memory usage

You should also measure performance separately for each variant type and size bin. A caller with excellent overall F1-score may have poor recall for insertions larger than 1 kilobase, which would matter if your study targets large insertions.

Step 4: Control for Confounding Variables

Run all callers on the same alignment file so that differences in output reflect caller behavior instead of alignment differences. Use the same reference genome build and the same read set. Record the exact version of each caller and the parameter settings you used. If you change any parameter, document the change and the reason.

If you have multiple samples, consider running the benchmark on a subset before scaling to the full dataset. This approach lets you identify problems early and adjust your workflow before committing substantial compute time.

Recording Benchmark Results in a Structured Format

Create a table with one row per caller and columns for each metric. Include the caller version, parameter settings, and the date of the run. Store this table alongside your analysis code and input file checksums so that the benchmark is fully reproducible.

MetricpbsvSniffles2cuteSV
Caller version
Reference build
Alignment file checksum
Precision (all SVs)
Recall (all SVs)
F1-score (all SVs)
Precision (deletions)
Recall (deletions)
Precision (insertions)
Recall (insertions)
Genotype concordance
Runtime (minutes)
Peak memory (GB)

Fill in this table for your specific data and use it as the primary evidence for your caller decision. The published benchmarks [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>] provide context, but your table reflects your actual conditions.

Interpreting Benchmark Results for Your Decision

When Results Are Close

If two callers produce similar F1-scores on your benchmark, choose based on secondary factors. Consider runtime if you process many samples. Consider population-level genotyping capability if you plan to scale beyond single samples. Sniffles2 was designed for family-level to population-level SV calling and produces fully genotyped VCF files [<a href="#ref-1">1</a>], which may save you a separate genotyping step. Consider the availability of support and documentation. pbsv benefits from PacBio ecosystem integration, which may simplify your workflow if you already use PacBio analysis tools.

When Results Favor One Caller for Specific Variant Types

If your benchmark shows that one caller has substantially better recall for the variant types central to your study, choose that caller even if its overall F1-score is slightly lower. For example, if you study insertions and one caller detects 20% more true insertions than the others, that advantage may outweigh a small deficit in deletion calling. Document this reasoning in your analysis report so that others understand why you made the choice.

When Results Are Inconclusive

If your benchmark produces noisy or inconsistent results, investigate before making a decision. Check whether the truth set has sufficient representation across variant types and size ranges. A truth set with only 20 deletions and 5 insertions cannot support a reliable comparison for insertions. Consider expanding your benchmark sample or adding simulated variants to increase statistical power.

Also check whether your alignment parameters match the recommendations for each caller. Some callers expect specific alignment tags or settings, and using the wrong alignment can depress performance for reasons unrelated to the caller itself. Revisit the documentation for each caller and verify that your alignment is compatible.

Common Benchmark Pitfalls and How to Avoid Them

Pitfall 1: Using a Truth Set That Does Not Match Your Variant Priorities

A genome-wide truth set may not contain enough variants in the size range or genomic context you care about. If your study targets SVs in the 50 to 10,000 base pair range, which are difficult to resolve with short reads but better handled by long-read callers [<a href="#ref-3">3</a>], ensure your truth set has adequate representation in that range. Stratify your results by size bin to confirm that your chosen caller performs well where it matters.

Pitfall 2: Comparing Callers on Different Alignments

If you align reads separately for each caller using different mappers or parameters, you cannot attribute performance differences to the caller alone. Use a single alignment file for all callers. This control is essential for a fair comparison.

Pitfall 3: Ignoring Genotype Accuracy

Detection and genotyping are separate challenges. A caller may detect an SV correctly but assign an incorrect genotype. The SVUPP study showed that incorporating phasing information into genotype likelihoods improves genotyping accuracy [<a href="#ref-2">2</a>]. If your downstream analysis depends on genotypes, measure genotype concordance in your benchmark and consider whether you need a phasing-based improvement step.

Pitfall 4: Overlooking Paralogous Regions

If your benchmark sample does not include paralogous regions, your results will overstate caller performance for studies that target those regions. The Paraphase evaluation showed that standard HiFi variant callers detected only 95 of 125 known clinically relevant variants across 11 paralogous loci, with the remaining 30 detected only by a dedicated haplotype-based caller [<a href="#ref-4">4</a>]. If your study includes segmental duplications or gene-pseudogene pairs, include such regions in your benchmark and plan for a specialized follow-up tool.

Pitfall 5: Failing to Document Parameter Settings

Benchmark results are only interpretable if you know the exact parameters used. Record the full command line for each caller, including all flags and values. Store this information with your benchmark results so that you can reproduce the comparison later or explain discrepancies if results change after a caller update.

Scaling From Benchmark to Production

Once your benchmark identifies a primary caller, run that caller on a small subset of your real data before processing the full dataset. Compare the output on the subset to your benchmark expectations. If the real data produces substantially different results, investigate the cause before scaling up. Differences may arise from coverage variation, sample quality, or genomic features not represented in your benchmark.

For production runs, maintain the same record-keeping discipline you used in the benchmark. Record the caller version, parameters, reference build, input file checksums, runtime, and memory usage for every run. This record supports reproducibility and helps you diagnose problems when results differ from expectations.

When to Repeat Your Benchmark

Repeat your benchmark when any of the following occur:

  • You update a caller to a new version
  • You change your alignment pipeline
  • You switch to a different reference genome build
  • You begin working with a new organism or sample type
  • You change your variant type priorities or size range of interest

Each of these changes can alter caller performance in ways that published benchmarks cannot predict. A fresh benchmark run on your current pipeline gives you reliable evidence for your decision.

Integrating Benchmark Results With Specialized Tool Decisions

Your benchmark may reveal that no general-purpose caller performs adequately for your regions of interest. If your study includes paralogous genes, the evidence is clear that standard callers will miss a substantial fraction of variants [<a href="#ref-4">4</a>]. In that case, your benchmark should include a comparison against a specialized tool such as Paraphase, which was shown to detect variants that standard callers missed [<a href="#ref-4">4</a>]. Similarly, if you need star-allele calls for pharmacogenomics, your benchmark should include a dedicated caller such as Aldy 4, which demonstrated near-perfect accuracy across sequencing technologies including PacBio HiFi [<a href="#ref-5">5</a>].

The benchmark framework described here is not a one-time exercise. It is a repeatable process that you can apply whenever your data, tools, or research questions change. The investment in building a benchmark pays off by giving you evidence-based confidence in your caller choice instead of relying on assumptions from published studies that may not reflect your conditions.

Frequently Asked Questions

What is the main difference between pbsv, Sniffles2, and cuteSV?

The main difference lies in their design priorities. pbsv is the PacBio-supported caller that integrates with the PacBio workflow ecosystem and serves as a production baseline. Sniffles2 emphasizes speed and accuracy, with repeat-aware clustering and coverage-adaptive filtering, and it supports population-level genotyping and mosaic SV detection [<a href="#ref-1">1</a>]. cuteSV offers substantial parameter control and can be paired with phasing-based genotyping improvement tools such as SVUPP [<a href="#ref-2">2</a>].

Which caller should I use for population-level SV genotyping?

Sniffles2 is the strongest choice for population-level genotyping because it was specifically designed to solve family-level to population-level SV calling and produces fully genotyped VCF files [<a href="#ref-1">1</a>]. It was evaluated across 11 probands and accurately identified causative SVs around MECP2, including complex alleles with multiple overlapping SVs [<a href="#ref-1">1</a>].

Can any of these callers detect mosaic structural variants?

Sniffles2 supports the detection of mosaic SVs in bulk long-read data [<a href="#ref-1">1</a>]. The published evaluation identified multiple mosaic SVs in brain tissue from a patient with multiple system atrophy, showing diversity within the cingulate cortex [<a href="#ref-1">1</a>]. pbsv and cuteSV do not have this capability as a primary design feature.

How does coverage depth affect caller choice?

Coverage depth significantly impacts SV calling performance. Sniffles2 was evaluated across 5 to 50x coverage and demonstrated improvements over prior methods across that range [<a href="#ref-1">1</a>]. Lower coverage reduces cost but also reduces sensitivity. A benchmark study showed that a hybrid approach using synthetic long reads achieved comparable F1-scores at 5x coverage to pbsv and Sniffles2 at 10x HiFi coverage [<a href="#ref-3">3</a>]. If your coverage is low, you should expect reduced sensitivity and plan accordingly.

What should I do if my study includes paralogous genes?

You should plan to use a dedicated haplotype-based caller in addition to your general-purpose SV caller. Standard HiFi variant callers detected only 95 of 125 known clinically relevant variants across 11 paralogous loci, while the remaining 30 were only identified by Paraphase [<a href="#ref-4">4</a>]. Combining standard callers with Paraphase detected all known variants, including SNVs, InDels, CNVs, SVs, and gene conversions [<a href="#ref-4">4</a>].

How can I improve genotyping accuracy for cuteSV?

You can pair cuteSV with a phasing-based genotyping improvement tool such as SVUPP. SVUPP incorporates read phasing information into genotype likelihoods and achieved higher accuracy than cuteSV2, Sniffles2, and kanpig for genotyping SVs without close neighbor SVs [<a href="#ref-2">2</a>]. SVUPP can take per-read phasing information from reference panel based phasing methods such as QUILT2 or reference-free phasing methods such as WhatsHap [<a href="#ref-2">2</a>].

What are the computational resource differences between the callers?

Sniffles2 was reported to be 11.8 times faster than prior state-of-the-art callers [<a href="#ref-1">1</a>], which suggests a substantial runtime advantage. However, you should measure runtime and memory for your specific dataset because results depend on coverage, parameters, and genomic region. Record these metrics for each run to inform future decisions.

When should I escalate to specialized tools beyond these three callers?

Escalate when your study involves paralogous regions, pharmacogenomics star-allele calling, mosaic SVs, or clinical adoption. For paralogous regions, use Paraphase [<a href="#ref-4">4</a>]. For pharmacogenomics, use a dedicated star-allele caller such as Aldy 4 [<a href="#ref-5">5</a>]. For mosaic SVs, use Sniffles2 [<a href="#ref-1">1</a>]. For clinical adoption, engage with clinical genomics experts and follow prospective clinical utility study requirements [<a href="#ref-4">4</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Detection of mosaic and population-level structural variants with Sniffles2.](https://pubmed.ncbi.nlm.nih.gov/38168980). Nature biotechnology, 2024. [2] [SVUPP: Pre-phasing long reads improves structural variant genotyping.](https://pubmed.ncbi.nlm.nih.gov/41134129). Bioinformatics (Oxford, England), 2022. [3] [Blackbird: structural variant detection using synthetic and low-coverage long-reads.](https://pubmed.ncbi.nlm.nih.gov/40630502). Bioinformatics advances, 2025. [4] [HiFi sequencing accurately identifies clinically relevant variants in paralogous genes.](https://pubmed.ncbi.nlm.nih.gov/42242190). American journal of human genetics, 2026. [5] [An efficient genotyper and star-allele caller for pharmacogenomics.](https://pubmed.ncbi.nlm.nih.gov/36657977). Genome research, 2023.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.