Detecting Insertions and Deletions in Long-Read Sequencing: A Guide to Identifying Novel and Complex Indels

By Dr. Zubair Khalid, DVM, MS, PhD ·

Detecting Insertions and Deletions in Long-Read Sequencing: A Guide to Identifying Novel and Complex Indels

Key Takeaways

  • Long-read sequencing technologies (Oxford Nanopore, PacBio HiFi) significantly improve insertion and deletion (indel) detection by spanning entire variant regions, resolving ambiguities that confound short-read platforms and enabling the identification of novel and complex indels, including structural variants up to kilobases in length.
  • The choice of sequencing platform is dictated by the target indel size range and required base accuracy; PacBio HiFi excels in detecting small indels (1-20 bp) due to its high consensus accuracy, while Oxford Nanopore is cost-effective for larger structural variants where per-base accuracy is less critical.
  • A robust long-read indel detection workflow necessitates careful alignment strategies using long-read specific aligners like Minimap2, appropriate variant calling tools (alignment-based or clustering-based), and stringent filtering based on mapping quality, read support, and strand bias to minimize false positives.
  • Visualization in genome browsers and manual inspection of read alignments are critical for distinguishing genuine indels from alignment artifacts, particularly in repetitive regions or genes with pseudogene homology, where breakpoint analysis is paramount.
  • Validation of clinically significant findings using orthogonal methods such as PCR amplification across breakpoints or optical genome mapping is essential to confirm indel calls and ensure diagnostic accuracy, especially for variants in known disease genes or those guiding treatment decisions.
  • Reproducibility in long-read indel detection relies on meticulous documentation of all pipeline parameters, tracking quality metrics at each stage, and maintaining detailed analysis logs, while common failure patterns include overfiltering true variants and ignoring reference genome limitations.

Insertions and deletions (indels) are among the most challenging variant classes to detect accurately in genomic sequencing data. Long-read sequencing technologies, including Oxford Nanopore and Pacific Biosciences HiFi, have substantially improved indel detection compared to short-read platforms because individual reads span entire repetitive regions and complex rearrangement breakpoints. This guide provides a systematic approach to indel detection with long reads, covering alignment strategies, variant calling, filtering, visualization, validation, and interpretation. The intended audience includes biology students, researchers, laboratory professionals, and life-science practitioners who need practical decision criteria for building and running indel detection pipelines.

The Indel Detection Problem in Genomic Sequencing

Indels range from single base pair changes to multi-kilobase structural variants. The conventional boundary between small indels and structural variants is approximately 50 base pairs, with larger polymorphisms classified as structural variations that include insertions, deletions, inversions, duplications, and translocations [<a href="#ref-1">1</a>]. This distinction matters for pipeline design because different bioinformatics tools are optimized for different size ranges.

Short-read sequencing platforms generate reads of 150 to 300 base pairs. When an indel exceeds the read length or occurs within a repetitive region, short reads often fail to map uniquely across the variant site. The result is either a false negative call or an alignment artifact that resembles a genuine variant. Long reads, typically 10 kilobases or more for Nanopore and 15 to 25 kilobases for HiFi, span entire variant regions and provide contiguous sequence context that resolves these ambiguities [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

The practical consequence is that long-read sequencing identifies indels that short-read approaches miss entirely. In a study of undiagnosed rare disease cases, HiFi long-read genome sequencing identified disease-causing variants in 11.8 percent of previously unsolved families, with indels and structural variants among the newly detected variant types [<a href="#ref-2">2</a>]. Another study demonstrated that a mobile element insertion in the IQCB1 gene, approximately 6.2 kilobases in length, remained undetected after routine exome and genome sequencing but was revealed through optical genome mapping and confirmed by long-read sequencing [<a href="#ref-3">3</a>]. These cases illustrate the diagnostic gap that long-read indel detection can close.

At a Glance: Long-Read Indel Detection Pipeline Decisions

The table below summarizes the key decisions in a long-read indel detection workflow. Each row corresponds to a pipeline stage where choices materially affect detection performance.

Pipeline StagePrimary OptionsKey Decision CriteriaCommon Pitfall
Sequencing platformOxford Nanopore, PacBio HiFi, optical genome mappingRead length, base accuracy, cost per genome, coverage depthChoosing platform without considering indel size distribution
Read alignmentMinimap2, other long-read alignersMapping quality, computational cost, ability to handle complex rearrangementsUsing short-read aligner parameters on long reads
Variant callingRead alignment-based callers, clustering-based callersIndel size range, coverage depth, genome complexityApplying one caller without benchmarking on similar data
FilteringMapping quality thresholds, read support counts, strand bias filtersFalse positive rate, coverage depth, repetitive regionsFiltering too aggressively and losing true variants
ValidationPCR amplification, orthogonal sequencing, optical mappingVariant size, breakpoint complexity, available samplesSkipping validation for clinically significant findings

Core Principles of Long-Read Indel Detection

Read Length Determines Detectable Variant Size

The fundamental advantage of long reads is that they can span an entire insertion or deletion. A deletion of 5 kilobases is invisible to short reads because no single read covers the breakpoint junction. Long reads of 10 to 20 kilobases easily span such deletions, producing alignment signatures that indicate the precise breakpoint location [<a href="#ref-1">1</a>][<a href="#ref-4">4</a>].

For insertions, the requirement is slightly different. The read must contain the inserted sequence plus sufficient flanking sequence on both sides to map uniquely. A 6.2-kilobase insertion requires reads of at least 8 to 10 kilobases for reliable detection, which explains why the IQCB1 insertion was missed by short-read sequencing [<a href="#ref-3">3</a>].

Base Accuracy Affects Small Indel Calling

Oxford Nanopore sequencing has historically had lower per-base accuracy than PacBio HiFi. This accuracy difference matters most for small indels of 1 to 20 base pairs, where a single sequencing error can create a false positive call. HiFi sequencing achieves high consensus accuracy through circular consensus sequencing, making it well suited for detecting small indels in clinical contexts [<a href="#ref-2">2</a>].

The choice between platforms therefore depends on the target variant size distribution. For projects focused on large structural variants, Nanopore at moderate coverage may suffice. For projects requiring accurate small indel detection, HiFi or high-accuracy Nanopore basecalling is preferable [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Coverage Depth Influences Detection Sensitivity

Coverage depth directly affects the confidence in indel calls. Low coverage reduces the number of reads supporting a variant, increasing the risk of false negatives. Benchmarking studies in crop plant genomes evaluated Oxford Nanopore indel detection at 5x, 10x, and 20x coverage, finding that detection performance varied with coverage and that lower coverage required more careful filtering [<a href="#ref-1">1</a>].

For clinical applications, 10-fold HiFi coverage was used in the Solve-RD study, which identified causative variants in previously undiagnosed rare disease families [<a href="#ref-2">2</a>]. This coverage level represents a practical balance between sequencing cost and detection sensitivity.

Building a Long-Read Indel Detection Workflow

Step 1: Define the Variant Size Range of Interest

Before selecting tools, define the indel size range that matters for the biological question. A project focused on rare disease diagnosis needs to detect variants from single base pairs to multi-kilobase insertions and deletions [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. A project studying crop genome structural variation may focus on variants larger than 50 base pairs [<a href="#ref-1">1</a>].

This definition determines sequencing platform choice, alignment parameters, and variant caller selection. Document the size range in the analysis plan so that downstream filtering decisions are consistent with the project goals.

Step 2: Select the Sequencing Platform and Coverage

Choose between Oxford Nanopore and PacBio HiFi based on the variant size range, required base accuracy, and available budget. For projects requiring detection of both small indels and large structural variants, HiFi provides the best combination of read length and accuracy [<a href="#ref-2">2</a>]. For projects focused primarily on large variants where per-base accuracy is less critical, Nanopore offers lower cost per gigabase [<a href="#ref-1">1</a>].

Set coverage depth based on the detection sensitivity required. Benchmarking data from crop genomes suggests that 10x to 20x coverage provides reliable indel detection for resequencing projects, while lower coverage requires more aggressive filtering and accepts reduced sensitivity [<a href="#ref-1">1</a>].

Step 3: Align Reads to the Reference Genome

Read alignment is the foundation of alignment-based indel detection. Long-read aligners are designed to handle the higher error rates and longer read lengths of Nanopore and HiFi data. The choice of reference genome also matters. The T2T-CHM13 reference genome has resolved many complex regions that were poorly assembled in earlier references, enabling accurate breakpoint characterization in regions that were previously inaccessible [<a href="#ref-5">5</a>].

For repetitive regions, consider whether the reference genome adequately represents the expected sequence. Pseudogene homology can create alignment artifacts that mimic genuine indels, as demonstrated in the PKD1 gene where six pseudogenes with high sequence similarity complicate variant detection [<a href="#ref-5">5</a>].

Step 4: Call Variants with Appropriate Tools

Alignment-based variant callers identify indels by analyzing read alignment patterns. These tools work well for most variant types but may struggle with complex rearrangements or chimeric reads. Clustering-based approaches, such as LcDel, first identify candidate deletion sites and then cluster reads to determine precise deletion boundaries [<a href="#ref-4">4</a>].

The choice of variant caller should be informed by benchmarking on data similar to the project. The crop genome benchmarking study evaluated multiple aligners and callers across different coverage levels, demonstrating that tool performance varies by genome complexity and coverage [<a href="#ref-1">1</a>]. Run a small validation set through candidate callers before committing to a full pipeline.

Step 5: Filter Candidate Variants

Raw variant calls contain false positives from sequencing errors, alignment artifacts, and repetitive regions. Apply filtering criteria that balance sensitivity and precision:

  • Mapping quality thresholds remove variants supported by poorly mapped reads
  • Read support counts require a minimum number of reads supporting each variant
  • Strand bias filters remove variants supported predominantly by one strand
  • Variant context filters flag variants in homopolymer runs or tandem repeats

The optimal thresholds depend on coverage depth and platform error rate. Higher coverage allows more stringent read support thresholds without losing true variants [<a href="#ref-1">1</a>].

Step 6: Visualize and Manually Inspect Candidate Variants

Visualization is essential for distinguishing genuine indels from alignment artifacts. View candidate variants in a genome browser with aligned reads displayed. Genuine indels show consistent read support across multiple reads, with breakpoints that align to the expected genomic coordinates. Artifacts often show inconsistent breakpoints, partial read support, or alignment patterns that suggest mismapping.

For complex variants, such as the LINE-1/ERV1 insertion in IQCB1, visualization reveals the inserted sequence structure and enables breakpoint characterization [<a href="#ref-3">3</a>]. Manual inspection is particularly important for variants in repetitive regions or genes with pseudogene homology [<a href="#ref-5">5</a>].

Step 7: Validate with Orthogonal Methods

Validation is required for clinically significant findings and recommended for research findings that will be reported. PCR amplification across the variant breakpoint provides direct evidence of the variant. Long-range PCR followed by sequencing can characterize large insertions and deletions [<a href="#ref-5">5</a>].

Orthogonal sequencing approaches provide independent confirmation. Optical genome mapping can detect large insertions and deletions without sequencing bias, making it useful for validating variants that are difficult to confirm by PCR [<a href="#ref-3">3</a>]. For small indels, Sanger sequencing across the variant site provides definitive validation.

Alignment Strategies for Complex Indel Detection

Choosing the Right Aligner

Long-read aligners differ in their handling of insertions, deletions, and complex rearrangements. The choice of aligner affects both sensitivity and precision of downstream variant calling. Benchmarking studies have shown that aligner performance varies by genome complexity, with polyploid genomes presenting additional challenges [<a href="#ref-1">1</a>].

For most projects, a single aligner with well-tested parameters is sufficient. However, for complex genomes or challenging variant types, running two aligners and comparing results can identify alignment-dependent artifacts.

Reference Genome Considerations

The reference genome provides the coordinate system for variant calling. Reference genome quality directly affects indel detection accuracy. Regions that are misassembled or missing from the reference produce spurious variant calls or hide genuine variants.

The T2T-CHM13 reference genome has improved indel detection in complex regions, including the PKD1 gene where breakpoints in an intronic AG-repeat could only be correctly characterized by aligning to this reference [<a href="#ref-5">5</a>]. For clinical applications, consider whether the reference genome adequately represents the population being studied.

Handling Repetitive Regions

Repetitive regions present the greatest challenge for indel detection. Reads from different repeat copies map equally well to multiple locations, creating ambiguity in variant calling. Long reads reduce this ambiguity by spanning entire repeat units, but complex repeat structures remain problematic.

For genes with pseudogene homology, such as PKD1, targeted approaches may be necessary. The CAPKD assay uses long-range PCR to amplify the gene of interest specifically, avoiding pseudogene contamination and enabling highly specific variant detection [<a href="#ref-5">5</a>].

Variant Calling Approaches and Their Tradeoffs

Alignment-Based Calling

Alignment-based variant callers analyze read alignments to identify positions where reads consistently disagree with the reference. These callers work well for small indels and simple structural variants. They require accurate alignments and sufficient read depth to distinguish genuine variants from sequencing errors.

The crop genome benchmarking study evaluated alignment-based callers for Oxford Nanopore data, finding that performance varied by coverage and genome complexity [<a href="#ref-1">1</a>]. For polyploid genomes, alignment-based calling requires careful handling of multi-mapping reads.

Clustering-Based Calling

Clustering-based approaches, such as LcDel, identify candidate variant sites and then cluster supporting reads to determine precise breakpoints. These methods can handle chimeric reads and complex variants that confuse alignment-based callers [<a href="#ref-4">4</a>].

The LcDel approach uses two clustering methods based on deletion length, followed by hierarchical clustering to determine deletion location and length [<a href="#ref-4">4</a>]. This multi-stage approach improves precision for deletion detection compared to single-stage methods.

Hybrid Approaches

Some pipelines combine multiple calling strategies to improve overall detection. For example, running both an alignment-based caller and a clustering-based caller, then intersecting or unioning the results, can improve sensitivity or precision depending on the intersection strategy.

The choice of combination strategy depends on the project goals. For clinical applications where false positives are costly, intersecting results from multiple callers reduces false positives at the cost of some sensitivity. For discovery projects where sensitivity is paramount, unioning results captures more variants but requires more extensive filtering.

Filtering Strategies to Reduce False Positives

Mapping Quality Filters

Mapping quality scores indicate the confidence that a read is correctly placed in the reference genome. Low mapping quality reads often produce spurious variant calls, particularly in repetitive regions. Set a minimum mapping quality threshold for reads supporting variant calls.

The optimal threshold depends on the genome complexity and the tolerance for false positives. In repetitive genomes, higher thresholds reduce false positives but may miss genuine variants in repeat regions.

Read Support Filters

Require a minimum number of reads supporting each variant call. The minimum depends on coverage depth and expected variant frequency. For germline variants at 10x coverage, at least 3 to 4 supporting reads provide reasonable confidence. For somatic variants or low-frequency variants, lower read support may be acceptable with appropriate statistical modeling.

The crop genome benchmarking study demonstrated that read support thresholds affect detection performance, with lower coverage requiring more careful threshold selection [<a href="#ref-1">1</a>].

Strand Bias Filters

Sequencing errors are often strand-specific, producing variants supported predominantly by reads from one strand. Genuine variants typically show support from both strands. Apply strand bias filters to remove variants with extreme strand imbalance.

Context-Specific Filters

Certain sequence contexts produce systematic errors in specific sequencing platforms. Homopolymer runs are problematic for Nanopore sequencing, producing false indel calls. Tandem repeats can cause alignment artifacts that mimic indels. Apply context-specific filters based on the known error profile of the sequencing platform.

Visualization and Manual Inspection

Genome Browser Visualization

Visualize candidate variants in a genome browser with aligned reads displayed. The Integrative Genomics Viewer and similar tools provide interactive visualization of read alignments. Genuine indels show consistent read support with clear breakpoints. Artifacts show inconsistent patterns that are difficult to interpret.

For complex variants, visualization reveals the structure of the variant and enables breakpoint characterization. The IQCB1 insertion was characterized by visualizing the inserted sequence and determining that it consisted of a LINE-1/ERV1 mobile element [<a href="#ref-3">3</a>].

Breakpoint Analysis

Examine the precise breakpoints of candidate variants. Genuine indels have breakpoints that are consistent across supporting reads. Artifacts often show variable breakpoints or breakpoints that fall in low-complexity sequence.

For deletions, the breakpoints define the deleted sequence. For insertions, the breakpoints define the insertion site and the inserted sequence can be extracted for further analysis. The CAPKD assay characterized breakpoints of large deletions and duplications in PKD1, including breakpoints in an intronic AG-repeat that required the T2T-CHM13 reference for correct characterization [<a href="#ref-5">5</a>].

Manual Review Criteria

Establish criteria for manual review of candidate variants. Variants that meet any of the following criteria warrant manual inspection:

  • Variants in known disease genes or genes of biological interest
  • Variants with borderline read support or mapping quality
  • Variants in repetitive regions or regions with pseudogene homology
  • Variants with complex breakpoint patterns
  • Variants that are candidates for clinical reporting

Validation Using PCR and Orthogonal Methods

PCR Validation

PCR amplification across the variant breakpoint provides direct evidence of the variant. Design primers that flank the predicted breakpoint and amplify the variant allele. The presence of an amplification product of the expected size confirms the variant.

For large insertions, long-range PCR can amplify the entire inserted sequence for characterization. The CAPKD assay uses long-range PCR to amplify PKD1 specifically, avoiding pseudogene contamination [<a href="#ref-5">5</a>].

Orthogonal Sequencing Validation

Sequencing the variant region with an independent technology provides strong validation. Sanger sequencing across small indels confirms the variant sequence. Short-read sequencing can validate variants that are detectable by both platforms.

For large structural variants, optical genome mapping provides an independent validation method that does not rely on sequencing [<a href="#ref-3">3</a>]. This approach is particularly useful for validating insertions and deletions that are difficult to amplify by PCR.

Validation Decision Criteria

Not all variants require validation. Establish criteria for when validation is necessary:

  • All variants that will be reported clinically
  • Variants that will guide treatment decisions
  • Variants in genes with established disease associations
  • Variants with unusual or complex structures
  • Variants that will be published or shared in public databases

For research projects where variants will not be reported, validation may be limited to a random sample for quality assessment.

Records and Measurements for Reproducible Indel Detection

Documenting Pipeline Parameters

Reproducible indel detection requires complete documentation of pipeline parameters. Record the following for each analysis:

  • Sequencing platform and basecalling version
  • Read alignment tool and version, with all parameters
  • Variant calling tool and version, with all parameters
  • Filtering thresholds and rationale
  • Reference genome version and source

The nf-core documentation provides standards for reproducible pipeline configuration, emphasizing the importance of version control and parameter documentation [<a href="#ref-6">6</a>]. Following these standards ensures that analyses can be reproduced and compared across projects.

Tracking Quality Metrics

Track quality metrics at each pipeline stage to identify problems early. Key metrics include:

  • Read length distribution and N50
  • Alignment rate and mapping quality distribution
  • Coverage depth and uniformity
  • Variant call count before and after filtering
  • Transition-transversion ratio for small variants
  • Indel size distribution

Compare these metrics to expected values for the sequencing platform and genome. Deviations from expected values indicate potential problems that require investigation.

Maintaining Analysis Logs

Maintain detailed logs of analysis runs, including software versions, parameters, and output file locations. The Carpentries lessons emphasize the importance of reproducible data analysis practices, including documentation and version control [<a href="#ref-7">7</a>]. These practices are essential for long-read indel detection where pipeline choices materially affect results.

Common Failure Patterns in Long-Read Indel Detection

Failure Pattern 1: Overfiltering True Variants

Aggressive filtering removes true variants along with false positives. This pattern is common when researchers apply thresholds optimized for short-read data to long-read data without adjustment. Long-read data has different error profiles and requires different filtering thresholds.

Prevention: Benchmark filtering thresholds on a validation set with known variants. Adjust thresholds based on the observed tradeoff between sensitivity and precision.

Failure Pattern 2: Underfiltering False Positives

Insufficient filtering produces high false positive rates that overwhelm downstream analysis. This pattern is common when researchers apply minimal filtering to maximize sensitivity.

Prevention: Apply multiple filtering criteria, including mapping quality, read support, and strand bias. Visualize a sample of calls to assess the false positive rate.

Failure Pattern 3: Ignoring Reference Genome Limitations

Reference genome gaps and errors produce spurious variant calls. This pattern is common in complex genomic regions that are poorly represented in the reference.

Prevention: Use the most complete reference genome available, such as T2T-CHM13 for human data [<a href="#ref-5">5</a>]. Consider whether reference genome limitations explain unusual variant patterns in specific regions.

Failure Pattern 4: Applying One Caller Without Benchmarking

Different variant callers have different strengths and weaknesses. Applying a single caller without benchmarking on similar data can miss variants that the caller does not detect well.

Prevention: Benchmark candidate callers on data similar to the project, as demonstrated in the crop genome study [<a href="#ref-1">1</a>]. Select callers based on observed performance, beyond popularity.

Failure Pattern 5: Skipping Validation for Clinically Significant Findings

Reporting unvalidated variants can lead to incorrect clinical decisions. This pattern is common when researchers prioritize throughput over accuracy.

Prevention: Establish validation criteria before beginning the analysis. Validate all variants that will be reported clinically or guide treatment decisions.

Limitations of Long-Read Indel Detection

Detection Limits in Complex Regions

Despite the advantages of long reads, some genomic regions remain difficult for indel detection. Highly repetitive regions, such as centromeres and segmental duplications, can produce ambiguous alignments even with long reads. Genes with pseudogene homology require targeted approaches for reliable variant detection [<a href="#ref-5">5</a>].

Coverage Requirements

Long-read sequencing at high coverage remains more expensive than short-read sequencing. The coverage required for reliable indel detection depends on the variant size range and the tolerance for false negatives. Projects with limited budgets may need to accept reduced sensitivity in exchange for lower coverage [<a href="#ref-1">1</a>].

Bioinformatics Expertise Requirements

Long-read indel detection requires substantial bioinformatics expertise. Researchers need proficiency in command-line tools, data management, and statistical interpretation. Training resources from EMBL-EBI and The Carpentries provide foundational skills for genomic data analysis [<a href="#ref-8">8</a>][<a href="#ref-7">7</a>].

Platform-Specific Error Profiles

Each sequencing platform has characteristic error profiles that affect indel detection. Oxford Nanopore has higher error rates in homopolymer regions, producing false indel calls. PacBio HiFi has lower error rates but may miss variants in regions with extreme GC content. Understanding platform-specific errors is essential for accurate variant calling.

Safety and Regulatory Context for Clinical Applications

Clinical Validation Requirements

Indel detection for clinical applications requires validation according to established standards. Variants that will be reported clinically must be confirmed by an orthogonal method. The CAPKD assay demonstrated this approach by comparing long-read sequencing results to next-generation sequencing and multiplex ligation-dependent probe amplification [<a href="#ref-5">5</a>].

Data Sharing and Privacy

Genomic data sharing must comply with applicable privacy regulations. Public databases such as NCBI provide resources for sharing and accessing genomic data [<a href="#ref-9">9</a>]. Researchers must ensure that data sharing complies with consent requirements and privacy protections.

Reporting Standards

Clinical reports should include sufficient information for interpretation, including the variant coordinates, variant type, validation status, and clinical significance. The Solve-RD study demonstrated the importance of clinical interpretation in converting sequencing findings into diagnoses [<a href="#ref-2">2</a>].

Professional Escalation Criteria

When to Consult a Bioinformatics Specialist

Seek specialist consultation when:

  • The analysis requires custom pipeline development beyond standard tools
  • Variant calls in clinically significant genes are ambiguous
  • The reference genome does not adequately represent the study population
  • Multiple callers produce discordant results that cannot be resolved by filtering

When to Escalate to a Clinical Genetics Team

Escalate to clinical genetics when:

  • A variant in a known disease gene is identified
  • A candidate variant requires clinical interpretation
  • Validation results are discordant with sequencing findings
  • The variant has implications for family members who may need testing

When to Consider Alternative Technologies

Consider alternative technologies when:

  • Long-read sequencing fails to resolve a suspected variant
  • The variant is in a region that is refractory to sequencing
  • Breakpoint characterization requires higher resolution than sequencing provides
  • Optical genome mapping or other orthogonal approaches may provide additional information [<a href="#ref-3">3</a>]

A Practical Decision Framework for Triaging Indel Candidates by Validation Priority

The preceding sections describe the technical components of long-read indel detection, but researchers often face a more immediate operational problem: how to allocate limited validation resources across dozens or hundreds of candidate variants. PCR validation, orthogonal sequencing, and manual inspection are time-consuming and expensive. A systematic triage framework helps prioritize which candidates warrant validation first, which require only computational scrutiny, and which can be safely deprioritized. This section provides a decision framework grounded in the evidence from clinical and agricultural long-read studies, along with a record system for tracking validation outcomes and troubleshooting methods for discordant results.

Establishing Validation Priority Tiers

Before examining individual variants, define three validation tiers that correspond to the consequences of a false positive or false negative call. Tier 1 includes variants in genes with established disease associations, variants that will guide treatment decisions, and variants that will be reported clinically. The Solve-RD study demonstrated that clinical interpretation and orthogonal validation of variants in known disease genes yielded 12 novel genetic diagnoses from de novo and rare inherited variants [<a href="#ref-2">2</a>]. These findings required rigorous confirmation because they directly affected patient care. Tier 2 includes variants in candidate genes or regions of biological interest where a false positive would waste downstream functional work but would not directly harm a patient. Tier 3 includes variants in intergenic regions, repetitive elements, or genes without known phenotypic associations where the cost of a false positive is minimal.

Assign each candidate variant to a tier before running validation experiments. This assignment prevents the common failure pattern where researchers validate an arbitrary subset of variants based on convenience instead of consequence. The tier assignment also determines the validation method. Tier 1 variants require orthogonal confirmation, meaning a method independent from the original sequencing platform. The CAPKD assay compared long-read sequencing results to next-generation sequencing and multiplex ligation-dependent probe amplification, providing a model for multi-platform confirmation [<a href="#ref-5">5</a>]. Tier 2 variants may be confirmed by PCR alone, while Tier 3 variants may require no wet-lab validation if computational evidence is strong.

Scoring Variants by Confidence and Consequence

Within each tier, score variants on two axes: call confidence and biological consequence. Call confidence incorporates mapping quality, read support count, strand balance, and breakpoint consistency across supporting reads. Biological consequence incorporates gene function, variant type, predicted protein effect, and population frequency. A variant with high call confidence and high biological consequence receives the highest validation priority. A variant with low call confidence and low biological consequence receives the lowest priority.

The scoring system should be documented before analysis begins. Use a simple numeric scale from 1 to 5 for each axis, with 5 representing the highest confidence or consequence. Multiply the two scores to produce a priority index from 1 to 25. Variants with a priority index above a threshold that you define based on your project goals proceed to validation. Variants below the threshold receive computational review only. This approach prevents the common failure pattern of overfiltering true variants, because the priority index separates the filtering decision from the validation decision.

Building a Validation Decision Tree

A decision tree provides a structured path from candidate variant to validated finding. Start with the variant size. For small indels under 50 base pairs, PCR amplification across the variant site followed by Sanger sequencing provides definitive confirmation. For larger indels, long-range PCR can amplify the entire variant region for characterization, as demonstrated in the CAPKD assay for PKD1 [<a href="#ref-5">5</a>]. For insertions larger than the PCR amplification limit, consider whether the inserted sequence can be characterized by targeted long-read sequencing or whether optical genome mapping provides a better validation path [<a href="#ref-3">3</a>].

The decision tree branches on variant location. Variants in unique regions are straightforward to validate by PCR because primers can be designed to flank the variant site specifically. Variants in repetitive regions or genes with pseudogene homology require additional care. The PKD1 gene has six pseudogenes with high sequence similarity, which complicates both detection and validation [<a href="#ref-5">5</a>]. For such regions, design primers that anneal to unique flanking sequence and verify primer specificity by in silico PCR against the reference genome before ordering primers.

The decision tree also branches on the availability of additional samples. If family members or biological replicates are available, segregation analysis provides powerful validation evidence. A variant that segregates with the phenotype across affected family members is more likely to be causal than a variant found in a single individual. The Solve-RD study included families with multiple affected individuals, enabling segregation analysis for candidate variants [<a href="#ref-2">2</a>].

Recording Validation Outcomes Systematically

A structured record system transforms validation from an informal process into an auditable workflow. For each candidate variant, record the following fields in a spreadsheet or database:

  • Variant identifier and genomic coordinates
  • Gene name and variant type
  • Validation tier and priority index
  • Validation method used
  • Validation date and operator
  • Primer sequences and PCR conditions
  • Validation result, including gel images or sequencing traces
  • Discordance notes if validation contradicts the original call
  • Final classification, such as confirmed, refuted, or inconclusive

The nf-core documentation emphasizes the importance of version control and parameter documentation for reproducible pipelines [<a href="#ref-6">6</a>]. Apply the same principle to validation records. Store primer sequences, PCR protocols, and analysis scripts in version control so that any validation result can be traced back to the exact conditions under which it was generated.

Track validation outcomes by tier and variant type. Calculate the confirmation rate for each tier, which is the proportion of validated variants that were confirmed as genuine. A low confirmation rate in Tier 1 indicates that the computational filtering is insufficient and that more stringent thresholds are needed before validation. A high confirmation rate in Tier 3 suggests that the pipeline may be missing true variants that are being filtered out before reaching the validation stage.

Troubleshooting Discordant Validation Results

Discordance between the computational call and the validation result requires systematic investigation instead of immediate acceptance of either outcome. The most common cause of discordance is primer design failure. Primers that anneal to repetitive sequence or that span the variant breakpoint may fail to amplify the variant allele. Check primer specificity by aligning primer sequences to the reference genome and to the variant sequence. Redesign primers in unique flanking regions if necessary.

A second cause of discordance is sample mix-up or contamination. If validation fails for a variant that had strong computational support, verify that the DNA sample used for validation matches the sample used for sequencing. Check sample tracking records and consider repeating the validation with a fresh aliquot.

A third cause is reference genome limitations. The T2T-CHM13 reference genome resolved complex regions that were poorly assembled in earlier references, enabling breakpoint characterization that was previously impossible [<a href="#ref-5">5</a>]. If validation fails in a region known to be problematic in the reference genome, consider whether the reference genome version used for alignment adequately represents the true sequence. Realigning to a different reference version may resolve the discordance.

A fourth cause is that the computational call is correct but the validation method is insufficiently sensitive. For example, PCR may fail to amplify a large insertion because the polymerase cannot process the entire region. In this case, consider an alternative validation method such as optical genome mapping, which detected a 6.2-kilobase insertion in IQCB1 that was missed by routine sequencing [<a href="#ref-3">3</a>].

Establishing Discordance Escalation Criteria

Define criteria for escalating discordant results to a specialist or to an alternative technology. Escalate when a Tier 1 variant shows discordance that cannot be resolved by primer redesign or sample verification. Escalate when multiple variants in the same genomic region show discordance, suggesting a regional problem instead of a variant-specific issue. Escalate when the discordance involves a variant that would change a clinical diagnosis or treatment decision.

The escalation pathway should include a bioinformatics specialist who can review the alignment and variant calling parameters, and a clinical genetics team if the variant has patient implications. The CAPKD study demonstrated the value of multi-disciplinary review by comparing results across multiple technologies and resolving discrepancies through careful analysis [<a href="#ref-5">5</a>].

Integrating the Decision Framework into the Pipeline

The decision framework operates after variant calling and filtering but before final reporting. Integrate it into the pipeline as a distinct stage with defined inputs and outputs. The input is the filtered variant call set. The output is a prioritized validation list with assigned tiers, priority scores, and validation methods.

Document the framework in the analysis plan so that all team members apply the same criteria. The Galaxy Training Network provides accessible workflow training that emphasizes the importance of structured analysis steps [<a href="#ref-10">10</a>]. Apply the same structured approach to the validation stage as to the alignment and calling stages.

Measuring Framework Performance

Track performance metrics for the decision framework itself. The primary metric is the validation confirmation rate by tier. A well-calibrated framework should show high confirmation rates in Tier 1 and progressively lower rates in Tiers 2 and 3. If Tier 1 confirmation rates are low, the computational filtering is insufficient. If Tier 3 confirmation rates are unexpectedly high, the framework may be assigning too many variants to low-priority tiers.

A secondary metric is the time from variant call to validated finding. The framework should reduce the time to validated findings by directing resources to the most consequential variants first. Track the median time to validation for each tier and adjust the framework if Tier 1 variants are not being validated promptly.

A third metric is the rate of discordance that requires escalation. A high escalation rate suggests that the validation methods are mismatched to the variant types being called. For example, if many large insertions require escalation to optical genome mapping because PCR fails, consider whether optical genome mapping should be a first-line validation method for large insertions instead of a rescue method [<a href="#ref-3">3</a>].

Common Failure Patterns in Validation Prioritization

The most common failure pattern is validating an arbitrary subset of variants instead of a prioritized subset. This pattern occurs when researchers lack a documented framework and default to validating variants that are easiest to amplify or that appear first in the call set. The result is that consequential variants may remain unvalidated while low-priority variants consume validation resources.

A second failure pattern is overvalidating Tier 3 variants. Researchers may feel compelled to validate every variant that passes filtering, even when the variant has no known biological consequence. This pattern wastes resources and delays validation of Tier 1 variants. The tier system exists to prevent this pattern by explicitly acknowledging that not all variants require the same level of confirmation.

A third failure pattern is ignoring discordance. When validation contradicts the computational call, researchers may accept the validation result without investigating the cause. This pattern can hide systematic problems in the pipeline, such as a reference genome artifact or a primer design issue that affects multiple variants. Investigate every discordance to determine whether it represents a single-variant problem or a systemic issue.

A fourth failure pattern is failing to update the framework based on validation outcomes. The framework should be a living document that evolves as validation data accumulate. If Tier 1 confirmation rates are consistently low, adjust the filtering thresholds or the tier assignment criteria. If certain variant types consistently fail validation, revise the validation method for those types.

Practical Implementation Steps

Implement the decision framework in six steps. First, define the validation tiers and document the assignment criteria. Second, establish the scoring system for call confidence and biological consequence. Third, build the decision tree that maps variant characteristics to validation methods. Fourth, create the record system for tracking validation outcomes. Fifth, define the escalation criteria for discordant results. Sixth, integrate the framework into the pipeline as a distinct stage with documented inputs and outputs.

The implementation should be tested on a small set of variants before full deployment. Run the framework on 10 to 20 variants that span the range of sizes, locations, and confidence levels expected in the full dataset. Verify that the framework assigns sensible priorities and that the validation methods are feasible for the available samples and equipment. Adjust the framework based on this pilot test before applying it to the full variant call set.

Relationship to Reproducible Analysis Standards

The decision framework aligns with reproducible analysis standards promoted by community resources. The nf-core documentation emphasizes version control and parameter documentation for reproducible pipelines [<a href="#ref-6">6</a>]. Apply the same standards to the validation stage by recording all validation parameters and outcomes. The Carpentries lessons emphasize foundational data analysis practices including documentation and version control [<a href="#ref-7">7</a>]. These practices are essential for the validation stage, where manual steps are more common and therefore more prone to undocumented variation.

The framework also aligns with training resources that emphasize structured analysis approaches. The EMBL-EBI training portal provides learning pathways for bioinformatics data analysis [<a href="#ref-8">8</a>]. The Galaxy Training Network provides accessible workflow training that emphasizes reproducible analysis steps [<a href="#ref-10">10</a>]. These resources can help team members understand the rationale for the decision framework and apply it consistently.

Limitations of the Decision Framework

The decision framework does not eliminate the need for expert judgment. The scoring system provides a structured starting point, but experienced researchers may adjust priorities based on knowledge that is not captured in the numeric scores. Document any such adjustments and the rationale so that the framework can be refined over time.

The framework assumes that validation resources are limited and that prioritization is necessary. Projects with abundant validation resources may not need the full tier system and can validate all candidates. However, even well-funded projects benefit from the record system and the troubleshooting methods, which improve the quality of validation regardless of the number of variants processed.

The framework does not address the choice of validation method for every possible variant type. The decision tree covers the common cases of small indels, large deletions, and large insertions, but unusual variants such as complex rearrangements or mobile element insertions may require bespoke validation approaches. The IQCB1 insertion required a combination of optical genome mapping, long-read sequencing, and targeted RNA sequencing for full characterization [<a href="#ref-3">3</a>]. Such cases require expert consultation and may not fit neatly into a standardized framework.

Professional Escalation Criteria for Validation Discordance

Escalate to a bioinformatics specialist when discordance persists after primer redesign and sample verification. Escalate when the discordance involves a variant in a gene with established disease associations and the validation result would change the clinical interpretation. Escalate when multiple variants in the same region show discordance, suggesting a regional artifact. Escalate when the discordance cannot be resolved with available validation methods and alternative technologies such as optical genome mapping may provide additional information [<a href="#ref-3">3</a>].

Escalate to a clinical genetics team when a Tier 1 variant shows discordance that affects patient care. The clinical team can assess whether the variant should be reported based on the totality of evidence, including the computational call, the validation result, and the clinical context. The Solve-RD study demonstrated the importance of clinical interpretation in converting sequencing findings into diagnoses [<a href="#ref-2">2</a>]. The same principle applies to discordant validation results, where clinical context may determine whether a variant is reported despite an inconclusive validation outcome.

Frequently Asked Questions

What size indels can long-read sequencing detect that short reads cannot?

Long reads can detect indels that exceed the read length of short-read platforms. A 6.2-kilobase insertion in the IQCB1 gene was detected by long-read sequencing after being missed by exome and genome sequencing [<a href="#ref-3">3</a>]. The practical limit depends on read length, with longer reads enabling detection of larger variants. For deletions, the read must span the entire deleted region. For insertions, the read must contain the inserted sequence plus sufficient flanking sequence for unique alignment.

How much coverage is needed for reliable indel detection with long reads?

Coverage requirements depend on the variant size range and the tolerance for false negatives. Benchmarking in crop genomes evaluated 5x, 10x, and 20x coverage, finding that detection performance varied with coverage [<a href="#ref-1">1</a>]. Clinical studies have used 10-fold HiFi coverage for variant detection in undiagnosed rare disease cases [<a href="#ref-2">2</a>]. Higher coverage improves sensitivity but increases cost, so the choice depends on project goals and budget.

Which is better for indel detection, Oxford Nanopore or PacBio HiFi?

The choice depends on the variant size range and accuracy requirements. HiFi provides higher per-base accuracy, making it better for small indel detection [<a href="#ref-2">2</a>]. Nanopore offers lower cost per gigabase and very long reads, making it suitable for large structural variant detection [<a href="#ref-1">1</a>]. For projects requiring both small and large indel detection, HiFi provides the best combination of read length and accuracy.

How do I distinguish genuine indels from alignment artifacts?

Genuine indels show consistent read support across multiple reads with clear breakpoints. Artifacts often show inconsistent breakpoints, strand bias, or support from poorly mapped reads. Visualization in a genome browser is essential for distinguishing genuine variants from artifacts. Apply filtering criteria including mapping quality, read support, and strand bias, then manually inspect candidate variants.

Why are indels in repetitive regions difficult to detect even with long reads?

Repetitive regions create ambiguity in read alignment because reads from different repeat copies map equally well to multiple locations. Long reads reduce this ambiguity by spanning entire repeat units, but complex repeat structures remain problematic. Genes with pseudogene homology, such as PKD1, require targeted approaches for reliable variant detection [<a href="#ref-5">5</a>].

What validation methods are appropriate for confirming long-read indel calls?

PCR amplification across the variant breakpoint provides direct evidence of the variant. Long-range PCR can amplify large insertions for characterization [<a href="#ref-5">5</a>]. Orthogonal sequencing approaches, including Sanger sequencing and short-read sequencing, provide independent confirmation. Optical genome mapping provides validation for large structural variants without sequencing bias [<a href="#ref-3">3</a>].

How do I choose the right variant caller for my long-read data?

Benchmark candidate callers on data similar to your project. The crop genome study evaluated multiple aligners and callers across different coverage levels, demonstrating that performance varies by genome complexity and coverage [<a href="#ref-1">1</a>]. Consider the variant size range of interest, the genome complexity, and the tolerance for false positives and false negatives.

What should I do if long-read sequencing does not identify a suspected variant?

Consider whether the variant is in a region that is refractory to sequencing or alignment. Evaluate whether the reference genome adequately represents the region. Consider alternative technologies, such as optical genome mapping, which can detect variants that sequencing approaches miss [<a href="#ref-3">3</a>]. Consult with a bioinformatics specialist or clinical genetics team for complex cases.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Benchmarking Oxford Nanopore read alignment-based insertion and deletion detection in crop plant genomes.](https://pubmed.ncbi.nlm.nih.gov/36988043). The plant genome, 2023. [2] [Unraveling undiagnosed rare disease cases by HiFi long-read genome sequencing.](https://pubmed.ncbi.nlm.nih.gov/40138663). Genome research, 2025. [3] [Long-read technologies identify a hidden LINE-1/ERV1 insertion in IQCB1 as causative variant for Senior-Løken syndrome.](https://pubmed.ncbi.nlm.nih.gov/40263280). NPJ genomic medicine, 2025. [4] [LcDel: deletion variation detection based on clustering and long reads.](https://pubmed.ncbi.nlm.nih.gov/38798694). Frontiers in genetics, 2024. [5] [Comprehensive Analysis of PKD1 and PKD2 by Long-Read Sequencing in Autosomal Dominant Polycystic Kidney Disease.](https://pubmed.ncbi.nlm.nih.gov/38527221). Clinical chemistry, 2024. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [8] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [9] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [10] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.