Copy Number Variant Calling from Long Reads: Strategies for Accurate Read-Depth and Split-Read Analysis

By Dr. Zubair Khalid, DVM, MS, PhD ·

Copy Number Variant Calling from Long Reads: Strategies for Accurate Read-Depth and Split-Read Analysis

Key Takeaways

  • Long-read sequencing offers superior resolution for Copy Number Variant (CNV) calling by spanning repetitive regions and capturing variant breakpoints within single molecules, a significant advantage over short-read methods.
  • A hybrid analytical approach, integrating both read-depth (for large events) and split-read (for precise breakpoints) strategies, is essential for comprehensive and accurate CNV detection across a broad size spectrum.
  • Read-depth analysis requires rigorous normalization to account for GC bias, mappability variations, and repeat content, with window size selection representing a critical trade-off between resolution and statistical power.
  • Split-read analysis leverages alignment signatures to identify exact breakpoints, offering nucleotide-level resolution, but is limited by the read length's ability to span the entire variant.
  • Reconciliation of read-depth and split-read outputs is crucial, with agreement between methods indicating high confidence, while discrepancies highlight variant complexity and method limitations requiring further investigation.
  • Systematic record-keeping of sample metadata, alignment statistics, tool versions, parameters, and call-level details is fundamental for reproducibility and troubleshooting in long-read CNV analysis.

Copy number variant (CNV) calling from long-read sequencing data requires a different analytical framework than the methods optimized for short-read platforms. Long reads provide two distinct advantages that change the analytical calculus: they span repetitive regions that confound short-read alignment, and they often contain the variant breakpoint within a single sequencing molecule. This article addresses the specific problem of how to combine read-depth and split-read strategies to achieve accurate CNV resolution from Pacific Biosciences HiFi and Oxford Nanopore data. The practical outcome is a workflow that leverages the complementary strengths of both signal types, with concrete guidance on normalization, breakpoint refinement, and quality control.

The core challenge for bioinformaticians working with long-read CNV data is that no single tool captures the full spectrum of variant sizes and types. Read-depth methods excel at detecting large copy number changes across broad genomic regions but provide poor breakpoint resolution. Split-read methods identify exact breakpoints but struggle with larger events where both breakpoints do not fall within a single read. A hybrid approach that runs both analyses and reconciles their outputs produces more reliable calls than either method alone. This article provides the analytical framework for implementing such a hybrid workflow, including specific tool configurations, normalization strategies, and validation steps.

At a Glance

The table below summarizes the key decisions required when building a long-read CNV calling workflow. These choices determine the sensitivity, specificity, and resolution of the final variant calls.

Workflow DecisionRead-Depth ApproachSplit-Read ApproachHybrid Integration
Primary signalCoverage counts across genomic windowsSplit alignments and discordant read pairsCombined evidence from both signals
Best variant size rangeLarge events, typically greater than 1 kbSmall to medium events, typically less than 10 kbFull spectrum with size-dependent weighting
Breakpoint resolutionLow, limited by window sizeHigh, nucleotide-level when breakpoints are spannedHigh for split-read events, window-limited for large events
Main normalization challengeGC bias, mappability variation, repeat contentAlignment artifacts, low-complexity sequenceCoordinating thresholds across tools
Recommended validationIndependent coverage confirmationPCR or targeted sequencing confirmationOrthogonal method confirmation for clinical reporting
Primary tool exampleCNVpytor with long-read modeSniffles2 with tandem repeat handlingCustom reconciliation script or manual curation

The hybrid approach requires running both analyses on the same alignment file, then comparing calls that overlap in genomic coordinates. When read-depth and split-read methods agree on a variant, confidence is high. When they disagree, the discrepancy itself provides useful information about the variant structure and the limitations of each method.

Long-Read Sequencing Context for CNV Analysis

Long-read sequencing platforms produce reads that range from 10 to 100 kb or more, fundamentally changing what is detectable in genomic data. The National Center for Biotechnology Information maintains the primary sequence databases and analysis resources that support variant interpretation, and their documentation describes the data structures and alignment formats that underpin CNV analysis workflows. Understanding these data resources is essential for researchers who need to access reference genomes, variant databases, and comparative genomics tools.

The practical advantage of long reads for CNV detection stems from their ability to span repetitive elements and structural variation breakpoints. A read that covers an entire duplicated region or that contains a deletion junction provides direct evidence of the variant structure. This contrasts with short-read approaches that must infer structural variants from coverage anomalies and paired-end distance signatures. The European Bioinformatics Institute offers training materials on sequence analysis that cover the alignment and variant calling concepts relevant to both short-read and long-read data.

For clinical applications, the improved resolution of long-read sequencing has demonstrated measurable benefits. A 2024 study in Genome Research examined 96 probands with rare diseases who had negative short-read sequencing results and found that long-read genome sequencing identified pathogenic or likely pathogenic variants in approximately 9.4% of cases. Among the variants only correctly interpreted in long-read data were copy number variants, an inversion, a mobile element insertion, two low-complexity repeat expansions, and a single base pair deletion. These findings illustrate that long-read data can resolve variant classes that short-read pipelines systematically miss or mischaracterize.

The same study noted that some variants were accurately called in both short-read and long-read data but required reinterpretation due to recently published gene-disease associations. This observation highlights that CNV calling accuracy depends on the sequencing platform and on the annotation and interpretation infrastructure that follows variant detection. Researchers should maintain current variant databases and periodically reanalyze existing data as new disease associations are published.

Core Principles of Read-Depth CNV Calling

Read-depth analysis for CNV detection operates on a simple premise: the number of sequencing reads mapping to a genomic region is proportional to the copy number of that region in the sample. A deletion reduces coverage by approximately half for heterozygous events and to near zero for homozygous events. A duplication increases coverage proportionally to the additional copy number. The analytical challenge lies in distinguishing true copy number changes from the substantial noise inherent in sequencing coverage.

Coverage Normalization Strategies

Raw coverage counts are influenced by multiple technical factors that have nothing to do with copy number. GC content bias remains the most significant confounder, as regions with extreme GC composition amplify and sequence less efficiently. Long-read platforms show different bias profiles than short-read platforms, but normalization remains essential. The Bioconductor project provides documented packages for genomic analysis that include normalization and segmentation functions applicable to coverage data.

Mappability variation also affects read-depth analysis. Regions with high sequence similarity to other genomic locations may have reduced unique mapping, artificially lowering coverage. Long reads reduce but do not eliminate this problem, as reads from near-identical paralogous regions may still map ambiguously. Repeat content creates similar issues, with highly repetitive regions showing coverage patterns that reflect alignment ambiguity instead of true copy number.

The standard normalization workflow involves several steps. First, divide the genome into windows of fixed size, typically 500 bp to 5 kb for long-read data. Second, count reads or bases mapping to each window. Third, correct for GC bias using a loess regression or similar approach that models the relationship between GC content and coverage. Fourth, correct for mappability by excluding or down-weighting windows with low unique mappability. Fifth, segment the normalized coverage to identify regions with statistically significant deviations from the sample median.

Window Size Selection

Window size represents a fundamental tradeoff between resolution and statistical power. Small windows provide finer breakpoint resolution but noisier coverage estimates. Large windows provide more stable coverage but blur breakpoints across broad regions. For long-read data, the higher per-base accuracy and longer read lengths permit smaller windows than would be practical with short reads.

A practical approach is to run the analysis at multiple window sizes and compare results. Large events should be detectable at both coarse and fine resolutions. Small events may only appear at fine resolution but will show noisier coverage profiles. The Galaxy Training Network provides accessible tutorials on coverage analysis and segmentation that demonstrate the impact of parameter choices on variant detection.

Segmentation and Copy Number Estimation

After normalization, the coverage profile must be segmented into regions of constant copy number. Segmentation algorithms identify change points where the mean coverage shifts significantly. The output is a series of segments with estimated copy number states. The copy number estimate for each segment derives from the ratio of observed coverage to the sample median coverage, adjusted for ploidy and sex chromosomes.

For diploid autosomes, a ratio of 0.5 indicates a heterozygous deletion, a ratio of 1.5 indicates a single copy gain, and a ratio of 2.0 indicates a homozygous duplication. The precision of these estimates depends on the number of reads in each segment and the uniformity of coverage. Low coverage regions produce noisy estimates that may not reach statistical significance even when true copy number changes exist.

Core Principles of Split-Read CNV Calling

Split-read analysis identifies structural variants by detecting reads that align to two or more non-contiguous genomic locations. A read spanning a deletion breakpoint will align partially to the reference on each side of the deletion, with the intervening sequence absent. A read spanning a duplication breakpoint will show a similar pattern but with the duplicated sequence present in the read. The Galaxy Training Network offers practical tutorials on structural variant calling that cover split-read alignment interpretation.

Alignment Signatures for Different Variant Types

Deletions produce the most straightforward split-read signature. A read that spans the deletion will have a portion aligning upstream of the breakpoint and a portion aligning downstream, with a gap corresponding to the deleted sequence. The alignment tool reports this as a split alignment with a large deletion in the read relative to the reference.

Duplications and insertions produce more complex signatures. A read spanning a tandem duplication breakpoint will show alignment to the reference that includes the duplicated region twice, which may appear as a split alignment with overlapping or rearranged segments. Inverted duplications and other complex rearrangements produce even more intricate patterns that require careful interpretation.

Breakpoint Resolution and Confidence

The primary advantage of split-read analysis is nucleotide-level breakpoint resolution. When a read spans the breakpoint, the exact genomic coordinates of the junction can be determined from the alignment. This precision is valuable for clinical reporting, primer design for validation, and understanding the functional consequences of the variant.

Breakpoint confidence depends on the number of reads supporting the junction and the quality of the alignments flanking the breakpoint. A single supporting read may represent a sequencing or alignment artifact. Multiple independent supporting reads with consistent breakpoint coordinates provide high confidence. The EMBL-EBI training materials on sequence analysis describe alignment quality metrics that inform breakpoint confidence assessment.

Limitations of Split-Read Analysis

Split-read methods have a critical limitation: they can only detect breakpoints that are spanned by individual reads. For large variants, the distance between breakpoints may exceed the read length, meaning no single read covers both junctions. In such cases, split-read analysis may detect one breakpoint but not the other, or may miss the variant entirely.

Long-read platforms mitigate this limitation by producing reads that span larger distances. HiFi reads typically range from 10 to 25 kb, while Oxford Nanopore reads can exceed 100 kb. However, even the longest reads cannot span very large CNVs, and the probability of a read spanning both breakpoints decreases as variant size increases. This limitation motivates the hybrid approach that combines split-read evidence with read-depth signals.

Hybrid Read-Depth and Split-Read Workflow

The hybrid workflow integrates read-depth and split-read analyses to leverage their complementary strengths. The goal is to produce a unified CNV call set that includes accurate breakpoints where split-read evidence exists and reliable copy number estimates for larger events where only read-depth signal is available.

Step 1: Alignment and Quality Control

The workflow begins with alignment of long reads to the reference genome. The choice of aligner and alignment parameters affects both read-depth and split-read analyses. Minimap2 is the most commonly used aligner for long-read data, with parameters optimized for HiFi or Nanopore data as appropriate. The National Center for Biotechnology Information provides reference genome sequences and documentation on alignment formats.

Quality control at the alignment stage includes checking mapping rates, coverage uniformity, and the distribution of alignment scores. Low mapping rates may indicate sample contamination, reference mismatch, or alignment parameter problems. Coverage uniformity across the genome provides a baseline for read-depth analysis. The Carpentries lessons on data processing provide foundational skills for managing and inspecting large bioinformatics files.

Step 2: Read-Depth Analysis with CNVpytor

CNVpytor is a widely used tool for read-depth CNV analysis that has been adapted for long-read data. The tool calculates coverage across genomic windows, applies GC correction, and segments the genome into copy number states. For long-read data, the tool should be configured with appropriate window sizes and coverage thresholds.

The CNVpytor workflow involves several commands that build progressively on previous outputs. The initial step calculates read depth per window. Subsequent steps apply corrections and perform segmentation. The final step generates a VCF file with CNV calls and supporting statistics. The Bioconductor project documentation describes similar segmentation approaches available in R packages.

Step 3: Split-Read Analysis with Sniffles2

Sniffles2 is a structural variant caller designed for long-read data that detects deletions, duplications, insertions, inversions, and translocations from split-read and coverage signals. The tool uses a combination of alignment signatures to identify variant breakpoints and estimate variant sizes.

For CNV analysis, Sniffles2 provides duplication and deletion calls with breakpoint coordinates. The tool also reports supporting read counts and other quality metrics. Configuration options include minimum supporting reads, minimum variant size, and tandem repeat handling. The nf-core documentation describes community standards for running structural variant callers within reproducible pipelines.

Step 4: Reconciliation and Merging

The reconciliation step compares CNV calls from read-depth and split-read analyses to produce a unified call set. Overlapping calls that agree on variant type and approximate coordinates receive high confidence. Calls from only one method require additional scrutiny.

A practical reconciliation approach uses genomic overlap to match calls between the two methods. For each read-depth call, check whether a split-read call overlaps the same region. For each split-read call, check whether the read-depth segmentation shows a corresponding copy number change. The output is a table with columns for variant coordinates, variant type, copy number estimate, breakpoint coordinates, and supporting evidence from each method.

Step 5: Annotation and Interpretation

The final step annotates the unified call set with gene information, population frequency data, and clinical significance where applicable. The NCBI maintains databases of genomic variation and clinical associations that support this interpretation. For research applications, annotation may focus on gene content and potential functional impact. For clinical applications, interpretation follows established standards for variant classification.

Tool Configuration and Parameter Selection

The performance of both read-depth and split-read tools depends heavily on parameter selection. The optimal parameters vary by sequencing platform, coverage depth, and the size distribution of expected variants. The following sections describe the key parameters for each tool and the rationale for their selection.

CNVpytor Configuration for Long Reads

CNVpytor requires specification of the reference genome, window size, and ploidy. For long-read data, a window size of 1 kb provides a reasonable balance between resolution and noise. The tool also requires a GC correction step that models the relationship between GC content and coverage.

The segmentation step uses a statistical model to identify change points in the coverage profile. The sensitivity of segmentation depends on the significance threshold, with more stringent thresholds producing fewer but more confident calls. The tool also supports specifying expected ploidy, which is important for sex chromosomes and for samples with known aneuploidies.

Sniffles2 Configuration for CNV Detection

Sniffles2 requires specification of the minimum supporting reads for a variant call. For high coverage data, a minimum of 5 to 10 supporting reads provides high confidence. For lower coverage data, lower thresholds may be necessary but produce more false positives.

The tool also has parameters for minimum and maximum variant size. Setting a minimum size filters out small indels that are not the focus of CNV analysis. The maximum size parameter is less critical because the tool naturally limits calls to sizes that can be supported by the data.

Tandem repeat handling is particularly important for CNV analysis because tandem repeats are prone to alignment artifacts that mimic structural variants. Sniffles2 includes options for filtering or flagging variants in tandem repeat regions. The nf-core documentation describes best practices for configuring structural variant callers within reproducible analysis pipelines.

Coverage Depth Considerations

Coverage depth fundamentally affects the sensitivity and specificity of both read-depth and split-read analyses. Higher coverage provides more statistical power for detecting small copy number changes and more supporting reads for split-read breakpoints. However, higher coverage also increases sequencing cost and computational requirements.

For read-depth analysis, the minimum detectable copy number change depends on the number of reads per window. A heterozygous deletion reduces coverage by 50%, which is detectable with approximately 20 reads per window at reasonable statistical power. For 1 kb windows, this corresponds to roughly 20x genome coverage. Higher coverage enables smaller windows or more sensitive detection of smaller copy number changes.

For split-read analysis, the probability of detecting a breakpoint depends on the probability that a read spans the junction. This probability increases with coverage and with read length. At 30x coverage with 15 kb reads, the expected number of reads spanning a given breakpoint is approximately 30 times the fraction of reads that cover the breakpoint region.

Normalization Challenges Specific to Long-Read Data

Long-read platforms present normalization challenges that differ from short-read platforms. Understanding these differences is essential for accurate read-depth analysis.

GC Bias in Long-Read Data

Long-read platforms show GC bias, but the pattern differs from short-read platforms. The bias arises from different mechanisms, including library preparation and sequencing chemistry. Empirical GC correction using the observed relationship between GC content and coverage in the sample remains the most reliable approach.

The GC correction model should be fit on the sample data instead of using a universal correction. This is because the bias pattern varies between platforms, library preparation methods, and even between sequencing runs. The Bioconductor project provides packages for fitting and applying GC correction models to coverage data.

Mappability and Repeat Content

Long reads improve mappability in repetitive regions because they can span repeat elements and align uniquely to flanking unique sequence. However, regions composed entirely of repeats with no unique flanking sequence remain problematic. These regions may show reduced coverage that is misinterpreted as copy number loss.

A practical approach is to compute mappability scores for the reference genome and exclude or down-weight low mappability regions from read-depth analysis. The NCBI provides reference genome annotations that include repeat content information useful for this purpose.

Coverage Uniformity

Long-read platforms show different coverage uniformity profiles than short-read platforms. Nanopore data may show systematic coverage variation related to pore performance and base calling quality. HiFi data generally shows more uniform coverage but may have platform-specific biases.

The segmentation algorithm should account for the expected noise level in the data. Tools that estimate the noise from the data itself, instead of assuming a fixed value, perform better across different platforms and coverage levels.

Breakpoint Refinement Strategies

Accurate breakpoint coordinates are essential for clinical reporting and for understanding the functional consequences of CNVs. The following strategies improve breakpoint resolution beyond what either method provides alone.

Using Split-Read Breakpoints to Refine Read-Depth Calls

When a read-depth call overlaps a split-read call, the split-read breakpoints provide the precise coordinates for the event. The read-depth call provides the copy number estimate and the overall size of the event. Combining these data produces a call with both accurate breakpoints and accurate copy number.

The refinement process involves replacing the read-depth segment boundaries with the split-read breakpoint coordinates. The copy number estimate from the read-depth analysis is retained. The final call includes both the precise breakpoints and the copy number state.

Local Assembly for Breakpoint Resolution

For variants where split-read evidence is insufficient, local assembly of reads spanning the breakpoint region can resolve the exact junction sequence. This approach collects all reads mapping to the breakpoint region and assembles them de novo. The assembled contig is then aligned to the reference to identify the exact breakpoint.

Local assembly is computationally intensive but provides the highest resolution breakpoint information. The EMBL-EBI training materials describe assembly concepts and tools applicable to this approach.

Validation of Breakpoints

Breakpoint validation is essential for clinical applications. PCR amplification across the predicted breakpoint followed by Sanger sequencing provides definitive confirmation. For large deletions, a PCR assay with primers flanking the deleted region will produce a product only if the deletion is present. For duplications, quantitative PCR or digital PCR can confirm the copy number change.

A 2025 study in Familial Cancer evaluated droplet digital PCR and targeted adaptive sampling long-read sequencing for VHL gene deletion analysis. The study found that adaptive sampling precisely identified the exact sequence alterations in cases of complex deletions. This approach demonstrates the clinical utility of long-read methods for breakpoint characterization.

Records and Measurements for CNV Calling

Systematic record keeping is essential for reproducible CNV analysis and for troubleshooting when results are unexpected. The following records should be maintained for each analysis.

Sample and Sequencing Metadata

Record the sample identifier, tissue type, DNA extraction method, sequencing platform, library preparation method, and sequencing run identifier. These metadata affect data quality and interpretation. The NCBI provides structured formats for submitting and accessing sequencing data with associated metadata.

Alignment Statistics

Record the number of reads sequenced, the number mapped, the mapping rate, the median coverage, and the coverage distribution. These statistics provide the context for interpreting CNV calls. Low mapping rates or unusual coverage distributions may indicate sample or data quality problems.

Tool Versions and Parameters

Record the exact versions of all software tools used in the analysis, including the aligner, read-depth caller, split-read caller, and any annotation tools. Also record all non-default parameters. This information is essential for reproducing the analysis and for understanding differences between runs.

Call-Level Records

For each CNV call, record the genomic coordinates, variant type, copy number estimate, breakpoint coordinates if available, supporting read counts, and quality scores. Also record which methods detected the variant and whether the methods agreed. The Galaxy Training Network provides examples of structured variant records and their interpretation.

Common Failure Patterns in Long-Read CNV Calling

Understanding common failure modes helps bioinformaticians troubleshoot unexpected results and design better workflows.

False Positive Calls in Repetitive Regions

Repetitive regions produce false positive CNV calls through multiple mechanisms. Alignment ambiguity can create coverage dips that mimic deletions. Tandem repeat expansions can create coverage spikes that mimic duplications. Split-read alignments in repetitive regions often represent alignment artifacts instead of true breakpoints.

The most effective mitigation is to annotate calls that fall in repetitive regions and apply additional scrutiny. Tools that provide repeat annotation or that flag calls in low complexity regions help prioritize review efforts.

False Negative Calls for Small Events

Small CNVs, particularly those below 1 kb, are difficult to detect with read-depth methods because the coverage change is distributed across too few windows. Split-read methods may detect these events if a read spans the breakpoint, but detection depends on read length and coverage.

For applications where small CNV detection is critical, targeted approaches may be necessary. A 2021 study in the International Journal of Molecular Sciences described a long-range PCR-based sequencing protocol that amplified entire genomic regions for a panel of inherited retinal disease loci. This approach detected copy number variants and characterized their breakpoints at nucleotide resolution, demonstrating the value of targeted amplification for small variant detection.

Discordant Calls Between Methods

When read-depth and split-read methods produce discordant calls, the discrepancy often indicates a complex variant structure. A read-depth call without split-read support may represent a large event that no single read spans. A split-read call without read-depth support may represent a small event below the resolution of the segmentation.

The reconciliation process should flag discordant calls for manual review instead of automatically discarding them. The review should examine the raw alignment data at the variant locus to determine which call is correct.

Coverage Artifacts from Sample Issues

Sample quality issues can create coverage artifacts that mimic CNVs. Degraded DNA produces uneven coverage with regions of low coverage that may be misinterpreted as deletions. Contamination from other samples creates mixed coverage patterns. The Carpentries lessons on data handling provide guidance on quality assessment that applies to sequencing data.

Limitations of Long-Read CNV Calling

Long-read CNV calling has limitations that should be acknowledged in any analysis report.

Size-Dependent Detection Sensitivity

The sensitivity of CNV detection depends on variant size. Small events below approximately 500 bp are difficult to detect with read-depth methods and may be missed by split-read methods if no read spans the breakpoint. Large events above approximately 100 kb may exceed the span of individual reads, limiting split-read breakpoint resolution.

The practical consequence is that the variant size spectrum should be reported with the expected sensitivity for each size range. A negative result does not exclude the presence of small CNVs below the detection threshold.

Platform-Specific Biases

HiFi and Nanopore data have different error profiles and bias patterns. HiFi data has higher per-base accuracy but shorter read lengths. Nanopore data has longer reads but higher error rates, particularly in homopolymer regions. These differences affect both read-depth normalization and split-read breakpoint detection.

The Galaxy Training Network provides platform-specific tutorials that describe the expected data characteristics and appropriate analysis parameters for each platform.

Interpretation Challenges

CNV interpretation requires distinguishing pathogenic variants from benign copy number polymorphisms. Population frequency data helps this distinction, but the databases are less complete for CNVs than for single nucleotide variants. The NCBI maintains databases of genomic variation that include CNV frequency data.

A 2023 review in Molecular Psychiatry discussed the genetic architecture of schizophrenia, noting that rare copy number variants contribute to disease risk alongside common alleles of small effect. The review emphasized that genomic studies need to address the Eurocentric bias in existing data to ensure benefits are distributed across diverse communities. This context is relevant for CNV interpretation, as population frequency data from diverse ancestries improves variant classification.

Quality Control and Validation Approaches

Quality control should occur at multiple stages of the CNV calling workflow.

Alignment-Level Quality Control

Inspect the alignment file for systematic problems before running variant callers. Check the distribution of alignment scores, the fraction of reads with supplementary alignments, and the coverage across the genome. The NCBI provides documentation on alignment formats and quality metrics.

Call-Level Quality Control

Each CNV call should have associated quality metrics that support or question its validity. For read-depth calls, the number of windows supporting the call and the magnitude of the coverage change provide confidence information. For split-read calls, the number of supporting reads and the consistency of breakpoint coordinates provide confidence information.

Orthogonal Validation

Orthogonal validation using an independent method provides the highest confidence in CNV calls. For clinical applications, validation is typically required before reporting. Methods include PCR-based approaches, digital PCR, and microarray analysis. A 2025 study in Familial Cancer demonstrated that droplet digital PCR and targeted adaptive sampling long-read sequencing accurately detected VHL gene deletions, including complex deletions that required precise breakpoint identification.

Benchmarking Against Known Variants

When available, samples with known CNVs provide valuable benchmarking data. These samples can be used to calibrate parameters and to measure the sensitivity and specificity of the workflow. The NCBI maintains reference materials and benchmarking datasets for variant calling.

Professional Escalation Criteria

Bioinformaticians should recognize situations that require escalation to more specialized expertise or additional analysis.

When to Consult a Clinical Geneticist

CNV calls that involve known disease genes, that affect dosage-sensitive regions, or that are being considered for clinical reporting should be reviewed by a clinical geneticist. The interpretation of CNV pathogenicity requires clinical context that extends beyond the bioinformatics analysis.

When to Request Additional Sequencing

Low coverage data that produces uncertain CNV calls may benefit from additional sequencing. The decision to sequence more should consider the cost of additional sequencing versus the cost of an incorrect call. For clinical applications, the cost of an incorrect call is high, justifying additional sequencing when confidence is low.

When to Use Specialized Analysis

Complex variants that produce discordant calls between methods may require specialized analysis approaches. Local assembly, targeted long-read sequencing, or optical mapping may resolve the structure of complex variants. A 2024 study in Genome Research found that long-read genome sequencing identified variants that were not called within short-read data, were represented by calls with incorrect sizes or structures, or failed quality control and filtration. These findings support the use of long-read methods when short-read results are ambiguous.

Decision Framework for Selecting CNV Calling Strategies by Variant Size and Data Type

The choice between read-depth, split-read, or hybrid analysis depends on the variant size spectrum of interest and the sequencing platform in use. A structured decision framework helps bioinformaticians allocate analytical effort where it produces the most reliable results. The framework below organizes the key considerations into a practical sequence of decisions that can be applied before running any CNV calling tool.

Step 1: Define the Variant Size Spectrum of Interest

The first decision is to specify the minimum and maximum variant sizes that the analysis must detect. This determination shapes every subsequent choice, including window size, tool parameters, and validation strategy. For clinical applications targeting known disease genes, the size spectrum may be narrow and defined by the gene structure. For research applications exploring genome-wide variation, the spectrum may span from a few hundred base pairs to several megabases.

The size spectrum directly determines which analytical method provides primary evidence. Variants below approximately 1 kb are poorly detected by read-depth methods because the coverage change distributes across too few windows. Variants above approximately 50 kb may exceed the span of individual HiFi reads, limiting split-read breakpoint resolution. The hybrid approach becomes most valuable when the size spectrum spans both regimes.

Step 2: Assess Data Characteristics Before Tool Selection

Before selecting tools, evaluate the alignment file for characteristics that affect CNV calling performance. Calculate the median coverage, the distribution of read lengths, and the fraction of reads with supplementary alignments. These metrics inform parameter choices and help predict which methods will perform reliably.

For coverage below 20x, read-depth analysis requires larger windows to achieve statistical power, which reduces breakpoint resolution. For coverage above 40x, smaller windows become feasible and split-read analysis gains sensitivity because more reads span each breakpoint. Read length distribution matters primarily for split-read analysis, as longer reads increase the probability of spanning both breakpoints of a variant.

The National Center for Biotechnology Information provides documentation on alignment formats and quality metrics that support this assessment. The European Bioinformatics Institute offers training materials on sequence analysis that cover the interpretation of alignment statistics in the context of variant calling.

Step 3: Match Method to Variant Size Range

For variants smaller than 1 kb, split-read analysis with Sniffles2 provides the primary detection signal. Read-depth analysis at this size range produces noisy coverage estimates that rarely reach statistical significance. The split-read approach detects these variants when a read spans the breakpoint, and the probability of spanning increases with coverage and read length.

For variants between 1 kb and 50 kb, both methods contribute useful evidence. Read-depth analysis detects the coverage change across multiple windows, while split-read analysis provides breakpoint coordinates when reads span the junctions. The hybrid reconciliation step becomes most valuable in this size range because the methods provide complementary information.

For variants larger than 50 kb, read-depth analysis provides the primary detection signal. Split-read analysis may detect one breakpoint if a read spans it, but the probability of spanning both breakpoints decreases as variant size increases. The read-depth call provides the copy number estimate and approximate boundaries, while any available split-read breakpoints refine the coordinates.

Step 4: Apply Platform-Specific Parameter Adjustments

HiFi and Nanopore data require different parameter configurations for optimal CNV calling. HiFi data has higher per-base accuracy, which reduces false positive split-read calls in repetitive regions. This accuracy permits more sensitive breakpoint detection with lower minimum supporting read thresholds. Nanopore data has longer reads but higher error rates, which increases the need for filtering tandem repeat artifacts and requires higher minimum supporting read thresholds for confident breakpoint calls.

The Galaxy Training Network provides platform-specific tutorials that describe expected data characteristics and appropriate analysis parameters. The nf-core documentation describes community standards for running structural variant callers within reproducible pipelines, including recommended parameter sets for different data types.

Step 5: Document the Decision Rationale

Record the rationale for each analytical decision, including the variant size spectrum, the data characteristics that informed parameter choices, and the expected sensitivity for each size range. This documentation supports interpretation of negative results and provides context for troubleshooting when calls appear unexpected.

The record should include the coverage threshold used to determine window size, the minimum supporting read threshold for split-read calls, and the criteria for reconciling discordant calls between methods. The Carpentries lessons on data management provide foundational practices for maintaining reproducible analysis records.

Comparison of Method Selection by Variant Size

The table below summarizes the recommended primary method, secondary method, and key parameter considerations for different variant size ranges. This comparison supports the decision framework by providing concrete guidance for workflow design.

Variant Size RangePrimary MethodSecondary MethodKey Parameter Considerations
Below 1 kbSplit-read with Sniffles2Local assembly for breakpoint resolutionMinimum supporting reads of 5 to 10, tandem repeat filtering enabled
1 kb to 10 kbHybrid read-depth and split-readRead-depth for copy number confirmationWindow size of 500 bp to 1 kb, minimum supporting reads of 5
10 kb to 50 kbHybrid read-depth and split-readSplit-read for breakpoint refinementWindow size of 1 kb to 5 kb, reconciliation threshold for overlap
Above 50 kbRead-depth with CNVpytorSplit-read for single breakpoint detectionWindow size of 5 kb to 10 kb, segmentation significance threshold

Troubleshooting Method Selection Failures

When the selected method fails to detect expected variants, the first troubleshooting step is to verify that the variant size falls within the detection range of the chosen method. A variant below 1 kb will not appear in read-depth analysis regardless of parameter adjustment. A variant above 100 kb may lack split-read support because no single read spans both breakpoints.

The second troubleshooting step is to examine the coverage profile at the expected variant location. If coverage appears uniform across the region, the variant may be absent or the window size may be too large to detect the change. Reducing the window size and rerunning the segmentation can reveal smaller coverage changes that the initial analysis missed.

The third troubleshooting step is to inspect the alignment file directly at the variant locus. The Integrative Genomics Viewer or similar visualization tools allow examination of individual read alignments. This inspection reveals whether split-read evidence exists but was filtered by quality thresholds, or whether alignment artifacts explain the absence of calls.

Escalation Criteria for Method Selection

Escalate to specialized analysis when the decision framework produces discordant results that cannot be resolved through parameter adjustment. Complex variants with multiple breakpoints, inverted duplications, or insertions at breakpoints may require local assembly or targeted long-read sequencing to resolve their structure.

A 2025 study in Familial Cancer demonstrated that targeted adaptive sampling long-read sequencing precisely identified the exact sequence alterations in cases of complex VHL gene deletions. This approach proved valuable when standard analysis methods could not resolve the variant structure. The study supports the use of targeted long-read approaches when initial CNV calling produces ambiguous or discordant results.

A 2024 study in Genome Research found that some variants were represented by calls with incorrect sizes or structures in short-read data, or failed quality control and filtration. The study demonstrated that long-read genome sequencing identified these variants correctly, supporting the use of long-read methods when short-read results are ambiguous or when the decision framework indicates that standard approaches may miss relevant variants.

For clinical applications, escalate to a clinical geneticist when the decision framework identifies a variant in a known disease gene but the breakpoint resolution is insufficient for clinical reporting. The geneticist can determine whether the available evidence meets the standards for clinical variant classification or whether additional analysis is required.

Frequently Asked Questions

What is the minimum coverage needed for reliable CNV calling from long reads?

The minimum coverage depends on the variant size and the detection method. For read-depth analysis, approximately 20x coverage provides sufficient reads per window to detect heterozygous deletions in 1 kb windows. For split-read analysis, higher coverage increases the probability that a read spans a breakpoint. The optimal coverage also depends on the expected variant size distribution and the tolerance for false positives and false negatives.

How do I choose between HiFi and Nanopore data for CNV analysis?

The choice depends on the specific application. HiFi data provides higher per-base accuracy, which improves split-read breakpoint resolution and reduces false positives in repetitive regions. Nanopore data provides longer reads, which improves the detection of large variants and the spanning of complex rearrangements. The optimal choice depends on the variant types of interest and the available sequencing infrastructure.

Why do read-depth and split-read methods produce different CNV calls?

The methods detect different signals and have different sensitivity profiles. Read-depth methods detect coverage changes across genomic windows and are sensitive to large events. Split-read methods detect alignment discontinuities and are sensitive to events with breakpoints that are spanned by individual reads. Discordant calls often indicate complex variant structures or events at the boundary of each method's detection capability.

How should I handle CNV calls in repetitive regions?

Calls in repetitive regions require additional scrutiny because alignment artifacts are more common in these regions. Annotate calls that overlap known repeat elements and examine the raw alignment data at the variant locus. If the call is supported by multiple independent reads with consistent breakpoints, it is more likely to be real. Consider orthogonal validation for calls that will be reported clinically.

What is the role of local assembly in CNV breakpoint resolution?

Local assembly collects reads spanning a breakpoint region and assembles them de novo to resolve the exact junction sequence. This approach provides nucleotide-level breakpoint resolution when split-read evidence is insufficient. Local assembly is computationally intensive but valuable for clinical reporting and for understanding the functional consequences of complex variants.

How do I validate CNV calls for clinical reporting?

Clinical validation requires an orthogonal method that independently confirms the variant. PCR amplification across the breakpoint followed by Sanger sequencing provides definitive confirmation for deletions and duplications with known breakpoints. Digital PCR provides quantitative copy number confirmation. The validation method should be documented in the clinical report.

What databases should I use for CNV interpretation?

The NCBI maintains databases of genomic variation, clinical associations, and population frequency data that support CNV interpretation. These resources provide information on known pathogenic CNVs, benign copy number polymorphisms, and population frequencies across diverse ancestries. The databases should be accessed through the official NCBI portal to ensure current data.

How often should I update my CNV calling pipeline?

The pipeline should be updated when new tool versions provide meaningful improvements in accuracy or when the sequencing platform changes. The nf-core community maintains versioned pipelines that incorporate current best practices for variant calling. Regular benchmarking against known variants helps identify when pipeline updates improve performance.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.