Variant Calling in Segmental Duplications: Strategies for Accurate Detection in Paralogous Regions
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Segmental duplications (SDs) pose a significant challenge in variant calling due to paralogous sequence variants (PSVs), which are fixed differences between duplicated DNA copies that can be misidentified as true biological variants by standard short-read pipelines. This leads to false positives and confounds interpretation, particularly in clinically relevant genes like SMN1 and SMN2.
- Specialized variant calling strategies are essential for accurate detection in SDs, including graph-based alignment refinement (e.g., DuploMap for long reads) that leverages PSVs to improve read mapping confidence, and multilocus callers (e.g., ParascopyVC) that jointly analyze reads across all repeat copies.
- Long-read sequencing technologies (PacBio, Oxford Nanopore) offer improved resolution in SDs due to longer read lengths spanning more of the duplicated regions, but even these require PSV-aware mapping refinement for high-identity duplications.
- For clinical applications, haplotype-based methods (e.g., Paraphase) using long-read HiFi data can provide accurate copy number determination and phased variant calls for genes with known paralogs, enabling precise risk assessment for conditions like spinal muscular atrophy.
- Rigorous filtering protocols are critical, focusing on mapping quality distributions, allele balance patterns that account for copy number, PSV overlap, and population frequencies to mitigate PSV artifacts and other false positives.
- Orthogonal validation methods such as Sanger sequencing with paralog-specific primers, digital droplet PCR, or targeted long-read sequencing are indispensable for confirming candidate variants in SDs before clinical reporting.
Researchers analyzing whole-genome sequencing data face a persistent problem when working with segmental duplications: distinguishing true biological variants from paralogous sequence variants (PSVs) that arise from sequence differences between duplicated copies. Standard short-read pipelines frequently misassign reads to the wrong paralog, producing false positive calls that waste validation resources and confound clinical interpretation. This article provides a practical framework for reducing PSV artifacts through graph-based alignment strategies, specialized variant callers, and rigorous filtering protocols. The approaches described here apply to germline and somatic variant calling workflows in research and diagnostic settings where segmental duplications are present in the target regions.
Understanding the Paralogous Mapping Problem
Segmental duplications, also known as low-copy repeats, are long stretches of duplicated DNA that share high sequence identity. These regions cover more than 5 percent of the human genome and present unique challenges for variant detection because short sequencing reads cannot always be assigned unambiguously to their true genomic origin. When a read maps equally well to two or more paralogous locations, aligners may place it at the wrong copy, and downstream variant callers then report differences between the read and the reference that are actually PSVs instead of genuine variants in the sample.
The core difficulty is that PSVs are fixed differences between paralogous copies that exist in the reference genome itself. When reads from one copy align to another copy, the caller interprets the PSV as a heterozygous or homozygous variant at that position. This artifact is particularly problematic in clinical contexts because many disease-associated genes reside within or near segmental duplications. The SMN1 and SMN2 genes, for example, share extremely high sequence identity, and accurate variant calling in these regions requires specialized approaches that account for their paralogous relationship.
Long-read sequencing technologies such as Pacific Biosciences and Oxford Nanopore can overcome some limitations of short reads because their longer read lengths span more of the duplicated region and provide more distinguishing information. However, long segmental duplications with high sequence identity still pose challenges for long-read mapping. Probabilistic methods that leverage PSVs to distinguish between multiple alignment locations have been developed to improve mapping accuracy in these regions. One such approach, DuploMap, analyzes reads mapped to segmental duplications using existing long-read aligners and uses PSVs to discriminate between candidate alignment positions. On simulated datasets, this method increased the percentage of correctly mapped reads with high confidence for multiple long-read aligners while maintaining high precision. Across whole-genome long-read datasets, DuploMap aligned an additional 8 to 21 percent of reads in segmental duplications with high confidence relative to Minimap2 alone.
The practical implication is that researchers must treat segmental duplications as a distinct analytical domain instead of applying generic variant calling pipelines without modification. The choice of aligner, variant caller, and filtering strategy materially affects the accuracy of variant calls in these regions.
Core Principles for Accurate Variant Detection in Duplicated Regions
Several principles guide successful variant calling in segmental duplications. First, the analysis must account for the copy number of the duplicated region, because the number of paralogous copies affects the expected allele balance and genotype likelihood calculations. Second, the method should use information from all copies jointly instead of treating each copy independently. Third, PSVs that differentiate paralogous copies should be identified and used to phase reads to their correct origin.
Standard variant callers that assume diploidy and unique mapping frequently fail in duplicated regions because they do not model the presence of multiple highly similar sequences. The ambiguity in read mapping means that reads with low mapping quality are often discarded, yet these reads may contain the only evidence for variants in the duplicated region. Specialized methods that utilize reads independent of mapping quality can recover this lost information.
A multilocus approach that performs variant calling jointly across all repeat copies has demonstrated substantial accuracy improvements in low-copy repeats. This method, ParascopyVC, aggregates reads mapped to different repeat copies and performs polyploid variant calling to identify candidate variants. It then uses population data to identify PSVs that differentiate repeat copies and estimates the genotype of variants for each copy. On simulated whole-genome sequence data, ParascopyVC achieved higher precision and recall than three state-of-the-art variant callers across 167 low-copy repeat regions. Benchmarking against genome-in-a-bottle high-confidence variant calls for the HG002 genome showed high precision and recall across low-copy repeat regions, significantly better than FreeBayes, GATK, and DeepVariant. Across seven human genomes, ParascopyVC demonstrated consistently higher accuracy than the other callers.
The key principle is that variant calling in segmental duplications requires methods that explicitly model the paralogous structure of the region. Generic pipelines that do not account for this structure will produce systematic errors that are difficult to detect without orthogonal validation.
At a Glance: Decision Table for Variant Calling Strategies
| Analysis Scenario | Recommended Approach | Key Considerations | Expected Outcome |
|---|---|---|---|
| Short-read WGS, duplicated region of interest | Use a multilocus caller such as ParascopyVC that aggregates reads across copies | Requires population data for PSV identification, validate with orthogonal methods | Higher precision and recall than generic callers in low-copy repeats |
| Long-read WGS, high-identity segmental duplications | Use PSV-aware mapping refinement such as DuploMap after initial alignment | Works with PacBio and Oxford Nanopore data, improves mapping confidence | Additional reads correctly mapped with high confidence, more mappable sequence |
| Clinical gene with known paralog (SMN1/SMN2) | Use haplotype-based method such as Paraphase with long-read HiFi data | Provides copy number and phased variants, enables silent carrier risk assessment | Accurate copy number calls and variant detection without pedigree information |
| Somatic variant calling in duplicated regions | Apply tumor-normal paired analysis with PSV filtering | Requires matched normal to distinguish germline PSVs from somatic variants | Reduced false positive somatic calls in paralogous regions |
Practical Workflow for Variant Calling in Segmental Duplications
Step 1: Define the Target Regions and Identify Segmental Duplications
Before running any variant caller, researchers should identify which regions of the genome contain segmental duplications relevant to their analysis. Public databases such as those maintained by the National Center for Biotechnology Information provide genomic annotation resources that include segmental duplication tracks. The UCSC Genome Browser and Ensembl also provide duplication annotations that can be downloaded as BED files for intersection with target regions.
For clinical applications, the target gene list should be cross-referenced with duplication annotations to identify genes that overlap segmental duplications. More than 150 genes overlapping low-copy repeats are associated with risk for human diseases, so this step is essential for prioritizing analysis strategies. Genes such as SMN1, CYP2D6, and STRC are well-known examples where paralogous copies complicate variant interpretation.
Step 2: Select Appropriate Sequencing Technology and Coverage
The choice between short-read and long-read sequencing affects the entire downstream analysis strategy. Short-read sequencing remains the most common approach for whole-genome and targeted sequencing due to cost and throughput. However, short reads are fundamentally limited in their ability to resolve high-identity segmental duplications because the read length may be shorter than the distinguishing PSV intervals.
Long-read sequencing technologies such as PacBio HiFi and Oxford Nanopore provide read lengths that can span larger portions of duplicated regions, enabling more confident assignment of reads to their correct paralogous copy. The ability to characterize repetitive regions of the human genome is limited by the read lengths of short-read sequencing technologies, and long-read technologies can potentially overcome this limitation. However, long segmental duplications with high sequence identity still pose challenges for long-read mapping, so specialized analysis methods remain necessary.
For clinical applications where accurate variant detection in a specific gene is critical, long-read sequencing with high coverage may be justified despite the higher cost. For research applications across the whole genome, a hybrid approach that combines short-read data with targeted long-read validation of candidate variants in duplicated regions may be more practical.
Step 3: Align Reads with Appropriate Parameters
The alignment step is the first point where PSV artifacts can be introduced or mitigated. Standard aligners such as BWA-MEM for short reads and Minimap2 for long reads are widely used, but their default parameters may not be optimal for segmental duplications.
For short-read data, the aligner should be configured to retain reads with multiple mapping locations instead of discarding them as ambiguous. Some multilocus variant callers require access to reads with low mapping quality because these reads may contain the only evidence for variants in duplicated regions. The alignment output should include all reads, including those with multiple reported alignments, so that downstream tools can use this information.
For long-read data, the initial alignment should be followed by a refinement step that uses PSVs to improve mapping accuracy. DuploMap is one such method that analyzes reads mapped to segmental duplications using existing long-read aligners and leverages PSVs to distinguish between multiple alignment locations. This approach increased the percentage of correctly mapped reads with high confidence for multiple long-read aligners including Minimap2 and BLASR on simulated datasets.
Step 4: Apply Specialized Variant Calling Methods
After alignment, the variant calling step should use methods that explicitly model the paralogous structure of duplicated regions. Generic callers such as GATK HaplotypeCaller and FreeBayes assume unique mapping and diploid genotypes, which leads to systematic errors in segmental duplications.
For short-read data, ParascopyVC performs variant calling jointly across all repeat copies and utilizes reads independent of mapping quality in low-copy repeats. The method aggregates reads mapped to different repeat copies and performs polyploid variant calling to identify candidate variants. It then uses population data to identify PSVs that differentiate repeat copies and estimates the genotype of variants for each repeat copy. This approach achieved higher precision and recall than GATK, FreeBayes, and DeepVariant in benchmark evaluations.
For long-read data, haplotype-based methods such as Paraphase can identify full-length haplotypes for specific genes, determine gene copy numbers, and call phased variants. Paraphase was developed for SMN1 and SMN2 analysis using long-read PacBio HiFi data and achieved copy-number calls highly concordant with orthogonal methods. The method identified major SMN1 and SMN2 haplogroups and characterized their co-segregation through pedigree-based analyses.
Step 5: Filter and Annotate Variants
Variant filtering is essential for removing PSV artifacts and other false positives. The filtering strategy should consider the following criteria:
- Mapping quality: Variants supported only by reads with low mapping quality should be flagged for review, although specialized callers may use these reads productively.
- Allele balance: In duplicated regions, the expected allele balance depends on copy number. A variant present in one copy of a duplicated region may appear at a lower allele fraction than a heterozygous variant in a unique region.
- PSV overlap: Variants that overlap known PSV positions should be treated with caution because they may represent differences between paralogous copies instead of true variants in the sample.
- Population frequency: Variants that appear at high frequency in population databases but are absent from validated variant sets in duplicated regions may represent mapping artifacts.
The annotation step should include information about whether the variant falls within a segmental duplication and whether it overlaps known PSVs. This annotation supports downstream interpretation and prioritization of candidate variants for validation.
Step 6: Validate Candidate Variants with Orthogonal Methods
Validation of candidate variants in segmental duplications is critical because even specialized callers produce false positives. Orthogonal validation methods include:
- Sanger sequencing with paralog-specific primers that amplify only one copy of the duplicated region
- Digital droplet PCR for copy number validation
- Long-read sequencing of the specific region to confirm variant presence and phase
- Linked-read sequencing data that provides long-range information about which reads originate from the same DNA molecule
The choice of validation method depends on the variant type, the genomic context, and the available resources. For clinical reporting, validation is typically required before a variant is reported to the patient or referring clinician.
Options and Tradeoffs in Analysis Strategies
Short-Read versus Long-Read Approaches
Short-read sequencing offers lower cost per base and established analysis pipelines, but it has fundamental limitations in segmental duplications. The ambiguity in read mapping leads to systematic false positives and false negatives that are difficult to eliminate entirely. Specialized callers such as ParascopyVC improve accuracy substantially, but they still achieve lower recall than long-read approaches in some regions.
Long-read sequencing provides the read length needed to resolve high-identity duplications, but it comes with higher cost and lower throughput. The analysis methods for long-read data are less mature than those for short-read data, although tools such as DuploMap and Paraphase have demonstrated substantial improvements. For clinical applications where accurate variant detection in a specific gene is critical, the additional cost of long-read sequencing may be justified.
Generic versus Specialized Variant Callers
Generic variant callers such as GATK and DeepVariant are well-validated in unique regions of the genome but perform poorly in segmental duplications. Benchmark evaluations have shown that ParascopyVC achieves higher precision and recall than these generic callers in low-copy repeat regions. However, generic callers remain useful for the majority of the genome that is not duplicated, and they provide a baseline for comparison.
The choice between generic and specialized callers depends on the analysis goals. For whole-genome analysis where the primary interest is in unique regions, a generic caller with appropriate filtering may be sufficient. For targeted analysis of genes within segmental duplications, a specialized caller is essential for accurate variant detection.
Reference-Based versus Graph-Based Approaches
Traditional variant calling uses a linear reference genome and aligns reads to it. Graph-based approaches represent the genome as a graph that includes alternative alleles and structural variation, which can improve read mapping in complex regions. However, graph-based methods are more computationally intensive and require careful construction of the graph to avoid introducing artifacts.
The multilocus approach used by ParascopyVC constructs a network of homologous sequences and realigns reads to each segmental duplication from its homologous counterparts. This approach is conceptually similar to graph-based methods but is specifically designed for the challenges of segmental duplications. The realignments are phased and assembled into haplotypes via graph-based algorithms, followed by integer linear programming to retain the two most plausible haplotypes.
Observations and Measurements for Quality Assessment
Mapping Quality Distributions
The distribution of mapping quality scores in segmental duplications differs markedly from unique regions. Reads in duplicated regions tend to have lower mapping quality because they align equally well to multiple locations. Researchers should examine the mapping quality distribution for their target regions to understand the extent of ambiguity.
A high proportion of reads with mapping quality near zero in a duplicated region indicates that the aligner cannot distinguish between paralogous copies. This observation suggests that a PSV-aware refinement step or a specialized caller is needed to extract information from these reads.
Allele Balance Patterns
In diploid unique regions, heterozygous variants typically show allele balance near 0.5. In duplicated regions, the expected allele balance depends on the copy number and the number of copies carrying the variant. A variant present in one copy of a two-copy duplication would be expected at an allele balance near 0.25, while a variant present in both copies would appear near 0.5.
Deviations from these expected patterns can indicate mapping artifacts. For example, a variant that appears homozygous in a duplicated region but is actually present in only one copy would show an allele balance near 1.0 instead of the expected 0.5. Examining allele balance distributions for candidate variants can help identify such artifacts.
Read Support and Strand Bias
The number of reads supporting a variant and the balance of forward and reverse strand support are standard quality metrics. In segmental duplications, these metrics can be misleading because reads from different paralogous copies may be counted together. A variant that appears to have strong read support may actually be supported by reads from multiple copies, with the variant present in only one copy.
Specialized callers that phase reads to their correct origin provide more reliable read support metrics. When using generic callers, researchers should examine the read alignments supporting each variant to determine whether the supporting reads all originate from the same paralogous copy.
Records and Documentation for Reproducible Analysis
Analysis Logs and Parameter Documentation
Reproducible variant calling in segmental duplications requires detailed documentation of all analysis parameters. The analysis log should record:
- Sequencing platform and version
- Aligner name, version, and all non-default parameters
- Variant caller name, version, and all non-default parameters
- Reference genome version and any masking or modification applied
- Filtering thresholds and their rationale
- Software container or environment versions
Workflow management systems such as those provided by the nf-core community support reproducible pipeline execution and documentation. The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context. Using such systems ensures that the analysis can be reproduced by other researchers or at a later time.
Version Control for Analysis Scripts
All custom analysis scripts should be maintained under version control using Git or a similar system. The Carpentries lessons provide foundational training in version control with Git, which is essential for tracking changes to analysis code and ensuring reproducibility. The lessons cover shell, Git, and programming skills that are directly applicable to bioinformatics analysis workflows.
Data Management and Storage
Raw sequencing data, aligned reads, and variant calls should be stored according to institutional data management policies. The NCBI provides data resources for sequence data deposition and retrieval, including the Sequence Read Archive for raw sequencing data and dbSNP for variant data. Depositing data in public repositories supports transparency and enables other researchers to reproduce or extend the analysis.
Common Failure Patterns and Troubleshooting
Failure Pattern 1: Excessive Variant Density in Duplicated Regions
A common observation is that variant calling produces an unusually high density of variants in segmental duplications compared to unique regions. This pattern indicates that PSVs are being called as variants because reads from one paralogous copy are aligning to another copy. The fix is to apply a specialized caller that models the paralogous structure or to filter variants that overlap known PSV positions.
Failure Pattern 2: Discordant Copy Number and Variant Calls
When copy number analysis and variant calling produce discordant results in a duplicated region, the cause is often that the variant caller assumes a diploid copy number that does not match the actual copy number. For example, a deletion of one copy of a duplicated region would reduce the expected allele balance for variants in the remaining copy. Specialized callers that estimate copy number as part of the variant calling process can resolve this discordance.
Failure Pattern 3: Variants That Do Not Validate by Sanger Sequencing
Sanger validation failures are common for variants in segmental duplications because the PCR primers may amplify multiple paralogous copies. The Sanger trace may show a mixture of sequences from different copies, making it impossible to confirm the variant. The fix is to design paralog-specific primers that amplify only the copy of interest, which requires identifying PSVs that distinguish the copies.
Failure Pattern 4: Inconsistent Results Across Analysis Pipelines
When different variant calling pipelines produce inconsistent results in a duplicated region, the cause is often that the pipelines use different alignment and calling strategies that handle paralogous mapping differently. The fix is to use a specialized caller that has been benchmarked in segmental duplications and to validate candidate variants with an orthogonal method.
Limitations of Current Approaches
Residual False Positives and False Negatives
Even the best current methods for variant calling in segmental duplications produce residual errors. ParascopyVC achieved high precision and recall in benchmark evaluations, but it did not achieve perfect accuracy. The remaining errors are concentrated in regions with complex duplication structure or high sequence identity where PSVs are insufficient to distinguish paralogous copies.
Dependence on Population Data for PSV Identification
Methods that use PSVs to differentiate repeat copies depend on population data to identify these variants. In populations that are underrepresented in reference datasets, the PSV catalog may be incomplete, leading to reduced accuracy. Researchers working with diverse populations should be aware of this limitation and may need to generate population-specific PSV data.
Computational Cost of Specialized Methods
Specialized variant callers for segmental duplications are more computationally intensive than generic callers because they perform additional analysis steps such as constructing networks of homologous sequences and phasing reads. The computational cost may be prohibitive for whole-genome analysis in large cohorts, and researchers may need to restrict specialized analysis to target regions of interest.
Incomplete Reference Genomes
Reference genomes are incomplete in some segmental duplications, particularly those with complex structure or very high sequence identity. The complete genome of a songbird, for example, added approximately 90 megabases of previously missing sequence, and 9 percent of previously unassembled or unannotated genes overlapped with segmental duplications. Incomplete reference sequence leads to reads that cannot be aligned, resulting in false negative variant calls in these regions.
Safety and Regulatory Context for Clinical Applications
Clinical Validation Requirements
Variant calling in segmental duplications for clinical applications requires validation according to regulatory standards. Laboratories performing clinical sequencing must validate their assays, including the bioinformatics pipeline, using appropriate reference materials and demonstrate acceptable accuracy, precision, and reproducibility. The validation should include specific assessment of performance in segmental duplications because these regions are known to be challenging.
Reporting Considerations
Variants in segmental duplications should be reported with appropriate caveats about the analytical limitations. If a variant is identified in a duplicated region using a method that cannot definitively assign it to a specific paralogous copy, the report should state this limitation. For genes such as SMN1 where the paralog SMN2 has clinical significance, the report should specify which gene the variant is in and the evidence supporting that assignment.
Incidental Findings and Carrier Screening
Analysis of segmental duplications may reveal variants relevant to carrier screening or incidental findings. For example, Paraphase analysis of SMN1 and SMN2 can identify haplotypes that confer silent carrier risk for spinal muscular atrophy. The method identified two SMN1 haplotypes that form a common two-copy SMN1 allele in African populations, and testing positive for these haplotypes in an individual with two copies of SMN1 gives a silent carrier risk substantially higher than currently used markers. Laboratories offering carrier screening should consider whether their analysis methods can detect such haplotypes.
Professional Escalation Criteria
When to Consult a Bioinformatics Specialist
Researchers should escalate to a bioinformatics specialist when they encounter any of the following situations:
- The target region has complex duplication structure that is not resolved by standard tools
- Variant calls in a duplicated region are inconsistent across multiple analysis pipelines
- The analysis requires integration of multiple data types such as short-read, long-read, and linked-read data
- The clinical significance of a variant depends on which paralogous copy it is in
When to Consider Orthogonal Validation
Orthogonal validation should be considered when:
- A variant in a duplicated region has potential clinical significance
- The variant is supported by reads with low mapping quality
- The variant allele balance is inconsistent with the expected pattern for the copy number
- The variant will be reported to a patient or referring clinician
When to Reanalyze with Different Methods
Reanalysis with different methods is appropriate when:
- Initial analysis produces an unexpectedly high number of variants in a duplicated region
- Copy number analysis and variant calling produce discordant results
- The analysis uses a generic variant caller that is known to perform poorly in segmental duplications
- New methods or reference data become available that could improve accuracy
Training and Skill Development for Laboratory Professionals
Foundational Bioinformatics Skills
Laboratory professionals who perform variant calling in segmental duplications should develop foundational bioinformatics skills including command-line operation, shell scripting, and version control. The Carpentries lessons provide training in these foundational computing and data skills, including shell, Git, and programming. These skills are essential for running analysis pipelines, troubleshooting errors, and maintaining reproducible workflows.
Domain-Specific Training
Training in genomic data analysis and variant calling is available through multiple sources. The EMBL-EBI Training program provides bioinformatics learning pathways and data-resource training that cover sequence analysis and variant interpretation. The Galaxy Training Network offers accessible workflow training and analysis tutorials that include hands-on exercises with real genomic data. These resources support skill development for researchers at all levels.
Reproducible Workflow Training
The nf-core documentation provides guidance on community pipeline standards, usage, and configuration for reproducible genomic analysis. Researchers who use nf-core pipelines should review this documentation to understand how to configure pipelines for their specific analysis needs and how to contribute improvements back to the community. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation for R-based analysis.
Decision Framework for Selecting Variant Calling Methods in Segmental Duplications
Choosing the right analysis strategy for segmental duplications requires a structured decision process that accounts for the specific genomic region, available data types, and clinical or research objectives. A practical decision framework helps researchers avoid the common error of applying a one-size-fits-all pipeline to regions with fundamentally different analytical requirements. The framework below organizes the key considerations into sequential decision points that can be applied before committing computational resources to a particular approach.
Decision Point 1: Characterize the Duplication Architecture
The first decision point requires determining whether the target region contains simple or complex duplication architecture. Simple duplications involve two paralogous copies with high sequence identity and a clear one-to-one relationship. Complex duplications involve multiple copies, nested duplications, or duplications that share sequence with more than one other genomic location. This distinction matters because methods optimized for simple duplications may fail when applied to complex architectures.
To characterize the architecture, researchers should examine the segmental duplication track for their target region in a genome browser and count the number of paralogous copies that share significant sequence identity. If the region contains more than two copies with greater than 95 percent identity, the analysis requires methods that can model multi-copy paralogy instead of simple diploid assumptions. The network-based approach used by SDrecall, which constructs a network of homologous sequences and realigns reads to each segmental duplication from its homologous counterparts, is designed for this more complex scenario.
Decision Point 2: Assess Available Sequencing Data Types
The second decision point evaluates whether the available sequencing data can support the required analytical approach. Short-read data alone can support multilocus variant calling methods such as ParascopyVC, which aggregates reads mapped to different repeat copies and performs polyploid variant calling. However, short-read data cannot resolve regions where the distinguishing PSV intervals are shorter than the read length or where the sequence identity between copies approaches 100 percent.
Long-read data from PacBio HiFi or Oxford Nanopore platforms provide the read length needed to span larger portions of duplicated regions. The ability to characterize repetitive regions of the human genome is limited by the read lengths of short-read sequencing technologies, and long-read technologies can potentially overcome this limitation. If long-read data are available, haplotype-based methods such as Paraphase become viable options for genes with known paralogs. If only short-read data are available, the analysis should use a multilocus caller and plan for orthogonal validation of any clinically significant findings.
Decision Point 3: Determine Whether the Target Gene Has Known Paralogs
The third decision point addresses whether the target gene has a well-characterized paralog with clinical significance. Genes such as SMN1 and SMN2 have extensively documented paralogous relationships, and specialized methods have been developed specifically for these genes. Paraphase identifies full-length SMN1 and SMN2 haplotypes, determines gene copy numbers, and calls phased variants using long-read PacBio HiFi data. The copy-number calls by Paraphase are highly concordant with orthogonal methods, achieving 99.2 percent concordance for SMN1 and 100 percent for SMN2.
For genes with known paralogs, researchers should use the specialized method if one exists instead of attempting to adapt a general-purpose caller. For genes without characterized paralogs, the analysis should use a general multilocus approach that can identify PSVs from population data and estimate genotypes for each repeat copy.
Decision Point 4: Evaluate the Required Resolution
The fourth decision point considers whether the analysis requires copy-number information, phased variants, or both. Some clinical applications require only the presence or absence of a pathogenic variant in a specific gene. Other applications require full haplotype information to assess carrier status or to determine whether two variants are on the same chromosome.
Paraphase extends beyond simple copy-number testing by detecting pathogenic variants and enabling potential haplotype-based screening of silent carriers through statistical phasing of haplotypes into alleles. The method identified two SMN1 haplotypes that form a common two-copy SMN1 allele in African populations, and testing positive for these haplotypes in an individual with two copies of SMN1 gives a silent carrier risk of 88.5 percent. This level of resolution is not available from standard variant calling pipelines.
Decision Point 5: Consider Computational Resources and Throughput
The fifth decision point addresses practical constraints on computational resources and sample throughput. Specialized methods for segmental duplications are more computationally intensive than generic callers because they perform additional analysis steps such as constructing networks of homologous sequences, phasing reads, and solving optimization problems. SDrecall uses integer linear programming to retain the two most plausible haplotypes after graph-based phasing and assembly, which adds computational overhead.
For large cohorts where whole-genome analysis is required, restricting specialized analysis to target regions of interest may be more practical than running specialized callers genome-wide. Researchers should benchmark the computational cost of specialized methods on a small subset of samples before committing to a large-scale analysis.
Decision Point 6: Plan for Validation and Confirmation
The final decision point requires planning for validation before the analysis begins. Variants in segmental duplications should be validated using orthogonal methods that can distinguish between paralogous copies. The choice of validation method depends on the variant type and the available resources. Sanger sequencing with paralog-specific primers can confirm variant presence in a specific copy, while digital droplet PCR can validate copy number. Long-read sequencing of the specific region provides the most definitive confirmation of variant presence and phase.
For clinical applications, the validation plan should be documented before the analysis begins so that resources are allocated appropriately. The validation results should be recorded and used to assess the performance of the analytical method in the specific genomic context.
Implementing the Decision Framework
The decision framework can be implemented as a structured assessment that is completed before the analysis pipeline is selected. The assessment should be documented in the analysis log so that the rationale for method selection is transparent and reproducible.
Step 1: Document the Target Region Characteristics
Record the genomic coordinates of the target region, the number of paralogous copies, the sequence identity between copies, and the presence of any known clinically significant genes. This documentation supports the selection of appropriate methods and provides context for interpreting results.
Step 2: Inventory Available Data and Resources
List the available sequencing data types, coverage levels, computational resources, and validation capabilities. This inventory determines which analytical approaches are feasible and which validation strategies can be implemented.
Step 3: Select the Primary Analysis Method
Based on the first five decision points, select the primary analysis method. The selection should be documented with the rationale for each decision point. For example, a researcher analyzing SMN1 with long-read HiFi data would select Paraphase because the gene has a known paralog, the data type supports haplotype-based analysis, and the clinical application requires copy-number and phased variant information.
Step 4: Define the Validation Strategy
Define the validation strategy before the analysis begins. The strategy should specify which variants will be validated, which validation methods will be used, and what criteria will determine whether a variant is confirmed or rejected.
Step 5: Record Outcomes and Refine the Approach
After the analysis is complete, record the outcomes including the number of variants identified, the validation results, and any discrepancies between the analytical and validation results. This information supports refinement of the approach for future analyses and contributes to the understanding of method performance in specific genomic contexts.
Comparison of Method Capabilities Across Decision Scenarios
The decision framework can be summarized by comparing method capabilities across the key decision scenarios. For short-read data in simple duplications, ParascopyVC provides a multilocus approach that aggregates reads across repeat copies and uses population data to identify PSVs. Benchmarking showed that ParascopyVC achieved higher precision and recall than GATK, FreeBayes, and DeepVariant across low-copy repeat regions, with a mean F1 score of 0.947 across seven human genomes.
For long-read data in simple duplications, DuploMap improves mapping accuracy by analyzing reads mapped to segmental duplications using existing long-read aligners and leveraging PSVs to distinguish between multiple alignment locations. On simulated datasets, DuploMap increased the percentage of correctly mapped reads with high confidence for multiple long-read aligners including Minimap2 and BLASR while maintaining high precision.
For long-read data in genes with known paralogs, Paraphase provides haplotype-based analysis that identifies full-length haplotypes, determines gene copy numbers, and calls phased variants. This method is particularly valuable for clinical applications such as carrier screening where haplotype information is essential.
For complex duplication architectures with multiple copies, SDrecall constructs a network of homologous sequences and realigns reads to each segmental duplication from its homologous counterparts. The realignments are phased and assembled into haplotypes via graph-based algorithms, followed by integer linear programming to retain the two most plausible haplotypes. Tested against long-read benchmarks, SDrecall achieved 95 percent sensitivity while maintaining manageable false positives for short variants.
Records for Method Selection and Performance Tracking
Maintaining records of method selection and performance is essential for improving variant calling in segmental duplications over time. The records should include the decision framework assessment for each target region, the selected method and its version, the validation results, and any discrepancies between expected and observed performance.
The analysis log should record the sequencing platform and version, aligner name and version with all non-default parameters, variant caller name and version with all non-default parameters, reference genome version, filtering thresholds, and software container or environment versions. Workflow management systems such as those provided by the nf-core community support reproducible pipeline execution and documentation of these parameters.
Version control for analysis scripts using Git or a similar system ensures that changes to the analysis approach are tracked and can be reviewed. The Carpentries lessons provide foundational training in version control with Git, which is essential for maintaining reproducible analysis workflows.
Common Mistakes in Method Selection
Mistake 1: Applying Generic Callers Without Modification
The most common mistake is applying a generic variant caller such as GATK or DeepVariant to segmental duplications without modification. These callers assume unique mapping and diploid genotypes, which leads to systematic errors in duplicated regions. Benchmark evaluations have shown that ParascopyVC achieves higher precision and recall than these generic callers in low-copy repeat regions.
Mistake 2: Ignoring Copy Number Variation
Researchers who ignore copy number variation in duplicated regions will misinterpret allele balance patterns. A variant present in one copy of a two-copy duplication appears at a lower allele balance than a heterozygous variant in a unique region. Variant callers that assume diploid copy number will produce incorrect genotype calls.
Mistake 3: Discarding Low Mapping Quality Reads
Standard pipelines often discard reads with low mapping quality, but these reads may contain the only evidence for variants in duplicated regions. Specialized methods such as ParascopyVC utilize reads independent of mapping quality in low-copy repeats, recovering information that generic pipelines lose.
Mistake 4: Failing to Plan for Validation
Researchers who do not plan for validation before the analysis begins may find that they lack the resources or methods needed to confirm variants in duplicated regions. Validation planning should be part of the initial analysis design, not an afterthought.
When to Escalate to Specialized Expertise
Researchers should escalate to a bioinformatics specialist when the decision framework identifies complex duplication architecture, when multiple specialized methods produce discordant results, or when the clinical significance of a variant depends on which paralogous copy it is in. Specialized expertise is also needed when integrating multiple data types such as short-read, long-read, and linked-read data for a single target region.
The decision framework provides a structured approach to method selection that reduces the risk of applying inappropriate analytical methods to segmental duplications. By documenting the rationale for each decision and tracking performance outcomes, researchers can build institutional knowledge that improves variant calling accuracy over time.
Frequently Asked Questions
What causes false positive variants in segmental duplications?
False positive variants in segmental duplications are primarily caused by reads that align to the wrong paralogous copy. When a read from one copy aligns to another copy, the sequence differences between the copies, known as paralogous sequence variants, are incorrectly interpreted as variants in the sample. Standard aligners and variant callers that assume unique mapping cannot distinguish between true variants and PSVs, leading to systematic false positives in duplicated regions.
How do paralogous sequence variants differ from true variants?
Paralogous sequence variants are fixed sequence differences between paralogous copies of a duplicated region that exist in the reference genome. They are not variants in the sample being analyzed. True variants are differences between the sample genome and the reference genome. The challenge is that when reads from one copy align to another copy, PSVs appear as variants because the read sequence differs from the reference at that location.
Can short-read sequencing accurately detect variants in segmental duplications?
Short-read sequencing can detect variants in segmental duplications when specialized analysis methods are used. Generic variant callers perform poorly in these regions, but methods such as ParascopyVC that aggregate reads across repeat copies and use PSVs to differentiate copies achieve substantially higher accuracy. However, short-read approaches still have limitations in regions with very high sequence identity or complex duplication structure, and long-read sequencing may be necessary for definitive results.
What is the advantage of long-read sequencing for segmental duplications?
Long-read sequencing provides read lengths that can span larger portions of duplicated regions, enabling more confident assignment of reads to their correct paralogous copy. The ability to characterize repetitive regions is limited by the read lengths of short-read sequencing technologies, and long-read technologies can overcome this limitation. However, long segmental duplications with high sequence identity still pose challenges for long-read mapping, and specialized methods such as DuploMap that use PSVs to refine mapping are needed for optimal accuracy.
How should I validate variants identified in segmental duplications?
Variants in segmental duplications should be validated using orthogonal methods that can distinguish between paralogous copies. Sanger sequencing with paralog-specific primers is a common approach, but the primers must be designed to amplify only the copy of interest. Digital droplet PCR can validate copy number, and long-read sequencing of the specific region can confirm variant presence and phase. Linked-read data can also support variant calls by providing long-range information about read origin.
What is the role of population data in variant calling in segmental duplications?
Population data are used to identify paralogous sequence variants that differentiate repeat copies. Methods such as ParascopyVC use population data to identify PSVs and estimate the genotype of variants for each repeat copy. The accuracy of these methods depends on the completeness of the PSV catalog, which may be limited in populations that are underrepresented in reference datasets.
How does copy number variation affect variant calling in duplicated regions?
Copy number variation affects the expected allele balance for variants in duplicated regions. A variant present in one copy of a two-copy duplication would be expected at a lower allele balance than a heterozygous variant in a unique region. Variant callers that assume diploid copy number will misinterpret variants in regions with different copy numbers. Specialized callers that estimate copy number as part of the variant calling process can resolve this issue.
When should I use a haplotype-based method such as Paraphase?
Haplotype-based methods such as Paraphase are appropriate when the analysis requires full-length haplotypes, gene copy numbers, and phased variants for a specific gene with known paralogs. Paraphase was developed for SMN1 and SMN2 analysis using long-read PacBio HiFi data and can identify silent carriers through statistical phasing of haplotypes into alleles. This approach is valuable for clinical applications such as carrier screening and molecular diagnosis of spinal muscular atrophy.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Object Detection in Medical Images: Algorithms, Datasets, and Validation Strategies
- From Raw Reads to Variants: A Diagnostic Blueprint for Next-Generation Sequencing (NGS) Workflows
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- Research Data Stewardship: Benefits and Implementation Strategies
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Sensitive alignment using paralogous sequence variants improves long-read mapping and variant calling in segmental duplications.. Nucleic acids research, 2020.
- Comprehensive SMN1 and SMN2 profiling for spinal muscular atrophy analysis using long-read PacBio HiFi sequencing.. American journal of human genetics, 2023.
- SDrecall: a sensitive approach for variant detection in segmental duplications.. Genome biology, 2026.
- The complete genome of a songbird.. bioRxiv : the preprint server for biology, 2025.
- A multilocus approach for accurate variant calling in low-copy repeats using whole-genome sequencing.. Bioinformatics (Oxford, England), 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.