# Paralogous Sequence Variants (PSVs) vs. True Variants: How to Distinguish Them in Segmental Duplications


## Key Takeaways

- Paralogous Sequence Variants (PSVs) are fixed sequence differences between duplicated genomic regions that are frequently misidentified as true heterozygous variants due to ambiguous read mapping. This artifact arises because short sequencing reads originating from one paralog can align with equal confidence to another, leading to a mixed allele signal at a single reference coordinate.

- Distinguishing PSVs from true variants relies on analyzing read mapping characteristics and population allele frequency patterns. PSVs are typically supported by reads with low mapping quality scores, indicating ambiguity in their origin, whereas true variants are supported by uniquely mapping reads with high confidence.

- Population allele frequency data is a critical filter: PSVs, being fixed differences, appear in nearly all individuals and thus exhibit near-fixed allele frequencies across population databases, unlike rare true variants. This high frequency in diverse cohorts strongly suggests a PSV artifact rather than a genuine biological variant.

- A systematic workflow for PSV filtering involves identifying variants within annotated segmental duplication regions, examining mapping quality distributions, checking allele balance consistency across multiple samples, querying population databases, and performing paralog-specific sequence comparisons.

- Orthogonal validation methods, such as targeted PCR amplification and sequencing of specific paralogs, are essential for confirming genuine variants and definitively ruling out PSV artifacts, especially for clinically significant findings in duplicated regions.

---

Segmental duplications are large, nearly identical blocks of genomic sequence that complicate variant calling because reads originating from one paralogous copy can align equally well to another copy. Paralogous sequence variants (PSVs) are fixed differences between duplicated copies, and they are frequently miscalled as true variants when they are actually artifacts of ambiguous read mapping. This article explains how to distinguish PSVs from genuine biological variants using allele frequency patterns, read mapping characteristics, and targeted validation strategies. The practical outcome is a filtering workflow that reduces false positives in germline and somatic variant calling pipelines when working with segmental duplications.

## The Core Problem: Why Segmental Duplications Break Standard Variant Calling

Standard variant calling assumes that each genomic region is unique and that reads mapping to a locus originate from that locus. Segmental duplications violate this assumption. These regions are defined as DNA segments larger than 1 kilobase that share high sequence identity, often exceeding 90 percent, between two or more locations in the genome. When a read originates from one copy of a duplicated region, aligners may map it to either copy with nearly equal confidence. The result is a pileup of reads at a position that mixes sequences from two different genomic locations.

PSVs are the sequence differences that distinguish one paralog from another. They are fixed in the population, meaning every individual carries the same difference between copies. When reads from two paralogs collapse into a single alignment position, the PSV appears as a heterozygous variant call. This is a false positive because the individual does not actually carry a variant at that position. The apparent heterozygosity is an artifact of comparing two different genomic locations against one reference coordinate.

The scale of this problem is substantial. A significant fraction of PSVs in segmental duplications overlaps with variant calls and adversely impacts short-read variant calling, as demonstrated in work on long-read mapping and variant calling in these regions [<a href="#ref-1">1</a>]. For researchers working with whole-genome or exome data, this means that a meaningful proportion of candidate variants in duplicated regions may be artifacts instead of true biological findings.

## At a Glance: PSV versus True Variant Decision Table

| Feature | Paralogous Sequence Variant (PSV) | True Heterozygous Variant | True Homozygous Variant |
| --- | --- | --- | --- |
| Allele frequency in population databases | Fixed or near-fixed in all individuals, appears as homozygous alternate in most samples | Variable, ranges from rare to common depending on the variant | Variable, ranges from rare to common depending on the variant |
| Allele balance within a single sample | Approximately 50 percent alternate allele fraction when two paralogs collapse into one position | Approximately 50 percent alternate allele fraction in germline heterozygotes | Near 100 percent alternate allele fraction |
| Read mapping characteristics | Reads supporting the alternate allele map with equal quality to multiple genomic locations, mapping quality scores are low or ambiguous | Reads supporting the alternate allele map uniquely to the variant locus with high mapping quality | Reads supporting the alternate allele map uniquely to the variant locus with high mapping quality |
| Consistency across samples | Present in nearly every sample at the same position with the same alternate allele | Present in a subset of samples, allele frequency varies by population | Present in a subset of samples, allele frequency varies by population |
| Validation by orthogonal method | PCR or capture-based enrichment followed by sequencing of a single copy confirms the alternate allele is not present | PCR or capture-based validation confirms the alternate allele is present in the original DNA | PCR or capture-based validation confirms the alternate allele is present in both alleles |

## The Biology of Segmental Duplications and PSVs

Segmental duplications arise through genomic rearrangements that copy large DNA blocks to new locations. These duplicated blocks can be arranged in tandem or dispersed across different chromosomes. Over evolutionary time, the copies accumulate mutations independently. PSVs are the mutations that became fixed in one copy but not the other. They serve as molecular signatures that distinguish paralogs from each other.

The human genome contains many gene families that exist as paralogs due to segmental duplications. The DExD/H-box RNA helicase family, for example, is a large paralogous gene family where pathogenic variation in a subset of members causes neurodevelopmental disorders and cancer [<a href="#ref-2">2</a>]. Similarly, the SMARCA4 and SMARCA2 genes are paralogs encoding catalytic subunits of chromatin remodeling complexes, and cells that lose SMARCA4 function depend on SMARCA2 for survival [<a href="#ref-3">3</a>]. These paralogous relationships create challenges for variant interpretation because a variant identified in one gene may actually be a PSV from its paralog.

The EP400 gene provides another example. EP400 encodes a core catalytic ATPase subunit of ATP-dependent chromatin remodeling complexes, and its paralogs are associated with neurodevelopmental disorders and epilepsy [<a href="#ref-4">4</a>]. When researchers perform exome sequencing to identify disease-causing variants in such genes, they must distinguish true variants from PSVs that originate from paralogous copies.

PSVs are not rare or exotic features. They are intrinsic to the structure of duplicated genomes. Any variant caller that processes reads from segmental duplications must account for them. The challenge is that most standard tools were designed for unique regions of the genome and do not model the possibility that reads at a single coordinate may come from multiple loci.

## How PSVs Create False Variant Calls

The mechanism by which PSVs become false variant calls involves read alignment and the reference genome. When sequencing reads are generated from a DNA sample, they are aligned to a reference genome. The reference genome contains one representation of each segmental duplication. If a read originates from a paralogous copy that differs from the reference at a PSV position, the read will show a mismatch at that position.

For a unique region, a mismatch between a read and the reference indicates either a sequencing error or a true variant. For a segmental duplication, the mismatch may indicate that the read comes from a different paralog. If reads from two paralogs both align to the same reference coordinate, the position will show a mix of reference and alternate alleles. This mix looks like a heterozygous variant.

The key distinction is that a true heterozygous variant occurs when an individual carries two different alleles at a single locus. A PSV artifact occurs when reads from two different loci are compared against a single reference coordinate. The individual may be homozygous at both loci, but the alignment makes it appear heterozygous.

This problem is compounded by the fact that many variant callers use allele frequency as a filter. A PSV artifact typically shows approximately 50 percent alternate allele fraction, which is the same pattern expected for a true germline heterozygous variant. Allele frequency alone cannot distinguish these cases.

## Read Mapping Quality as a Distinguishing Feature

Read mapping quality is one of the most useful features for distinguishing PSVs from true variants. Mapping quality reflects the confidence that a read is placed at the correct genomic location. When a read maps equally well to multiple locations, the mapping quality is low. This is exactly what happens with reads from segmental duplications.

A true variant in a unique region is supported by reads with high mapping quality. These reads align to only one location in the genome, and the aligner assigns them high confidence scores. A PSV artifact is supported by reads that align to multiple locations with similar scores. The aligner cannot determine which location is correct, so it assigns low mapping quality.

The practical implication is that variant calls supported predominantly by low-mapping-quality reads should be treated with suspicion. This is particularly true when the variant is located within a known segmental duplication. Variant callers that incorporate mapping quality into their probabilistic models will assign lower confidence to variants supported by ambiguous reads.

Long-read sequencing technologies can partially overcome these limitations because longer reads span more of the duplicated region and can contain multiple PSVs that uniquely identify the paralog of origin. Work on long-read mapping in segmental duplications has shown that leveraging PSVs improves mapping accuracy. A probabilistic method called DuploMap analyzes reads mapped to segmental duplications and uses PSVs to distinguish between multiple alignment locations. On simulated datasets, this approach increased the percentage of correctly mapped reads with high confidence for multiple long-read aligners while maintaining high precision [<a href="#ref-1">1</a>].

## Allele Frequency Patterns Across Populations

Population allele frequency is a powerful filter for identifying PSV artifacts. Because PSVs are fixed differences between paralogs, they appear in every individual. A variant that is called in nearly 100 percent of samples in a population database is unlikely to be a rare true variant. It is more likely a PSV artifact or a common polymorphism.

The distinction between a PSV and a common polymorphism requires additional evidence. A common polymorphism is a true variant that segregates in the population at high frequency. A PSV artifact is not a true variant at all. The difference can be resolved by examining whether the variant is consistently called in all samples and whether the supporting reads have ambiguous mapping.

Population databases such as those maintained by NCBI provide allele frequency information for variants across diverse populations [<a href="#ref-5">5</a>]. When a candidate variant in a segmental duplication shows near-fixed allele frequency in these databases, it should be flagged for further investigation. The variant may be a PSV artifact instead of a true biological variant.

For somatic variant calling, the situation is different. Somatic variants are not expected to appear in population databases because they arise in individual tumors. However, PSV artifacts will still appear in somatic calls because they are present in the germline DNA of every sample. A somatic variant caller that does not account for segmental duplications will call PSV artifacts as somatic variants in every tumor sample.

## Practical Workflow for Distinguishing PSVs from True Variants

The following workflow provides a systematic approach to filtering PSV artifacts from variant calls in segmental duplications. This workflow assumes that raw sequencing data and aligned reads are available for analysis.

### Step 1: Identify Variant Calls in Segmental Duplication Regions

The first step is to determine which variant calls fall within segmental duplications. Segmental duplication annotations are available from genome browsers and annotation databases. NCBI provides sequence resources and analysis services that include genomic annotation data [<a href="#ref-5">5</a>]. Overlap your variant calls with segmental duplication coordinates to create a list of candidate PSV artifacts.

### Step 2: Examine Mapping Quality Distributions

For each candidate variant, examine the mapping quality of reads supporting the alternate allele. Extract the reads that support the alternate allele and inspect their mapping quality scores. If the majority of supporting reads have low or ambiguous mapping quality, the variant is likely a PSV artifact. If the supporting reads have high mapping quality and map uniquely, the variant may be genuine.

### Step 3: Check Allele Balance Across Multiple Samples

PSV artifacts show consistent allele balance across all samples. If you have access to multiple samples from the same population, check whether the variant is called at approximately 50 percent alternate allele fraction in every sample. True rare variants will appear in only a subset of samples. PSV artifacts will appear in nearly all samples.

### Step 4: Query Population Databases

Submit the variant coordinates to population databases to check allele frequency. NCBI maintains databases that aggregate variant information across large cohorts [<a href="#ref-5">5</a>]. If the variant shows near-fixed allele frequency, it is likely a PSV artifact. If the variant is absent or rare in population databases, it may be a true variant.

### Step 5: Apply Paralog-Specific Analysis

For genes with known paralogs, compare the candidate variant against the paralogous sequence. If the alternate allele matches the sequence of the paralog at the corresponding position, the variant is almost certainly a PSV artifact. This analysis requires a multiple sequence alignment of the gene family.

### Step 6: Validate with Orthogonal Methods

For variants that remain ambiguous after filtering, orthogonal validation is necessary. PCR amplification using primers specific to one paralog, followed by sequencing, can determine whether the variant is present in the original DNA. This approach requires careful primer design to avoid amplifying both paralogs.

## Tools and Resources for PSV Analysis

Several categories of tools and resources support PSV analysis. The choice of tools depends on the sequencing platform and the specific analysis goals.

### Read Alignment and Mapping Quality

Read aligners differ in how they handle multi-mapping reads. Some aligners report all possible mapping locations, while others report only the best location. The choice of aligner affects the mapping quality scores available for downstream analysis. For long-read data, aligners such as Minimap2 and BLASR are commonly used, and methods like DuploMap can improve mapping accuracy in segmental duplications by leveraging PSVs [<a href="#ref-1">1</a>].

### Variant Calling with Duplication Awareness

Some variant callers incorporate models for segmental duplications and can assign lower confidence to variants in these regions. The choice of variant caller should be informed by the known limitations of each tool. No variant caller perfectly resolves segmental duplications, so downstream filtering remains necessary.

### Population Frequency Databases

Population frequency data are essential for distinguishing PSVs from true variants. NCBI provides access to databases that aggregate variant information across large cohorts [<a href="#ref-5">5</a>]. These resources allow researchers to check whether a candidate variant is present at high frequency in the population, which would suggest a PSV artifact.

### Training and Educational Resources

Bioinformatics training resources provide practical guidance for working with genomic data. EMBL-EBI offers training on bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover variant calling and filtering [<a href="#ref-7">7</a>]. The Carpentries offers foundational computing and data skills that support reproducible analysis workflows [<a href="#ref-8">8</a>]. These resources are valuable for researchers who need to build or refine their variant calling pipelines.

### Reproducible Workflow Frameworks

Reproducibility is critical for variant calling because different tools and parameters produce different results. Community pipeline standards provide structured approaches to building reproducible analysis workflows. The nf-core documentation describes community pipeline standards, usage, and configuration for reproducible workflows [<a href="#ref-9">9</a>]. Bioconductor provides official package and workflow documentation for reproducible genomic analysis [<a href="#ref-10">10</a>]. Adopting these frameworks ensures that PSV filtering steps are documented and repeatable.

## Variant Calling Workflow Considerations

The choice of variant calling workflow affects how PSVs are handled. Germline and somatic variant calling have different requirements and different approaches to filtering.

### Germline Variant Calling

Germline variant calling identifies variants present in the inherited genome. These variants are present in all cells of the body and are expected to show approximately 50 percent alternate allele fraction for heterozygotes. PSV artifacts mimic this pattern, making them difficult to distinguish without additional evidence.

For germline variant calling, population frequency is a powerful filter. A variant that appears in nearly all samples in a population database is unlikely to be a rare disease-causing variant. It is more likely a PSV artifact or a common polymorphism. The distinction requires examining mapping quality and paralog-specific sequences.

Family-based analysis can also help. If a variant is called in all family members, including those who do not share the disease phenotype, it may be a PSV artifact. True disease-causing variants should segregate with the phenotype in families where the variant is causal.

### Somatic Variant Calling

Somatic variant calling identifies variants that arise in tumors or other tissues. These variants are not present in the germline and are not expected to appear in population databases. However, PSV artifacts will appear in somatic calls because they are present in the germline DNA of every sample.

Somatic variant callers typically compare tumor and normal samples from the same individual. Variants present in both samples are filtered out as germline. However, PSV artifacts may be called differently in tumor and normal samples due to differences in coverage and allele balance. This can lead to PSV artifacts being called as somatic variants.

For somatic variant calling, the key filter is whether the variant is present in the matched normal sample. If a variant is called in the tumor but not the normal sample, it may be a true somatic variant or a PSV artifact that was not called in the normal due to lower coverage. Examining the raw reads in the normal sample can resolve this ambiguity.

### Variant Filtering Strategies

Variant filtering is the process of removing false positives from variant calls. For segmental duplications, the following filters are particularly useful:

- Mapping quality filter: Remove variants supported predominantly by low-mapping-quality reads.
- Allele balance filter: Remove variants with allele balance inconsistent with the expected pattern for the sample type.
- Population frequency filter: Remove variants with near-fixed allele frequency in population databases.
- Segmental duplication overlap filter: Flag variants that fall within known segmental duplications for additional scrutiny.
- Paralog match filter: Remove variants where the alternate allele matches the paralogous sequence.

These filters should be applied in combination instead of individually. A variant that passes all filters is more likely to be genuine than a variant that passes only one or two.

## Common Failure Patterns in PSV Filtering

Several common failure patterns lead to PSV artifacts being reported as true variants. Recognizing these patterns helps researchers avoid them.

### Failure Pattern 1: Ignoring Mapping Quality

The most common failure is ignoring mapping quality when filtering variants. Many researchers filter on depth and allele frequency but do not examine mapping quality. This allows PSV artifacts to pass through the pipeline because they have adequate depth and the expected allele balance.

The solution is to incorporate mapping quality into the filtering process. Variants supported by reads with low mapping quality should be flagged or removed. This is particularly important for variants in segmental duplications.

### Failure Pattern 2: Relying Solely on Population Frequency

Population frequency is a useful filter, but it is not sufficient on its own. Some true variants are common in the population, and some PSV artifacts may not be present in population databases if the databases were built using methods that filtered them out.

The solution is to use population frequency in combination with mapping quality and paralog-specific analysis. A variant with high population frequency and low mapping quality is almost certainly a PSV artifact. A variant with high population frequency and high mapping quality may be a true common polymorphism.

### Failure Pattern 3: Using Inappropriate Reference Sequences

The choice of reference genome affects PSV detection. If the reference genome contains only one copy of a segmental duplication, reads from the other copy will show mismatches that look like variants. If the reference genome contains both copies, reads can be aligned to the correct copy.

The solution is to be aware of the reference genome limitations and to use alternative references or pangenome graphs when working with segmental duplications. These approaches can improve mapping accuracy and reduce PSV artifacts.

### Failure Pattern 4: Applying the Same Filters to All Genomic Regions

Variant calling pipelines often apply the same filters to all genomic regions. This approach fails for segmental duplications because these regions have different error profiles than unique regions.

The solution is to apply region-specific filters. Variants in segmental duplications should be subjected to stricter filtering than variants in unique regions. This may involve requiring higher mapping quality, lower allele balance deviation, or orthogonal validation.

### Failure Pattern 5: Not Validating Candidate Variants

The final common failure is not validating candidate variants with orthogonal methods. Even with careful filtering, some PSV artifacts will pass through. Orthogonal validation using PCR or capture-based enrichment can confirm whether a variant is genuine.

The solution is to validate clinically significant variants in segmental duplications before reporting them. This is particularly important for variants that will be used for clinical decision-making.

## Records and Measurements for PSV Analysis

Maintaining detailed records of the PSV filtering process is essential for reproducibility and for building confidence in variant calls. The following records should be maintained for each analysis.

### Variant Call Records

For each variant call, record the genomic coordinates, reference and alternate alleles, allele frequency, mapping quality distribution, and the results of each filtering step. This information allows the filtering process to be audited and repeated.

### Filtering Decision Logs

Record the rationale for each filtering decision. If a variant was removed because of low mapping quality, document the mapping quality threshold used. If a variant was flagged for orthogonal validation, document the reason for the flag.

### Sample Metadata

Record the sample type, sequencing platform, coverage, and any relevant clinical information. This metadata is essential for interpreting variant calls and for comparing results across samples.

### Pipeline Configuration

Record the exact tools and parameters used for alignment, variant calling, and filtering. This information is essential for reproducing the analysis and for troubleshooting unexpected results.

## Quality Controls for PSV Filtering

Quality controls ensure that the filtering process is working correctly and that the final variant calls are reliable.

### Positive Controls

Include samples with known variants in segmental duplications to verify that the pipeline can detect true variants in these regions. These controls should be processed through the entire pipeline to confirm that the filtering steps do not remove genuine variants.

### Negative Controls

Include samples with no expected variants in segmental duplications to verify that the pipeline does not produce false positives. These controls should be processed through the entire pipeline to confirm that PSV artifacts are being filtered correctly.

### Replicate Analysis

Process replicate samples to assess the reproducibility of variant calls. Variants that are called consistently across replicates are more reliable than variants that appear in only one replicate.

### Cross-Platform Validation

If possible, validate a subset of variants using a different sequencing platform or a different analysis pipeline. Cross-platform validation provides strong evidence that a variant is genuine.

## Limitations of PSV Filtering Approaches

PSV filtering approaches have inherent limitations that researchers must understand.

### Incomplete Segmental Duplication Annotations

Segmental duplication annotations are not complete. Some duplicated regions may not be annotated, and variants in these regions will not be flagged for additional scrutiny. This limitation means that some PSV artifacts will pass through even well-designed filtering pipelines.

### Reference Genome Bias

Reference genomes contain one representation of each segmental duplication. This representation may not match either paralog perfectly, and it may contain errors. Reference genome bias can cause both false positives and false negatives in variant calling.

### Short-Read Limitations

Short-read sequencing cannot span entire segmental duplications. This limitation means that short reads cannot always be assigned to the correct paralog, even with sophisticated alignment algorithms. Long-read sequencing can overcome this limitation but is more expensive and less widely available.

### Population Database Gaps

Population databases may not include all populations or all variant types. A variant that is absent from a population database may still be a common polymorphism in an underrepresented population. This limitation means that population frequency filters must be applied with caution.

### Somatic Variant Complexity

Somatic variant calling in segmental duplications is particularly challenging because tumor samples may have copy number alterations that affect allele balance. These alterations can make PSV artifacts appear as somatic variants or can mask true somatic variants.

## Safety and Regulatory Context for Clinical Reporting

When variant calls are used for clinical decision-making, additional considerations apply. PSV artifacts that are reported as true variants can lead to incorrect diagnoses and inappropriate treatments.

### Clinical Validation Requirements

Variants reported for clinical use should be validated using orthogonal methods. This is particularly important for variants in segmental duplications, where the risk of PSV artifacts is high. Clinical laboratories should have documented validation procedures that include testing for PSV artifacts.

### Reporting Standards

Clinical reports should clearly indicate when a variant is located in a segmental duplication and when the evidence for the variant is limited by mapping ambiguity. This information allows clinicians to interpret the variant appropriately.

### Professional Escalation Criteria

When a variant in a segmental duplication is identified as a candidate for clinical action, the case should be escalated to a molecular pathologist or clinical geneticist for review. These professionals can assess the evidence and determine whether the variant should be reported.

### Data Sharing and Reinterpretation

Variant data from segmental duplications should be shared with appropriate databases to improve population frequency estimates and to support reinterpretation as new evidence becomes available. NCBI provides resources for data sharing and analysis [<a href="#ref-5">5</a>].

## Professional Escalation Criteria

The following criteria indicate when a variant call in a segmental duplication should be escalated for expert review:

- The variant is in a clinically significant gene and falls within a segmental duplication.
- The variant has been flagged as a potential PSV artifact but cannot be definitively classified.
- The variant is being considered for clinical action and has not been validated by an orthogonal method.
- The variant shows discordant results across different analysis pipelines or sequencing platforms.
- The variant is in a gene with known paralogs and the alternate allele matches the paralogous sequence.

In these cases, consultation with a molecular geneticist, bioinformatician, or clinical laboratory director is appropriate before the variant is reported.

## Building a Reproducible PSV Filtering Pipeline

Reproducibility is essential for variant calling because different tools and parameters produce different results. A reproducible pipeline ensures that the same input data produces the same output, regardless of who runs the analysis.

### Version Control

Use version control for all pipeline code and configuration files. This allows the exact pipeline used for each analysis to be identified and reproduced.

### Containerization

Use containers to package the pipeline and its dependencies. Containers ensure that the pipeline runs in the same environment regardless of the host system.

### Documentation

Document all pipeline steps, including tool versions, parameters, and filtering thresholds. This documentation should be sufficient for another researcher to reproduce the analysis.

### Workflow Management

Use workflow management tools to orchestrate the pipeline. These tools track the execution of each step and provide logs that can be used for troubleshooting.

The nf-core community provides documentation on pipeline standards, usage, and configuration that supports reproducible workflow development [<a href="#ref-9">9</a>]. Bioconductor provides official documentation for reproducible genomic analysis using R packages [<a href="#ref-10">10</a>]. The Galaxy Training Network provides tutorials on building and running analysis workflows [<a href="#ref-7">7</a>]. These resources support the development of reproducible PSV filtering pipelines.

## Training and Skill Development for PSV Analysis

PSV analysis requires a combination of genomics knowledge, computational skills, and biological interpretation. Training resources are available for each of these areas.

### Genomics and Variant Calling Fundamentals

EMBL-EBI training provides learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>]. These resources cover the fundamentals of variant calling and genomic analysis.

### Computational Skills

The Carpentries offers lessons on foundational computing, data, shell, Git, and programming [<a href="#ref-8">8</a>]. These skills are essential for building and running reproducible analysis pipelines.

### Workflow Development

The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-7">7</a>]. These tutorials cover the practical aspects of building and running genomic analysis workflows.

### Community Standards

The nf-core documentation describes community pipeline standards for reproducible workflows [<a href="#ref-9">9</a>]. These standards support the development of high-quality analysis pipelines.

## Interpretation of Results and Reporting

The interpretation of variant calls in segmental duplications requires careful consideration of the evidence. A variant that passes all filtering steps is more likely to be genuine, but it is not guaranteed to be genuine. The following considerations apply to interpretation and reporting.

### Evidence Integration

Integrate all available evidence when interpreting a variant. This includes mapping quality, allele frequency, population frequency, paralog sequence comparison, and any orthogonal validation results. A variant supported by multiple lines of evidence is more reliable than a variant supported by only one.

### Uncertainty Communication

Communicate uncertainty clearly in reports. If a variant is in a segmental duplication and the evidence is limited, state this limitation explicitly. This allows downstream users to interpret the variant appropriately.

### Variant Classification

Use standard variant classification frameworks when classifying variants. These frameworks consider population frequency, segregation data, functional evidence, and computational predictions. PSV artifacts should be classified as benign or likely benign because they are not true variants.

### Follow-Up Testing

Recommend follow-up testing when a variant is ambiguous. This may include orthogonal validation, family segregation studies, or functional studies. Follow-up testing can resolve ambiguity and provide definitive evidence.

## Frequently Asked Questions

### What is the difference between a PSV and a common polymorphism?

A PSV is a fixed sequence difference between two paralogous copies of a duplicated region. It is present in every individual because it distinguishes one genomic location from another. A common polymorphism is a true variant that segregates in the population at high frequency. The distinction matters because a PSV is not a variant at all, while a common polymorphism is a genuine biological variant. The two can be distinguished by examining mapping quality and paralog-specific sequences. A PSV artifact will be supported by reads with ambiguous mapping, while a true common polymorphism will be supported by uniquely mapping reads.

### How does mapping quality help identify PSV artifacts?

Mapping quality reflects the confidence that a read is placed at the correct genomic location. Reads from segmental duplications often map equally well to multiple locations, so aligners assign them low mapping quality scores. A true variant in a unique region is supported by reads with high mapping quality because those reads align to only one location. When a variant call is supported predominantly by low-mapping-quality reads, it is likely a PSV artifact. Variant callers that incorporate mapping quality into their models will assign lower confidence to such variants.

### Can long-read sequencing solve the PSV problem completely?

Long-read sequencing can substantially reduce PSV artifacts because longer reads span more of the duplicated region and can contain multiple PSVs that uniquely identify the paralog of origin. Methods like DuploMap leverage PSVs to improve long-read mapping accuracy in segmental duplications [<a href="#ref-1">1</a>]. However, long-read sequencing does not completely solve the problem. Even long reads may not span an entire segmental duplication, and some duplicated regions are too similar for any read length to resolve. Long-read data also require careful alignment and variant calling to avoid introducing new artifacts.

### Why do PSV artifacts appear as heterozygous variants?

PSV artifacts appear as heterozygous variants because reads from two different paralogous copies align to the same reference coordinate. The reference genome contains one representation of each segmental duplication. When reads from both copies align to that single representation, the position shows a mix of reference and alternate alleles. This mix looks like a heterozygous variant, even though the individual may be homozygous at both loci. The apparent heterozygosity is an artifact of comparing two different genomic locations against one reference coordinate.

### What filters are most effective for removing PSV artifacts?

The most effective filters combine mapping quality, population frequency, and paralog-specific sequence comparison. Mapping quality filters remove variants supported by ambiguous reads. Population frequency filters remove variants that appear in nearly all samples, which is the pattern expected for PSVs. Paralog match filters remove variants where the alternate allele matches the sequence of a known paralog. These filters should be applied in combination because no single filter is sufficient on its own.

### How should somatic variant callers handle segmental duplications?

Somatic variant callers should compare tumor and normal samples from the same individual and filter out variants present in both. However, PSV artifacts may be called differently in tumor and normal samples due to differences in coverage and allele balance. Examining the raw reads in the normal sample can resolve this ambiguity. If a variant is called in the tumor but not the normal sample, it may be a true somatic variant or a PSV artifact that was not called in the normal due to lower coverage. Somatic variant callers should also apply mapping quality and segmental duplication overlap filters.

### When should orthogonal validation be performed?

Orthogonal validation should be performed for any variant that will be used for clinical decision-making, particularly when the variant falls within a segmental duplication. Validation is also appropriate when a variant passes all filtering steps but remains ambiguous due to limited evidence. PCR amplification using primers specific to one paralog, followed by sequencing, can determine whether the variant is present in the original DNA. This approach requires careful primer design to avoid amplifying both paralogs.

### What should be done when a variant cannot be definitively classified?

When a variant cannot be definitively classified as a true variant or a PSV artifact, the case should be escalated for expert review. Consultation with a molecular geneticist, bioinformatician, or clinical laboratory director is appropriate. The variant should be reported with clear communication of the uncertainty, including the fact that it is located in a segmental duplication and that the evidence is limited by mapping ambiguity. Follow-up testing, including orthogonal validation and family segregation studies, may resolve the ambiguity.

## Related Bioinformatics Guides

- [Genomic Data Visualization Tools: Choosing and Using Them Effectively](/knowledge/bioinformatics/genomic-data-visualization-tools-choosing-and-using-them-effectively)
- [Volcano Plot Proteomics: How to Create and Interpret Them Effectively](/knowledge/bioinformatics/volcano-plot-proteomics-how-to-create-and-interpret-them-effectively)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Ancestral Sequence Reconstruction for Viral Evolution](/knowledge/bioinformatics/ancestral-sequence-reconstruction-for-viral-evolution)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Sensitive alignment using paralogous sequence variants improves long-read mapping and variant calling in segmental duplications.](https://pubmed.ncbi.nlm.nih.gov/33035301). Nucleic acids research, 2020.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Monoallelic variation in DHX9, the gene encoding the DExH-box helicase DHX9, underlies neurodevelopment disorders and Charcot-Marie-Tooth disease.](https://pubmed.ncbi.nlm.nih.gov/37467750). American journal of human genetics, 2023.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Functional characterization of SMARCA4 variants identified by targeted exome-sequencing of 131,668 cancer patients.](https://pubmed.ncbi.nlm.nih.gov/33144586). Nature communications, 2020.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Variants in EP400, encoding a chromatin remodeler, cause epilepsy with neurodevelopmental disorders.](https://pubmed.ncbi.nlm.nih.gov/39708813). American journal of human genetics, 2025.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.