# Single-Individual Haplotype Phasing with Long Reads: Overcoming the Challenges of De Novo Phasing

Researchers who need phased haplotypes from a single individual's long-read sequencing data face a distinct problem: without parental genotypes or pedigree information, traditional trio-based phasing approaches are unavailable, and the computational strategies that remain must contend with switch errors, missing variants, and fragmented haplotype blocks. This article explains why single-individual phasing is fundamentally harder than pedigree-based phasing, what specific error modes emerge when using long reads alone, and which practical strategies, including ultra-long reads and linked-read data, can improve accuracy. The scope covers laboratory professionals and bioinformatics analysts who have already generated or plan to generate Oxford Nanopore Technologies (ONT) or Pacific Biosciences (PacBio) sequencing data and need to produce phased haplotypes without access to family samples.

## The Core Problem: Reconstructing Two Haplotypes from One Genome

A diploid genome contains two copies of each chromosome, one inherited from each parent. The set of alleles on a single chromosome copy is called a haplotype. Phasing is the process of assigning heterozygous variants to their respective parental chromosomes. When researchers have sequencing data from an individual plus both parents, phasing becomes a matter of comparing allele inheritance patterns. Without parental data, the only signal available is the physical co-occurrence of variants on the same sequencing read or on the same long DNA molecule.

Long-read sequencing platforms produce reads that span multiple heterozygous variants, allowing direct observation of which alleles travel together. This physical linkage is the foundation of single-individual phasing. However, the accuracy and completeness of the resulting haplotypes depend on read length, sequencing depth, variant density, and the algorithms used to assemble haplotypes from read evidence.

The fundamental limitation is that a single read can only phase variants that it physically covers. If two heterozygous variants are separated by a distance greater than the read length, no single read can connect them. The haplotype block, defined as a contiguous stretch of variants that have been assigned to haplotypes, ends at that gap. Researchers must therefore choose strategies that maximize the span of physical connectivity across the genome.

## At a Glance: Single-Individual Phasing Strategies Compared

| Strategy | Data Requirement | Typical Haplotype Block Length | Primary Limitation | Best Use Case |
| --- | --- | --- | --- | --- |
| Standard long reads (ONT or PacBio) | 30x to 60x coverage | 10 kb to several hundred kb | Read length limits connectivity across repetitive or low-complexity regions | Targeted regions, gene-level phasing, clinical amplicon panels |
| Ultra-long reads (ONT) | 30x to 60x coverage with size-selected DNA | Hundreds of kb to Mb scale | Requires high-molecular-weight DNA extraction and library preparation expertise | Whole-genome phasing, structural variant resolution, repeat expansion loci |
| Linked-read sequencing (barcode partitioning) | Standard short-read library with long-molecule barcoding | 100 kb to several Mb | Requires specialized instrumentation and reagents | Scalable phasing without long-read infrastructure |
| Combined long reads plus linked reads | Both data types from same individual | Chromosome scale | Higher cost and computational burden | Highest accuracy requirements, clinical applications, complex genomic regions |

The choice among these strategies depends on the biological question, the genomic region of interest, available infrastructure, and the tolerance for switch errors in downstream analyses.

## Why Single-Individual Phasing Differs from Pedigree-Based Approaches

Pedigree-based phasing uses Mendelian inheritance patterns to assign alleles to haplotypes. When genotypes from parents and offspring are available, the transmission of alleles can be tracked across generations, and haplotype assignment becomes a logical deduction problem. This approach does not require long reads and works with standard short-read sequencing data.

Single-individual phasing removes this information source entirely. The researcher must rely on the physical evidence within the sequencing data itself. This distinction has practical consequences:

- **Error propagation differs.** In pedigree phasing, a genotyping error in one family member can often be detected through Mendelian inconsistencies. In single-individual phasing, a variant calling error creates a false heterozygous site that the phasing algorithm must place on one haplotype or the other, potentially introducing a switch error that affects all downstream variants.
- **Coverage requirements increase.** Pedigree phasing can work with modest coverage per individual because the statistical signal comes from allele transmission patterns across multiple family members. Single-individual phasing requires sufficient coverage to confidently call heterozygous variants and to observe multiple reads spanning each variant pair.
- **Completeness is harder to achieve.** Even with deep coverage, regions of the genome that are difficult to sequence, such as segmental duplications, centromeres, and telomeres, will have gaps in haplotype assignment. Pedigree approaches can sometimes bridge these gaps through inheritance patterns, but single-individual approaches cannot.

The practical implication is that researchers planning single-individual phasing must design their sequencing strategy with the phasing problem in mind from the start. Read length and coverage decisions made at the library preparation stage directly determine the quality of the final haplotypes.

## Core Principles of Long-Read Phasing

### Physical Linkage as the Primary Signal

Long-read phasing works by identifying reads that contain two or more heterozygous variants. Each such read provides evidence that the alleles it carries belong to the same haplotype. When multiple reads support the same allele combination, the phasing algorithm can assign those variants to a haplotype with confidence.

The strength of this signal depends on the density of heterozygous variants in the region and the length of the reads. In a region with one heterozygous variant every 1 kb, a 10 kb read will typically span ten variants, providing substantial phasing information. In a region with sparse variation, such as a gene-poor stretch of the genome, a read may span only one or two variants, and the phasing connections will be weaker.

### Switch Errors and Their Consequences

A switch error occurs when the phasing algorithm incorrectly assigns alleles to haplotypes, effectively switching the haplotype labels for all variants downstream of the error point. For example, if variants 1 through 5 are correctly phased as haplotype A and haplotype B, but variant 6 is incorrectly assigned, then variants 7 through 20 will also be incorrectly assigned relative to the true haplotypes.

Switch errors are particularly problematic because they are invisible in the final output unless the researcher has an independent method to validate the phasing. The error rate is typically reported as the number of switch errors per megabase or as the switch error rate, which is the proportion of heterozygous sites that are incorrectly phased relative to a truth set.

The consequences of switch errors depend on the downstream application. For allele-specific expression analysis, a switch error can cause the researcher to attribute expression to the wrong haplotype. For structural variant analysis, a switch error can break a large deletion or insertion into two smaller events. For clinical interpretation, a switch error can place a pathogenic variant on the wrong haplotype, leading to incorrect compound heterozygosity assessment.

### Missing Variants and Their Impact on Phasing

Phasing algorithms can only phase variants that have been identified. If a heterozygous variant is missed during variant calling, it creates a gap in the haplotype. The surrounding variants may still be phased, but the missing variant cannot be assigned to a haplotype.

Missing variants are particularly common in regions with low complexity, high GC content, or repetitive sequence. These regions are difficult for both short-read and long-read aligners, and variant callers may fail to identify true heterozygous sites. The result is a haplotype that is incomplete, with gaps that may correspond to biologically important regions.

The interaction between missing variants and switch errors is important. A missing variant does not directly cause a switch error, but it reduces the density of phased sites, making it harder for the phasing algorithm to detect and correct errors. In regions with sparse variant density, a single misaligned read can cause a switch error that would have been corrected if additional variant sites had been available.

## Practical Workflow for Single-Individual Phasing

### Step 1: Assess the Biological Question and Region of Interest

Before designing the sequencing experiment, define what the phased haplotypes will be used for. The requirements differ substantially by application:

- **Targeted gene phasing.** If the goal is to phase variants within a single gene or a small genomic region, amplicon-based long-read sequencing may be sufficient. A proof-of-concept study using Oxford Nanopore Technologies data demonstrated reliable single-nucleotide variant calling and phasing from as little as 60 reads in a 12 kb region of the APOE locus, with full haplotype reconstruction achieved using 600 reads when variants were first identified with a highly accurate platform [8].
- **Whole-genome phasing.** If the goal is chromosome-scale haplotypes, the sequencing strategy must maximize read length and coverage across the entire genome. The DipAsm approach, which uses long accurate reads and long-range conformation data, generated chromosome-scale phased assemblies for single individuals with NG50 values up to 25 Mb and phased approximately 99.5 percent of heterozygous sites at 98 to 99 percent accuracy [11].
- **Repeat expansion analysis.** For repeat expansion disorders, the repeat length often exceeds the capacity of short-read sequencing, and targeted long-read approaches are required. Full-length transcript phasing methods have been developed for the HTT gene that can be adapted for other repeat expansion disorders [9].

### Step 2: Choose the Sequencing Platform and Library Preparation

The choice between ONT and PacBio platforms affects read length, accuracy, and cost. Both platforms can produce reads long enough to span multiple heterozygous variants, but the practical details differ:

- **ONT offers ultra-long reads.** With appropriate DNA extraction and library preparation, ONT can produce reads exceeding 100 kb, and some reads can reach megabase lengths. These ultra-long reads dramatically improve phasing connectivity. The tradeoff is lower per-base accuracy compared to PacBio HiFi reads, although the accuracy is sufficient for variant calling and phasing in most applications.
- **PacBio HiFi reads offer high accuracy.** PacBio circular consensus sequencing produces reads of 10 to 25 kb with accuracy exceeding 99.9 percent. The high accuracy simplifies variant calling and reduces the risk of false heterozygous sites, but the shorter read length limits phasing connectivity compared to ultra-long ONT reads.

For single-individual phasing, the optimal strategy may involve combining both platforms. The high accuracy of PacBio HiFi reads can be used for variant calling, while ultra-long ONT reads provide the long-range connectivity needed for phasing. This combined approach was used in the DipAsm method, which employed long accurate reads plus long-range conformation data [11].

### Step 3: Generate Sufficient Coverage

Coverage requirements for phasing are higher than for standard variant calling. The phasing algorithm needs multiple reads spanning each pair of adjacent heterozygous variants to establish phase with confidence. Low coverage results in fragmented haplotypes with many small blocks.

For targeted amplicon approaches, the coverage requirement is modest. The APOE locus study demonstrated that reliable phasing could be achieved with 600 reads covering a 12 kb region [8]. For whole-genome approaches, coverage of 30x to 60x is typically recommended, with higher coverage providing more opportunities to observe variant pairs on the same read.

The relationship between coverage and phasing completeness is not linear. Increasing coverage from 10x to 30x produces a substantial improvement in haplotype block length, but the improvement from 30x to 60x is smaller. The optimal coverage depends on the read length distribution and the variant density in the region of interest.

### Step 4: Perform Variant Calling with Appropriate Tools

Variant calling is a prerequisite for phasing. The phasing algorithm needs a list of heterozygous variants with their genomic positions and alleles. The accuracy of this variant list directly affects phasing quality.

For ONT data, variant calling can be performed with tools designed for nanopore data, which account for the specific error profile of the platform. For PacBio HiFi data, the high accuracy allows the use of standard variant callers with minimal filtering.

The APOE locus study found that ONT data allowed reliable single-nucleotide variant calling and phasing, but the recognition of indels was less efficient [8]. This finding suggests that researchers should consider using a highly accurate sequencing platform for initial variant identification, then use long reads for phasing. The study identified the best combination of ONT read sets and software as BWA or Minimap2 for alignment and HapCUT2 for phasing, enabling full haplotype reconstruction when both SNVs and indels had been identified previously using a highly accurate platform [8].

### Step 5: Run Phasing Algorithms

Several phasing algorithms are available for long-read data. The choice of algorithm depends on the data type and the desired output:

- **HapCUT2** is a widely used phasing tool that works with both short and long reads. It builds a fragment graph from read alignments and uses max-cut approximation to assign variants to haplotypes. The APOE study identified HapCUT2 as part of the best combination for ONT-based phasing [8].
- **WhatsHap** is another popular phasing tool that uses a dynamic programming approach to find the maximum likelihood haplotype assignment given the read data.
- **DipAsm** is a diploid assembly approach that uses long accurate reads and long-range conformation data to generate chromosome-scale phased assemblies [11].

The output of these algorithms is a set of haplotype blocks, each containing a contiguous stretch of phased variants. The blocks are separated by gaps where no read spans the intervening variants.

### Step 6: Evaluate Phasing Quality

Phasing quality should be evaluated using multiple metrics:

- **Haplotype block length.** The N50 of haplotype blocks, which is the length such that 50 percent of the phased genome is in blocks of at least that length, provides a summary of phasing completeness.
- **Switch error rate.** If a truth set is available, such as from a pedigree or from an independent phasing method, the switch error rate can be calculated directly.
- **Phasing completeness.** The proportion of heterozygous variants that are assigned to a haplotype, also called the phasing rate, indicates how much of the genome was successfully phased.

For the DipAsm method, the reported metrics were NG50 up to 25 Mb and approximately 99.5 percent of heterozygous sites phased at 98 to 99 percent accuracy [11]. These numbers provide a benchmark for what is achievable with current technology.

## Options and Tradeoffs in Data Acquisition

### Standard Long Reads versus Ultra-Long Reads

The most direct way to improve single-individual phasing is to increase read length. Ultra-long reads from ONT can span hundreds of kilobases, connecting variants that would be in separate haplotype blocks with standard-length reads.

The tradeoff is that ultra-long read library preparation requires high-molecular-weight DNA. DNA extraction methods that shear DNA into fragments shorter than 50 kb will not produce ultra-long reads, regardless of the sequencing platform. Researchers must use gentle extraction methods, such as agarose plug lysis or magnetic bead-based approaches designed for high-molecular-weight DNA.

The sequencing yield of ultra-long reads is also lower than standard reads. A typical ONT flow cell may produce fewer ultra-long reads than standard-length reads, requiring more flow cells to achieve the same coverage. The cost per gigabase is therefore higher for ultra-long read strategies.

### Linked-Read Sequencing as an Alternative

Linked-read sequencing offers a different approach to single-individual phasing. In this method, long DNA molecules are partitioned and barcoded such that all sequencing reads originating from the same long molecule share a common barcode. The barcode information provides long-range connectivity without requiring long sequencing reads.

The CPTv2-seq method demonstrated barcode partitioning of long DNA molecules in a single tube using on-bead barcoded tagmentation. This approach transferred homogeneous populations of barcodes from beads to individual long DNA molecules that were fragmented at the same time, producing sequencing libraries where all reads from each long DNA molecule shared a common barcode [10]. The method provided a barcode-linked read structure that revealed long-range molecular contiguity and enabled accurate haplotype-resolved sequencing and phasing of structural variants [10].

Linked-read approaches have the advantage of using standard short-read sequencing infrastructure, which is widely available and less expensive than long-read platforms. The tradeoff is that the barcode information is probabilistic instead of deterministic. If two variants are on the same long DNA molecule but the molecule is not fully sequenced, the barcode may not connect them.

### Combining Data Types for Maximum Accuracy

The highest quality single-individual phasing results come from combining multiple data types. The DipAsm method used long accurate reads plus long-range conformation data to achieve chromosome-scale phasing [11]. The long-range conformation data, which can come from Hi-C or other proximity ligation methods, provides connectivity across distances that exceed even ultra-long read lengths.

The practical question is whether the additional cost and complexity of multiple data types is justified by the improvement in phasing quality. For research applications where haplotype block length is not critical, standard long reads alone may be sufficient. For clinical applications or studies of structural variation, the higher accuracy of combined approaches may be necessary.

## Observations and Measurements for Quality Control

### Monitoring Read Length Distributions

The read length distribution is the single most important quality metric for long-read phasing. A sequencing run that produces mostly short reads will yield fragmented haplotypes regardless of the phasing algorithm used.

Before proceeding with phasing, check the read length N50, which is the length such that 50 percent of the sequenced bases are in reads of at least that length. For phasing applications, a read length N50 of at least 10 kb is a reasonable minimum, with 50 kb or higher preferred for whole-genome phasing.

The read length distribution should be examined for each sample and each sequencing run. Batch effects can cause substantial variation in read length, and a run that produced poor read lengths should be repeated or supplemented before proceeding with phasing.

### Tracking Variant Density and Heterozygosity

The density of heterozygous variants in the region of interest determines how much phasing information each read provides. In a region with one heterozygous variant per kilobase, a 10 kb read will typically span ten variants. In a region with one heterozygous variant per 100 kb, the same read will span only a fraction of a variant, providing no phasing information.

For whole-genome phasing, the average heterozygosity of the individual is an important consideration. Individuals with low heterozygosity, such as those from isolated populations or with consanguineous parents, will have fewer heterozygous variants and therefore less phasing information per read. In these cases, longer reads or higher coverage may be needed to achieve the same phasing completeness.

### Measuring Haplotype Block Statistics

After phasing, calculate the haplotype block N50 and the phasing rate. These metrics provide a quantitative assessment of phasing quality that can be compared across samples and across different phasing strategies.

The haplotype block N50 is particularly informative for whole-genome phasing. A block N50 of 1 Mb indicates that half of the phased genome is in blocks of at least 1 Mb, which is sufficient for most gene-level analyses. A block N50 of 10 Mb or higher is needed for chromosome-scale analyses, such as studying the distribution of variants across entire chromosome arms.

The phasing rate, or the proportion of heterozygous variants assigned to a haplotype, should be reported alongside the block N50. A high block N50 with a low phasing rate indicates that the phasing algorithm is producing long blocks but missing many variants. A low block N50 with a high phasing rate indicates that most variants are phased but the blocks are fragmented.

## Common Failure Patterns and How to Address Them

### Fragmented Haplotype Blocks

The most common failure pattern in single-individual phasing is fragmented haplotype blocks. The genome is phased in many small blocks, each containing only a few variants, with no connectivity between blocks.

This pattern typically results from insufficient read length or insufficient coverage. The fix is to generate longer reads, increase coverage, or use a linked-read approach to provide connectivity across gaps.

For targeted regions, amplicon-based approaches can produce long reads that span the entire region of interest. The APOE locus study demonstrated that amplicon-based ONT sequencing could phase a 12 kb region with 600 reads [8]. For larger regions, the amplicon approach becomes impractical, and whole-genome or targeted capture approaches are needed.

### High Switch Error Rates

A high switch error rate indicates that the phasing algorithm is assigning alleles to the wrong haplotypes. This pattern can result from:

- **Low variant density.** In regions with sparse heterozygous variants, the phasing algorithm has little information to distinguish correct from incorrect phase assignments.
- **Alignment errors.** Reads that are misaligned can create false variant pairs that support incorrect phase assignments.
- **Variant calling errors.** False heterozygous sites create noise that can confuse the phasing algorithm.

The fix for high switch error rates depends on the cause. If variant density is low, longer reads or linked-read data can provide connectivity across larger distances. If alignment errors are suspected, more stringent alignment parameters or a different aligner may help. If variant calling errors are the issue, using a more accurate variant caller or filtering low-quality variants before phasing can reduce the error rate.

### Missing Variants in Difficult Regions

Regions with low complexity, high GC content, or repetitive sequence are prone to missing variants. These regions are difficult for aligners, and variant callers may fail to identify true heterozygous sites.

The consequence is that the haplotype is incomplete in these regions. The surrounding variants may be phased correctly, but the missing variants cannot be assigned to a haplotype.

For clinical applications, missing variants in medically important regions can be a serious problem. The human leukocyte antigen (HLA) and killer cell immunoglobulin-like receptor (KIR) regions are highly polymorphic and medically important, and they are also difficult to sequence and phase [11]. Researchers studying these regions should use targeted approaches designed to capture and sequence the full region, including the difficult subregions.

### Inconsistent Results Across Samples

When phasing multiple samples from the same study, inconsistent results can indicate technical problems. One sample may have long haplotype blocks while another has fragmented blocks, or one sample may have a high switch error rate while another has a low rate.

These inconsistencies often result from differences in DNA quality or sequencing depth. A sample with degraded DNA will produce shorter reads and therefore more fragmented haplotypes. A sample with lower coverage will have less phasing information and therefore more switch errors.

The fix is to standardize the DNA extraction and library preparation protocols across all samples and to monitor read length and coverage metrics for each sample. Samples that fall below quality thresholds should be resequenced or excluded from the analysis.

## Limitations of Single-Individual Phasing

### Inability to Phase All Variants

Even with the best available technology, single-individual phasing cannot phase all heterozygous variants. Some variants will be in regions that are impossible to sequence with current platforms, such as highly repetitive satellite DNA. Other variants will be in regions where the read length is insufficient to connect them to neighboring variants.

The practical consequence is that the phased haplotypes will have gaps. The researcher must decide whether these gaps are acceptable for the intended application. For gene-level analyses, gaps within the gene of interest may be unacceptable. For genome-wide analyses, gaps in non-genic regions may be tolerable.

### Difficulty in Validating Phasing Accuracy

Without parental data, validating the accuracy of single-individual phasing is challenging. The researcher cannot compare the phased haplotypes to a known truth set derived from inheritance patterns.

Several validation strategies are available:

- **Comparison to an independent phasing method.** If both long-read and linked-read data are available for the same individual, the two phasing results can be compared. Regions where the two methods agree are likely to be correct.
- **Comparison to population reference panels.** If the individual's population is represented in a reference panel with phased haplotypes, the single-individual phasing can be compared to the population haplotypes. This comparison is not definitive, because the individual may have rare haplotypes not present in the reference panel, but it can identify gross errors.
- **Experimental validation.** For specific variants of interest, experimental methods such as allele-specific PCR or Sanger sequencing of cloned alleles can confirm the phase assignment. This approach is practical for a small number of variants but not for whole-genome validation.

### Computational Resource Requirements

Single-individual phasing with long reads requires substantial computational resources. The alignment of long reads to the reference genome is computationally intensive, and the phasing algorithms can require large amounts of memory.

For whole-genome phasing, the computational requirements can exceed the capacity of a standard desktop workstation. Researchers may need access to a high-performance computing cluster or a cloud computing environment. The nf-core community provides standardized pipelines for bioinformatics analysis that can be configured to run on clusters or in the cloud, with documentation covering usage and configuration [5].

For researchers who are new to long-read analysis, the Galaxy Training Network offers accessible workflow training and analysis tutorials that cover the practical aspects of working with sequencing data [4]. The European Bioinformatics Institute provides training materials for bioinformatics data resources and practical analysis education [2].

## Safety and Regulatory Context for Clinical Applications

### Accuracy Requirements for Clinical Phasing

When phased haplotypes are used for clinical decision-making, the accuracy requirements are substantially higher than for research applications. A switch error in a clinically relevant region can lead to incorrect interpretation of compound heterozygosity, which can affect diagnosis and treatment decisions.

The APOE locus study described a streamlined proof-of-concept workflow for variant calling and phasing based on ONT data in a clinically relevant 12 kb region, with the goal of establishing standards for variant phasing in ONT-based targeted resequencing [8]. The study noted that robust standards for variant phasing in ONT-based targeted resequencing efforts were not yet available, highlighting the need for careful validation in clinical applications [8].

Researchers and laboratory professionals who plan to use phased haplotypes for clinical purposes should:

- Validate the phasing accuracy using an independent method before reporting results.
- Document the phasing algorithm, parameters, and quality metrics for each sample.
- Report the limitations of the phasing approach, including the haplotype block N50 and the phasing rate.
- Consult with clinical genetics professionals about the appropriate use of phased haplotypes in diagnostic testing.

### Data Management and Reproducibility

Reproducibility is a critical concern in bioinformatics analysis. The phasing results should be reproducible given the same input data and the same analysis parameters.

Several resources support reproducible bioinformatics analysis:

- The Bioconductor project provides packages for the analysis and comprehension of high-throughput genomic data, with documentation covering installation and reproducible analysis workflows [3].
- The nf-core community provides standards for pipeline development and usage, with documentation covering configuration and reproducible workflow context [5].
- The Carpentries offers lessons on foundational computing, data, shell, Git, and programming training that support reproducible research practices [6].

For clinical applications, the data management requirements may be governed by regulatory frameworks that vary by jurisdiction. Researchers should be aware of the applicable regulations and ensure that their data management practices comply.

## Professional Escalation Criteria

### When to Seek Additional Expertise

Single-individual phasing with long reads is a specialized analysis that requires expertise in both laboratory techniques and bioinformatics. Researchers should consider escalating to additional expertise in the following situations:

- **Unexpectedly fragmented haplotypes.** If the haplotype block N50 is substantially lower than expected given the read length and coverage, the problem may be in the library preparation, the sequencing run, or the phasing algorithm. A specialist in long-read sequencing can help diagnose the issue.
- **High switch error rates that persist after parameter optimization.** If the switch error rate remains high after trying different phasing algorithms and parameters, the problem may be in the variant calling or the alignment. A bioinformatics specialist can help identify the source of the errors.
- **Clinical applications with stringent accuracy requirements.** If the phased haplotypes will be used for clinical decision-making, the analysis should be reviewed by a clinical genetics professional who can assess the accuracy and limitations of the results.
- **Regions of the genome that are difficult to phase.** If the region of interest includes highly polymorphic or repetitive regions, such as the HLA or KIR regions, specialized approaches may be needed. The DipAsm study demonstrated the importance of chromosome-scale phased assemblies for the discovery of structural variants in these regions [11].

### Documentation for Escalation

When escalating a phasing problem, provide the following documentation:

- The sequencing platform, library preparation method, and read length distribution for each sample.
- The variant calling pipeline, including the aligner, variant caller, and filtering parameters.
- The phasing algorithm and parameters used.
- The phasing quality metrics, including haplotype block N50, phasing rate, and switch error rate if available.
- The specific region of interest and the phasing results in that region.

This documentation allows the specialist to identify the source of the problem and recommend appropriate solutions.

## Frequently Asked Questions

### What is the minimum read length needed for single-individual phasing?

The minimum read length depends on the variant density in the region of interest. To phase two adjacent heterozygous variants, a single read must span both variants. In a region with one heterozygous variant per kilobase, a read length of 2 to 3 kb is sufficient to phase adjacent variants. However, longer reads produce longer haplotype blocks because they connect more distant variants. For whole-genome phasing, reads of at least 10 kb are recommended, and ultra-long reads exceeding 100 kb substantially improve phasing completeness. The APOE locus study demonstrated reliable phasing of a 12 kb region with ONT reads, with full haplotype reconstruction achieved using 600 reads [8].

### How much sequencing coverage is needed for phasing?

Coverage requirements for phasing are higher than for standard variant calling. The phasing algorithm needs multiple reads spanning each pair of adjacent heterozygous variants to establish phase with confidence. For targeted amplicon approaches, 600 reads covering the target region were sufficient for full haplotype reconstruction in the APOE locus study [8]. For whole-genome phasing, coverage of 30x to 60x is typically recommended. Higher coverage provides more opportunities to observe variant pairs on the same read, but the improvement diminishes beyond approximately 60x.

### Can linked-read sequencing replace long-read sequencing for phasing?

Linked-read sequencing can provide long-range connectivity without long sequencing reads. The CPTv2-seq method demonstrated barcode partitioning of long DNA molecules in a single tube, producing sequencing libraries where all reads from each long DNA molecule shared a common barcode [10]. This approach enabled accurate haplotype-resolved sequencing and phasing of structural variants [10]. However, linked-read approaches have limitations. The barcode information is probabilistic, and the connectivity depends on the length of the original DNA molecules and the completeness of sequencing. For the highest quality phasing, combining linked-read data with long-read data may be necessary.

### What causes switch errors in single-individual phasing?

Switch errors occur when the phasing algorithm incorrectly assigns alleles to haplotypes, effectively switching the haplotype labels for all variants downstream of the error point. Common causes include low variant density, alignment errors, and variant calling errors. In regions with sparse heterozygous variants, the phasing algorithm has little information to distinguish correct from incorrect phase assignments. Misaligned reads can create false variant pairs that support incorrect phase assignments. False heterozygous sites from variant calling errors create noise that can confuse the phasing algorithm.

### How can I validate phasing accuracy without parental data?

Several validation strategies are available when parental data are not available. Comparison to an independent phasing method, such as using both long-read and linked-read data for the same individual, can identify regions where the two methods agree. Comparison to population reference panels can identify gross errors, although the individual may have rare haplotypes not present in the reference panel. Experimental validation using allele-specific PCR or Sanger sequencing of cloned alleles can confirm the phase assignment for specific variants of interest. The choice of validation strategy depends on the application and the tolerance for errors.

### What is the difference between haplotype phasing and diploid assembly?

Haplotype phasing assigns heterozygous variants to haplotypes without necessarily assembling the full haplotype sequence. The output is a set of phased variants, typically represented as variant call format (VCF) files with phase information. Diploid assembly goes further, producing two complete haplotype sequences. The DipAsm method is a diploid assembly approach that uses long accurate reads and long-range conformation data to generate chromosome-scale phased assemblies [11]. Diploid assembly is more computationally intensive than phasing but provides a more complete picture of the genome, including structural variants and regions that are difficult to phase.

### Can single-individual phasing be used for repeat expansion disorders?

Repeat expansion disorders present a special challenge for phasing because the repeat length often exceeds the capacity of short-read sequencing. Full-length transcript phasing methods have been developed for repeat expansion disorders, with a method for targeted haplotype phasing of the HTT gene that can be adopted for other repeat expansion disorders [9]. These methods use long-read sequencing to span the repeat region and determine the haplotype of the expansion-containing allele. The approach is important because repeat expansion disorders are typically characterized by a heterozygous expansion locus associated with a single haplotype, and precision genetic medicines can be used to selectively target expansion-containing sequences in a haplotype-specific manner [9].

### What are the computational requirements for single-individual phasing?

Single-individual phasing with long reads requires substantial computational resources. The alignment of long reads to the reference genome is computationally intensive, and the phasing algorithms can require large amounts of memory. For whole-genome phasing, access to a high-performance computing cluster or a cloud computing environment is often necessary. The nf-core community provides standardized pipelines that can be configured to run on clusters or in the cloud [5]. The Galaxy Training Network offers accessible workflow training for researchers who are new to long-read analysis [4]. The Bioconductor project provides packages for reproducible genomic analysis [3].

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data](/knowledge/bioinformatics/long-read-metagenome-assembly-overcoming-challenges-with-nanopore-and-pacbio-data)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Haplotype phasing in single-cell DNA-sequencing data.](https://pubmed.ncbi.nlm.nih.gov/29950014). Bioinformatics (Oxford, England), 2018.
- [A Long-Read Sequencing Approach for Direct Haplotype Phasing in Clinical Settings.](https://pubmed.ncbi.nlm.nih.gov/33271988). International journal of molecular sciences, 2020.
- [Full-Length Transcript Phasing with Third-Generation Sequencing.](https://pubmed.ncbi.nlm.nih.gov/36335491). Methods in molecular biology (Clifton, N.J.), 2023.
- [Haplotype phasing of whole human genomes using bead-based barcode partitioning in a single tube.](https://pubmed.ncbi.nlm.nih.gov/28650462). Nature biotechnology, 2017.
- [Chromosome-scale, haplotype-resolved assembly of human genomes.](https://pubmed.ncbi.nlm.nih.gov/33288905). Nature biotechnology, 2021.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.