# Haplotyping in De Novo Assembly: How to Phase Heterozygous Genomes Without a Reference


## Key Takeaways

- Haplotyping in de novo assembly aims to resolve the two parental chromosome copies into distinct sequences without a reference genome, overcoming the tendency of standard assemblers to collapse heterozygous alleles into a single consensus.
- Long, high-fidelity (HiFi) reads are critical for read-based phasing, as their length spans multiple heterozygous sites, providing direct phase links, while their accuracy distinguishes true variants from sequencing errors for robust graph construction.
- Graph-based haplotyping, exemplified by tools like hifiasm, constructs a phased assembly graph that explicitly preserves all haplotype paths, enabling the extraction of contiguous haplotype sequences even in highly heterozygous or polyploid genomes.
- Complementary phasing signals such as Hi-C (chromatin conformation) and Strand-seq can provide long-range phase information, bridging gaps where read-based phasing is insufficient, particularly when parental samples are unavailable.
- Quality control for haplotype-resolved assemblies necessitates evaluating contiguity (N50/NG50), completeness (gene content), phasing accuracy (switch errors), and read mapping uniformity to ensure the faithful representation of both parental genomes.

---

## Direct Answer and Scope

Haplotyping in de novo assembly is the process of separating the two parental copies of a diploid genome into distinct contiguous sequences without relying on a reference genome. Researchers who attempt to assemble heterozygous genomes often find that standard assemblers collapse divergent alleles into a single consensus sequence or produce fragmented contigs that switch between haplotypes. The practical solution involves using long high-fidelity reads, graph-based assembly algorithms, and additional phasing signals such as parental data, chromatin conformation, or strand-specific sequencing. This article explains the principles of read-based phasing and graph-based haplotyping and provides a practical framework for integrating tools like hifiasm, Hi-C, Strand-seq, and long-read phasing into assembly pipelines. The target reader is a biology student, researcher, or laboratory professional who needs to produce haplotype-resolved assemblies and interpret their quality correctly.

## The Problem of Heterozygous Genome Assembly

### Why Heterozygosity Breaks Conventional Assemblers

A diploid genome contains two copies of each chromosome, one inherited from each parent. These copies differ at millions of positions, including single nucleotide variants, insertions, deletions, and larger structural rearrangements. When an assembler encounters a heterozygous site, it must decide whether to represent both alleles or merge them into one. Many early assemblers chose to collapse heterozygous alleles into a single consensus copy, which loses the phase information that links variants along each parental chromosome. This collapse creates a mosaic sequence that does not correspond to either real haplotype and can obscure the true variant structure of the genome.

The consequences of collapse are measurable. Variants that are heterozygous but physically close on the same chromosome become scrambled, and the resulting assembly cannot be used to determine which alleles travel together. For researchers studying gene expression, disease association, or population genetics, this loss of phase information is a critical limitation. The problem becomes more severe in complex regions such as the major histocompatibility complex, where high polymorphism and structural variation make collapse almost certain with short reads alone.

### The Shift from Consensus Assembly to Haplotype Assembly

The field has moved from producing a single consensus genome to producing two or more haplotype-resolved assemblies. A haplotype-resolved assembly preserves the sequence of each parental chromosome separately, allowing researchers to study allele-specific expression, compound heterozygosity, and structural variation that spans heterozygous sites. This shift was enabled by two developments: long-read sequencing technologies that produce reads long enough to span multiple heterozygous sites, and assembly algorithms that represent haplotypes explicitly in a graph structure.

Long high-fidelity reads, such as those produced by PacBio HiFi, provide both length and accuracy. These reads can span several kilobases and carry base-level error rates low enough for confident variant detection. When such reads cover a heterozygous site, the assembler can use the read sequence to determine which allele is present and link that allele to other variants on the same read. This read-based phasing is the foundation of modern haplotype-resolved assembly. The hifiasm assembler, for example, uses long high-fidelity reads to preserve the contiguity of all haplotypes in a phased assembly graph, instead of maintaining only one haplotype as earlier graph-based assemblers did. This design choice is described in the original [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886), which demonstrates improved haplotype-resolved assembly across human and nonhuman datasets, including a hexaploid California redwood genome of roughly 30 gigabases.

## Core Principles of Phasing Without a Reference

### Read-Based Phasing and the Phased Assembly Graph

Read-based phasing relies on the simple observation that a single sequencing read originates from one physical DNA molecule and therefore carries alleles from only one haplotype. If a read covers two or more heterozygous positions, the alleles observed on that read are phased with respect to each other. The challenge is to extend this local phasing across longer distances, which requires reads that span multiple heterozygous sites or linking information from other sources.

The phased assembly graph is the computational structure that makes this extension possible. In a standard assembly graph, nodes represent sequence and edges represent overlaps between reads. In a phased assembly graph, the graph is constructed so that each haplotype follows a distinct path. The hifiasm assembler builds such a graph by using the error patterns and coverage of high-fidelity reads to distinguish true allelic variation from sequencing errors. Once the graph is constructed, the assembler traverses it to extract contiguous haplotype sequences. The key innovation is that the graph preserves the contiguity of all haplotypes, beyond one, which allows the assembler to produce complete phased assemblies even in highly heterozygous genomes.

### The Role of Long High-Fidelity Reads

Long high-fidelity reads are the primary input for read-based phasing because they provide two properties that short reads cannot. First, their length allows a single read to span multiple heterozygous sites, creating direct phase links. Second, their accuracy allows the assembler to distinguish true heterozygous variants from sequencing errors, which is essential for building a correct phased graph.

The practical implication is that sequencing strategy determines phasing success. A researcher planning a haplotype-resolved assembly should prioritize read length and accuracy over raw throughput. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) demonstrates that high-fidelity reads enable haplotype-resolved assembly even in complex genomes, and the same principle applies to single-cell sequencing approaches. A [study using single-cell long-read sequencing](https://pubmed.ncbi.nlm.nih.gov/35819189) on PacBio HiFi and Oxford Nanopore Technologies platforms completed human genome assemblies with high continuity from individual cells, showing that read-based phasing can work even when DNA input is limited and heterogeneous.

### Graph-Based Haplotyping and the Limits of Local Phasing

Graph-based haplotyping extends read-based phasing by representing all possible haplotype paths in a single structure. This approach is necessary because reads are finite in length, and local phase information must be connected across gaps. The graph encodes the relationships between reads and variants, and the assembler uses graph traversal algorithms to find the most likely haplotype paths.

The limitation of graph-based haplotyping is that it depends on the coverage and distribution of reads. Regions with low coverage may lack the reads needed to connect phase information across long distances, resulting in phase breaks. Regions with high heterozygosity may create complex graph structures that are difficult to resolve unambiguously. These limitations are constraints of the input data. A researcher who understands these constraints can design sequencing strategies that address them, such as increasing coverage in difficult regions or adding complementary phasing data.

## At a Glance

| Phasing Strategy | Input Data | Strengths | Limitations | Best Use Case |
| --- | --- | --- | --- | --- |
| Read-based phasing with HiFi reads | PacBio HiFi or equivalent long high-fidelity reads | Direct phase links across heterozygous sites, high accuracy, works without parental data | Requires high coverage, phase breaks in low-complexity or low-coverage regions | Standard diploid genomes, including human and animal genomes |
| Graph-based haplotyping with hifiasm | HiFi reads, optional parental short reads for trio binning | Preserves contiguity of all haplotypes, handles complex and polyploid genomes | Computationally intensive, requires careful parameter tuning | Highly heterozygous genomes, polyploid genomes, pangenome construction |
| Chromatin conformation phasing with Hi-C | Hi-C reads plus an initial assembly | Provides long-range phase links across megabases, works without parental data | Requires additional library preparation, resolution limited by chromatin structure | Scaffolding and phasing across large genomic distances |
| Strand-specific phasing with Strand-seq | Strand-seq libraries | Provides genome-wide phase information without parental data, useful for structural variant analysis | Requires specialized protocol, lower throughput than standard sequencing | Diverse human genomes, structural variation studies |
| Parental trio binning | Short reads from both parents plus long reads from the offspring | Assigns haplotypes to parental origin, resolves phase unambiguously | Requires parental samples, not applicable when parents are unavailable | Family-based studies, clinical genetics |

## Practical Workflow for Haplotype-Resolved Assembly

### Step 1: Assess the Genome and Define the Phasing Goal

Before selecting tools and designing the sequencing strategy, define what the assembly must achieve. A researcher studying a single gene region may need only local phasing across that region. A researcher building a pangenome resource needs complete haplotype-resolved assemblies for many individuals. A researcher studying structural variation needs phased assemblies that integrate all forms of genetic variation, including complex loci that are difficult to assemble with short reads.

The genome itself imposes constraints. Genome size, heterozygosity rate, ploidy, and the presence of repetitive regions all affect the choice of sequencing platform and assembly algorithm. A highly heterozygous genome requires an assembler that preserves haplotype contiguity, such as hifiasm. A polyploid genome requires an assembler that can represent more than two haplotypes. A genome with large repeat arrays requires long reads that can span those repeats.

### Step 2: Design the Sequencing Strategy

The sequencing strategy must provide sufficient coverage of long high-fidelity reads for read-based phasing. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) demonstrates successful assembly with high-fidelity reads across multiple human and nonhuman datasets, and the [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) shows that even relatively modest coverage can produce contiguous assemblies when reads are long and accurate.

Consider whether additional phasing data is needed. If parental samples are available, trio binning can assign haplotypes to parental origin and resolve phase unambiguously. If parental samples are not available, Hi-C or Strand-seq can provide long-range phase information. The choice depends on the research question and the availability of samples.

### Step 3: Select the Assembly Tool and Configure Parameters

The primary tool for haplotype-resolved assembly with high-fidelity reads is hifiasm. The assembler constructs a phased assembly graph and extracts haplotype sequences from it. Configuration parameters include coverage thresholds, error rate settings, and options for trio binning or Hi-C integration.

For researchers who prefer workflow-based approaches, community pipelines provide standardized implementations. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflow configuration, and the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training for assembly analysis. These resources are useful for researchers who want to avoid manual configuration errors and ensure reproducibility.

### Step 4: Run the Assembly and Evaluate Output

After the assembly completes, evaluate the output for contiguity, completeness, and phasing accuracy. Contiguity is measured by metrics such as N50, which indicates the contig length at which half the assembly is contained. Completeness is assessed by comparing the assembly to expected gene content or by mapping reads back to the assembly to check coverage. Phasing accuracy is assessed by examining switch errors, which occur when the assembly switches from one haplotype to the other.

The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) reports that its assemblies frequently deliver better contiguity than existing tools, and the [diverse human genome study](https://pubmed.ncbi.nlm.nih.gov/33632895) reports average minimum contig length needed to cover 50 percent of the genome of 26 million base pairs across 64 assembled haplotypes. These benchmarks provide context for evaluating your own assembly results.

### Step 5: Polish and Validate the Assembly

Assembly polishing corrects residual errors in the assembled sequence. For haplotype-resolved assemblies, polishing must be performed carefully to avoid collapsing haplotypes that were correctly separated. The [template mutagenesis approach](https://pubmed.ncbi.nlm.nih.gov/35822882) described in the literature achieves ultra-low error rates by imprinting template strands with mutation patterns and using unmutated libraries for correction, but this approach is specialized and not applicable to all projects.

Validation involves checking that the assembly correctly represents both haplotypes. This can be done by mapping reads back to the assembly and verifying that heterozygous sites are represented in both haplotypes, or by comparing the assembly to known variant data if available.

## Tools and Their Tradeoffs

### Hifiasm for Graph-Based Haplotyping

Hifiasm is the most widely used assembler for haplotype-resolved assembly with high-fidelity reads. Its design philosophy is to preserve the contiguity of all haplotypes in the phased assembly graph, which distinguishes it from earlier graph-based assemblers that aimed to maintain only one haplotype. The [publication describing hifiasm](https://pubmed.ncbi.nlm.nih.gov/33526886) demonstrates its performance on three human and five nonhuman datasets, including a hexaploid genome, and reports that it consistently outperforms existing tools on haplotype-resolved assembly.

The tradeoff is computational cost. Building and traversing a phased assembly graph for a large genome requires substantial memory and processing time. Researchers working with large genomes should plan for adequate computational resources or use cloud-based infrastructure.

### Hi-C for Long-Range Phasing

Hi-C captures chromatin conformation by crosslinking DNA in the nucleus, digesting it, and sequencing the resulting ligation junctions. The frequency of contacts between genomic regions reflects their physical proximity in the nucleus, which correlates with their linear distance along the chromosome. This information provides long-range phase links that can connect haplotypes across megabases.

Hi-C phasing is particularly useful when parental samples are unavailable. The [diverse human genome study](https://pubmed.ncbi.nlm.nih.gov/33632895) used long-read and strand-specific sequencing technologies together to assemble haplotype-resolved human genomes without parent-child trio data, demonstrating that Hi-C and Strand-seq can substitute for parental information.

The tradeoff is resolution. Hi-C contact frequencies are noisy and depend on chromatin structure, which varies across cell types and conditions. The phase information is statistical instead of deterministic, and errors can occur in regions with unusual chromatin conformation.

### Strand-Seq for Genome-Wide Phasing

Strand-seq is a specialized sequencing technique that preserves the template strand of each DNA molecule. By sequencing single strands, researchers can determine which parental chromosome each read originated from, providing genome-wide phase information without parental data.

The [diverse human genome study](https://pubmed.ncbi.nlm.nih.gov/33632895) used Strand-seq in combination with long-read sequencing to assemble 64 haplotypes from 32 diverse human genomes. The study identified 107,590 structural variants, of which 68 percent were not discovered with short-read sequencing, demonstrating the value of phased assemblies for structural variation analysis.

The tradeoff is throughput and complexity. Strand-seq requires specialized library preparation and produces data that require dedicated analysis tools. It is not a drop-in replacement for standard sequencing but rather a complementary data source for phasing.

### Trio Binning with Parental Data

Trio binning uses short reads from both parents to assign offspring reads to parental haplotypes before assembly. This approach resolves phase unambiguously because each read is assigned to the haplotype of the parent from which it was inherited. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) describes a graph trio binning algorithm that advances over standard trio binning by integrating parental information into the phased assembly graph.

The tradeoff is sample availability. Trio binning requires DNA from both parents, which is not always available. For nonmodel organisms, parental samples may be difficult or impossible to obtain. In such cases, Hi-C or Strand-seq are the practical alternatives.

### Single-Cell Assembly Approaches

Single-cell genome assembly is an emerging approach that addresses the challenge of cell heterogeneity. Most genome assemblies require large amounts of DNA from homogeneous cell lines, which does not reflect the heterogeneity present in real tissues. The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) used SMOOTH-seq to sequence individual cells on PacBio HiFi and Oxford Nanopore Technologies platforms and completed human genome assemblies with high continuity from individual cells.

The tradeoff is coverage. Single-cell sequencing produces limited DNA per cell, so achieving sufficient coverage requires sequencing many cells. The study used 95 individual K562 cells to achieve an NG50 of approximately 2 megabases, and 30 diploid HG002 cells at an average coverage of 41.7 percent to achieve an NG50 of over 1.3 megabases on the Oxford Nanopore platform.

## Observations and Measurements for Quality Control

### Contiguity Metrics

Contiguity metrics describe how continuous the assembly is. The N50 statistic indicates the contig length at which half the assembly is contained in contigs of that length or longer. A higher N50 indicates a more contiguous assembly. The NG50 statistic is the same metric applied to the genome size instead of the assembly size, which is useful for comparing assemblies of different sizes.

The [diverse human genome study](https://pubmed.ncbi.nlm.nih.gov/33632895) reports an average minimum contig length needed to cover 50 percent of the genome of 26 million base pairs across 64 assembled haplotypes. This benchmark provides a reference point for evaluating human genome assemblies. For nonhuman genomes, the expected N50 depends on genome size, heterozygosity, and repeat content.

### Completeness Assessment

Completeness is assessed by checking whether the assembly contains expected genes and functional elements. This can be done by comparing the assembly to a curated gene set or by using completeness assessment tools that search for conserved single-copy genes. The [zebrafish genome study](https://pubmed.ncbi.nlm.nih.gov/41332591) used long-read sequencing and assembly algorithms to produce complete genome assemblies that incorporated 7 percent more genomic sequence than the previous reference and added 130 million bases of previously unassembled sequence.

For haplotype-resolved assemblies, completeness must be assessed for each haplotype separately. A haplotype that is missing genes or contains gaps may indicate that the phasing process failed in that region.

### Phasing Accuracy and Switch Errors

Phasing accuracy is measured by switch errors, which occur when the assembly switches from one haplotype to the other. A switch error means that the phase information is incorrect at that point, and variants on either side of the switch are assigned to the wrong haplotype.

Switch errors can be detected by comparing the assembly to known phase information, such as parental genotypes or independent phasing data. In the absence of such data, switch errors can be inferred from read mapping patterns. Reads that map to both haplotypes in a region may indicate a switch error.

### Read Mapping and Coverage Uniformity

Mapping reads back to the assembly provides a direct check on assembly quality. Coverage should be relatively uniform across the assembly, with no regions of zero coverage that indicate collapsed or missing sequence. Regions with unusually high coverage may indicate collapsed repeats or misassembled haplotypes.

The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) demonstrates that coverage requirements vary by platform and assembly strategy. The study achieved contiguous assemblies with an average coverage of 41.7 percent on the Oxford Nanopore platform, showing that high coverage is not always necessary when reads are long and accurate.

## Records and Documentation for Reproducibility

### Recording Assembly Parameters

Reproducible assembly requires recording all parameters used in the assembly process. This includes the assembler version, the input data and its quality metrics, the parameter settings, and the computational environment. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that emphasize reproducibility, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that includes documentation practices.

A practical approach is to maintain an assembly log that records each step of the process. The log should include the date, the software version, the input files, the parameters, and the output files. This log serves as the primary record for reproducing the assembly and for troubleshooting if problems arise.

### Version Control for Assembly Scripts

Assembly scripts and configuration files should be under version control. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in version control with Git, which is essential for tracking changes to assembly workflows. Version control allows researchers to revert to previous versions, compare different parameter settings, and share workflows with collaborators.

### Data Management for Large Assembly Files

Assembly files are large, and data management is a practical concern. Raw sequencing data, intermediate files, and final assemblies should be stored in a structured directory hierarchy with clear naming conventions. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of databases and search systems for sequence data, which are useful for depositing and retrieving assembly data.

### Documentation of Quality Control Results

Quality control results should be documented alongside the assembly. This includes the contiguity metrics, completeness assessment results, and phasing accuracy measurements. The documentation should note any regions where the assembly is uncertain, such as low-coverage regions or regions with high switch error rates.

## Common Failure Patterns and Their Causes

### Haplotype Collapse in Heterozygous Regions

Haplotype collapse occurs when the assembler merges two parental haplotypes into a single consensus sequence. This failure is most common in regions with high heterozygosity, where the assembler cannot distinguish true allelic variation from sequencing errors. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) notes that existing algorithms either collapse heterozygous alleles into one consensus copy or fail to cleanly separate the haplotypes, which is the problem that hifiasm was designed to solve.

The practical consequence of collapse is that the assembly contains only one copy of the region, and variants that are heterozygous in the original genome appear homozygous in the assembly. This can lead to incorrect variant calls and misinterpretation of the genome structure.

### Phase Breaks in Low-Complexity Regions

Phase breaks occur when the assembler cannot connect phase information across a region. This is common in low-complexity regions, such as homopolymer runs or simple repeats, where reads do not contain enough variation to establish phase links. Phase breaks result in haplotypes that are fragmented, with each fragment phased correctly but the relationship between fragments unknown.

The [template mutagenesis approach](https://pubmed.ncbi.nlm.nih.gov/35822882) described in the literature addresses this problem for targeted regions by imprinting template strands with dense mutation patterns, which creates the variation needed for phasing. However, this approach is not applicable to whole-genome assembly.

### Misassembly at Repeat Boundaries

Misassembly occurs when the assembler incorrectly joins sequences from different genomic locations. This is most common at repeat boundaries, where identical or nearly identical sequences create ambiguous overlaps. Misassembly can create chimeric contigs that combine sequences from different chromosomes or different haplotypes.

Long reads reduce but do not eliminate misassembly at repeat boundaries. The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) demonstrates that long-read sequencing can produce contiguous assemblies even from individual cells, but the study also notes that assembly quality depends on the assembler and sequencing strategy.

### Switch Errors from Incorrect Graph Traversal

Switch errors occur when the assembler traverses the phased assembly graph incorrectly, switching from one haplotype path to another. This can happen when the graph is ambiguous, such as in regions with low coverage or high similarity between haplotypes. Switch errors are particularly problematic because they are difficult to detect without independent phasing data.

### Coverage Bias and Uneven Sequencing

Coverage bias occurs when some regions of the genome are sequenced more or less than others. This can result from library preparation artifacts, amplification bias, or the physical properties of the genome. Uneven coverage can cause phase breaks in low-coverage regions and misassembly in high-coverage regions.

## Limitations and Interpretation Boundaries

### The Limits of Read-Based Phasing

Read-based phasing is limited by read length and coverage. A read can only phase variants that it physically spans, so phase information is local unless reads are long enough to span multiple heterozygous sites. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) demonstrates that high-fidelity reads provide sufficient length for haplotype-resolved assembly in many genomes, but the approach has limits in genomes with very long stretches of low heterozygosity.

### The Challenge of Polyploid Genomes

Polyploid genomes present a greater challenge than diploid genomes because they have more than two haplotypes to resolve. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) demonstrates successful assembly of a hexaploid California redwood genome of approximately 30 gigabases, showing that graph-based haplotyping can handle polyploidy. However, the computational cost and complexity increase with ploidy, and the phase information becomes more ambiguous.

### The Problem of Cell Heterogeneity

Cell heterogeneity affects haplotype assembly because different cells may have different genomes due to somatic mutation or structural variation. The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) notes that cell heterogeneity could profoundly affect haplotype assembly results, which is why most genome assemblies use homogeneous cell lines. Single-cell assembly approaches address this problem but require specialized sequencing strategies.

### The Interpretation of Structural Variants

Haplotype-resolved assemblies enable the identification of structural variants that are invisible to short-read sequencing. The [diverse human genome study](https://pubmed.ncbi.nlm.nih.gov/33632895) identified 107,590 structural variants, of which 68 percent were not discovered with short-read sequencing. However, structural variant calls from assemblies require careful validation because assembly errors can create false structural variants.

### The Need for Independent Validation

Haplotype-resolved assemblies should be validated with independent data whenever possible. This can include parental genotypes, independent phasing data, or experimental validation of specific variants. The [template mutagenesis study](https://pubmed.ncbi.nlm.nih.gov/35822882) demonstrates that ultra-low error rates are achievable with specialized protocols, but standard assemblies have higher error rates that require validation.

## Safety and Regulatory Context

### Data Privacy for Human Genome Assemblies

Human genome assemblies contain sensitive genetic information that requires careful handling. Researchers working with human data must comply with applicable privacy regulations and institutional review board requirements. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of databases and search systems that include data access controls for sensitive human data.

### Ethical Considerations for Nonmodel Organisms

Assembling genomes of nonmodel organisms raises ethical considerations related to sample collection, species conservation, and benefit sharing. Researchers should ensure that samples were collected legally and ethically and that the resulting data are shared in accordance with community standards.

### Reproducibility Standards for Published Assemblies

Published assemblies should include sufficient documentation for reproduction. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that emphasize reproducibility, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that includes reproducibility context. Journals increasingly require assembly data and scripts to be deposited in public repositories.

### Professional Escalation Criteria

Researchers should seek professional assistance when assembly problems exceed their expertise. This includes situations where the assembly fails repeatedly, where quality metrics are consistently poor, or where the assembly is intended for clinical or regulatory use. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) provides bioinformatics learning pathways that can help researchers build the skills needed to address assembly challenges.

## A Decision Framework for Selecting Phasing Signals and Allocating Sequencing Budget

### The Core Decision: Which Phasing Signal Matches Your Constraints

Researchers often choose a phasing strategy based on what is familiar instead of what the genome and project constraints require. This leads to wasted sequencing budget, failed assemblies, or phase information that cannot answer the biological question. A structured decision framework should be applied before any sequencing is ordered, because the choice of phasing signal determines library preparation, sequencing depth, and computational requirements.

The first decision point is sample availability. If parental DNA is accessible, trio binning provides the most unambiguous phase assignment because each offspring read is assigned to the haplotype of the parent from which it was inherited. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) describes a graph trio binning algorithm that advances over standard trio binning by integrating parental information into the phased assembly graph. This approach is particularly valuable for clinical genetics and family-based studies where parental origin of alleles matters. However, trio binning requires that both parents are available and that their DNA can be sequenced to sufficient depth for accurate variant calling. For nonmodel organisms, parental samples may be difficult or impossible to obtain, and this option is immediately eliminated.

The second decision point is the required phase contiguity. If the research question requires linking variants across megabases, such as for studying compound heterozygosity or allele-specific expression across large genes, then Hi-C or Strand-seq is necessary because read-based phasing alone cannot span these distances. The [diverse human genome study](https://pubmed.ncbi.nlm.nih.gov/33632895) used long-read and strand-specific sequencing technologies together to assemble haplotype-resolved human genomes without parent-child trio data, demonstrating that these approaches can substitute for parental information. If the research question only requires local phasing across a few kilobases, such as for resolving a single gene region, then read-based phasing with high-fidelity reads alone may be sufficient.

The third decision point is the heterozygosity rate of the genome. Highly heterozygous genomes require an assembler that preserves haplotype contiguity, such as hifiasm, and may benefit from additional phasing signals to resolve complex graph structures. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) demonstrates successful assembly of a hexaploid California redwood genome of approximately 30 gigabases, showing that graph-based haplotyping can handle extreme heterozygosity and polyploidy. Low-heterozygosity genomes present a different challenge: the lack of variation means that reads may not contain enough polymorphic sites to establish phase links, and additional phasing signals may be needed to bridge long stretches of identical sequence.

### A Practical Decision Matrix for Phasing Signal Selection

The following decision matrix organizes the selection process by sample availability, required phase contiguity, and genome complexity. This matrix is intended to be used during the experimental design phase, before sequencing is ordered.

| Scenario | Sample Availability | Required Phase Contiguity | Recommended Phasing Signal | Primary Tool or Method |
| --- | --- | --- | --- | --- |
| Family-based study, clinical genetics | Both parents available | Any | Trio binning plus HiFi reads | hifiasm with graph trio binning |
| Nonmodel organism, no parental samples | No parents | Local, within read length | HiFi reads only | hifiasm standard mode |
| Nonmodel organism, no parental samples | No parents | Long-range, across megabases | HiFi reads plus Hi-C | hifiasm with Hi-C integration |
| Diverse human genomes, no trios | No parents | Long-range, genome-wide | HiFi reads plus Strand-seq | hifiasm with Strand-seq integration |
| Single-cell study, limited DNA | No parents | Local to moderate | HiFi or ONT reads from single cells | hifiasm or single-cell assembly workflow |
| Targeted region, ultra-low error required | Any | Local, 10 kb or longer | Template mutagenesis with short reads | Specialized bench protocol and assembly algorithm |

The [template mutagenesis approach](https://pubmed.ncbi.nlm.nih.gov/35822882) occupies a distinct niche in this matrix. It uses short reads only and achieves per-base error rates below 10 to the negative ninth power for regions 10 kb and longer by imprinting template strands with dense mutation patterns. This approach is appropriate when ultra-low error is required for a targeted region and when long-read sequencing is unavailable or impractical. It is not appropriate for whole-genome assembly because the protocol is designed for targeted regions.

### Allocating Sequencing Budget Across Phasing Signals

Once the phasing signal is selected, the sequencing budget must be allocated across the required data types. The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) provides a useful reference point for coverage requirements. The study achieved an NG50 of approximately 2 megabases using 95 individual K562 cells on PacBio HiFi and Oxford Nanopore Technologies platforms, and an NG50 of over 1.3 megabases using 30 diploid HG002 cells at an average coverage of 41.7 percent on the Oxford Nanopore platform. These results demonstrate that contiguous assemblies are achievable with moderate coverage when reads are long and accurate.

For a standard diploid genome with HiFi reads, plan for at least 30-fold coverage of the nuclear genome for read-based phasing. This coverage provides enough overlapping reads to establish phase links across heterozygous sites and to distinguish true variation from sequencing errors. If Hi-C is added for long-range phasing, plan for an additional 30 to 50-fold coverage of Hi-C reads, which are typically shorter and require more depth to produce reliable contact maps. If Strand-seq is used, the coverage requirements depend on the number of single-strand libraries sequenced, which is typically determined by the desired density of phase-informative reads.

For trio binning, the offspring HiFi reads remain the primary assembly input, but parental short reads must be sequenced to sufficient depth for accurate variant calling. A common target is 30-fold coverage of each parent with short reads, which provides enough depth to identify the parental alleles that are used to bin the offspring reads. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) describes the graph trio binning algorithm that uses parental information to improve the phased assembly graph, and this algorithm requires parental variant calls as input.

### Recording the Decision Rationale and Budget Allocation

A record of the phasing signal selection and budget allocation is essential for reproducibility and for troubleshooting if the assembly fails. The record should include the following elements:

1. The research question and the required phase contiguity
2. The sample availability and whether parental samples were obtained
3. The estimated heterozygosity rate of the genome, based on prior knowledge or a pilot sequencing run
4. The selected phasing signal and the rationale for that selection
5. The sequencing platform, read type, and coverage target for each data type
6. The assembler and version, and the specific parameters used
7. The expected quality metrics, such as N50 and switch error rate, based on published benchmarks

This record should be maintained in a version-controlled repository alongside the assembly scripts. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in version control with Git, which is essential for tracking changes to assembly workflows. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that emphasize reproducibility, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that includes documentation practices.

### Troubleshooting When the Selected Strategy Fails

When an assembly fails to meet quality thresholds, the first step is to determine whether the failure is due to insufficient data, incorrect parameter settings, or a mismatch between the phasing signal and the genome. The following troubleshooting sequence addresses the most common failure patterns.

If the assembly shows haplotype collapse in heterozygous regions, the cause is likely insufficient read length or coverage to distinguish true allelic variation from sequencing errors. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) notes that existing algorithms either collapse heterozygous alleles into one consensus copy or fail to cleanly separate the haplotypes, which is the problem that hifiasm was designed to solve. If hifiasm was used and collapse still occurred, increase coverage or verify that the input reads meet the quality thresholds for high-fidelity reads.

If the assembly shows phase breaks in low-complexity regions, the cause is likely insufficient variation to establish phase links. This is a fundamental limitation of read-based phasing and cannot be fully resolved by increasing coverage alone. The [template mutagenesis approach](https://pubmed.ncbi.nlm.nih.gov/35822882) addresses this problem for targeted regions by imprinting template strands with dense mutation patterns, but this approach is not applicable to whole-genome assembly. For whole-genome assembly, consider adding Hi-C or Strand-seq to provide long-range phase information that bridges low-complexity regions.

If the assembly shows switch errors, the cause is likely incorrect graph traversal in regions where the phased assembly graph is ambiguous. Switch errors can be detected by comparing the assembly to known phase information, such as parental genotypes or independent phasing data. If switch errors are concentrated in specific regions, examine those regions for low coverage or high similarity between haplotypes, and consider whether additional sequencing or a different phasing signal would resolve the ambiguity.

If the assembly shows misassembly at repeat boundaries, the cause is likely ambiguous overlaps between identical or nearly identical sequences. Long reads reduce but do not eliminate this problem. The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) demonstrates that long-read sequencing can produce contiguous assemblies even from individual cells, but the study also notes that assembly quality depends on the assembler and sequencing strategy. If misassembly persists, consider whether the repeat content of the genome requires even longer reads or a different assembly approach.

### Professional Escalation Criteria for Phasing Strategy Failures

Researchers should seek professional assistance when the troubleshooting sequence does not resolve the failure. Specific escalation criteria include:

1. The assembly fails repeatedly with the same error pattern despite parameter adjustments
2. The quality metrics are consistently below published benchmarks for similar genomes
3. The assembly is intended for clinical or regulatory use and must meet specific quality standards
4. The genome has unusual features, such as extreme polyploidy or large repeat arrays, that exceed the researcher's experience
5. The computational requirements exceed available resources and a different approach is needed

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) provides bioinformatics learning pathways that can help researchers build the skills needed to address assembly challenges. The [Bioconductor Project](https://bioconductor.org/) offers official package and workflow documentation for reproducible genomic analysis, which may be useful for downstream validation and interpretation of phased assemblies. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of databases and search systems for sequence data, which are useful for depositing and retrieving assembly data and for comparing your assembly to existing resources.

## Frequently Asked Questions

### What is the difference between haplotyping and genome assembly?

Genome assembly is the process of reconstructing a genome sequence from sequencing reads. Haplotyping is the process of separating the two parental copies of a diploid genome into distinct sequences. A standard genome assembly produces one consensus sequence that may collapse heterozygous alleles. A haplotype-resolved assembly produces two sequences, one for each parental chromosome, that preserve the phase information linking variants along each chromosome.

### Why do standard assemblers collapse heterozygous alleles?

Standard assemblers collapse heterozygous alleles because they are designed to produce a single consensus sequence. When an assembler encounters a heterozygous site, it must decide whether to represent both alleles or merge them. Many assemblers merge the alleles to simplify the assembly graph, which loses the phase information. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) describes this problem and presents a solution that preserves the contiguity of all haplotypes in a phased assembly graph.

### What is a phased assembly graph?

A phased assembly graph is a computational structure that represents all possible haplotype paths in a single graph. Nodes represent sequence, and edges represent overlaps between reads. In a phased assembly graph, each haplotype follows a distinct path, allowing the assembler to extract contiguous haplotype sequences. The hifiasm assembler constructs such a graph using high-fidelity reads and preserves the contiguity of all haplotypes.

### How do Hi-C reads help with phasing?

Hi-C reads capture chromatin conformation by crosslinking DNA in the nucleus, digesting it, and sequencing the resulting ligation junctions. The frequency of contacts between genomic regions reflects their physical proximity in the nucleus, which correlates with their linear distance along the chromosome. This information provides long-range phase links that can connect haplotypes across megabases, which is useful when parental samples are unavailable.

### What is trio binning and when should it be used?

Trio binning uses short reads from both parents to assign offspring reads to parental haplotypes before assembly. This approach resolves phase unambiguously because each read is assigned to the haplotype of the parent from which it was inherited. Trio binning should be used when parental samples are available and when unambiguous phase assignment is critical. The [hifiasm publication](https://pubmed.ncbi.nlm.nih.gov/33526886) describes a graph trio binning algorithm that advances over standard trio binning.

### What are switch errors and how can they be detected?

Switch errors occur when the assembly switches from one haplotype to the other, meaning that variants on either side of the switch are assigned to the wrong haplotype. Switch errors can be detected by comparing the assembly to known phase information, such as parental genotypes or independent phasing data. In the absence of such data, switch errors can be inferred from read mapping patterns, where reads that map to both haplotypes in a region may indicate a switch error.

### Can haplotype-resolved assembly be done from single cells?

Yes, haplotype-resolved assembly can be done from single cells using long-read sequencing technologies. The [single-cell assembly study](https://pubmed.ncbi.nlm.nih.gov/35819189) used SMOOTH-seq to sequence individual cells on PacBio HiFi and Oxford Nanopore Technologies platforms and completed human genome assemblies with high continuity from individual cells. The study demonstrated that cell heterogeneity, which could affect haplotype assembly results, can be addressed with single-cell approaches.

### What is the role of polishing in haplotype-resolved assembly?

Polishing corrects residual errors in the assembled sequence. For haplotype-resolved assemblies, polishing must be performed carefully to avoid collapsing haplotypes that were correctly separated. The [template mutagenesis approach](https://pubmed.ncbi.nlm.nih.gov/35822882) achieves ultra-low error rates by imprinting template strands with mutation patterns and using unmutated libraries for correction, but this approach is specialized and not applicable to all projects.

## Related Bioinformatics Guides

- [Transcriptome Assembly Without a Reference Genome](/knowledge/bioinformatics/transcriptome-assembly-without-a-reference-genome)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm.](https://pubmed.ncbi.nlm.nih.gov/33526886). Nature methods, 2021.
- [De novo assembly of human genome at single-cell levels.](https://pubmed.ncbi.nlm.nih.gov/35819189). Nucleic acids research, 2022.
- [Complete de novo assembly and re-annotation of the zebrafish genome.](https://pubmed.ncbi.nlm.nih.gov/41332591). bioRxiv : the preprint server for biology, 2025.
- [Targeted de novo phasing and long-range assembly by template mutagenesis.](https://pubmed.ncbi.nlm.nih.gov/35822882). Nucleic acids research, 2022.
- [Haplotype-resolved diverse human genomes and integrated analysis of structural variation.](https://pubmed.ncbi.nlm.nih.gov/33632895). Science (New York, N.Y.), 2021.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.