Graph-Based Haplotyping in De Novo Assembly: Reconstructing Diploid Genomes without a Reference
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Graph-based haplotyping reconstructs diploid genomes de novo by leveraging assembly graphs, where heterozygous sites manifest as "bubbles" representing alternative allele paths. Standard assemblers often collapse these bubbles into a single consensus, necessitating specialized algorithms that preserve haplotype contiguity.
- Long, high-fidelity reads (e.g., PacBio HiFi, Oxford Nanopore) are crucial for graph-based haplotyping as they can span multiple heterozygous sites, directly linking variants into haplotype blocks and enabling the resolution of complex genomic regions and structural variations.
- Trio binning, utilizing parental sequencing data, offers the most robust method for haplotype resolution by enabling accurate classification of offspring reads into maternal and paternal origins before assembly, thereby minimizing switch errors.
- When parental data are unavailable, single-sample haplotyping relies on long reads to infer haplotype structure by identifying reads that traverse heterozygous sites, though completeness of phasing is dependent on read length and coverage density.
- Hi-C data serve to scaffold contigs into chromosome-level assemblies and phase distant variants, while linked reads offer a cost-effective alternative for phasing within the 10-100 kb range, bridging the gap between short and long reads.
- Assembly quality assessment for phased genomes must include contiguity (N50/NG50 for each haplotype), completeness (BUSCO scores per haplotype), base accuracy, and critically, the switch error rate, which quantifies inter-haplotype misassignments.
Direct Answer and Scope
Researchers who need to phase variants and reconstruct haplotypes from assembly graphs face a distinct methodological problem: most assemblers collapse heterozygous alleles into a single consensus copy, and separating the two parental haplotypes requires a clear understanding of graph structures, data types, and algorithmic choices. This article explains how assembly graphs encode haplotype information, compares the practical inputs of Hi-C, linked reads, and long reads for graph-based haplotyping, and provides concrete workflow decisions, quality checks, and limitations for diploid genome reconstruction without a reference genome. The content is written for biology students, researchers, laboratory professionals, and life-science practitioners who need actionable guidance on selecting sequencing strategies, running assemblers, evaluating phased assemblies, and interpreting results within the bounds of current evidence.
The Problem of Diploid Assembly and Haplotype Collapse
Why Reference-Based Phasing Fails for Novel Genomes
Reference-based phasing aligns sequencing reads to an existing genome and uses variant calls to assign alleles to parental chromosomes. This approach fails when no closely related reference exists, when the organism has high structural divergence from the reference, or when the genomic region of interest is absent from the reference assembly. For non-model organisms, agricultural species with complex genomes, or populations with high structural variation, the reference itself may contain collapsed repeats, misassembled regions, or missing sequences that bias every downstream phasing decision.
De novo assembly removes the dependency on a reference by reconstructing the genome from the reads themselves. However, standard de novo assemblers that produce a single haploid consensus sequence cannot represent the two distinct haplotypes present in a diploid organism. Heterozygous alleles are either collapsed into one sequence or represented as ambiguous paths in the assembly graph. The result is a mosaic of the two parental genomes that obscures true haplotype structure and complicates variant interpretation.
The Assembly Graph as a Haplotype Substrate
Assembly graphs, including de Bruijn graphs and overlap graphs, represent reads as nodes or edges and connections between overlapping sequences as graph topology. In a diploid genome, heterozygous sites create bubbles in the graph where two alternative sequences diverge and then reconverge. These bubbles are the raw material for haplotyping. A phased assembly graph preserves the contiguity of all haplotypes instead of collapsing them into one path.
The hifiasm assembler explicitly takes advantage of long high-fidelity sequence reads to represent haplotype information in a phased assembly graph. Unlike graph-based assemblers that aim to maintain the contiguity of only one haplotype, hifiasm preserves the contiguity of all haplotypes, which enables the development of graph trio binning algorithms that advance over standard trio binning approaches. This design choice matters for researchers because the assembler's objective function determines whether the output contains one sequence per locus or two distinct haplotype sequences.
Heterozygosity and Sequencing Coverage Requirements
The success of graph-based haplotyping depends on the density and distribution of heterozygous sites. Regions with low heterozygosity provide few graph bubbles to distinguish haplotypes, while regions with high heterozygosity create complex graph structures that require deep coverage to resolve. Long reads that span multiple heterozygous sites allow the assembler to link variants into haplotype blocks. The read length determines how far apart two heterozygous sites can be and still be connected by a single read.
For single-cell genome assembly, sequencing coverage per cell is a limiting factor. A study using SMOOTH-seq on PacBio HiFi and Oxford Nanopore Technologies platforms completed human genome assembly with high continuity using 95 individual K562 cells, and with sequencing data from 30 diploid individual HG002 cells at average coverage of approximately 41.7 percent on the ONT platform, the NG50 reached over 1.3 Mb. These results demonstrate that single-cell de novo assembly is feasible but requires careful attention to coverage depth and platform choice.
At a Glance: Data Inputs for Graph-Based Haplotyping
| Data Type | Read Length | Phasing Range | Key Strength | Primary Limitation | Best Use Case |
|---|---|---|---|---|---|
| Hi-C | Short reads with long-range ligation contacts | Chromosome scale | Links distant loci across megabases | Requires high-quality chromatin preparation | Scaffolding haplotigs into chromosome-level assemblies |
| Linked reads | Short reads with shared barcodes | 10 to 100 kb | Cost-effective phasing of common variants | Barcode density limits resolution in low-heterozygosity regions | Phasing variants in projects with limited long-read budget |
| Long reads (HiFi or ONT) | 10 to 100 kb | Read length dependent | Directly spans heterozygous sites and complex repeats | Higher cost per base and higher DNA input requirements | De novo assembly and haplotype-resolved assembly of complex genomes |
Core Principles of Graph-Based Haplotyping
Phased Assembly Graphs versus Standard Assembly Graphs
A standard assembly graph contains all reads and their overlaps, but the output is a single consensus sequence per locus. A phased assembly graph explicitly represents alternative paths that correspond to different haplotypes. The distinction matters for downstream analysis because a phased graph allows direct extraction of haplotype sequences, whereas an unphased graph requires additional computational steps to separate haplotypes after assembly.
The hifiasm approach strives to preserve the contiguity of all haplotypes in the graph itself. This design enables the development of graph trio binning, which uses parental sequencing data to assign reads to maternal or paternal haplotypes before assembly. The graph structure then reflects the two haplotypes as separate paths instead of collapsed consensus.
Trio Binning and Its Graph-Based Extension
Standard trio binning uses short reads from both parents to classify offspring reads as maternal, paternal, or ambiguous based on which parental genome they match more closely. The classified reads are then assembled separately to produce two haploid assemblies. Graph trio binning extends this concept by using the phased assembly graph to improve read classification and to resolve regions where parental reads do not provide clear discrimination.
The practical advantage of graph trio binning is that it produces two assemblies that are each derived from a single haplotype, reducing the chance of switching errors where the assembled sequence jumps from one haplotype to the other. For researchers with access to parental DNA, trio binning is the most direct route to haplotype-resolved assembly.
Single-Sample Haplotyping without Trio Data
When parental samples are unavailable, haplotyping must rely on the reads themselves. Long reads that span multiple heterozygous sites provide the linkage information needed to phase variants. The assembler identifies bubbles in the graph and uses reads that traverse the bubbles to determine which alleles belong to the same haplotype.
A study of 32 diverse human genomes demonstrated that long-read and strand-specific sequencing technologies together facilitate the de novo assembly of high-quality haplotype-resolved human genomes without parent-child trio data. The study produced 64 assembled haplotypes with an average minimum contig length needed to cover 50 percent of the genome of 26 million base pairs. This result shows that single-sample haplotyping is feasible with sufficient long-read coverage and appropriate assembly algorithms.
Practical Workflow for Graph-Based Haplotyping
Step 1: Define the Biological Question and Genome Characteristics
Before selecting a sequencing strategy, determine whether the goal is a single haploid reference-quality assembly, a fully phased diploid assembly, or a pangenome resource that represents multiple individuals. The answer determines the required data types and coverage. Also assess the genome size, ploidy level, heterozygosity rate, and repeat content. A hexaploid genome such as California redwood, with an estimated size of approximately 30 Gb, presents different challenges than a diploid human genome of 3 Gb.
Step 2: Select Sequencing Platforms and Coverage Targets
For haplotype-resolved assembly, long high-fidelity reads are the primary input. The hifiasm assembler was designed for PacBio HiFi reads and delivers better assemblies than existing tools on human and nonhuman datasets. Oxford Nanopore Technologies long reads are also viable, particularly when ultra-long reads are needed to span complex repeats or structural variation.
Coverage targets depend on the genome complexity and the desired haplotype resolution. Higher coverage improves the chance of spanning heterozygous sites and resolving graph bubbles, but also increases cost. For single-cell assembly, coverage per cell is limited by the amount of DNA available from one cell, and the NG50 of the resulting assembly depends on both the number of cells and the per-cell coverage.
Step 3: Generate or Obtain Parental Data for Trio Binning
If parental DNA is available, generate short-read or long-read data from both parents. The parental data are used to classify offspring reads into maternal and paternal bins. The quality of the parental data directly affects the accuracy of read classification and the completeness of the resulting haplotypes. For species where parental samples are difficult to obtain, consider whether single-sample haplotyping with long reads is sufficient for the research question.
Step 4: Run the Assembler with Haplotype-Aware Parameters
Use an assembler that explicitly supports phased assembly graphs. Configure the assembler to output both the primary assembly and the haplotype-resolved assemblies. For hifiasm, the default output includes the primary assembly and the two haplotype assemblies. Review the assembler documentation for parameters that control the tradeoff between contiguity and haplotype separation.
Step 5: Evaluate Assembly Quality with Multiple Metrics
Assembly quality is not a single number. Assess contiguity with NG50 and N50 statistics, completeness with BUSCO scores, and base accuracy with read mapping and variant calling. For haplotype-resolved assemblies, also evaluate the switch error rate, which measures how often the assembled haplotype switches from one parental chromosome to the other. A low switch error rate indicates that the haplotypes are cleanly separated.
Step 6: Validate Haplotype Assignments with Independent Data
If possible, validate the phased haplotypes using an independent method such as Hi-C data, which provides long-range chromatin interaction information that can confirm the assignment of contigs to chromosomes and haplotypes. Alternatively, use PCR-free short-read sequencing to genotype known variants and check that the phased haplotypes are consistent with the observed allele combinations.
Data Requirements and Tradeoffs
Hi-C Data for Chromosome-Scale Phasing
Hi-C sequencing captures physical contacts between distant genomic loci, providing information about the three-dimensional organization of chromosomes. In the context of graph-based haplotyping, Hi-C data can be used to scaffold haplotype-resolved contigs into chromosome-level assemblies and to phase variants that are too far apart to be connected by long reads alone.
The limitation of Hi-C is that the contact maps are noisy and require substantial sequencing depth to produce reliable phasing decisions. The quality of the Hi-C library preparation is critical, and cross-linking artifacts can create spurious contacts that mislead scaffolding algorithms. Researchers should expect to generate at least 30 to 50 million read pairs for a mammalian-sized genome, with higher coverage needed for complex genomes.
Linked Reads for Cost-Effective Phasing
Linked reads are produced by partitioning long DNA molecules into many small aliquots, tagging each aliquot with a unique barcode, and sequencing the fragments with short-read technology. Reads that share a barcode originate from the same long DNA molecule and therefore provide linkage information across the length of that molecule.
The practical advantage of linked reads is cost. Short-read sequencing is substantially cheaper per base than long-read sequencing, and the barcode information provides phasing information that is not available from standard short reads. However, the phasing range is limited by the length of the original DNA molecules, typically 10 to 100 kb, and the barcode diversity limits the number of distinct molecules that can be tracked in a single experiment.
Long Reads for Direct Haplotype Spanning
Long reads provide the most direct evidence for haplotyping because a single read can span multiple heterozygous sites. PacBio HiFi reads offer high accuracy with read lengths of 10 to 25 kb, while Oxford Nanopore reads can be much longer but with lower per-base accuracy. The choice between the two platforms depends on the specific requirements of the assembly project.
For graph-based haplotyping, the key metric is the number of heterozygous sites that a single read can span. Longer reads reduce the number of reads needed to connect distant variants and improve the contiguity of the resulting haplotype blocks. However, longer reads also require more DNA input and may have higher error rates that complicate graph construction.
Single-Cell Sequencing for Heterogeneous Samples
Most genome assembly projects require large amounts of DNA from homogeneous cell lines, which does not preserve cell heterogeneity. Single-cell genome long-read sequencing addresses this limitation by sequencing the genome of individual cells. A study using SMOOTH-seq completed human genome assembly with high continuity using 95 individual K562 cells, and the resulting assembly enabled the identification of more complete and accurate insertion events and complex structural variations compared to bulk assembly.
The tradeoff is that single-cell sequencing provides limited coverage per cell, and the assembly quality depends on the number of cells sequenced and the per-cell coverage. Researchers studying heterogeneous tissues or tumors may need single-cell approaches, but should expect lower contiguity than bulk assembly with equivalent total sequencing output.
Assembly Quality Assessment and Records
Contiguity Metrics
Contiguity metrics describe how much of the genome is contained in the largest contigs or scaffolds. The N50 statistic is the length of the shortest contig such that contigs of that length or longer cover at least 50 percent of the assembly. The NG50 statistic is the same calculation applied to the known or estimated genome size instead of the assembly size. For haplotype-resolved assemblies, report N50 and NG50 separately for the primary assembly and for each haplotype assembly.
The zebrafish genome assembly project used homozygous fish from two lab strains for de novo genome assemblies and produced assemblies that incorporated 7 percent more genomic sequence than the previous reference assembly GRCz11, with an additional 130 million bases of previously unassembled sequence. This example illustrates that contiguity improvements come from both better sequencing data and better assembly algorithms.
Completeness Metrics
Completeness metrics assess whether the assembly contains all expected genes or conserved genomic elements. BUSCO scores measure the fraction of conserved single-copy orthologs that are present in the assembly. For haplotype-resolved assemblies, completeness should be assessed for each haplotype separately, because a haplotype that is missing a conserved gene may indicate a collapse or misassembly.
Base Accuracy and Switch Error Rate
Base accuracy is measured by mapping reads back to the assembly and counting mismatches, or by comparing the assembly to a known reference when one is available. For haplotype-resolved assemblies, the switch error rate is a critical metric that measures how often the assembled haplotype switches from one parental chromosome to the other. A high switch error rate indicates that the haplotypes are not cleanly separated and that downstream variant analysis may be unreliable.
Record Keeping for Reproducibility
Maintain detailed records of the sequencing platforms, coverage levels, assembler versions, and parameter settings used for each assembly. Record the input data file names and checksums, the exact commands used to run the assembler, and the versions of all software dependencies. This information is essential for reproducing the assembly and for troubleshooting when quality issues arise.
Reproducible workflow tools can help standardize the assembly process. The nf-core documentation describes community pipeline standards for usage, configuration, and reproducible workflow context, and the Galaxy Training Network provides accessible workflow training and analysis tutorials. The Carpentries lessons offer foundational computing, data, shell, Git, and programming training that is useful for researchers who need to build the computational skills required for genome assembly.
Common Failure Patterns in Graph-Based Haplotyping
Haplotype Collapse in Low-Heterozygosity Regions
When heterozygous sites are sparse, the assembly graph may not contain enough bubbles to distinguish the two haplotypes. The assembler may collapse the two haplotypes into a single consensus sequence, producing a haploid assembly that does not represent either parental chromosome accurately. This failure pattern is more common in inbred organisms or in regions of the genome with low diversity.
Switching Errors in Complex Repeat Regions
Repetitive regions create complex graph structures with many alternative paths. The assembler may switch from one haplotype to the other at a repeat boundary, producing a chimeric haplotype that combines sequences from both parental chromosomes. Switching errors are difficult to detect without independent validation data and can propagate errors into downstream variant analysis.
Coverage Imbalance Between Haplotypes
If one haplotype is sequenced more deeply than the other, the assembler may preferentially follow the higher-coverage path and discard the lower-coverage haplotype. This imbalance can arise from biases in DNA extraction, library preparation, or sequencing. The result is an assembly that represents one haplotype well and the other poorly.
Misassembly at Structural Variation Breakpoints
Structural variants create large differences between haplotypes that are difficult to represent in a single assembly graph. The assembler may misjoin sequences at the breakpoints of deletions, insertions, inversions, or translocations, producing contigs that do not correspond to any true genomic sequence. A study of 32 diverse human genomes identified 107,590 structural variants, of which 68 percent were not discovered with short-read sequencing, highlighting the importance of long reads for resolving structural variation.
Limitations and Interpretation Boundaries
What Graph-Based Haplotyping Cannot Resolve
Graph-based haplotyping cannot resolve haplotypes in regions where the sequencing data do not provide sufficient linkage information. If no read spans two heterozygous sites, the phase between those sites remains unknown. Similarly, if a region is not covered by any reads, the assembly will contain a gap that cannot be filled without additional sequencing.
The Impact of Sequencing Errors on Graph Topology
Sequencing errors create spurious branches in the assembly graph that must be distinguished from true heterozygous variants. High-accuracy reads such as PacBio HiFi reduce the number of spurious branches, but errors still occur and can lead to incorrect graph topology. The assembler's error correction and graph simplification steps are critical for producing a clean graph that accurately represents the true haplotypes.
The Role of the Reference in Validation
Even when the goal is reference-free assembly, a reference genome can be useful for validation. Comparing the de novo assembly to a reference can identify misassemblies, measure base accuracy, and assess completeness. However, the reference may contain errors or may not represent the population being studied, so validation against a reference should be complemented by read-based validation.
Pangenome Resources and Their Limitations
Pangenome resources that represent multiple individuals provide a broader view of genomic diversity than a single reference. The zebrafish genome project generated 40 draft haplotypes to create a zebrafish pangenome resource and demonstrated its utility for variant analysis. Pangenome graphs can represent variation more completely than a linear reference, but they require substantial computational resources and careful interpretation.
Safety and Regulatory Context for Genome Assembly Data
Data Management and Privacy Considerations
Genome assembly projects generate large amounts of sequence data that may include sensitive information about individuals, particularly for human studies. Researchers must comply with applicable data protection regulations and institutional review board requirements. Raw sequencing data should be stored securely, and access should be restricted to authorized personnel.
Data Sharing and Database Submission
Many funding agencies and journals require that genome assembly data be deposited in public databases. The National Center for Biotechnology Information provides official descriptions of NCBI databases, search systems, sequence resources, and analysis services, and the European Bioinformatics Institute offers bioinformatics learning pathways and data-resource training. Researchers should plan for data submission early in the project to ensure that the data are formatted correctly and that all required metadata are collected.
Ethical Considerations for Human Genome Assembly
Human genome assembly projects raise ethical considerations related to consent, privacy, and the potential for re-identification. Researchers should ensure that participants have provided informed consent for the specific analyses being conducted and that the data are managed in accordance with ethical guidelines. The potential for incidental findings should be addressed in the study protocol.
Professional Escalation Criteria
When to Seek Additional Expertise
Researchers should seek additional expertise when the assembly quality metrics fall below the thresholds required for the intended downstream analysis, when the assembly graph contains complex structures that cannot be resolved with the available data, or when the computational resources required for the assembly exceed the capacity of the local computing environment.
When to Generate Additional Data
Additional sequencing data may be needed when the assembly contains excessive gaps, when the haplotype switch error rate is too high, or when the coverage in specific genomic regions is insufficient. The decision to generate additional data should be based on the quality metrics and the specific requirements of the research question.
When to Consult a Bioinformatics Core Facility
A bioinformatics core facility can provide expertise in assembly algorithms, quality assessment, and troubleshooting. Researchers should consider consulting a core facility when the assembly project is large or complex, when the research team lacks experience with graph-based haplotyping, or when the project timeline is constrained.
Decision Framework for Selecting a Graph-Based Haplotyping Strategy
Step 1: Inventory Available Biological Materials and Data
Before choosing an assembly strategy, document what biological materials are accessible. The presence or absence of parental DNA is the single largest determinant of the haplotyping approach. If parental samples exist, graph trio binning with hifiasm delivers cleaner haplotype separation than any single-sample method, because the parental reads provide an external reference for classifying offspring reads into maternal and paternal bins before assembly begins. The hifiasm assembler explicitly enables graph trio binning, which advances over standard trio binning by using the phased assembly graph to improve read classification and resolve regions where parental reads do not provide clear discrimination.
If parental DNA is unavailable, the decision shifts to whether the research question requires fully phased haplotypes or whether a primary assembly with partial phasing is acceptable. Single-sample haplotyping with long reads can produce haplotype-resolved assemblies, as demonstrated in a study of 32 diverse human genomes that used long-read and strand-specific sequencing technologies to assemble 64 haplotypes without parent-child trio data. However, the completeness of single-sample phasing depends on read length and coverage, and some regions will remain unphased.
Also inventory the amount of DNA available. Standard bulk assembly requires microgram quantities of high-molecular-weight DNA. If only single cells are available, the strategy changes fundamentally. A study using SMOOTH-seq completed human genome assembly with high continuity using 95 individual K562 cells on PacBio HiFi and Oxford Nanopore Technologies platforms, and with 30 diploid individual HG002 cells at average coverage of approximately 41.7 percent on the ONT platform, the NG50 reached over 1.3 Mb. These results establish that single-cell assembly is feasible but requires sequencing many cells to compensate for the limited coverage per cell.
Step 2: Classify the Genome Complexity Profile
Record the genome size, ploidy, heterozygosity rate, and repeat content before selecting a sequencing platform. These parameters determine the coverage needed and the expected difficulty of graph construction. A hexaploid genome such as California redwood, with an estimated size of approximately 30 Gb, presents fundamentally different challenges than a diploid human genome of 3 Gb. The hifiasm assembler consistently outperformed other tools on both human and nonhuman datasets, including the California redwood hexaploid genome, which demonstrates that haplotype-aware assemblers can scale to complex polyploid genomes.
For heterozygosity, estimate the density of heterozygous sites per kilobase. This estimate can come from a small pilot sequencing run or from published data on closely related species. Low heterozygosity regions provide few graph bubbles to distinguish haplotypes, and the assembler may collapse the two haplotypes into a single consensus sequence. High heterozygosity regions create complex graph structures that require deep coverage to resolve. The read length determines how far apart two heterozygous sites can be and still be connected by a single read, so the interaction between heterozygosity and read length is the central constraint on phasing contiguity.
Step 3: Match Data Type to the Phasing Distance Requirement
The required phasing distance is the genomic span over which variants must be connected into haplotype blocks. This distance is determined by the downstream analysis. If the goal is to phase common single-nucleotide variants within genes, a phasing distance of 10 to 100 kb may suffice, and linked reads can provide this at lower cost than long reads. If the goal is to resolve structural variation or to produce chromosome-scale haplotypes, long reads combined with Hi-C scaffolding are necessary.
Use the following decision rules for data type selection. Choose Hi-C when the primary need is chromosome-scale scaffolding of haplotype-resolved contigs or when variants must be phased across megabase distances that exceed the span of any single read. Choose linked reads when the budget is constrained, the genome is small, and the phasing distance requirement is within the 10 to 100 kb range of the original DNA molecules. Choose long reads when the genome has complex repeats, when structural variation is the focus, or when single-sample haplotyping without trio data is required. Long reads provide the most direct evidence for haplotyping because a single read can span multiple heterozygous sites, and the hifiasm assembler was designed specifically for PacBio HiFi reads.
Step 4: Determine Coverage Targets from Genome Parameters
Coverage targets should be calculated from the genome size, heterozygosity, and the desired haplotype resolution instead of adopted from published defaults. For a diploid genome with moderate heterozygosity, the coverage must be sufficient for the assembler to distinguish true heterozygous bubbles from sequencing errors. Higher coverage improves the chance of spanning heterozygous sites and resolving graph bubbles, but also increases cost and computational requirements.
For single-cell assembly, the coverage calculation is different because the total sequencing output is distributed across many individual cells. The NG50 of the resulting assembly depends on both the number of cells sequenced and the per-cell coverage. The SMOOTH-seq study demonstrated that 95 individual K562 cells produced an assembly with NG50 of approximately 2 Mb, while 30 individual HG002 cells at average coverage of approximately 41.7 percent produced an NG50 over 1.3 Mb. These numbers provide a starting point for estimating the number of cells needed, but the actual requirement depends on the genome complexity and the desired contiguity.
Step 5: Select the Assembly Algorithm and Configure Parameters
Choose an assembler that explicitly supports phased assembly graphs. The hifiasm assembler preserves the contiguity of all haplotypes in the graph, unlike other graph-based assemblers that aim to maintain the contiguity of only one haplotype. This design choice is not a minor implementation detail. It determines whether the output contains one sequence per locus or two distinct haplotype sequences.
Configure the assembler to output both the primary assembly and the haplotype-resolved assemblies. For hifiasm, the default output includes the primary assembly and the two haplotype assemblies. Review the assembler documentation for parameters that control the tradeoff between contiguity and haplotype separation. Some parameters favor longer contigs at the expense of clean haplotype separation, while others favor clean separation at the expense of contiguity. The correct setting depends on the downstream analysis. If the goal is to identify structural variants, clean haplotype separation is more important than maximum contiguity. If the goal is to produce a reference-quality assembly for gene annotation, contiguity may take priority.
Step 6: Build a Validation Plan Before Running the Assembly
Plan the validation strategy before the assembly run begins, because some validation methods require additional data that must be generated in the same sequencing campaign. If Hi-C data will be used for validation, the Hi-C library must be prepared from the same DNA sample or from a sample from the same individual. If PCR-free short-read data will be used to genotype known variants and check phase consistency, these reads must be generated from the same individual.
The validation plan should specify the independent data types, the quality metrics that will be computed, and the thresholds that will trigger additional sequencing or a change in assembly parameters. A switch error rate above the threshold required for the downstream analysis should trigger either additional sequencing coverage or a re-run with different assembler parameters. A BUSCO completeness score below the expected range for the species should trigger an investigation into whether the assembly collapsed haplotypes or missed genomic regions.
Record System for Haplotyping Projects
Required Records for Each Assembly Attempt
Maintain a structured record for every assembly attempt, whether it succeeds or fails. The record should include the sample identifier, the DNA extraction method, the sequencing platform, the read length distribution, the total sequencing output in gigabases, the estimated coverage, the assembler name and version, the exact command line used, the parameter settings, the date of the run, and the computing environment including the operating system and available memory.
Record the input data file names and checksums so that the exact input data can be retrieved if the assembly needs to be reproduced. Record the versions of all software dependencies, because assembler behavior can change between versions and a reproducible assembly requires identical software versions. The nf-core documentation describes community pipeline standards for usage, configuration, and reproducible workflow context, and the Galaxy Training Network provides accessible workflow training and analysis tutorials that can help standardize the assembly process.
Quality Metric Log
Create a table that records the quality metrics for each assembly attempt. Include the N50 and NG50 for the primary assembly and for each haplotype assembly, the BUSCO completeness scores, the base accuracy measured by read mapping, and the switch error rate for the phased haplotypes. Record the date each metric was computed and the software version used for the computation.
The quality metric log serves two purposes. First, it allows comparison across assembly attempts to identify which parameter settings produce the best results. Second, it provides the documentation needed for publication and for data submission to public databases. The National Center for Biotechnology Information provides official descriptions of NCBI databases, search systems, sequence resources, and analysis services, and the European Bioinformatics Institute offers bioinformatics learning pathways and data-resource training that can help researchers prepare the required metadata.
Decision Log
Record the rationale for each major decision in the haplotyping project. This includes the choice of sequencing platform, the coverage target, the assembler selection, the parameter settings, and the validation strategy. The decision log should note the alternatives that were considered and the reason each alternative was rejected. This documentation is valuable when the project is reviewed by collaborators, when the results are challenged, or when the project is extended to additional samples.
The decision log also helps identify systematic errors in the decision-making process. If every assembly attempt uses the same sequencing platform and the same assembler parameters, the log will reveal that the project has not explored the parameter space and may be missing better assembly configurations.
Troubleshooting Method for Assembly Failures
Diagnose Low Contiguity
When the N50 or NG50 of the assembly falls below the expected range for the genome size and data type, check the coverage first. Low coverage is the most common cause of low contiguity because the assembler cannot extend contigs through regions where reads are sparse. Calculate the actual coverage from the read count and the estimated genome size, and compare it to the target coverage. If the actual coverage is below the target, additional sequencing is required.
If the coverage is adequate, check the read length distribution. Short reads produce fragmented assemblies because they cannot span repetitive regions or long heterozygous stretches. If the read length distribution is shorter than expected for the platform, the DNA may have been sheared during extraction or library preparation. Check the DNA quality metrics from the sequencing facility and consider re-extracting DNA if the fragment length is insufficient.
Diagnose High Switch Error Rate
A high switch error rate indicates that the assembled haplotypes are not cleanly separated. The most common cause is insufficient linkage information between heterozygous sites. If the reads are too short to span multiple heterozygous sites, the assembler cannot determine which alleles belong to the same haplotype. Increasing the read length or adding Hi-C data can improve the linkage information.
Another cause of high switch error rate is the collapse of haplotypes in low-heterozygosity regions. When heterozygous sites are sparse, the assembler may collapse the two haplotypes into a single consensus sequence, and the resulting assembly switches between haplotypes when it encounters a region with sufficient heterozygosity to separate them. This failure pattern is more common in inbred organisms or in regions of the genome with low diversity.
Diagnose Missing Sequence
When the assembly is missing sequence that is expected to be present, check whether the missing sequence is in a repetitive region. Repetitive regions create complex graph structures with many alternative paths, and the assembler may fail to resolve these structures, leaving gaps in the assembly. Long reads that span the entire repeat unit can resolve these regions, but the reads must be long enough to cover the repeat and the flanking unique sequence.
The zebrafish genome assembly project demonstrated that new assemblies can incorporate substantially more sequence than previous references. The new zebrafish assemblies incorporated 7 percent more genomic sequence than the previous reference genome GRCz11, with an additional 130 million bases of previously unassembled sequence. This example illustrates that missing sequence is often present in the reads but was not assembled by earlier algorithms or was not covered by earlier sequencing data.
Diagnose Coverage Imbalance Between Haplotypes
When one haplotype is sequenced more deeply than the other, the assembler may preferentially follow the higher-coverage path and discard the lower-coverage haplotype. This imbalance can arise from biases in DNA extraction, library preparation, or sequencing. Check the coverage distribution across the genome to identify regions where one haplotype has substantially lower coverage than the other. If the imbalance is consistent across the genome, the bias is likely in the library preparation. If the imbalance is localized, the bias may be in a specific genomic region such as a repeat or a GC-rich area.
Escalation Criteria for Professional Consultation
When to Generate Additional Data
Generate additional sequencing data when the assembly contains excessive gaps, when the haplotype switch error rate is too high for the downstream analysis, or when the coverage in specific genomic regions is insufficient. The decision to generate additional data should be based on the quality metrics and the specific requirements of the research question. If the assembly is intended for structural variant discovery, a higher switch error rate may be acceptable than if the assembly is intended for allele-specific expression analysis.
When to Change the Assembly Strategy
Change the assembly strategy when the current approach cannot achieve the required quality metrics even with additional data. For example, if single-sample haplotyping with long reads produces a switch error rate that is too high, consider whether parental DNA can be obtained for trio binning. If linked reads cannot provide the required phasing distance, consider whether the budget can accommodate long-read sequencing.
When to Consult a Bioinformatics Core Facility
Consult a bioinformatics core facility when the assembly project is large or complex, when the research team lacks experience with graph-based haplotyping, or when the project timeline is constrained. A core facility can provide expertise in assembly algorithms, quality assessment, and troubleshooting. The Carpentries lessons offer foundational computing, data, shell, Git, and programming training that can help researchers build the computational skills required for genome assembly, and the Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can support the downstream analysis of assembled genomes.
When to Reconsider the Research Question
Reconsider the research question when the required quality metrics cannot be achieved with any available data type or assembly strategy. If the genome is extremely large, highly repetitive, or has very low heterozygosity, fully phased haplotype-resolved assembly may not be feasible with current technology. In these cases, a primary assembly with partial phasing may be sufficient for the research question, or the question may need to be reframed to focus on regions of the genome that can be assembled and phased reliably.
Frequently Asked Questions
What is the difference between a primary assembly and a haplotype-resolved assembly?
A primary assembly is a single sequence per locus that represents a consensus of the two haplotypes. A haplotype-resolved assembly contains two distinct sequences per locus, one for each parental chromosome. The primary assembly is useful for many applications, but it obscures the true diploid structure and cannot be used to study allele-specific expression or to identify compound heterozygous variants.
How much sequencing coverage is needed for graph-based haplotyping?
The required coverage depends on the genome size, heterozygosity, and repeat content. For human genomes, coverage of 30 to 60 times with PacBio HiFi reads is typically sufficient for haplotype-resolved assembly. For single-cell assembly, the per-cell coverage is limited by the amount of DNA in a single cell, and the overall assembly quality depends on the number of cells sequenced.
Can graph-based haplotyping be done without parental data?
Yes, single-sample haplotyping is possible when long reads span multiple heterozygous sites. The assembler uses the linkage information in the reads to phase variants. However, the completeness and accuracy of the haplotypes depend on the read length and coverage, and some regions may remain unphased.
What is the switch error rate and why does it matter?
The switch error rate measures how often the assembled haplotype switches from one parental chromosome to the other. A high switch error rate indicates that the haplotypes are not cleanly separated, which can lead to incorrect variant calls and misinterpretation of allele-specific effects.
How do Hi-C data improve graph-based haplotyping?
Hi-C data provide long-range chromatin interaction information that can be used to scaffold haplotype-resolved contigs into chromosome-level assemblies and to phase variants that are too far apart to be connected by long reads. Hi-C is particularly useful for assigning contigs to chromosomes and for validating the overall structure of the assembly.
What are the main limitations of linked reads for haplotyping?
Linked reads provide phasing information across the length of the original DNA molecules, typically 10 to 100 kb. The phasing range is limited by the DNA molecule length, and the barcode diversity limits the number of distinct molecules that can be tracked. Linked reads are less effective in low-heterozygosity regions where there are few variants to phase.
How does single-cell sequencing differ from bulk sequencing for genome assembly?
Single-cell sequencing preserves cell heterogeneity but provides limited coverage per cell. Bulk sequencing uses DNA from many cells and provides higher coverage but averages over the cell population. Single-cell assembly is useful for studying heterogeneous tissues or tumors, but the assembly contiguity is generally lower than bulk assembly with equivalent total sequencing output.
What quality metrics should be reported for a haplotype-resolved assembly?
Report contiguity metrics such as N50 and NG50 for the primary assembly and for each haplotype, completeness metrics such as BUSCO scores, base accuracy measured by read mapping or comparison to a reference, and the switch error rate for the phased haplotypes. Also report the sequencing platforms, coverage levels, assembler versions, and parameter settings used for the assembly.
Related Bioinformatics Guides
- Transcriptome Assembly Without a Reference Genome
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm.. Nature methods, 2021.
- De novo assembly of human genome at single-cell levels.. Nucleic acids research, 2022.
- Complete de novo assembly and re-annotation of the zebrafish genome.. bioRxiv : the preprint server for biology, 2025.
- Targeted de novo phasing and long-range assembly by template mutagenesis.. Nucleic acids research, 2022.
- Haplotype-resolved diverse human genomes and integrated analysis of structural variation.. Science (New York, N.Y.), 2021.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.