# Heterozygosity and Assembly Quality: How to Assess and Improve Haplotype-Resolved Genomes


## Key Takeaways

-   **Heterozygosity necessitates specialized assembly:** Standard genome assemblers collapse divergent haplotypes into a single consensus sequence, leading to false structural variants, collapsed repeats, and missing genes, particularly problematic for outbred animals, cultivated plants, and wild populations.
-   **K-mer spectra are critical diagnostics:** Analyzing k-mer frequency distributions reveals heterozygosity levels, estimates genome size, and identifies sequencing errors, guiding the decision to pursue haplotype phasing when heterozygosity exceeds approximately 1%.
-   **Merqury quantifies assembly quality and haplotype issues:** This tool assesses completeness, accuracy (QV), and crucially, haplotype duplication by comparing assembly k-mers to read k-mers, providing a reference-free evaluation essential for non-model organisms.
-   **Long-read sequencing and advanced phasing are paramount:** Technologies like PacBio and Nanopore are essential for spanning heterozygous regions, while methods such as trio binning (using parental DNA) or Hi-C chromatin conformation capture are required to resolve individual haplotypes across entire chromosomes.
-   **Per-haplotype metrics are vital for evaluation:** Assessing assembly size, contiguity (N50), completeness (BUSCO, k-mer based), and structural variation independently for each haplotype is crucial to identify and rectify assembly imbalances or errors.

---

Genome assembly from highly heterozygous organisms presents a distinct challenge: standard assembly algorithms collapse divergent haplotypes into a single mosaic sequence, producing errors that appear as false structural variants, collapsed repeats, and missing genes. For researchers working with outbred animals, cultivated plants, or wild populations, the central question is whether an assembly faithfully represents both parental chromosome copies or merely a chimeric approximation. This article explains how to evaluate haplotype representation using k-mer spectra and tools such as Merqury, how to interpret metrics including haplotype duplication and completeness, and how to adjust assembly strategies when heterozygosity compromises quality.

The practical outcome for readers is a decision framework: when to accept a collapsed assembly, when to pursue haplotype phasing, which quality metrics to report, and how to recognize when an assembly requires rework. The guidance applies to researchers generating new assemblies, laboratory professionals evaluating assembly pipelines, and students learning genome assembly quality assessment.

## The Heterozygosity Problem in Genome Assembly

### Why Standard Assemblies Fail on Heterozygous Genomes

Most genome assemblers were designed under the assumption that a diploid genome contains two nearly identical copies of each chromosome. When sequence divergence between haplotypes exceeds a few percent, assemblers face a choice at each genomic region: merge the two haplotypes into one consensus sequence or attempt to represent them separately. The default behavior of many assemblers is to collapse, producing a haploid representation that mixes alleles from both parents.

The consequences of collapse are measurable. Genes present in only one haplotype may appear missing. Paralogous gene families can be incorrectly counted as duplications when haplotype variants are assembled as separate loci. Structural variants between haplotypes, such as inversions or translocations, become invisible or appear as assembly errors. The Vertebrate Genomes Project reported that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly, based on lessons from generating assemblies for 16 vertebrate species [<a href="#ref-1">1</a>].

For highly heterozygous species, the problem intensifies. The blue bat star genome showed a k-mer based heterozygosity rate of 5.16%, higher than any previously reported echinoderm species [<a href="#ref-2">2</a>]. At this level of divergence, standard assembly approaches produce fragmented, chimeric results. The cultivated octoploid strawberry, with high heterozygosity at most loci, required dedicated phasing to produce separate haplotypes of 825 Mb and 808 Mb with contig N50 values above 26 Mb [<a href="#ref-3">3</a>].

### What Haplotype-Resolved Assembly Means

A haplotype-resolved assembly represents each chromosome copy separately. For a diploid organism, this means two assemblies: one for the maternal haplotype and one for the paternal haplotype. For polyploids, the number of haplotypes equals the ploidy level. The octoploid strawberry work produced 56 chromosomes total, representing eight haplotypes across seven chromosome groups [<a href="#ref-3">3</a>].

Haplotype resolution differs from simple phasing. Phasing assigns variants to parental chromosomes across short regions. Haplotype-resolved assembly extends this concept to entire chromosomes, producing contiguous sequences for each homolog. The distinction matters for downstream analysis: variant calling against a collapsed reference misses heterozygous sites, while analysis against a haplotype-resolved assembly can identify allele-specific expression, parent-of-origin effects, and structural differences between haplotypes.

### When Haplotype Resolution Is Necessary

Not every genome project requires haplotype resolution. For a first-pass reference genome of a homozygous inbred line, a collapsed assembly may suffice. For clinical diagnostics targeting known variants, a high-quality collapsed reference with accurate variant calls can be adequate. The decision depends on the biological questions:

- Population genetics studies require accurate allele frequencies, which collapsed assemblies distort
- Comparative genomics of structural variants requires true haplotype representation
- Gene family analysis requires distinguishing paralogs from alleles
- Functional studies of allele-specific expression require phased haplotypes
- Breeding programs tracking trait-linked haplotypes require haplotype continuity

The Lactarius hatsudake genome provides a contrasting example. This mycorrhizal mushroom was assembled with a genome size of 76.7 Mb and scaffold N50 of 223.2 kb, with heterozygosity estimated at 1.14% [<a href="#ref-4">4</a>]. The assembly was not haplotype-resolved, yet it served its purpose for functional genomics and breeding applications. The researchers prioritized contiguity and completeness over haplotype separation because the biological questions did not require allele-level resolution.

## K-Mer Spectra as a Diagnostic Tool

### Principles of K-Mer Frequency Analysis

K-mer spectra provide a genome-wide view of sequence composition without alignment. A k-mer is a sequence of length k, typically 17 to 31 bases. Counting all k-mers in sequencing reads produces a frequency distribution that reflects genome structure. The spectrum contains distinct peaks corresponding to different sequence categories:

- Low-frequency k-mers represent sequencing errors, appearing once or twice
- The main peak represents homozygous regions where both haplotypes share identical sequence
- A secondary peak at approximately half the main peak frequency indicates heterozygous regions where the two haplotypes differ
- Very high-frequency k-mers represent repetitive sequences

The ratio of the heterozygous peak to the homozygous peak estimates heterozygosity. The Lactarius hatsudake study used k-mer frequency distribution to estimate genome size at 63.84 Mb and heterozygosity at 1.14% [<a href="#ref-4">4</a>]. The blue bat star study used k-mer analysis to estimate 5.16% heterozygosity [<a href="#ref-2">2</a>]. The Isodon rubescens f. lushanensis genome showed 1.7% heterozygosity with 83.43% repeat content [<a href="#ref-5">5</a>].

### Interpreting K-Mer Spectra for Assembly Decisions

Before assembly, k-mer spectra guide sequencing strategy. A spectrum with a prominent heterozygous peak signals that standard assembly will collapse divergent regions. The expected collapse rate correlates with the proportion of k-mers in the heterozygous peak. When heterozygosity exceeds approximately 1%, haplotype phasing should be considered.

After assembly, k-mer spectra assess assembly quality by comparing k-mers in the assembly to k-mers in the reads. This comparison reveals:

- Missing sequence: read k-mers absent from the assembly indicate gaps or collapsed regions
- Duplicated sequence: assembly k-mers present at higher frequency than read k-mers indicate false duplications
- Error k-mers: assembly k-mers absent from reads indicate base errors introduced during assembly

The Genome 10K consortium identified haplotype heterozygosity as a major source of assembly error when not handled correctly [<a href="#ref-1">1</a>]. K-mer analysis provides the quantitative basis for detecting this problem before it propagates through downstream analysis.

### Limitations of K-Mer Analysis

K-mer spectra have resolution limits. Very low heterozygosity below 0.1% may not produce a visible heterozygous peak. High repeat content complicates interpretation because repetitive k-mers appear at multiple frequencies. The Isodon rubescens genome, with 83.43% repeat content, required careful interpretation of k-mer spectra to distinguish heterozygous peaks from repeat peaks [<a href="#ref-5">5</a>].

K-mer analysis also cannot distinguish between different types of heterozygosity. A single nucleotide polymorphism and a large structural variant both reduce shared k-mer content between haplotypes, but they have different assembly implications. Complementary approaches, including alignment-based methods and long-read mapping, provide additional information.

## Merqury and Assembly Quality Metrics

### How Merqury Works

Merqury evaluates assembly quality by comparing k-mers in the assembly to k-mers in the sequencing reads. The tool requires a k-mer database built from the read set, typically using meryl or similar software. The assembly is then evaluated against this database to produce three core metrics:

- Completeness: the fraction of read k-mers present in the assembly
- Accuracy: the fraction of assembly k-mers present in the reads
- Haplotype duplication: the proportion of assembly k-mers present at twice the expected frequency

These metrics provide a reference-free assessment that does not depend on alignment to an external genome. This independence matters for non-model organisms where no closely related reference exists.

### Completeness and Its Interpretation

Completeness measures how much of the sequenced genome is represented in the assembly. A completeness value of 95% means 5% of k-mers in the reads are missing from the assembly. Missing k-mers can indicate:

- Unassembled regions, often repetitive or high-GC content
- Collapsed haplotypes where divergent sequence was merged
- Sequencing depth insufficient to cover all regions

The BUSCO assessment provides a complementary completeness measure based on conserved single-copy orthologs. The Lactarius hatsudake genome achieved 89.0% BUSCO completeness [<a href="#ref-4">4</a>]. For haplotype-resolved assemblies, completeness should be evaluated separately for each haplotype, because collapsing can inflate apparent completeness by merging divergent alleles into one sequence.

### Haplotype Duplication and False Duplications

Haplotype duplication occurs when both haplotypes are assembled as separate sequences but the assembler also produces a collapsed copy. This creates three representations of a region that should have two. The Genome 10K consortium identified false gene duplications as a common error in assemblies that did not properly handle heterozygosity [<a href="#ref-1">1</a>].

Merqury detects haplotype duplication by examining k-mer copy numbers. A region with true haplotype representation has k-mers at the expected diploid frequency. A region with duplication has k-mers at higher frequency, indicating that the assembler produced redundant sequence. The QV (quality value) metric derived from k-mer analysis provides an overall accuracy estimate, with higher values indicating fewer errors.

### Using Merqury in Practice

Merqury analysis requires a k-mer database from the read set and the assembly fasta file. The workflow follows these steps:

1. Build a k-mer database from the sequencing reads using meryl
2. Run Merqury with the assembly and the k-mer database
3. Examine the completeness, QV, and duplication metrics
4. Generate the k-mer spectrum plot showing the assembly k-mers overlaid on the read k-mer spectrum
5. Compare metrics across haplotypes and across assembly versions

The k-mer spectrum plot provides visual confirmation of assembly behavior. A well-assembled haploid genome shows a single peak matching the read spectrum. A haplotype-resolved assembly shows two peaks, one for each haplotype. A collapsed assembly shows a single peak at higher frequency than the read spectrum, indicating that divergent haplotypes were merged.

## Assembly Strategies for Heterozygous Genomes

### Long-Read Sequencing Requirements

Long-read sequencing is essential for haplotype-resolved assembly. The Genome 10K consortium confirmed that long-read sequencing technologies are essential for maximizing genome quality [<a href="#ref-1">1</a>]. Short reads cannot span heterozygous regions reliably because the reads are shorter than the distance between variants.

The sequencing depth required for haplotype resolution exceeds that for collapsed assembly. The Isodon rubescens f. lushanensis project used 139.07 Gb of PacBio and Nanopore data, achieving approximately 328x sequencing depth for a 349 Mb genome [<a href="#ref-5">5</a>]. The octoploid strawberry project used single molecule real-time sequencing combined with Hi-C for chromosome-scale phasing [<a href="#ref-3">3</a>].

For most diploid genomes, 30x to 60x coverage of long reads provides sufficient depth for haplotype phasing. Higher heterozygosity and higher repeat content require more depth. The decision should be based on k-mer spectrum analysis of initial sequencing data, not on a fixed coverage target.

### Trio Binning and Parental Phasing

Trio binning uses sequencing data from both parents to phase the offspring genome. Short reads from each parent are used to classify offspring long reads as maternal or paternal. The classified reads are then assembled separately, producing two haplotypes without the need for variant-based phasing.

This approach requires:

- DNA from both parents and the offspring
- Sufficient short-read coverage of each parent for k-mer classification
- Long-read coverage of the offspring for assembly

Trio binning works well for species where parental samples are available. It fails for wild-caught individuals, ancient DNA, or species where parents cannot be sampled. For these cases, algorithmic phasing based on variant linkage must be used.

### Hi-C and Chromatin Conformation Phasing

Hi-C sequencing captures chromatin interactions that occur more frequently within chromosomes than between chromosomes. This signal provides long-range phasing information that can separate haplotypes across entire chromosomes. The octoploid strawberry assembly used Hi-C to anchor and phase the genome into 56 chromosomes [<a href="#ref-3">3</a>].

Hi-C phasing works by:

1. Assembling contigs without phasing
2. Mapping Hi-C reads to the assembly
3. Using interaction frequency patterns to assign contigs to haplotypes
4. Ordering and orienting contigs within each haplotype

Hi-C phasing requires high-quality Hi-C libraries and sufficient sequencing depth. The interaction signal must be strong enough to distinguish within-haplotype from between-haplotype interactions. For highly heterozygous genomes, the Hi-C signal can be confounded by structural differences between haplotypes.

### Single-Cell and Strand-Sequence Approaches

Single-cell sequencing and strand-seq provide alternative phasing approaches. Single-cell sequencing amplifies one chromosome copy at a time, allowing direct haplotype separation. Strand-seq exploits the strand-specific inheritance of DNA to phase variants.

These approaches are technically demanding and expensive at scale. They are most useful for validating haplotype-resolved assemblies or for species where other phasing methods fail. Most genome projects should first attempt trio binning or Hi-C phasing before considering single-cell approaches.

## Evaluating Haplotype-Resolved Assemblies

### Per-Haplotype Quality Metrics

Haplotype-resolved assemblies require quality assessment for each haplotype separately. The octoploid strawberry assembly reported Hap1 at 825 Mb with contig N50 of 26.70 Mb and Hap2 at 808 Mb with contig N50 of 27.51 Mb [<a href="#ref-3">3</a>]. These metrics show that both haplotypes achieved similar contiguity, indicating balanced assembly quality.

Per-haplotype metrics to report include:

- Assembly size and contig N50
- Completeness based on k-mer analysis and BUSCO
- QV or base accuracy
- Number of gaps and unplaced contigs
- Structural variant calls between haplotypes

Discrepancies between haplotypes indicate assembly problems. A substantially smaller haplotype may have collapsed regions. A haplotype with lower BUSCO completeness may have missing genes. These discrepancies should be investigated before downstream analysis.

### Comparing Haplotypes for Structural Variants

Haplotype-resolved assemblies enable direct comparison of structural variation between chromosome copies. The octoploid strawberry assembly identified a 10 Mb inversion and translocation on chromosome 2-1 [<a href="#ref-3">3</a>]. The Isodon rubescens f. lushanensis assembly revealed genomic structure variations between two forms of Isodon rubescens [<a href="#ref-5">5</a>].

Structural variant detection between haplotypes requires:

1. Aligning one haplotype to the other
2. Identifying breakpoints in the alignment
3. Validating breakpoints with read coverage and junction reads
4. Classifying variants as inversions, translocations, duplications, or deletions

The biological significance of structural variants between haplotypes depends on the species and the genes affected. In the strawberry example, the inversion and translocation may affect gene expression and recombination patterns [<a href="#ref-3">3</a>]. In medicinal plants, structural variants can affect biosynthetic gene clusters and chemotype differences [<a href="#ref-5">5</a>].

### Gene Annotation Across Haplotypes

Gene annotation of haplotype-resolved assemblies presents unique challenges. The octoploid strawberry assembly annotated 104,957 and 102,356 protein-coding genes in Hap1 and Hap2, respectively [<a href="#ref-3">3</a>]. The difference of approximately 2,600 genes between haplotypes requires interpretation: some differences represent true presence-absence variation, while others reflect annotation errors.

When annotating haplotypes:

- Use the same annotation pipeline for both haplotypes to enable comparison
- Validate gene models with transcriptome data from multiple tissues
- Distinguish paralogs from alleles using phylogenetic analysis
- Report presence-absence variation separately from copy number variation

The Isodon rubescens f. lushanensis assembly predicted 34,865 protein-coding genes [<a href="#ref-5">5</a>]. For this diploid species, the single assembly represents one haplotype, and the gene count reflects that representation. A haplotype-resolved version would likely show similar gene counts per haplotype with allelic variation between them.

## At a Glance: Assembly Quality Assessment Decision Table

| Scenario | Recommended Action | Key Metrics to Monitor | Escalation Criterion |
| --- | --- | --- | --- |
| Heterozygosity below 0.5% | Proceed with standard collapsed assembly | Assembly size vs k-mer estimate, BUSCO completeness | Completeness below 90% or unexpected duplication |
| Heterozygosity 0.5% to 2% | Consider haplotype phasing with trio binning or Hi-C | Per-haplotype contig N50, Merqury duplication, haplotype size balance | Size difference between haplotypes exceeds 10% |
| Heterozygosity above 2% | Haplotype phasing strongly recommended with multiple long-read platforms | K-mer spectrum shape, phasing continuity, structural variant validation | Chimeric contigs or alternating haplotype coverage |
| High repeat content above 60% | Increase sequencing depth and use ultra-long reads | Repeat composition, k-mer peak separation, gap distribution | Unresolved repeat arrays or fragmented centromeres |
| Polyploid genome | Use polyploid-aware assemblers and Hi-C for chromosome assignment | Per-subgenome completeness, chromosome number verification | Missing chromosomes or unbalanced subgenome representation |

## Practical Workflow for Assembly Quality Assessment

### Step 1: Pre-Assembly K-Mer Analysis

Before assembling, analyze the k-mer spectrum of the sequencing reads to estimate genome size, heterozygosity, and repeat content. This analysis guides assembly strategy and identifies potential problems early.

Actions based on results:

- Heterozygosity below 0.5%: standard assembly may suffice
- Heterozygosity 0.5% to 2%: consider haplotype phasing
- Heterozygosity above 2%: haplotype phasing is strongly recommended
- High repeat content: increase sequencing depth and use ultra-long reads

The blue bat star genome at 5.16% heterozygosity required specialized approaches [<a href="#ref-2">2</a>]. The Lactarius hatsudake genome at 1.14% heterozygosity was assembled without phasing [<a href="#ref-4">4</a>]. These examples illustrate the range of strategies appropriate for different heterozygosity levels.

### Step 2: Assembly and Initial Quality Metrics

Run the assembly and compute initial quality metrics including contig N50, assembly size, and BUSCO completeness. Compare these metrics to the k-mer based genome size estimate. Large discrepancies indicate assembly problems.

The Lactarius hatsudake assembly illustrates this comparison: k-mer based genome size was 63.84 Mb, but the final assembly was 76.7 Mb [<a href="#ref-4">4</a>]. The difference likely reflects repetitive sequences that were not fully represented in the k-mer estimate. The assembly size exceeding the k-mer estimate can also indicate haplotype duplication.

### Step 3: Merqury Evaluation

Run Merqury to obtain completeness, QV, and duplication metrics. Generate the k-mer spectrum plot and examine the distribution of assembly k-mers relative to read k-mers.

Interpretation guidelines:

- Completeness above 95%: acceptable for most applications
- Completeness 90% to 95%: investigate missing regions
- Completeness below 90%: assembly requires improvement
- QV above 40: high base accuracy
- QV 30 to 40: moderate accuracy, check for systematic errors
- QV below 30: assembly requires polishing or rework

### Step 4: Haplotype-Specific Assessment

For haplotype-resolved assemblies, evaluate each haplotype separately. Compare assembly sizes, contiguity, completeness, and gene content between haplotypes. Investigate any substantial discrepancies.

The strawberry assembly showed Hap1 at 825 Mb and Hap2 at 808 Mb, a difference of 17 Mb [<a href="#ref-3">3</a>]. This difference could reflect true size variation between haplotypes or assembly artifacts. The contig N50 values of 26.70 Mb and 27.51 Mb indicate similar contiguity, suggesting the size difference is biological instead of technical.

### Step 5: Validation and Polishing

Validate the assembly with independent data. Map reads back to the assembly and check for coverage uniformity. Use polishing tools to correct base errors. Re-run Merqury after polishing to confirm improvement.

Polishing decisions:

- Low QV with high coverage: run multiple polishing rounds
- Low QV with low coverage: consider additional sequencing
- Localized errors: polish specific regions instead of the whole assembly
- Structural errors: cannot be fixed by polishing, require reassembly

### Step 6: Documentation and Reporting

Document all assembly parameters, quality metrics, and validation results. Report the k-mer based heterozygosity estimate, the assembly strategy, and the final quality metrics. This documentation enables reproducibility and provides context for downstream users.

The nf-core documentation emphasizes reproducible workflow standards for genomic analysis [<a href="#ref-6">6</a>]. Following community standards for assembly reporting facilitates comparison across projects and enables meta-analyses of assembly quality.

## Common Failure Patterns in Haplotype-Resolved Assembly

### Collapsed Haplotypes

Collapsed haplotypes occur when the assembler merges divergent alleles into a single sequence. This failure produces:

- Assembly size smaller than expected
- Reduced heterozygosity in the assembly compared to reads
- False homozygous variants in downstream analysis
- Missing allele-specific expression

Detection: compare assembly size to k-mer based genome size estimate. Run Merqury and examine the k-mer spectrum for evidence of collapse. The assembly k-mer peak will appear at higher frequency than the read k-mer peak.

### Haplotype Duplication

Haplotype duplication occurs when the assembler produces both haplotypes plus a collapsed copy. This failure produces:

- Assembly size larger than expected
- Duplicated genes in annotation
- False copy number variants
- Problems with variant calling due to multi-mapping reads

Detection: Merqury duplication metrics will show elevated copy numbers. BUSCO analysis will show increased duplication rates. The k-mer spectrum will show assembly k-mers at twice the expected frequency.

### Chimeric Haplotypes

Chimeric haplotypes occur when the assembler joins sequences from different haplotypes into one contig. This failure produces:

- Contigs with alternating haplotype origin
- False structural variants
- Broken gene models
- Incorrect phasing

Detection: map reads from each parent to the assembly and check for alternating coverage patterns. Use Hi-C data to check for consistent chromatin interaction patterns along contigs.

### Unbalanced Haplotype Representation

Unbalanced representation occurs when one haplotype is well assembled but the other is fragmented or incomplete. This failure produces:

- Large size difference between haplotypes
- Different contig N50 values between haplotypes
- Different BUSCO completeness between haplotypes
- Missing genes in one haplotype

Detection: compare all quality metrics between haplotypes. Investigate regions present in one haplotype but absent from the other.

## Records and Measurements for Assembly Projects

### Essential Records

Maintain detailed records of all assembly decisions and measurements. The following records support reproducibility and troubleshooting:

- Sequencing platform, chemistry version, and basecalling model
- Read length distributions and coverage calculations
- K-mer analysis parameters and results
- Assembly software versions and parameters
- Phasing method and parameters
- Polishing rounds and changes made
- Quality metrics at each assembly stage

The Galaxy Training Network provides accessible workflow training for genomic analysis [<a href="#ref-7">7</a>]. Following structured training materials helps ensure that assembly records follow community standards.

### Quality Metrics to Track

Track these metrics throughout the assembly process:

- Read N50 and read length distribution
- K-mer based genome size estimate
- Heterozygosity estimate
- Assembly size and contig N50
- Scaffold N50 after scaffolding
- BUSCO completeness and duplication
- Merqury completeness, QV, and duplication
- Per-haplotype metrics for phased assemblies

The EMBL-EBI training resources provide guidance on data resources and analysis education [<a href="#ref-8">8</a>]. These resources help researchers understand which metrics are most informative for different assembly types.

### Version Control and Reproducibility

Use version control for assembly scripts and parameters. Record software versions and container images. The nf-core documentation describes standards for reproducible workflow configuration [<a href="#ref-6">6</a>]. The Carpentries lessons provide foundational training in version control with Git [<a href="#ref-9">9</a>].

Reproducibility requirements:

- Record all software versions
- Use containerized workflows where possible
- Document parameter choices and rationale
- Archive intermediate files and final assemblies
- Store raw sequencing data in public repositories

The NCBI provides data resources for sequence storage and analysis [<a href="#ref-10">10</a>]. Depositing raw data and assemblies in public databases supports reproducibility and enables community validation.

## Limitations and Interpretation Boundaries

### What K-Mer Metrics Cannot Tell You

K-mer based metrics have inherent limitations. They cannot:

- Distinguish between different types of errors
- Identify the location of errors
- Detect errors in repetitive regions
- Validate structural variant calls
- Assess biological correctness of phasing

These limitations mean that k-mer metrics must be complemented with alignment-based methods, manual inspection, and biological validation. The Genome 10K consortium emphasized that high-quality assemblies require multiple lines of evidence [<a href="#ref-1">1</a>].

### Heterozygosity Estimation Uncertainty

K-mer based heterozygosity estimates have uncertainty, particularly for:

- Genomes with high repeat content
- Species with variable ploidy
- Samples with contamination
- Sequencing with biased error profiles

The Isodon rubescens genome, with 83.43% repeat content, required careful interpretation of k-mer spectra [<a href="#ref-5">5</a>]. The blue bat star genome, with 5.16% heterozygosity, showed that high heterozygosity can be estimated reliably but requires sufficient sequencing depth [<a href="#ref-2">2</a>].

### Ploidy Complications

K-mer analysis assumes diploidy. Polyploid genomes produce more complex spectra with multiple peaks. The octoploid strawberry genome required specialized analysis to interpret the k-mer spectrum [<a href="#ref-3">3</a>]. Researchers working with polyploids should use tools designed for polyploid analysis or consult with specialists.

### Reference Bias in Validation

When validating assemblies against a reference genome, reference bias can obscure true differences. A haplotype that is more similar to the reference will appear better assembled than a divergent haplotype. This bias affects:

- BUSCO completeness scores
- Read mapping rates
- Variant calling accuracy
- Structural variant detection

For haplotype-resolved assemblies, validate each haplotype independently and avoid comparing haplotypes to a single reference.

## Professional Escalation Criteria

### When to Seek Specialist Consultation

Consult with genome assembly specialists when:

- K-mer spectra show unexpected patterns that cannot be interpreted
- Merqury metrics indicate problems that persist after multiple assembly attempts
- Haplotype phasing produces unbalanced or chimeric results
- Polyploid genomes require specialized analysis
- Structural variant validation requires cytogenetic confirmation

The complexity of haplotype-resolved assembly means that some problems require specialized expertise. The Vertebrate Genomes Project brought together international experts to address assembly challenges across vertebrate lineages [<a href="#ref-1">1</a>].

### When to Consider Reassembly

Reassembly should be considered when:

- Completeness is below 90% and cannot be improved by polishing
- Haplotype duplication affects more than 5% of the genome
- Chimeric haplotypes are detected in multiple contigs
- Structural variant calls conflict with independent evidence
- Gene annotation reveals systematic errors in gene models

Reassembly is expensive and time-consuming. Before deciding to reassemble, consider whether additional sequencing, different assembly parameters, or targeted polishing could resolve the problems.

### When to Change Sequencing Strategy

Changes to sequencing strategy may be necessary when:

- Initial sequencing depth is insufficient for phasing
- Read length is too short to span heterozygous regions
- Hi-C libraries have low interaction signal
- Parental samples are unavailable for trio binning

The Isodon rubescens project used both PacBio and Nanopore platforms to achieve 328x coverage [<a href="#ref-5">5</a>]. The strawberry project used single molecule real-time sequencing and Hi-C [<a href="#ref-3">3</a>]. These examples show that multiple sequencing strategies may be needed for challenging genomes.

## Safety and Ethical Context

### Data Management and Privacy

Genome assembly projects involve sensitive data, particularly for endangered species, agricultural cultivars, or human-associated organisms. Follow institutional data management policies and applicable regulations. The NCBI provides data resources with controlled access options for sensitive data [<a href="#ref-10">10</a>].

### Sample Provenance and Permits

For wild-collected samples, ensure that collection permits and export-import regulations are followed. The blue bat star genome project involved a species from Korean waters, requiring appropriate collection authorization [<a href="#ref-2">2</a>]. The Lactarius hatsudake project involved a mycorrhizal mushroom from pine forests, requiring collection permissions [<a href="#ref-4">4</a>].

### Ethical Use of Genomic Data

Genomic data from agricultural species may have commercial value. The strawberry genome project involved a cultivated cultivar, "Yanli", with potential breeding applications [<a href="#ref-3">3</a>]. Researchers should clarify data sharing and intellectual property arrangements before starting assembly projects.

### Biosafety Considerations

For pathogenic or toxin-producing organisms, follow biosafety regulations for sample handling and data management. The Isodon rubescens project involved a medicinal plant with bioactive diterpenoids [<a href="#ref-5">5</a>]. While the plant itself is not hazardous, researchers should be aware of any regulatory requirements for medicinal plant research.

## A Practical Decision Framework for Haplotype Assembly Strategy

### Establishing a Structured Triage System

Researchers often struggle to translate heterozygosity estimates into concrete assembly decisions. A structured triage system based on measurable thresholds provides clarity. The framework below uses k-mer based heterozygosity as the primary input, with secondary considerations for repeat content and biological questions. This system mirrors the decision process used in published genome projects, where heterozygosity directly influenced assembly strategy.

The Lactarius hatsudake genome at 1.14% heterozygosity was assembled without phasing and served its purpose for functional genomics [<a href="#ref-4">4</a>]. The blue bat star at 5.16% heterozygosity required specialized approaches [<a href="#ref-2">2</a>]. The Isodon rubescens f. lushanensis genome at 1.7% heterozygosity with 83.43% repeat content needed a different strategy than a low-repeat genome at the same heterozygosity [<a href="#ref-5">5</a>]. These examples show that heterozygosity alone does not determine the approach.

### Triage Level 1: Low Heterozygosity Below 0.5%

For genomes with heterozygosity below 0.5%, a standard collapsed assembly is usually sufficient. The primary risk is unnecessary expenditure on phasing that the biological questions do not require. Proceed with standard assembly but verify that the k-mer spectrum shows no hidden heterozygous peak that the initial estimate missed.

Record the following at this level:

- K-mer based genome size estimate
- Heterozygosity percentage with k-mer size used
- Assembly size compared to k-mer estimate
- BUSCO completeness and duplication rates
- Merqury completeness and QV scores

Escalate to Level 2 if the assembly size differs from the k-mer estimate by more than 10% or if BUSCO duplication exceeds 5%. These signs indicate that the heterozygosity estimate may have been too low or that repeat content is confounding the analysis.

### Triage Level 2: Moderate Heterozygosity 0.5% to 2%

Genomes in this range require a deliberate decision about phasing. The Lactarius hatsudake project at 1.14% heterozygosity chose a collapsed assembly because the biological questions centered on functional genomics and breeding applications [<a href="#ref-4">4</a>]. The Isodon rubescens project at 1.7% heterozygosity produced a single reference genome without haplotype separation, focusing instead on gene content and biosynthetic pathways [<a href="#ref-5">5</a>].

The decision at this level depends on the intended analyses:

- Population genetics requiring allele frequencies demands phasing
- Gene family analysis requires distinguishing paralogs from alleles
- Structural variant detection between haplotypes requires phasing
- Functional genomics of a single reference may proceed without phasing

If phasing is chosen, trio binning works when parental samples are available. Hi-C phasing works when parents cannot be sampled. The strawberry project combined single molecule real-time sequencing with Hi-C to achieve chromosome-scale phasing [<a href="#ref-3">3</a>].

### Triage Level 3: High Heterozygosity Above 2%

Heterozygosity above 2% makes collapsed assembly unreliable. The blue bat star at 5.16% heterozygosity required specialized assembly approaches [<a href="#ref-2">2</a>]. At this level, haplotype phasing is not optional for most biological questions because collapsed assemblies produce chimeric sequences that corrupt downstream analysis.

The sequencing strategy must change at this level:

- Increase long-read coverage beyond the standard 30x to 60x range
- Consider multiple long-read platforms for complementary error profiles
- Use ultra-long reads to span heterozygous regions
- Plan for Hi-C or trio binning from the start

The Isodon rubescens project used both PacBio and Nanopore platforms to achieve 328x coverage for a 349 Mb genome [<a href="#ref-5">5</a>]. This depth supported both assembly and phasing in a genome with high repeat content.

### Triage Level 4: High Repeat Content Above 60%

Repeat content complicates heterozygosity interpretation because repetitive k-mers appear at multiple frequencies. The Isodon rubescens genome at 83.43% repeat content required careful k-mer spectrum interpretation to distinguish heterozygous peaks from repeat peaks [<a href="#ref-5">5</a>].

For high-repeat genomes:

- Increase sequencing depth to resolve repeat arrays
- Use ultra-long reads to span repetitive regions
- Validate assembly with independent methods beyond k-mer analysis
- Expect lower BUSCO completeness due to unresolved repeats

The Lactarius hatsudake genome achieved 89.0% BUSCO completeness [<a href="#ref-4">4</a>]. For high-repeat genomes, completeness below this level may reflect repeat content instead of assembly failure. Document the repeat content estimate alongside completeness metrics.

### Triage Level 5: Polyploid Genomes

Polyploid genomes require specialized analysis because k-mer spectra show multiple peaks corresponding to different ploidy levels. The octoploid strawberry genome required dedicated phasing to produce eight haplotypes across 56 chromosomes [<a href="#ref-3">3</a>]. Standard diploid assembly tools will fail on polyploid genomes.

For polyploid genomes:

- Use polyploid-aware assemblers designed for multiple haplotypes
- Plan for Hi-C to assign sequences to chromosomes
- Expect higher sequencing depth requirements
- Validate chromosome number against karyotype data

The strawberry project produced Hap1 at 825 Mb and Hap2 at 808 Mb with contig N50 values above 26 Mb [<a href="#ref-3">3</a>]. These metrics demonstrate that polyploid phasing can achieve contiguity comparable to diploid assemblies when the strategy is appropriate.

### Implementing the Decision Framework

The triage system translates into a practical workflow with defined decision points. At each point, record the evidence that supports the decision. This documentation enables troubleshooting if the assembly fails quality checks.

Step 1: Run k-mer analysis on initial sequencing data. Record genome size, heterozygosity, and repeat content estimates.

Step 2: Apply the triage level based on heterozygosity and repeat content. Document the rationale for the chosen strategy.

Step 3: Execute the assembly strategy appropriate for the triage level. Record all parameters and software versions.

Step 4: Run Merqury and BUSCO quality assessment. Compare results to the thresholds in the triage table.

Step 5: If quality metrics fall below thresholds, escalate to the next triage level. Document the failure pattern and the revised strategy.

Step 6: For phased assemblies, evaluate each haplotype separately. Compare size, contiguity, and completeness between haplotypes.

### Records and Measurements for Triage Decisions

Maintain a structured record for each assembly project that captures the triage decision and its outcomes. The following fields support reproducibility and troubleshooting:

- Species and sample identifier
- Sequencing platforms and chemistry versions
- Total sequencing output and coverage calculation
- K-mer size used for heterozygosity estimation
- Heterozygosity estimate with confidence interval if available
- Repeat content estimate
- Triage level assigned
- Assembly strategy chosen
- Assembly size and contig N50
- Merqury completeness, QV, and duplication metrics
- BUSCO completeness and duplication
- Per-haplotype metrics for phased assemblies
- Escalation events and their outcomes

The nf-core documentation emphasizes reproducible workflow standards for genomic analysis [<a href="#ref-6">6</a>]. Following structured records supports comparison across projects and enables meta-analyses of assembly quality.

### Common Failure Patterns in Triage Implementation

Failure to apply the triage system correctly produces recognizable patterns. The most common failure is underestimating heterozygosity because the k-mer analysis used insufficient sequencing depth. Low coverage produces a noisy spectrum where the heterozygous peak is not visible.

Another common failure is ignoring repeat content when interpreting heterozygosity. High repeat content can mask the heterozygous peak or produce false peaks that are misinterpreted. The Isodon rubescens genome at 83.43% repeat content required specialized interpretation [<a href="#ref-5">5</a>].

A third failure pattern is proceeding with a collapsed assembly at moderate heterozygosity without documenting the decision. If downstream analysis later requires allele-level information, the assembly cannot provide it. Document the biological rationale for the assembly strategy at the time of the decision.

### Professional Escalation Criteria

Escalate to specialist consultation when:

- K-mer spectra show patterns that cannot be interpreted with standard tools
- Heterozygosity estimates vary substantially across k-mer sizes
- Assembly quality metrics do not improve after multiple strategy adjustments
- Polyploid genomes require analysis beyond standard diploid tools
- Structural variant validation requires cytogenetic confirmation

The Vertebrate Genomes Project brought together international experts to address assembly challenges across vertebrate lineages [<a href="#ref-1">1</a>]. For challenging genomes, specialist consultation can prevent wasted effort on ineffective strategies.

### Validation of Triage Decisions

Validate triage decisions by comparing assembly outcomes across projects with similar characteristics. The strawberry genome at high heterozygosity required phasing [<a href="#ref-3">3</a>]. The Lactarius hatsudake genome at 1.14% heterozygosity did not [<a href="#ref-4">4</a>]. These examples provide reference points for decision making.

For each triage level, document the expected quality metrics and compare actual outcomes. If outcomes consistently fall below expectations, revise the thresholds for future projects. This iterative refinement improves the triage system over time.

### Integration with Existing Quality Assessment

The triage framework complements the k-mer and Merqury based quality assessment described in the main workflow. The triage system determines which assembly strategy to use. The quality metrics determine whether the strategy succeeded. Together, they provide a complete decision framework for haplotype-resolved genome assembly.

The Galaxy Training Network provides accessible workflow training for genomic analysis [<a href="#ref-7">7</a>]. The EMBL-EBI training resources offer guidance on data resources and analysis education [<a href="#ref-8">8</a>]. These resources support researchers in implementing structured assembly strategies.

## Frequently Asked Questions

### What is the difference between a collapsed assembly and a haplotype-resolved assembly?

A collapsed assembly merges the two parental chromosome copies into a single consensus sequence. This approach loses allele-specific information and can create false structural variants. A haplotype-resolved assembly represents each parental chromosome separately, preserving allele-specific sequence and enabling analysis of heterozygosity, structural variation, and allele-specific expression. The choice depends on the biological questions and the heterozygosity level of the species.

### How much heterozygosity requires haplotype phasing?

There is no universal threshold, but heterozygosity above approximately 1% generally warrants consideration of phasing. The Lactarius hatsudake genome at 1.14% heterozygosity was assembled without phasing [<a href="#ref-4">4</a>]. The blue bat star genome at 5.16% heterozygosity required specialized approaches [<a href="#ref-2">2</a>]. The decision should be based on the k-mer spectrum and the intended downstream analyses.

### What does Merqury measure and how should I interpret the results?

Merqury measures assembly completeness, accuracy, and haplotype duplication by comparing k-mers in the assembly to k-mers in the sequencing reads. Completeness indicates the fraction of read k-mers present in the assembly. QV estimates base accuracy. Duplication metrics indicate whether haplotypes were correctly separated or redundantly assembled. These metrics should be interpreted together with assembly size, contiguity, and BUSCO results.

### How do I know if my assembly has collapsed haplotypes?

Signs of collapsed haplotypes include assembly size smaller than the k-mer based genome size estimate, reduced heterozygosity in the assembly compared to reads, and Merqury spectra showing assembly k-mers at higher frequency than read k-mers. Comparing the assembly to parental reads, when available, can confirm collapse by showing mixed parental coverage in single contigs.

### What is the role of Hi-C in haplotype-resolved assembly?

Hi-C provides long-range chromatin interaction data that can phase contigs into haplotypes and order them into chromosomes. The octoploid strawberry assembly used Hi-C to produce 56 chromosomes across eight haplotypes [<a href="#ref-3">3</a>]. Hi-C is particularly useful when parental samples are unavailable for trio binning.

### How should I report heterozygosity in my genome paper?

Report the k-mer based heterozygosity estimate with the k-mer size and analysis method used. Include the k-mer spectrum plot as supplementary data. State whether the assembly is collapsed or haplotype-resolved and describe the phasing method. Report per-haplotype quality metrics for phased assemblies, as done for the strawberry genome [<a href="#ref-3">3</a>].

### What BUSCO completeness should I expect for a haplotype-resolved assembly?

BUSCO completeness for haplotype-resolved assemblies should be comparable to or better than collapsed assemblies. The Lactarius hatsudake genome achieved 89.0% BUSCO completeness [<a href="#ref-4">4</a>]. For haplotype-resolved assemblies, report BUSCO for each haplotype separately and for the combined assembly. Lower completeness in one haplotype may indicate assembly problems.

### When should I consider reassembly instead of polishing?

Polishing corrects base errors but cannot fix structural problems such as collapsed haplotypes, chimeric contigs, or missing regions. If Merqury or BUSCO indicate structural problems, reassembly should be considered. If only base accuracy is low, polishing may suffice. The decision should be based on the specific quality metrics and the intended use of the assembly.

## Related Bioinformatics Guides

- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Towards complete and error-free genome assemblies of all vertebrate species.](https://pubmed.ncbi.nlm.nih.gov/33911273). Nature, 2021.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [The first high-quality genome assembly and annotation of Patiria pectinifera.](https://pubmed.ncbi.nlm.nih.gov/37730712). Scientific data, 2023.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [High-quality haplotype-resolved genome assembly of cultivated octoploid strawberry.](https://pubmed.ncbi.nlm.nih.gov/37077373). Horticulture research, 2023.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [A high-quality genome assembly of Lactarius hatsudake strain JH5.](https://pubmed.ncbi.nlm.nih.gov/36171643). G3 (Bethesda, Md.), 2022.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [High-quality assembly of the T2T genome for Isodon rubescens f. lushanensis reveals genomic structure variations between 2 typical forms of Isodon rubescens.](https://pubmed.ncbi.nlm.nih.gov/39388604). GigaScience, 2024.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.