Long-Read Sequencing for De Novo Plant and Animal Genomes: Overcoming Size and Heterozygosity Challenges
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Long-read sequencing (PacBio HiFi, Oxford Nanopore Technologies) is essential for de novo assembly of plant and animal genomes with high repetitive content (>60%) or significant heterozygosity (>1.41%), as short reads fail to span repetitive elements and resolve parental haplotypes, leading to fragmented assemblies.
- Hybrid sequencing strategies, combining long reads for contiguity and short reads for base-level polishing, are a robust workflow for achieving high-accuracy assemblies, particularly for complex genomes where ultralong reads may be required to span extremely long repeat arrays.
- Coverage depth is critical, with recommendations ranging from 30-40x for HiFi reads in diploid genomes to 50-100x for ONT reads in highly repetitive genomes, and higher coverage (60-80x) is often necessary for polyploid species to accurately distinguish homeologous chromosomes.
- Assembly quality is assessed using BUSCO for gene completeness, mapping rates of sequencing reads to the assembly, and structural validation via methods like Hi-C chromatin conformation capture to confirm chromosome-level contiguity and telomere placement.
- Polyploidy and high heterozygosity necessitate specialized assembly parameters and potentially trio binning for haplotype separation, while ultralong reads are crucial for achieving telomere-to-telomere assemblies by resolving complex centromeric and telomeric repeat regions.
Researchers assembling genomes for plant and animal species with large, complex, or highly heterozygous genomes face obstacles that short-read platforms cannot resolve. This article provides a practical framework for designing long-read sequencing projects, selecting platforms and coverage depths, managing assembly quality, and interpreting results within the constraints of current bioinformatics tools. The guidance applies to laboratory professionals, graduate students, and established researchers who need concrete decision criteria for their assembly projects.
The Core Problem: Why Short Reads Fail on Complex Genomes
Short-read sequencing platforms generate fragments typically 150 to 300 base pairs in length. These reads are highly accurate, but their limited length creates fundamental difficulties when assembling genomes that contain long repetitive elements, segmental duplications, or high heterozygosity between parental haplotypes. When a genome contains repeated sequences longer than the read length, the assembler cannot determine whether two reads originate from the same copy of a repeat or from different copies. This ambiguity produces fragmented assemblies with gaps, misjoins, and collapsed regions.
The practical consequences for genome assembly quality are measurable. Assemblies built exclusively from short reads often contain thousands of contigs, with N50 values in the tens of kilobases for complex genomes. This fragmentation obscures gene order, disrupts structural variant detection, and prevents the resolution of telomere-to-telomere structure. For agricultural species where breeding decisions depend on accurate genomic information, these limitations translate directly into reduced power for trait mapping and marker development.
Long-read platforms address this problem by producing reads that span repetitive regions and extend across heterozygous sites. Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) instruments generate reads from tens of kilobases to megabases in length. When these reads cover an entire repeat unit or span the junction between unique and repetitive sequence, the assembler can resolve the correct genomic arrangement. The tradeoff is that long-read platforms historically produced lower per-base accuracy than short-read platforms, requiring careful polishing strategies to achieve base-level precision.
A hybrid sequencing strategy combines the strengths of both approaches. Long reads provide contiguity and structural resolution, while short reads contribute high per-base accuracy for polishing. This workflow has been applied successfully to viral genomes with complex internal repeat structures, where marker-based assessments of partial genes proved insufficient for resolving genome architecture. The combination of ONT long reads for assembly contiguity and Illumina short reads for single-nucleotide polishing generated complete de novo assemblies that supported detection of structural inversions consistent with genome isomerization.
Genome Complexity Factors That Drive Platform Choice
Genome Size and Repetitive Content
Genome size alone does not determine sequencing strategy, but it establishes the scale of the project. A bacterial genome of five megabases requires different coverage calculations than a plant genome of 500 megabases or a fish genome of 700 megabases. The total sequencing output needed scales linearly with genome size for a given coverage target, which directly affects cost and instrument time.
Repetitive content is a more important determinant of platform choice than size alone. Plant genomes frequently contain 60 to 80 percent repetitive DNA, including transposable elements, ribosomal RNA gene clusters, and satellite repeats. The enset genome, an Ethiopian orphan staple crop, displays 64.64 percent repetitive DNA content across its 540 megabase assembly. This level of repetition means that most short reads map to multiple locations, and assembly algorithms must rely on read depth and paired-end information to resolve copy numbers. Long reads that span entire repeat units or extend into flanking unique sequence resolve these regions unambiguously.
The distribution of repeats matters as well. Tandem repeats, where identical or similar sequences are arranged consecutively, create particular difficulties because the repeat unit length determines the read length needed to span it. Interspersed repeats, such as transposable elements, require reads that span the junction between the repeat and its insertion site. Both patterns are common in agricultural species, and both favor long-read approaches.
Heterozygosity and Haplotype Phasing
Heterozygosity describes the proportion of genomic positions where the two parental copies differ. Highly heterozygous genomes, common in outcrossing plants and many animal species, present a specific assembly challenge. When the two haplotypes diverge sufficiently, the assembler may attempt to assemble them separately, producing two copies of each region that do not correspond to the true genome structure. Alternatively, the assembler may collapse the haplotypes into a single consensus sequence that represents neither parental copy accurately.
The enset genome illustrates this challenge with 1.41 percent heterozygosity, a level that complicates assembly but remains tractable with appropriate strategies. Higher heterozygosity, particularly above two percent, requires careful attention to assembly parameters and may benefit from trio binning or other phasing approaches that separate haplotypes before assembly.
Long reads help manage heterozygosity because they can span heterozygous sites and link them into haplotype-specific contigs. When a single read covers multiple heterozygous positions, the phasing information is preserved across the entire read length. This property enables haplotype-resolved assembly, where both parental copies are reconstructed separately, which is valuable for breeding applications that need to distinguish alleles.
Polyploidy and Whole-Genome Duplication
Polyploid species carry more than two complete sets of chromosomes, multiplying the complexity of assembly. Whole-genome duplication creates paralogous regions that resemble segmental duplications but involve entire chromosomes or chromosome sets. The Australian burrowing frogs of the genus Neobatrachus include diploid and tetraploid species, with tetraploids showing selection on genes involved in chromosome segregation and crossover distribution. Assembling a tetraploid genome requires distinguishing homeologous chromosomes that share high sequence identity while maintaining the correct copy number for each locus.
Long-read platforms provide the read length needed to span the conserved regions between homeologous chromosomes and identify diagnostic variants that distinguish them. However, polyploid assembly remains computationally demanding, and current tools handle tetraploid and higher ploidy levels less reliably than diploid assembly. Researchers working on polyploid species should expect to invest substantial time in parameter optimization and manual curation.
Long-Read Platform Options and Their Tradeoffs
PacBio HiFi Sequencing
PacBio HiFi reads are produced by circular consensus sequencing, where the same molecule is read multiple times to generate a highly accurate consensus sequence. HiFi reads typically range from 10 to 25 kilobases with accuracy above 99.9 percent. This combination of length and accuracy makes HiFi reads suitable for both assembly and variant detection without separate polishing steps.
The mandarin fish genome assembly used PacBio HiFi reads as a primary data source, integrating them with ONT ultralong reads and Hi-C chromatin conformation capture to produce a near-complete telomere-to-telomere genome. The assembly spanned 24 chromosomes with telomeric repeats detected at both ends of 20 chromosomes and at one end of the remaining four. BUSCO evaluation against the Actinopterygii database showed 98.7 percent genome completeness, and alignment analyses demonstrated mapping rates above 97 percent for all three data types.
HiFi reads excel at resolving heterozygous regions because the high accuracy allows confident identification of variants within individual reads. The main limitation is read length, which may be insufficient to span the largest repeat arrays or structural variants. For genomes with very long repeats, HiFi reads alone may leave gaps that require ultralong reads to bridge.
Oxford Nanopore Technologies Sequencing
ONT sequencing produces reads that can exceed one megabase in length, with typical read lengths in the 20 to 100 kilobase range depending on library preparation and basecalling settings. The ultralong reads generated by ONT are particularly valuable for resolving complex repeat structures and providing long-range contiguity. The mandarin fish project used ONT ultralong reads specifically to complement HiFi data and close gaps in the assembly.
The tradeoff for ONT reads is lower per-base accuracy compared to HiFi, although recent basecalling improvements have narrowed this gap. Current ONT platforms can achieve median read accuracy above 99 percent with appropriate basecalling models, but systematic errors in homopolymer regions and methylation-sensitive motifs remain a concern. These errors are typically corrected during polishing, either with the same ONT data using appropriate tools or with short-read data in a hybrid approach.
ONT sequencing offers flexibility in throughput and cost, with flow cells that can be scaled to project needs. The technology also supports direct RNA sequencing and methylation detection, which may be valuable for projects that extend beyond genome assembly into functional genomics.
Hybrid Approaches
Hybrid sequencing combines long reads for assembly contiguity with short reads for polishing accuracy. This strategy was used effectively for infectious laryngotracheitis virus vaccine strains, where ONT long reads improved assembly contiguity and Illumina short reads provided single-nucleotide-level polishing. The resulting assemblies showed high sequence concordance with targeted regions validated by Sanger sequencing, and the workflow supported detection of a structural inversion in the unique short region of one strain.
For large plant and animal genomes, hybrid approaches remain common because they balance cost and quality. The long-read component provides the scaffold that determines assembly structure, while the short-read component corrects residual base errors. This division of labor is particularly valuable when the long-read platform produces lower accuracy reads, as with standard ONT protocols, or when the genome contains homopolymer regions that challenge specific sequencing chemistries.
The choice between a pure HiFi approach and a hybrid approach depends on the specific genome and project goals. HiFi-only assemblies can achieve high quality for many genomes, but the addition of ultralong reads or short-read polishing may be necessary for genomes with extreme repeat content or for applications requiring the highest possible base accuracy.
Coverage Depth: Calculating What You Need
Minimum Coverage for Assembly
Coverage depth is the average number of times each genomic position is represented in the sequencing data. For de novo assembly, coverage must be sufficient for the assembler to distinguish true sequence from sequencing errors and to resolve repetitive regions. The required coverage depends on the platform, the genome complexity, and the assembly strategy.
For HiFi reads, coverage of 30 to 40 times the genome size is commonly recommended for diploid genomes. This depth provides enough overlapping reads to build accurate consensus while allowing for the detection of heterozygous sites. For ONT reads used in assembly, coverage of 50 to 100 times is often recommended, with the higher end of this range needed for genomes with high repeat content or when using lower-accuracy basecalling.
The coverage calculation must account for the genome size, not the sequencing output. A project targeting 40 times coverage of a 500 megabase genome requires 20 gigabases of HiFi sequence. The same coverage of a 5 gigabase genome, common in some plant families, requires 200 gigabases. This scaling drives both cost and instrument time, and it should be calculated explicitly before project initiation.
Coverage for Polishing and Validation
Polishing coverage is distinct from assembly coverage. When using short reads to polish a long-read assembly, coverage of 50 to 100 times is typically recommended to ensure that every position is covered by multiple high-quality reads. Lower coverage may leave systematic errors uncorrected, particularly in regions with extreme GC content or other sequencing biases.
Validation coverage, used to assess assembly quality, can be lower than polishing coverage. Mapping a subset of reads back to the assembly and calculating mapping rates provides a quality check without requiring full coverage. The mandarin fish project demonstrated mapping rates above 97 percent for ONT ultralong reads, PacBio HiFi reads, and Hi-C data, providing confidence that the assembly represents the sequenced individual.
Adjusting Coverage for Genome Complexity
Genomes with high heterozygosity or polyploidy may require higher coverage than simple diploid genomes. The extra coverage helps the assembler distinguish true heterozygous sites from sequencing errors and provides the read depth needed to separate haplotypes. For tetraploid genomes, coverage of 60 to 80 times HiFi data may be necessary to achieve comparable assembly quality to a diploid genome at 40 times coverage.
Repeat content also influences coverage requirements. Highly repetitive genomes may have regions that are underrepresented in sequencing libraries due to biases in library preparation or amplification. Higher total coverage helps ensure that even underrepresented regions receive sufficient reads for assembly. However, coverage alone cannot overcome the fundamental limitation that repeats longer than the read length remain unresolvable without ultralong reads.
At a Glance: Platform Selection and Coverage Decisions
| Genome Feature | Recommended Platform Strategy | Coverage Guidance | Primary Quality Concern |
|---|---|---|---|
| Diploid, moderate repeat content (under 50 percent) | PacBio HiFi alone or HiFi plus short-read polishing | 30 to 40 times HiFi | Base accuracy in GC-rich regions |
| Diploid, high repeat content (over 60 percent) | HiFi plus ONT ultralong reads, with optional Hi-C scaffolding | 40 to 60 times HiFi, 50 to 100 times ONT | Repeat resolution and gap closure |
| Highly heterozygous or polyploid | HiFi plus ONT ultralong reads, trio binning if parents available | 60 to 80 times HiFi for tetraploids | Haplotype separation and copy number accuracy |
Assembly Workflow: From Raw Reads to Finished Genome
Read Quality Assessment and Filtering
The assembly workflow begins with quality assessment of the raw sequencing data. For ONT reads, this includes evaluating read length distribution, per-read accuracy estimates from the basecaller, and the presence of adapter contamination. For HiFi reads, quality assessment focuses on read length distribution and the accuracy scores assigned during circular consensus calling.
Read filtering removes low-quality reads, adapter sequences, and reads that fail internal quality thresholds. The specific thresholds depend on the platform and the downstream assembler. Some assemblers accept raw reads and perform their own filtering, while others expect pre-filtered input. The choice of filtering strategy should be documented and reported to ensure reproducibility.
Assembly Algorithms and Parameter Selection
Long-read assemblers use different algorithms for constructing the assembly graph and resolving repeats. Some assemblers are optimized for HiFi reads, others for ONT reads, and some handle both. The choice of assembler and its parameters has a substantial effect on assembly quality, and parameter optimization is often necessary for complex genomes.
Key parameters include the minimum read length for assembly, the overlap thresholds for read alignment, and the strategies for resolving the assembly graph. For heterozygous genomes, parameters that control haplotype separation are critical. For polyploid genomes, parameters that handle more than two haplotypes may be required, though support for higher ploidy remains limited in many tools.
The Galaxy Training Network provides accessible workflow training for genome assembly and related analyses, offering tutorials that walk through the practical steps of assembly and quality assessment. These resources are valuable for researchers who are new to long-read assembly or who need to refresh their skills with specific tools.
Scaffolding and Chromosome-Level Assembly
Contig-level assemblies represent the first output of the assembly process, but chromosome-level assemblies require additional data. Hi-C chromatin conformation capture provides long-range information about the physical proximity of genomic regions within the nucleus, enabling the ordering and orientation of contigs into chromosome-scale scaffolds.
The mandarin fish genome assembly integrated Hi-C data to achieve chromosome-level resolution, with the final assembly spanning 24 chromosomes. This approach is now standard for agricultural species where chromosome-level assemblies are needed for breeding applications. The Hi-C data also provides a quality check, as the expected interaction patterns between chromosome arms can be compared to the observed patterns in the assembled genome.
Telomere-to-telomere assembly represents the highest standard of genome completeness, with all chromosomes assembled from one telomere to the other without gaps. This standard requires ultralong reads to span the repetitive telomeric and centromeric regions that resist assembly from shorter reads. The mandarin fish assembly achieved telomeric repeats at both ends of 20 chromosomes and at one end of the remaining four, approaching but not fully achieving the telomere-to-telomere standard.
Polishing Strategies
Polishing corrects residual errors in the assembly after the initial assembly step. The polishing strategy depends on the data available and the accuracy requirements of the project. For HiFi-only assemblies, polishing may be performed with the same HiFi reads using tools designed for this purpose. For hybrid assemblies, short reads provide the polishing data.
The infectious laryngotracheitis virus study demonstrated the value of hybrid polishing, using Illumina short reads to achieve single-nucleotide-level accuracy after ONT-based assembly. This approach corrected the systematic errors characteristic of ONT data while preserving the contiguity provided by the long reads.
Polishing should be evaluated by comparing the polished assembly to independent data, such as Sanger sequencing of targeted regions or mapping of reads that were not used in the assembly. The concordance between the assembly and independent validation data provides confidence in the final sequence.
Quality Assessment and Validation
BUSCO Completeness Assessment
BUSCO (Benchmarking Universal Single-Copy Orthologs) provides a standardized assessment of assembly completeness by searching for a set of genes expected to be present in single copy in the target lineage. The mandarin fish assembly achieved 98.7 percent completeness against the Actinopterygii database, indicating that nearly all expected genes were present in the assembly.
BUSCO results are reported as the percentage of complete single-copy genes, complete duplicated genes, fragmented genes, and missing genes. High duplication rates may indicate haplotype separation issues, where both parental copies are assembled separately. High fragmentation or missing rates may indicate assembly gaps or collapsed regions.
The choice of BUSCO database is important. The database must match the lineage of the target species, and different databases may produce different completeness scores. The database version should be reported alongside the BUSCO results to ensure comparability across studies.
Mapping Rate Assessment
Mapping rates measure the proportion of sequencing reads that align to the assembled genome. High mapping rates indicate that the assembly represents the sequenced individual well, while low mapping rates may indicate assembly errors, contamination, or the presence of sequences that were not assembled.
The mandarin fish project demonstrated mapping rates above 97 percent for ONT ultralong reads, PacBio HiFi reads, and Hi-C data. This consistency across data types provides strong evidence that the assembly is complete and accurate. Mapping rates below 90 percent warrant investigation, as they may indicate systematic assembly problems.
Structural Validation
Structural validation confirms that the assembly has the expected large-scale features, such as chromosome number, chromosome size distribution, and the presence of telomeric and centromeric sequences. For species with known karyotypes, the assembly should produce the expected number of chromosomes with sizes consistent with cytogenetic observations.
The mandarin fish assembly detected telomeric repeats at the expected chromosome ends, providing structural validation of the assembly. The presence of telomeric repeats at both ends of most chromosomes indicates that those chromosomes were assembled to completion. Chromosomes with telomeric repeats at only one end may have gaps at the other end that require additional sequencing or assembly effort.
Comparison to Independent Data
Independent validation data, such as Sanger sequencing of targeted regions or genetic maps, provides the strongest evidence of assembly accuracy. The infectious laryngotracheitis virus study validated targeted regions by Sanger sequencing, confirming the sequence concordance of the assembly.
For agricultural species, genetic maps or linkage maps can validate the ordering and orientation of scaffolds. If the assembly places markers in an order that conflicts with the genetic map, the assembly likely contains misjoins that require correction.
Case Studies in Complex Genome Assembly
The Mandarin Fish Genome
The mandarin fish (Siniperca scherzeri) is a commercially significant aquaculture species in China, valued for its flesh quality, disease resistance, and domestication adaptability. The genome assembly project integrated PacBio HiFi long-read sequencing, ONT ultralong-read sequencing, and Hi-C chromatin conformation capture to produce a near-complete telomere-to-telomere genome.
The assembly spanned 24 chromosomes with telomeric repeats detected at both ends of 20 chromosomes and at only one end of the remaining four. BUSCO evaluation against the Actinopterygii database revealed 98.7 percent genome completeness. Alignment analyses using minimap2 demonstrated mapping rates above 97 percent for ONT ultralong reads, PacBio HiFi reads, and Hi-C data against the assembled genome.
The project annotated 23,296 protein-coding genes, establishing a genomic resource for evolutionary biology and molecular breeding strategies. The integration of multiple long-read platforms was essential for achieving this assembly quality, as HiFi reads provided accurate sequence while ultralong reads resolved the largest repetitive regions.
The Enset Genome
Enset (Ensete ventricosum), known as the tree against hunger, plays a key role in Ethiopian food security and farming systems, feeding more than 20 million people. The genome of enset landrace Mazia was sequenced to support breeding programs for this drought-tolerant crop, which is partially resistant to Xanthomonas wilt.
The Mazia assembly was 540.14 megabases, more complete than the previously published genome assembly of landrace Bedadeti at 451.28 megabases. The assembly displayed 1.41 percent heterozygosity and 64.64 percent repetitive DNA content. Comparative analyses with the Bedadeti assembly and chromosome-level genome sequences of the two main banana progenitors revealed that approximately 25 percent of the Mazia genome is unique to enset.
Gene Ontology and sequence similarity searches identified enset-specific protein-coding genes with functions related to DNA integration, carbohydrate metabolism, disease resistance, and transcriptional regulation. These findings support the distinct lifestyle, adaptation, and corm productive quality of enset. The improved genomic resources are intended to supplement conventional clonal selection-based breeding programs.
The Neobatrachus Frog Polyploidy Study
The Australian burrowing frogs of the genus Neobatrachus include three polyploid species: N. aquilonius, N. kunapalari, and N. sudellae. Researchers assembled a reference genome for the diploid N. pictus and sequenced 96 individuals from all nine Neobatrachus species to investigate how animals adapt to polyploidy.
The study found tetraploid-specific selection on genes with meiotic roles in synaptonemal complex and crossover distribution, including SYCE2 and PRR19, and genes involved in chromosome and spindle size, including Condensin-2 and KifC1. These changes may represent genetic adaptation in vertebrate polyploids with tetrasomic and mixed inheritance that raises crossover interference and scales chromosome and spindle size to ensure successful chromosomal segregation.
The study also showed that adaptive alleles are shared between the tetraploids via interspecific introgression. This finding has implications for understanding how polyploid species maintain fertility and fitness, which is relevant for agricultural species where polyploidy is common.
The Infectious Laryngotracheitis Virus Vaccine Study
The infectious laryngotracheitis virus (ILTV) vaccine study demonstrated the value of hybrid whole-genome sequencing for organisms with complex genome structures. The ILTV genome contains long internal inverted repeats that can give rise to genomic isomers, complicating short-read assembly and accurate resolution of genome structure.
The study used a hybrid whole-genome sequencing strategy, combining ONT long reads to improve assembly contiguity with Illumina short reads for high-accuracy polishing at the single-nucleotide level. This approach generated complete de novo genome assemblies for the commercial Serva and Salsbury vaccine strains.
The assemblies showed high sequence concordance with targeted regions validated by Sanger sequencing. Whole-genome analysis enabled detection and independent validation of a structural inversion in the unique short region of the Salsbury strain, consistent with herpesvirus genome isomerization. The study also performed a pangenome-based analysis to define a conserved core-genome dataset that robustly resolved vaccine-associated lineages.
The study was cross-sectional, with two strains and a single lot per strain, so it cannot distinguish between strain-specific and lot-specific variation. This limitation highlights the importance of understanding the scope of conclusions that can be drawn from a given dataset.
Pangenomics and the Shift Beyond Single Reference Genomes
The Rationale for Pangenome Resources
A single reference genome represents one individual of a species, capturing only a fraction of the genetic diversity present in the population. Pangenomics addresses this limitation by representing genetic diversity across multiple individuals, providing access to structural variants, copy number variations, presence or absence variations, and non-reference regulatory or coding sequences.
In agricultural species, pangenome resources support the discovery of hidden variants that may contribute to domestication, adaptation, and breeding traits. The review of pangenomics for agricultural breeding describes three major technical routes: variant integration, reference-guided iterative graph construction, and reference-free graph construction. Each route has distinct tradeoffs in accuracy, scalability, coordinate consistency, reference bias, computational demand, annotation transfer, and suitability for downstream breeding questions.
Pangenome Construction Strategies
Variant integration approaches start with a linear reference genome and add variants identified from additional individuals. This approach is computationally efficient and maintains coordinate consistency with the reference, but it may miss variants in regions that are absent from the reference or that are too complex to align.
Reference-guided iterative graph construction builds a graph by iteratively adding new individuals to an existing graph. This approach captures more diversity than variant integration but requires careful handling of coordinate systems and may introduce reference bias.
Reference-free graph construction builds a graph directly from all input genomes without a reference. This approach minimizes reference bias but is computationally demanding and produces graphs that are difficult to coordinate with existing resources.
Applications and Limitations of Pangenomes
Pangenome resources support hidden variant discovery, QTL and GWAS interpretation, environmental adaptation analysis, and multi-omics-based candidate prioritization. However, the review highlights unresolved limitations, including graph complexity, pipeline-dependent structural variant calls, incomplete functional annotation, weak cross-study comparability, and the difficulty of distinguishing causal variants from linked or neutral variation.
Pangenome studies should be treated as connected but non-equivalent evidence. Resource-building studies establish representational breadth, method papers define technical feasibility, and trait-focused studies provide varying levels of biological support. Apparent inconsistencies among studies may reflect differences in methods, data, or biological context instead of true contradictions.
For researchers working on agricultural species, pangenome resources can complement long-read assemblies by providing population-level context for the variants identified in individual genomes. The combination of high-quality reference assemblies and pangenome resources supports more accurate variant discovery and trait mapping.
Data Management and Reproducibility
Documentation Standards
Reproducible genome assembly requires comprehensive documentation of every step in the workflow. This includes the version of each software tool, the parameters used, the input data and its quality metrics, and the intermediate and final outputs. Without this documentation, other researchers cannot reproduce the assembly or assess its quality.
The nf-core documentation provides standards for community pipelines, including usage, configuration, and reproducibility context. These pipelines offer pre-built workflows for common bioinformatics tasks, reducing the burden of workflow development and ensuring that best practices are followed.
Version Control and Workflow Management
Version control systems track changes to analysis scripts and configuration files, providing a record of how the analysis evolved. Workflow management systems automate the execution of analysis steps, ensuring that each step runs with the correct inputs and parameters.
The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible research practices. These skills are essential for researchers who need to manage complex analysis workflows and collaborate with others.
Data Storage and Archiving
Genome assembly projects generate large volumes of data, including raw sequencing reads, intermediate files, and final assemblies. Storage requirements can reach terabytes for large genomes, and data management plans should address storage capacity, backup, and long-term archiving.
Raw sequencing data should be archived in public repositories such as those maintained by NCBI, which provides official descriptions of databases, search systems, sequence resources, and analysis services. Public archiving ensures that the data remain available for reanalysis and validation by other researchers.
Common Failure Patterns and Troubleshooting
Fragmented Assemblies
Fragmented assemblies, characterized by a large number of contigs and low N50 values, typically result from insufficient coverage, inadequate read length, or assembly parameters that are too conservative. The first diagnostic step is to assess read length distribution and coverage depth. If reads are shorter than expected, library preparation or sequencing conditions may need adjustment.
If coverage is adequate but the assembly remains fragmented, the genome may contain repeats longer than the read length. In this case, ultralong reads or Hi-C data may be needed to bridge the gaps. Alternatively, assembly parameters that control repeat resolution may need adjustment.
Haplotype Collapse or Separation
Haplotype collapse occurs when the assembler merges the two parental haplotypes into a single consensus sequence, losing heterozygous information. Haplotype separation occurs when the assembler creates two copies of each region, inflating the assembly size and creating false duplications.
BUSCO results can indicate these problems. High duplication rates suggest haplotype separation, while high missing rates may indicate collapse. The choice of assembly parameters that control heterozygosity handling is critical, and some assemblers offer specific modes for heterozygous genomes.
Misjoins and Structural Errors
Misjoins occur when the assembler incorrectly connects sequences that are not adjacent in the genome. These errors are difficult to detect from the assembly alone but may be identified by comparing the assembly to genetic maps, Hi-C data, or independent assemblies.
Hi-C data provides a powerful tool for detecting misjoins because the expected interaction patterns between genomic regions can be compared to the observed patterns. Regions that show unexpected interaction patterns may contain misjoins that require correction.
Base Errors After Polishing
Residual base errors after polishing may result from systematic errors in the sequencing platform that are not corrected by the polishing reads. Homopolymer regions are a common source of errors in ONT data, and these regions may require targeted validation.
Sanger sequencing of targeted regions provides independent validation of base accuracy. The infectious laryngotracheitis virus study used this approach to confirm sequence concordance in targeted regions, providing confidence in the overall assembly quality.
Records and Measurements for Assembly Projects
Essential Records
Every genome assembly project should maintain records of the following items: the source organism and sample identifier, the DNA extraction method and quality metrics, the sequencing platform and chemistry version, the basecalling model and version, the read quality metrics including length distribution and accuracy estimates, the coverage depth calculation, the assembler and version, the assembly parameters, the polishing tools and parameters, the BUSCO database and results, the mapping rate results, and the final assembly statistics including N50, contig count, and total length.
These records support reproducibility and provide the basis for quality assessment. They also enable comparison with other assemblies and identification of systematic issues.
Quality Metrics to Track
The key quality metrics for genome assembly include N50 and L50 values, which describe the contiguity of the assembly, BUSCO completeness scores, which describe the representation of expected genes, mapping rates, which describe how well the assembly represents the sequenced reads, and assembly size compared to the expected genome size based on flow cytometry or other estimates.
These metrics should be tracked throughout the assembly process, from the initial contig assembly through polishing and scaffolding. Improvements in one metric may come at the cost of another, and the final assembly represents a balance of competing quality considerations.
Professional Escalation Criteria
Researchers should seek additional expertise or escalate to more experienced colleagues when the assembly quality metrics fall below acceptable thresholds, when the assembly process produces unexpected results that cannot be explained by known genome features, when the computational requirements exceed available resources, or when the project timeline is at risk due to assembly difficulties.
Specific escalation triggers include BUSCO completeness below 90 percent for a diploid genome, mapping rates below 90 percent, assembly size that differs from the expected genome size by more than 20 percent, or the presence of structural errors that cannot be resolved with available data.
Limitations and Interpretation Constraints
Cross-Sectional Study Limitations
Genome assemblies represent the sequenced individual at the time of sampling. They do not capture the full genetic diversity of a species, and they cannot distinguish between variation that is specific to the sequenced individual and variation that is shared across the population.
The infectious laryngotracheitis virus study illustrates this limitation. The study was cross-sectional, with two strains and a single lot per strain, so it cannot distinguish between strain-specific and lot-specific variation. Researchers should be careful not to overgeneralize from a single assembly to the entire species.
Pangenome Interpretation Constraints
Pangenome resources provide a broader view of genetic diversity than a single reference genome, but they have their own limitations. The review of pangenomics for agricultural breeding highlights the difficulty of distinguishing causal variants from linked or neutral variation, the weak cross-study comparability of structural variant calls, and the incomplete functional annotation of pangenome resources.
Researchers using pangenome resources should treat them as complementary to reference assemblies instead of replacements. The combination of a high-quality reference assembly and a pangenome resource provides the most complete picture of genetic diversity.
Computational Resource Constraints
Long-read assembly is computationally demanding, requiring substantial memory and processing time. The computational requirements scale with genome size and complexity, and projects on large or polyploid genomes may require access to high-performance computing resources.
The Bioconductor project provides official documentation for packages, workflows, installation, and reproducible genomic analysis, which can help researchers manage the computational aspects of their projects. The EMBL-EBI Training resources provide learning pathways for bioinformatics data resources and practical analysis education.
Safety and Ethical Considerations
Sample Collection and Biosecurity
Sample collection for genome sequencing should follow appropriate biosafety and biosecurity protocols. For agricultural species, this includes following quarantine regulations, preventing cross-contamination between samples, and documenting the provenance of all samples.
For pathogens, such as the infectious laryngotracheitis virus, additional precautions are necessary to prevent the release of infectious material. Researchers should follow institutional biosafety guidelines and national regulations for work with pathogenic organisms.
Data Sharing and Intellectual Property
Genome sequence data may have commercial value, particularly for agricultural species with breeding applications. Researchers should be aware of intellectual property considerations and data sharing agreements before depositing data in public repositories.
Public data sharing is generally encouraged for research purposes, but the terms of data sharing should be documented and agreed upon by all parties before the project begins. The NCBI provides official descriptions of databases and search systems that support data sharing and access.
Ethical Use of Genomic Information
Genome sequence data can reveal information about the sequenced individual, including traits that may be sensitive or commercially valuable. Researchers should consider the ethical implications of their work and ensure that data are used appropriately.
For agricultural species, genomic information may be used for breeding decisions that affect farmers and communities. The enset genome project, for example, has implications for food security in Ethiopia, and the research should be conducted in a way that benefits the communities that depend on the crop.
Frequently Asked Questions
What coverage depth is recommended for a diploid plant genome with high repeat content?
For a diploid plant genome with high repeat content, coverage of 40 to 60 times with PacBio HiFi reads is a reasonable starting point. The higher end of this range is appropriate when the genome has very high repeat content or when the assembly will be used for applications requiring high accuracy. Coverage should be calculated based on the genome size, and the actual coverage achieved should be verified after sequencing. If the genome has regions that are underrepresented in the sequencing library, higher total coverage may be needed to ensure adequate representation of all regions.
How do I decide between PacBio HiFi and Oxford Nanopore Technologies for my project?
The choice between PacBio HiFi and ONT depends on the specific requirements of your project. HiFi reads provide higher per-base accuracy and are well suited for variant detection and polishing. ONT reads can be much longer, which is valuable for resolving very long repeats and providing long-range contiguity. Many projects use both platforms, with HiFi reads providing accurate sequence and ONT ultralong reads providing the longest-range information. The mandarin fish genome project used this combined approach to achieve near-complete telomere-to-telomere assembly.
What is the role of Hi-C data in long-read genome assembly?
Hi-C data provides long-range information about the physical proximity of genomic regions within the nucleus. This information is used to order and orient contigs into chromosome-scale scaffolds, producing a chromosome-level assembly. Hi-C data also provides a quality check, as the expected interaction patterns between genomic regions can be compared to the observed patterns in the assembled genome. The mandarin fish project used Hi-C data to achieve chromosome-level assembly spanning 24 chromosomes.
How do I assess the completeness of my genome assembly?
BUSCO completeness assessment is the standard method for evaluating assembly completeness. BUSCO searches for a set of genes expected to be present in single copy in the target lineage and reports the percentage of complete single-copy genes, complete duplicated genes, fragmented genes, and missing genes. The mandarin fish assembly achieved 98.7 percent completeness against the Actinopterygii database. The choice of BUSCO database is important, and the database version should be reported alongside the results.
What should I do if my assembly has high BUSCO duplication rates?
High BUSCO duplication rates may indicate haplotype separation, where both parental copies of the genome are assembled separately. This can occur in heterozygous genomes when the assembler fails to merge the two haplotypes. Options include adjusting assembly parameters to handle heterozygosity differently, using trio binning to separate haplotypes before assembly, or accepting the haplotype-resolved assembly if it is appropriate for your research questions. The interpretation of duplication rates depends on the ploidy of the species and the goals of the project.
How does polyploidy affect genome assembly strategy?
Polyploidy multiplies the complexity of genome assembly because the assembler must distinguish homeologous chromosomes that share high sequence identity while maintaining the correct copy number for each locus. The Neobatrachus frog study provides an example of polyploid species with tetraploid-specific selection on genes involved in chromosome segregation. Polyploid assembly requires higher coverage and careful parameter optimization, and current tools handle tetraploid and higher ploidy levels less reliably than diploid assembly.
What is the difference between a reference genome and a pangenome?
A reference genome represents one individual of a species, while a pangenome represents genetic diversity across multiple individuals. Pangenomes provide access to structural variants, copy number variations, presence or absence variations, and non-reference regulatory or coding sequences that may contribute to domestication, adaptation, and breeding traits. The review of pangenomics for agricultural breeding describes three major technical routes for pangenome construction, each with distinct tradeoffs in accuracy, scalability, and suitability for downstream breeding questions.
How should I validate my assembly beyond BUSCO and mapping rates?
Beyond BUSCO and mapping rates, validation can include comparison to independent data such as Sanger sequencing of targeted regions, comparison to genetic maps or linkage maps, and assessment of structural features such as telomeric repeats and chromosome number. The infectious laryngotracheitis virus study used Sanger sequencing to validate targeted regions, and the mandarin fish project detected telomeric repeats at the expected chromosome ends. Independent validation data provides the strongest evidence of assembly accuracy.
Related Bioinformatics Guides
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Long-Read Sequencing Cost and Market: What to Expect
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Hybrid Whole-Genome Sequencing for Genetic Stability Assessment of Infectious Laryngotracheitis Virus Vaccine Strains.. 2026.
- Pangenomics for Agricultural Breeding: Construction Strategies, Evidence Integration, and Translational Constraints.. 2026.
- Genetic adaptation to polyploidy in animals: a case study in Australian burrowing frogs Neobatrachus. 2026.
- What makes a banana false? How the genome of Ethiopian orphan staple Ensete ventricosum differs from the banana A and B sub-genomes. 2026.
- Telomere-to-telomere gapless genome assembly of Siniperca scherzeri.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.