Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Section: Infrastructure, Cloud & Policy

Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices

Long-read sequencing has become the standard approach for assembling complex genomes that resist short-read methods. This article explains how long-read technologies enable high-quality de novo assembly of genomes with segmental duplications, high heterozygosity, and abundant repeats, using published case studies to extract practical best practices. Readers will learn how to choose assembly strategies, evaluate assembly quality, and avoid common failure modes when working with complex genomes.

At a Glance

The table below summarizes the key decisions and expected outcomes when assembling complex genomes with long-read sequencing.

Assembly Challenge Recommended Approach Expected Outcome
Segmental duplications and complex repeats Use uncollapsed assembly methods with haplotype-resolved output Complete haplotypes reconstructed across duplicated regions
High heterozygosity in plant or animal genomes Use HiFi reads with assemblers designed for haplotype separation Phased assemblies that represent both haplotypes accurately
Repetitive content exceeding 45% of the genome Combine ultra-long reads with Hi-C scaffolding Chromosome-scale contiguity with N50 values above 40 Mb
Evaluating assembly correctness Use reference-free evaluators such as Inspector Precise identification of structural errors and their locations

Why Complex Genomes Require Long Reads

Short-read sequencing cannot produce complete genome assemblies when the genome contains high duplication and multiple heterozygosity. Genome-wide short-read approaches struggle to resolve repetitive regions and closely related haplotypes, leaving gaps and collapsed sequences in the final assembly. High-quality chromosome-scale sequences provide an important basis for downstream analysis, including genome annotation, mutation detection, evolutionary analysis, gene function research, and comparative genomics. The emergence of long-read sequencing technology has greatly improved the integrity of complex genome assembly by allowing reads to span repeats and structural variants that break short-read assemblies.

Long reads produced by Oxford Nanopore and PacBio technologies can span structural variants and resolve complex repetitive regions such as centromeres, unlocking previously inaccessible genomic information. The ability to sequence DNA directly from samples without amplification also supports applications in metagenomics and direct RNA sequencing. Nanopore sequencing has advanced rapidly over the past decade, transitioning from an emerging technology to a significant instrument in genomic sequencing, with applications in telomere-to-telomere genome assembly, direct RNA sequencing, and metagenomics.

Core Principles of Long-Read Assembly

Read Length and Accuracy Tradeoffs

Long-read technologies present different tradeoffs between read length and accuracy. PacBio HiFi sequencing yields highly accurate long-read datasets with read lengths averaging 10 to 25 kb and accuracies greater than 99.5%. These accurate long reads improve results for complex applications such as single nucleotide and structural variant detection, genome assembly, assembly of difficult polyploid or highly repetitive genomes, and assembly of metagenomes. Oxford Nanopore sequencing produces ultra-long reads that can span entire repetitive regions, though raw nanopore reads have higher error rates than HiFi reads.

Third-generation sequencers are known to produce more error-prone reads than short-read platforms, generating new challenges for assembly algorithms and pipelines. The introduction of HiFi reads with substantially reduced error rates has provided a promising solution for more accurate assembly outcomes. Researchers must choose the correct assembler for their projects based on the read type and genome characteristics.

Assembly Strategies for Complex Genomes

Computational methods for complex genome assembly fall into three categories: collapsed, semi-collapsed, and uncollapsed assemblers. Collapsed assemblers merge similar sequences into a single representation, which loses haplotype information. Semi-collapsed assemblers partially separate haplotypes. Uncollapsed assembly is the most correct and complete way to represent genomes because it preserves distinct haplotypes. Genome assembly is closely related to haplotype reconstruction, where uncollapsed assembly realizes haplotype reconstruction and haplotype reconstruction promotes uncollapsed assembly.

The goal of gapless, telomere-to-telomere, and accurate assembly of complex genomes can be achieved routinely using long-read data, though current methods may require multiple tools and careful parameter selection.

Case Study: Haplotype-Resolved Human Genomes

The Human Pangenome Reference Consortium demonstrated the power of long-read assembly for complex human genomes. Long-read and strand-specific sequencing technologies together facilitated the de novo assembly of high-quality haplotype-resolved human genomes without parent-child trio data. The project produced 64 assembled haplotypes from 32 diverse human genomes. These highly contiguous haplotype assemblies achieved an average minimum contig length needed to cover 50% of the genome of 26 million base pairs. The assemblies integrated all forms of genetic variation, even across complex loci.

The study identified 107,590 structural variants, of which 68% were not discovered with short-read sequencing. The authors characterized 278 structural variant hotspots spanning megabases of gene-rich sequence and found that 63% of all structural variants arise through homology-mediated mechanisms. This resource enabled reliable graph-based genotyping from short reads of up to 50,340 structural variants, resulting in the identification of 1,526 expression quantitative trait loci.

This case study illustrates that long-read assembly can resolve complex loci that short-read methods miss, and that haplotype-resolved assemblies provide a foundation for understanding structural variation in the human genome.

Case Study: Pangenome Graph of Japanese Haplotypes

A recent study generated 20 near-complete haplotypes from 10 Japanese male individuals using three complementary long-read and long-range datasets and constructed a pangenome graph from these haplotype-resolved assemblies. All haplotypes achieved an N50 value exceeding 100 Mbp for gapless contigs. The study substantially improved the average reconstruction rate of complete haplotypes from 46.8% and 52.8% in two previous pangenome graphs to 91.2% within 30 segmentally duplicated complex regions.

The researchers identified complete minor haplotypes in the KIR and SMN regions that were absent in previous pangenome graphs. They found putatively biased gene conversion events occurring in only one direction around the SMN and beta-defensin genes, implying non-random evolution in these regions. This work demonstrates that repetitive regions are challenging to assemble yet medically important, and that pangenome graphs built from haplotype-resolved assemblies offer a more refined view of human genomic diversity involving complex segmental duplications.

Case Study: Chromosome-Scale Assembly of the Southern White Rhinoceros

The southern white rhinoceros genome provides an example of how long-read sequencing can transform a fragmented assembly into a chromosome-scale resource. The fragmented southern white rhinoceros genome assembly had limited chromosome-scale structural and evolutionary comparisons with the functionally extinct northern subspecies. Researchers produced a chromosome-scale genome assembly by integrating Oxford Nanopore Technology long-read sequencing, Illumina short-read polishing, and high-throughput chromosome conformation capture scaffolding.

The final assembly spans 2.48 Gb and achieves a contig N50 of 42.06 Mb, representing a 452-fold improvement in contiguity over the previous assembly. In total, 2.46 Gb of sequence was anchored to 40 autosomes plus the X and Y chromosomes. Genome annotation identified 1.13 Gb of repetitive elements, which is 45.7% of the assembly, 22,593 protein-coding genes, and 100.68 Mb of segmental duplications. Inspection of the major histocompatibility complex class II gene region supported the local assembly and annotation reliability, revealing conserved gene composition and order between the southern and northern white rhinoceroses.

Whole-genome comparison with the northern white rhinoceros assembly indicated extensive chromosome-scale synteny along with localized structural variants between the two subspecies, including 111 inversions spanning 33.48 Mb and 497 translocations spanning 36.48 Mb. This case study shows that integrating multiple data types can produce chromosome-scale assemblies for genomes with high repeat content.

Case Study: Tamarind Genome Assembly

The tamarind genome represents a plant genome assembly that integrated long-read data with transcriptomic and metabolomic analysis. Tamarindus indica is the sole member of the genus Tamarindus of the Leguminosae family and is a multipurpose horticultural plant. Researchers reported the first high-quality genome assembly of T. indica anchored to 12 chromosomes with an N50 of 56.6 Mb. Supported by comprehensive transcriptome data, they reported 48,867 protein-coding genes.

Through phylogenetic and evolutionary analysis, the study uncovered an independent whole-genome duplication event in T. indica and highlighted the expression divergence of segmentally duplicated genes and their role in the better adaptivity of the plant. The researchers observed a high expansion of the terpene cyclase mutase gene family and identified nine oxidosqualene cyclases and their putative functions. By employing integrated genomic, transcriptomic, and metabolomic analysis, they identified the putative L-Idonate dehydrogenase gene and provided evidence about its possible role in the accumulation of high tartaric acid content.

This case study demonstrates that long-read assembly provides an important resource for future genetic and biotechnological studies to understand essential pathways and assist breeding programs for trait enhancements.

Case Study: Immune Gene Regions in Cynomolgus Macaque

The major histocompatibility complex and killer cell immunoglobulin-like receptors are key regulators of immune responses. The cynomolgus macaque, an Old World monkey species, serves as an important preclinical model for studying human diseases. Several MHC-KIR combinations have been associated with either a poor or good prognosis, so macaques with a well-characterized immunogenetic profile may improve drug evaluation and speed up vaccine development.

A complete overview of the MHC and KIR haplotype organizations in cynomolgus macaques was lacking because characterization by conventional techniques is hampered by the extensive expansion of the macaque MHC-B region that complicates the discrimination between genes and alleles. Researchers assembled complete MHC and KIR genomic regions of cynomolgus macaque using third-generation long-read sequencing. This was the first physical mapping of complete MHC and KIR gene regions in a Vietnamese cynomolgus macaque. They identified four functional Mafa-B loci and showed that alleles of the Mafa-I01, B056, B034, and B001 functional lineages are highly frequent in the Vietnamese cynomolgus macaque population.

This case study shows that long-read assembly can resolve immune gene regions with extensive expansion and high sequence similarity, providing insights into haplotype organizations and diversity that refine the selection of animals with specific genetic markers for medical research.

Case Study: Endemic Betta Fishes

The genus Betta comprises over 70 species, many of which are endemic to Southeast Asia and highly vulnerable to habitat loss. Most wild Betta species remain poorly characterized at the genomic level. Researchers conducted an integrated meristic and genomic comparison of three endemic Bangka Island Betta species. Specimens were collected from peatland waters in Bangka, and high-molecular-weight DNA was extracted and sequenced using Oxford Nanopore PromethION technology, followed by de novo assembly and reference-guided scaffolding using the Betta splendens genome.

Genome assemblies were highly contiguous and complete with BUSCO scores above 97%. The species with the largest genome size also showed the highest scaffold N50 and elevated retrotransposon content. Gene duplication analysis revealed dispersed duplications as the dominant category across all genomes, with variation in tandem and proximal duplicates. Comparative genomic analysis demonstrated conserved collinearity, with two species showing the closest relationship while the third diverged earlier.

This case study demonstrates that long-read assembly can be applied to non-model organisms with limited prior genomic resources, producing assemblies suitable for comparative genomics and evolutionary analysis.

Case Study: Nippostrongylus brasiliensis

The de novo assembly of the complex genome of Nippostrongylus brasiliensis using MinION long reads provides an early example of nanopore assembly for a parasitic nematode with a complex genome. The study used Oxford Nanopore MinION long reads to assemble the genome of this gastrointestinal parasite of rodents, which serves as a model for human hookworm infection. The assembly demonstrated that nanopore sequencing could produce contiguous assemblies for genomes that were difficult to assemble with short reads alone.

Practical Workflow for Complex Genome Assembly

Step 1: Define Assembly Goals and Data Requirements

Before starting an assembly project, define the biological questions that the assembly must answer. If the goal is to identify structural variants or resolve segmental duplications, haplotype-resolved assembly is necessary. If the goal is to produce a reference genome for gene annotation, a collapsed assembly may suffice. The choice of sequencing platform and coverage depends on these goals.

Deep coverage HiFi datasets are available for complex samples including inbred model genomes, octoploid strawberry, and a diploid frog species. These datasets can be used without restriction to develop new algorithms and explore complex genome structure and evolution. Researchers can use such public datasets to benchmark assembly tools before committing to their own sequencing runs.

Step 2: Choose the Sequencing Platform

PacBio HiFi sequencing produces reads with accuracies greater than 99.5% and lengths averaging 10 to 25 kb. These accurate long reads are well suited for assembly of difficult polyploid or highly repetitive genomes. Oxford Nanopore sequencing offers ultra-long reads that can span entire repetitive regions, though raw read accuracy is lower. Some projects combine both platforms to leverage the strengths of each.

The choice between platforms also affects cost and turnaround time. Nanopore sequencing provides a faster, more cost-effective approach with extended read lengths, demonstrating significant potential for complex genome assembly. The openness and versatility of nanopore sequencing have established it as a preferred option for an increasing number of research teams.

Step 3: Select the Assembler

Assembly algorithm choice should be a deliberate, well-justified decision when researchers create genome assemblies for eukaryotic organisms from third-generation sequencing technologies. A benchmark of state-of-the-art long-read de novo assemblers used 12 real and 64 simulated datasets from different eukaryotic genomes with different read length distributions, imitating PacBio continuous long-read, PacBio HiFi, and Oxford Nanopore sequencing.

For Oxford Nanopore and PacBio continuous long-read data, the benchmark included Canu, Flye, Miniasm, Raven, and wtdbg2. For PacBio HiFi reads, the benchmark included HiCanu, Flye, Hifiasm, LJA, and MBG. Evaluation categories addressed reference-based metrics, assembly statistics, misassembly count, BUSCO completeness, runtime, and RAM usage. The study also investigated the effect of increased read length on assembly quality.

Canu, a successor of Celera Assembler, was specifically designed for noisy single-molecule sequences. Canu introduced support for nanopore sequencing, halved depth-of-coverage requirements, and improved assembly continuity while simultaneously reducing runtime by an order of magnitude on large genomes versus Celera Assembler 8.2. These advances resulted from new overlapping and assembly algorithms, including an adaptive overlapping strategy based on tf-idf weighted MinHash and a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes. Canu can reliably assemble complete microbial genomes and near-complete eukaryotic chromosomes using either PacBio or Oxford Nanopore technologies and achieves a contig NG50 of greater than 21 Mbp on both human and Drosophila melanogaster PacBio data sets.

For assembly structures that cannot be linearly represented, Canu provides graph-based assembly outputs in graphical fragment assembly format for analysis or integration with complementary phasing and scaffolding techniques. The combination of such highly resolved assembly graphs with long-range scaffolding information promises the complete and automated assembly of complex genomes.

Step 4: Scaffold to Chromosome Scale

Contig-level assemblies can be scaffolded to chromosome scale using high-throughput chromosome conformation capture data. The southern white rhinoceros assembly integrated Oxford Nanopore long-read sequencing, Illumina short-read polishing, and Hi-C scaffolding to anchor 2.46 Gb of sequence to 40 autosomes plus the X and Y chromosomes. The tamarind genome was anchored to 12 chromosomes with an N50 of 56.6 Mb.

Hi-C scaffolding uses the spatial proximity of DNA sequences in the nucleus to order and orient contigs along chromosomes. This approach is particularly valuable for genomes where genetic maps are unavailable or incomplete.

Step 5: Polish and Evaluate the Assembly

Assembly polishing uses raw reads to correct errors in the assembled sequence. Polishing can be performed with long reads, short reads, or reference genomes. Several methods for generating genome assemblies from error-prone long reads have been developed, complemented by various tools for assembly polishing. End users are left with a plethora of possible combinations of programs for obtaining a final trusted assembly, so there is a need to measure the completeness and accuracy of such assemblies.

Workflow management systems such as Snakemake can automate the running of multiple genome assembly and evaluation programs at once. Two workflows provide end users with an easy-to-run solution for testing various genome assemblies from their sequencing data. Both workflows use the conda packaging system, so there is no need for manual installation of each program. The workflows are available as open source software under the MIT license.

Evaluating Assembly Quality

Reference-Free Evaluation with Inspector

Long-read de novo genome assembly continues to advance rapidly, but there is a lack of effective tools to accurately evaluate assembly results, especially for structural errors. Inspector is a reference-free long-read de novo assembly evaluator that faithfully reports types of errors and their precise locations. Inspector can correct assembly errors based on consensus sequences derived from raw reads covering erroneous regions. Based on in silico and long-read assembly results from multiple long-read data and assemblers, Inspector can accurately identify both large-scale and small-scale assembly errors in addition to providing generic metrics.

Reference-free evaluation is essential when no closely related reference genome is available. Inspector uses the raw reads themselves as the ground truth for evaluating assembly correctness, making it applicable to any organism.

Metrics for Assembly Quality

Common metrics for evaluating assembly quality include contig N50, which is the minimum contig length needed to cover 50% of the genome, and BUSCO completeness, which measures the presence of conserved single-copy orthologs. The Vertebrate Genomes Project confirmed that long-read sequencing technologies are essential for maximizing genome quality, and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly.

The project generated assemblies for 16 species representing six major vertebrate lineages and found that their assemblies corrected substantial errors, added missing sequence in some of the best historical reference genomes, and revealed biological discoveries. These included the identification of many false gene duplications, increases in gene sizes, chromosome rearrangements that are specific to lineages, a repeated independent chromosome breakpoint in bat genomes, and a canonical GC-rich pattern in protein-coding genes and their regulatory regions.

Records and Measurements to Maintain

For each assembly project, maintain records of the following measurements:

Measurement Purpose Recording Frequency
Read length distribution and N50 Assess whether reads can span repeats and structural variants After each sequencing run
Depth of coverage Ensure sufficient coverage for assembly and polishing After each sequencing run
Contig N50 and scaffold N50 Track assembly contiguity After each assembly and scaffolding step
BUSCO completeness score Measure gene space completeness After each assembly iteration
Misassembly count Identify structural errors After each assembly iteration
Runtime and RAM usage Plan computational resources After each assembly run

Common Failure Patterns

Collapsed Repeats and Haplotypes

The most common failure in complex genome assembly is the collapse of diverged repeats and haplotypes into a single sequence. This occurs when the assembler cannot distinguish between similar sequences that are actually distinct loci or alleles. The Vertebrate Genomes Project found that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Canu was designed with a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes.

Insufficient Read Length

If reads are too short to span repetitive regions, the assembly will contain gaps at those loci. Ultra-long reads from Oxford Nanopore can span entire repetitive regions, but generating sufficient ultra-long reads requires careful DNA extraction and library preparation. The human genome assembly with ultra-long reads demonstrated that nanopore sequencing and assembly of a human genome with ultra-long reads is feasible.

Overpolishing

Polishing with short reads can introduce errors in repetitive regions where short reads map incorrectly. The southern white rhinoceros assembly used Illumina short-read polishing, but this approach requires careful validation to avoid introducing errors. Some assembly pipelines use long-read polishing exclusively to avoid this problem.

Computational Resource Limitations

Assembly of complex eukaryotic genomes requires substantial computational resources. Runtime and RAM usage vary widely among assemblers, and the choice of assembler may be constrained by available hardware. The benchmark of long-read assemblers evaluated runtime and RAM usage as part of the assessment, recognizing that these practical considerations affect assembler choice.

Limitations and Interpretation Constraints

Assembly Is Not Complete

Even with long-read sequencing, most genome assemblies are not complete. Telomere-to-telomere assemblies that cover every chromosome end to end remain challenging for complex genomes. The Vertebrate Genomes Project aims to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species, but this effort is ongoing.

Haplotype Representation

Collapsed assemblies represent a single consensus sequence that may not correspond exactly to any real haplotype. Haplotype-resolved assemblies are more accurate but require higher coverage and more complex assembly strategies. The choice between collapsed and haplotype-resolved assembly depends on the biological questions being addressed.

Annotation Dependence

Genome annotation depends on the quality of the assembly. The tamarind genome study used comprehensive transcriptome data to support gene annotation, identifying 48,867 protein-coding genes. Without transcriptome support, gene models in repetitive regions may be incomplete or incorrect.

Data Sharing and Reproducibility

Genomic data sharing is governed by policies that vary by funding agency and jurisdiction. The National Institutes of Health Genomic Data Sharing Policy outlines expectations for data sharing in NIH-funded research. Researchers should consult their funding agency policies and institutional requirements before depositing or sharing genomic data.

The FAIR Guiding Principles provide a framework for making data findable, accessible, interoperable, and reusable. Applying these principles to genome assemblies and associated data supports reproducibility and enables downstream analysis by other researchers.

Safety and Regulatory Context

Data Privacy and Consent

Human genome assemblies contain sensitive genetic information. Researchers working with human data must ensure that consent covers the intended uses of the data and that data sharing complies with applicable regulations. The Japanese pangenome study and the Human Pangenome Reference Consortium both addressed these considerations in their data release plans.

Controlled Data Access

Some genomic datasets are available through controlled access mechanisms that require approval from data access committees. The NCBI provides data resources for depositing and accessing genomic data, including assemblies and raw sequencing reads. Researchers should familiarize themselves with the data access policies for the repositories they use.

Ethical Use of Non-Human Data

For non-human genomes, ethical considerations focus on sample collection and species conservation. The Betta fish study collected specimens from peatland waters in Bangka, and the white rhinoceros study involved an endangered subspecies. Researchers should ensure that sample collection complies with relevant permits and conservation regulations.

Professional Escalation Criteria

When to Seek Additional Expertise

Consult a bioinformatics specialist or assembly expert when any of the following conditions apply:

Condition Action
Assembly N50 falls below expectations for the read length and coverage used Review assembler parameters and consider alternative assemblers
BUSCO completeness is below 90% for a eukaryotic genome Investigate whether reads cover the full gene space and consider additional sequencing
Structural errors are detected by Inspector or other evaluation tools Correct errors using Inspector or reassemble with different parameters
Computational resources are insufficient for the chosen assembler Select a more memory-efficient assembler or use cloud computing resources
The genome contains extreme repeat content or polyploidy Consult with researchers who have assembled similar genomes

When to Reassemble

Reassembly may be necessary when evaluation reveals widespread structural errors, when additional data types become available, or when the biological questions require higher resolution than the current assembly provides. The southern white rhinoceros assembly represented a 452-fold improvement in contiguity over the previous assembly, demonstrating that reassembly with better data and methods can transform a genome resource.

Frequently Asked Questions

What is the difference between collapsed and uncollapsed assembly?

Collapsed assembly merges similar sequences into a single representation, which loses haplotype information and can obscure structural variation. Uncollapsed assembly preserves distinct haplotypes and is the most correct and complete way to represent genomes. Uncollapsed assembly realizes haplotype reconstruction, and haplotype reconstruction promotes uncollapsed assembly.

How much coverage is needed for long-read assembly of a complex genome?

Coverage requirements vary by genome complexity and assembly strategy. Canu halves depth-of-coverage requirements compared to earlier assemblers, and HiFi assembly typically requires lower coverage than continuous long-read assembly because of higher per-read accuracy. The benchmark of long-read assemblers used datasets with different read length distributions and coverage levels to evaluate assembler performance.

What is the role of Hi-C data in genome assembly?

Hi-C data provides long-range information about the spatial proximity of DNA sequences in the nucleus. This information is used to order and orient contigs along chromosomes, producing chromosome-scale scaffolds. The southern white rhinoceros assembly used Hi-C scaffolding to anchor 2.46 Gb of sequence to 40 autosomes plus the X and Y chromosomes.

How does Inspector evaluate assembly quality without a reference genome?

Inspector uses the raw reads themselves as the ground truth for evaluating assembly correctness. It reports types of errors and their precise locations and can correct assembly errors based on consensus sequences derived from raw reads covering erroneous regions. Inspector can accurately identify both large-scale and small-scale assembly errors.

What are the main sources of assembly error in complex genomes?

Unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Collapsed repeats and haplotypes produce false gene duplications and missing sequence. The Vertebrate Genomes Project found that their assemblies corrected substantial errors and added missing sequence in some of the best historical reference genomes.

Can nanopore sequencing alone produce high-quality complex genome assemblies?

Nanopore sequencing can produce high-quality assemblies, especially when ultra-long reads are generated. The human genome assembly with ultra-long reads demonstrated the feasibility of nanopore-only assembly. However, some projects combine nanopore data with other data types, such as HiFi reads or Hi-C data, to achieve chromosome-scale contiguity and high base accuracy.

How do I choose between PacBio HiFi and Oxford Nanopore sequencing?

The choice depends on the genome characteristics and the biological questions. PacBio HiFi reads have accuracies greater than 99.5% and lengths averaging 10 to 25 kb, making them well suited for assembly of difficult polyploid or highly repetitive genomes. Oxford Nanopore offers ultra-long reads and a faster, more cost-effective approach with extended read lengths. Some projects use both platforms to leverage the strengths of each.

What should I do if my assembly has structural errors?

Use Inspector or a similar reference-free evaluator to identify the types and locations of errors. Inspector can correct assembly errors based on consensus sequences derived from raw reads covering erroneous regions. If errors are widespread, consider reassembling with different parameters or a different assembler. The choice of assembler should be a deliberate, well-justified decision based on the read type and genome characteristics.

Related Bioinformatics Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.