How to Choose Between Long-Read-Only and Hybrid Assembly: A Decision Framework for Your Genome Project
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Long-read-only assembly is often sufficient for complete, circular bacterial genomes with low repetitive content, utilizing ligation libraries for high output and low noise, thus avoiding additional short-read sequencing costs.
- Hybrid assembly, incorporating short reads, is crucial for resolving complex genomes like large eukaryotes or polyploids, and for accurately characterizing bacterial plasmids and mobile elements by correcting homopolymer errors and verifying replicon content.
- Short reads provide essential high-accuracy polishing for long-read assemblies, particularly for correcting systematic errors like homopolymers common in Nanopore data, and for resolving highly repetitive regions that long reads alone may misjoin.
- Library preparation choice significantly impacts assembly quality; ligation libraries offer the highest output and lowest artifactual content for long reads, while tagmentation is suitable for limited DNA input, and amplification libraries should be avoided for de novo assembly due to high noise and short read lengths.
- Cost considerations for hybrid assembly must account for both long-read and short-read library preparation and sequencing runs, which can double the expense but often yields a finished genome more efficiently than extensive manual finishing of a fragmented long-read-only assembly.
- Metagenomic projects with high population diversity benefit from hybrid assembly to improve rare species detection, as short reads offer deep coverage that long reads might miss due to sampling stochasticity, though this increases overall sequencing depth and cost.
Direct Answer and Scope
The central decision in genome assembly is whether to generate only long-read sequencing data or to combine long reads with short reads in a hybrid approach. Long-read-only assembly uses data from Oxford Nanopore Technologies (ONT) or Pacific Biosciences (PacBio) platforms to construct genomes from reads that span thousands of bases. Hybrid assembly adds short-read data, typically from Illumina platforms, to correct errors in long reads and resolve repetitive regions. The choice depends on genome size, complexity, budget, and the biological questions the assembly must answer. For small bacterial genomes with modest repetitive content, long-read-only assembly often produces finished genomes with minimal additional sequencing cost. For large eukaryotic genomes, polyploid samples, or metagenomes with high population diversity, hybrid assembly frequently yields more accurate and contiguous results. This article provides a decision framework grounded in published benchmarking evidence and practical laboratory considerations.
The framework applies to researchers, laboratory professionals, and life-science practitioners who must allocate sequencing budgets and choose analysis workflows. It covers data inputs, workflow choices, quality controls, reproducibility, interpretation limits, and reporting standards. The guidance draws on peer-reviewed comparisons of library preparation strategies, benchmarking studies of assembly tools, and official documentation from bioinformatics training and workflow resources.
At a Glance: Decision Table for Assembly Strategy
| Genome or Sample Type | Recommended Strategy | Primary Rationale | Key Cost Consideration |
|---|---|---|---|
| Small bacterial genome, low repetitive content | Long-read-only with ligation library | Ligation produces high output, low noise, and long reads suitable for complete circular assemblies | One ONT flow cell plus ligation kit, no additional short-read sequencing needed |
| Bacterial genome with plasmids or mobile elements | Hybrid assembly with short-read polishing | Short reads resolve plasmid sequences and detect read leakage artifacts that distort replicon counts | Add one Illumina lane or run, short-read data also supports variant confirmation |
| Large eukaryotic genome with high repeat content | Hybrid assembly with coverage optimization | Short reads correct homopolymer errors and resolve segmental duplications that long reads alone may misjoin | Higher total sequencing cost, requires both platform types |
| Metagenome or soil community sample | Long-read shotgun with optional short-read integration | Long reads recover longer contigs but population diversity biases assembly, short reads help quantify rare taxa | Cost scales with depth needed for rare species detection |
| Transcriptome without reference genome | Long-read-only for isoform discovery | Long reads generate longer assembled transcripts than short reads for reference-free analysis | No short-read data required unless differential expression accuracy is critical |
Core Principles of Assembly Strategy Selection
What Long Reads Provide That Short Reads Cannot
Long-read sequencing platforms generate reads that span repetitive elements, structural variants, and entire mobile genetic elements. This read length advantage directly improves assembly contiguity because assemblers can bridge regions that short reads cannot resolve uniquely. The Oxford Nanopore platform offers several library preparation strategies, each with distinct output characteristics. A 2023 comparison of ONT library strategies for bacterial genomics found that ligation-based libraries produced the largest output at 33.62 gigabases, followed by tagmentation at 11.72 gigabases and amplification at 4.79 gigabases. Average read lengths for tagmentation and ligation exceeded 5,000 base pairs, while amplification produced average reads below 1,100 base pairs. The same study reported that amplification libraries generated the most sequencing noise, with only 22.7 percent of reads mappable to curated genomes, compared to 92.9 percent for ligation and 87.3 percent for tagmentation. Artifactual tandem content was most abundant in amplification libraries at 22.5 percent, while ligation and tagmentation produced only 0.9 percent and 2.2 percent respectively. These findings demonstrate that library preparation choice within a single platform can affect assembly quality as much as the platform selection itself.
What Short Reads Add to a Long-Read Assembly
Short-read data contributes two distinct benefits to genome assembly. First, short reads provide high per-base accuracy that can correct systematic errors in long-read data, particularly homopolymer length errors common in nanopore sequencing. Second, short reads offer deep coverage of repetitive regions at low cost, which helps assemblers distinguish between nearly identical copies of repeated elements. Hybrid assembly strategies use short reads both before and after long-read assembly. Pre-assembly correction uses short reads to polish long reads before the assembly step, while post-assembly polishing applies short reads to the assembled contigs. The choice of when to apply short-read correction affects computational cost and final accuracy.
The Cost Structure of Each Approach
Long-read-only assembly requires one sequencing platform and one library preparation. The per-sample cost depends on throughput, with bacterial genomes typically multiplexed across a single flow cell. Hybrid assembly requires both platforms, doubling the library preparation cost and adding a second sequencing run. The cost difference narrows when considering the cost of finishing an assembly. A long-read-only assembly that produces fragmented contigs may require additional sequencing or manual finishing, while a hybrid assembly often produces a finished genome in one workflow. Budget planning should include the cost of repeated sequencing runs, beyond the initial library preparation.
Genome Size and Complexity as Primary Decision Drivers
Bacterial Genomes and Small Plasmids
Bacterial genomes range from roughly 0.5 to 15 megabases, with most pathogenic and environmental isolates falling between 2 and 7 megabases. The repetitive content of bacterial genomes varies widely. Some species contain few repeated elements and assemble readily from long reads alone. Others contain multiple copies of ribosomal RNA operons, insertion sequences, and prophages that create assembly ambiguities. The 2023 ONT library comparison demonstrated that ligation-based long-read assembly produced high-quality bacterial assemblies with low artifactual content. However, the same study warned that users should not accept assembly results at face value without careful replicon verification, including detection of plasmids assembled from leaked reads. Read leakage occurs when barcoded samples are incorrectly assigned during demultiplexing, causing sequences from one sample to appear in another sample's assembly. This artifact can create false plasmids or merge distinct replicons. Hybrid assembly with short reads provides an independent check on plasmid content because short-read coverage of plasmid sequences should match the expected copy number.
Eukaryotic Genomes and Polyploid Samples
Eukaryotic genomes present greater assembly challenges due to their size, repetitive content, and ploidy. A typical plant or animal genome ranges from 100 megabases to several gigabases, with repetitive elements often constituting more than half of the genome. Long-read-only assembly of these genomes produces highly contiguous contigs but may contain base-level errors in homopolymer runs and tandem repeats. Hybrid assembly adds short-read polishing that corrects these errors but requires sufficient short-read coverage across the entire genome. Polyploid samples add another layer of complexity because assemblers must distinguish between homologous chromosomes and paralogous sequences. The choice between long-read-only and hybrid assembly for polyploid genomes depends on whether the research question requires haplotype resolution or a collapsed consensus assembly.
Metagenomes and Population Diversity
Metagenomic samples contain multiple species at vastly different abundances, creating assembly challenges that single-genome projects do not face. A 2026 review of soil microbiome sequencing methods found that short-read 16S amplicon sequencing, full-length ONT 16S sequencing, and long-read shotgun metagenomics each have distinct biases that shape the recovered community. The review noted that long-read and short-read 16S approaches generally converge on dominant taxa and between-sample differences, but disagree substantially on alpha diversity estimates, rare taxon detection, and the relative abundances of entire phyla. Shotgun metagenomics reveals systematic biases in both short and long-read assembly that depend on population diversity within the sample. For metagenomes, hybrid assembly can improve the recovery of rare species because short reads provide deep coverage that long reads may miss due to sampling stochasticity. However, the cost of deep short-read sequencing across a complex metagenome can exceed the cost of additional long-read sequencing.
Library Preparation Strategies and Their Tradeoffs
Ligation-Based Libraries
Ligation-based library preparation is the standard approach for ONT sequencing. The 2023 comparison found that ligation produced the largest output, the most homogeneous output per channel, and the lowest artifactual tandem content at 0.9 percent. Ligation libraries also produced the highest mappability at 92.9 percent, meaning the vast majority of sequenced reads aligned to the expected genome. The main disadvantage of ligation is the requirement for high-molecular-weight DNA, which demands careful extraction and purification. Samples with degraded DNA may produce shorter reads and lower output with ligation than with amplification-based approaches.
Tagmentation-Based Libraries
Tagmentation combines fragmentation and adapter ligation in a single enzymatic step, reducing hands-on time and DNA input requirements. The 2023 comparison found that tagmentation produced intermediate output at 11.72 gigabases, with mappability of 87.3 percent and artifactual tandem content of 2.2 percent. Output per channel was intermediate in variability. Tagmentation is suitable for samples with limited DNA quantity, but the reduced output compared to ligation may require additional flow cells for large genomes or high-coverage projects.
Amplification-Based Libraries
Amplification-based libraries use PCR to amplify the library, which allows sequencing from very small DNA inputs. The 2023 comparison found that amplification produced the smallest output at 4.79 gigabases, the shortest average read lengths below 1,100 base pairs, and the highest sequencing noise with only 22.7 percent mappability. Artifactual tandem content reached 22.5 percent, the highest of the three strategies. The study also found that basecalling and demultiplexing of barcoded libraries resulted in approximately 20 percent data loss as unclassified reads and 1.5 percent read leakage. Amplification is appropriate when DNA quantity is severely limited, but the reduced read length and increased noise make it a poor choice for de novo assembly projects where long reads are essential.
Platform-Specific Considerations for PacBio
Pacific Biosciences platforms offer an alternative long-read technology with different error profiles. PacBio HiFi reads provide high accuracy in a single pass, reducing the need for short-read polishing. The choice between ONT and PacBio affects the hybrid assembly decision because PacBio HiFi reads may not require short-read correction to the same degree as ONT reads. However, PacBio platforms typically have higher per-run costs and longer turnaround times than ONT. The decision framework should consider the specific error profile of the chosen platform, beyond the read length.
Practical Workflow for Assembly Strategy Selection
Step 1: Define the Biological Question
The assembly strategy must serve the research question. A project aimed at identifying antimicrobial resistance genes in bacterial pathogens requires complete plasmid sequences and accurate gene annotation. A project aimed at characterizing transcript isoforms in a non-model organism requires long reads that span full-length transcripts. A project aimed at describing soil microbial community composition requires methods that detect rare taxa without systematic bias. Write the biological question in explicit terms before selecting sequencing platforms.
Step 2: Estimate Genome Size and Complexity
Use published genome sizes for the target species or closely related organisms. For metagenomes, estimate the total community genome size and the expected abundance range of target species. Assess repetitive content by reviewing published assemblies or running k-mer analysis on pilot sequencing data. Genomes with high repeat content benefit from hybrid assembly because short reads help resolve repeated elements.
Step 3: Assess DNA Quantity and Quality
High-molecular-weight DNA is essential for ligation-based long-read libraries. Measure DNA integrity using pulsed-field gel electrophoresis or an automated fragment analyzer. Samples with DNA fragments below 20 kilobases will produce shorter reads and may require amplification-based libraries, which have known quality tradeoffs. For hybrid assembly, confirm that sufficient DNA remains for both library preparations.
Step 4: Calculate Coverage Requirements
Long-read assembly typically requires 30 to 50-fold coverage for bacterial genomes and 30 to 60-fold coverage for eukaryotic genomes, depending on the platform and error rate. Short-read polishing requires 50 to 100-fold coverage of short-read data. Calculate the total output required and compare with the expected output of each library strategy. The 2023 ONT comparison provides expected output ranges for ligation, tagmentation, and amplification libraries.
Step 5: Compare Costs Across Scenarios
Build a cost model that includes library preparation kits, flow cells or chips, sequencing runs, and computational resources. Include the cost of repeated runs if the initial assembly fails quality checks. Compare the total cost of long-read-only assembly with the total cost of hybrid assembly, including the cost of short-read sequencing. For small genomes, the cost difference may be small enough that hybrid assembly is the safer choice.
Step 6: Select Assembly Tools and Workflows
Choose assembly tools that are maintained and documented. The Galaxy Training Network provides accessible workflow training and analysis tutorials for genome assembly. The nf-core documentation describes community pipeline standards for reproducible workflow configuration. Bioconductor offers official package and workflow documentation for genomic analysis in R. Select tools that have been benchmarked on data similar to the project's target genome type.
Step 7: Define Quality Metrics Before Assembly
Establish quality thresholds before running the assembly. Common metrics include contig N50, number of contigs, completeness based on Benchmarking Universal Single-Copy Orthologs (BUSCO), and base-level accuracy estimated by mapping reads back to the assembly. For bacterial genomes, circularity of the chromosome and plasmids is a key quality indicator. For metagenomes, the recovery of expected species and the absence of chimeric contigs are important checks.
Records and Measurements for Assembly Projects
Sequencing Run Records
Maintain a record of each sequencing run, including the library preparation method, flow cell or chip type, run duration, total output in gigabases, read length distribution, and mappability to a reference genome if available. The 2023 ONT library comparison provides baseline expectations for ligation, tagmentation, and amplification libraries. Record the proportion of unclassified reads after demultiplexing and the estimated read leakage rate. These records allow comparison across runs and early detection of protocol drift.
Assembly Quality Records
Record assembly metrics for each assembly attempt, including the assembler used, version number, parameter settings, and computational resources required. Track contig N50, total assembly size, number of contigs, and completeness scores. For bacterial assemblies, record whether the chromosome and each plasmid assembled as circular molecules. For hybrid assemblies, record the short-read polishing method and the number of polishing rounds.
Sample Metadata
Link every assembly to complete sample metadata, including species identification, source, collection date, DNA extraction method, and library preparation details. The NCBI Data Resources provide official descriptions of sequence databases and search systems that support metadata standards. Proper metadata ensures that assemblies can be interpreted in context and reproduced by other researchers.
Version Control and Reproducibility
Use version control for analysis scripts and workflow configurations. The Carpentries lessons provide foundational training in shell, Git, and programming that supports reproducible analysis. The nf-core documentation describes community standards for pipeline usage and configuration. Record the exact software versions and parameter settings for every assembly step to ensure that results can be reproduced or audited.
Common Failure Patterns and How to Avoid Them
Accepting Assembly Results Without Replicon Verification
The 2023 ONT library comparison explicitly warned that users should not accept assembly results at face value without careful replicon verification. Plasmids assembled from leaked reads can create false positives in antimicrobial resistance surveillance. Verify every circular contig by checking read coverage consistency and by mapping short reads to the assembly when available. For bacterial projects, compare the expected plasmid profile with the assembled replicons.
Choosing Amplification Libraries for De Novo Assembly
Amplification libraries produce short reads, high noise, and high artifactual tandem content. Using amplification libraries for de novo assembly of unknown genomes will likely produce fragmented and error-prone assemblies. Reserve amplification libraries for samples with extremely limited DNA and use hybrid assembly with short-read correction to compensate for the quality deficits.
Ignoring Read Leakage in Multiplexed Runs
Read leakage between barcoded samples can create false sequences in assemblies. The 2023 ONT comparison reported approximately 1.5 percent read leakage in barcoded libraries. For projects with many multiplexed samples, include negative controls and verify that control samples contain no reads from other samples. Consider using unique dual barcodes or increasing the stringency of demultiplexing parameters.
Underestimating the Impact of Population Diversity in Metagenomes
Metagenomic assembly quality depends on the population diversity within the sample. The 2026 soil microbiome review found that shotgun metagenomics reveals systematic biases in both short and long-read assembly that depend on population diversity. High-diversity samples produce fragmented assemblies because closely related strains cannot be distinguished. For such samples, consider assembly binning strategies that separate contigs by coverage and composition.
Skipping Short-Read Polishing for Nanopore Assemblies
Nanopore sequencing has systematic errors in homopolymer regions that persist in assembled contigs. Hybrid assembly with short-read polishing corrects these errors and improves base-level accuracy. Skipping short-read polishing to save cost may produce assemblies with error rates that affect downstream analysis, particularly variant calling and gene annotation.
Quality Controls and Validation Steps
Read-Level Quality Checks
Before assembly, assess read quality using platform-specific metrics. For ONT data, check the read length distribution, quality scores, and the proportion of reads that pass the chosen quality threshold. The 2023 ONT comparison provides expected mappability rates for each library strategy. Low mappability indicates contamination, adapter issues, or sequencing noise that will degrade assembly quality.
Assembly-Level Quality Checks
After assembly, run multiple quality checks. Map reads back to the assembly to estimate base-level accuracy. Run BUSCO to assess completeness against expected single-copy orthologs. For bacterial genomes, check that the chromosome and plasmids are circular and that coverage is uniform across the genome. For eukaryotic genomes, check for unexpected duplications or collapses in repetitive regions.
Cross-Platform Validation
When hybrid assembly is used, validate the final assembly by mapping both long and short reads to the assembled contigs. Discrepancies between the two data types may indicate assembly errors. For metagenomes, compare the taxonomic composition derived from the assembly with the composition derived from amplicon sequencing of the same sample. The 2026 soil microbiome review found that long-read and short-read 16S approaches converge on dominant taxa but disagree on rare taxon detection, so disagreements in rare taxa may reflect method bias instead of assembly error.
Replicon Verification for Bacterial Assemblies
For bacterial genomes, verify every circular contig independently. Check that the coverage of each replicon is consistent with its expected copy number. Plasmids present at high copy number should show higher coverage than the chromosome. The 2023 ONT comparison detected plasmids assembled from leaked reads, demonstrating that replicon verification is essential for accurate plasmid reporting.
Interpretation Limits and Reporting Standards
What Assembly Quality Metrics Do Not Capture
Assembly quality metrics measure contiguity and completeness but do not capture all biological features. A highly contiguous assembly may still contain base-level errors in repetitive regions. A complete BUSCO score does not guarantee that structural variants are correctly assembled. Report the limitations of the assembly alongside the quality metrics so that downstream users can interpret results appropriately.
Reporting Requirements for Public Databases
Deposit assemblies in public databases with complete metadata. The NCBI Data Resources provide official descriptions of sequence databases, search systems, and analysis services that support assembly submission and retrieval. Include the assembly method, sequencing platforms, coverage, and quality metrics in the submission. The EMBL-EBI Training resources describe data-resource training and practical analysis education that supports proper data submission.
Reproducibility Standards
Publish the complete analysis workflow, including software versions, parameter settings, and computational environment. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration. The Galaxy Training Network provides accessible workflow training and analysis tutorials that demonstrate reproducible analysis practices. Bioconductor offers official package and workflow documentation for reproducible genomic analysis in R.
Welfare and Safety Context for Veterinary and Clinical Applications
Antimicrobial Resistance Surveillance
Long-read and hybrid assembly play a critical role in antimicrobial resistance surveillance. A 2026 study of carbapenem-resistant Acinetobacter species from companion animals in Japan used whole-genome sequencing to identify resistance genes and their genomic contexts. The study found that a plasmid containing blaNDM-1 and blaOXA-58 was present in an A. johnsonii isolate, representing the first report of blaNDM-1-harboring Acinetobacter from companion animals in Japan. This finding demonstrates that assembly quality directly affects the detection of mobile genetic elements carrying resistance genes. Hybrid assembly with short-read validation provides confidence that plasmid-borne resistance genes are correctly localized.
Veterinary Pathogen Detection
Nanopore sequencing has emerged as a tool for veterinary pathogen detection, offering portability, real-time long-read data, and minimal infrastructure requirements. A 2026 review of nanopore sequencing in veterinary pathogen detection noted that for bacterial pathogens, long-read sequencing enables near-complete genome assembly and identification of plasmid-borne antimicrobial resistance genes. The review also acknowledged current limitations in accuracy and host-DNA interference. For veterinary diagnostics, the choice between long-read-only and hybrid assembly affects the speed of results and the confidence in resistance gene detection. Hybrid assembly provides higher accuracy but requires additional sequencing time, which may delay clinical decisions.
One Health Surveillance
One Health surveillance integrates human, animal, and environmental health data. Assembly quality affects the interpretation of pathogen transmission chains and resistance gene spread. The 2026 veterinary nanopore review emphasized the role of nanopore sequencing in One Health surveillance of emerging zoonoses. For surveillance projects, hybrid assembly provides the accuracy needed for confident strain typing and resistance gene characterization, while long-read-only assembly offers speed and portability for field settings. The choice depends on whether the project prioritizes speed or accuracy.
Professional Escalation Criteria
When to Seek Additional Expertise
Escalate to a bioinformatics specialist or sequencing facility when the assembly fails quality thresholds after two attempts, when the genome contains unexpected structural features that cannot be resolved, or when the assembly produces conflicting results across different assemblers. Also escalate when the project requires regulatory submission, clinical reporting, or public health action, because these contexts demand higher validation standards.
When to Add Short-Read Data
Add short-read data to a long-read-only project when the assembly contains unresolved repetitive regions, when base-level accuracy is insufficient for variant calling, or when plasmid content cannot be verified. The 2023 ONT comparison demonstrated that even high-quality ligation libraries benefit from careful replicon verification. If the long-read assembly produces fragmented contigs in specific genomic regions, short-read data can often resolve those regions.
When to Change Library Preparation
Change library preparation strategy when the current strategy produces insufficient output, excessive noise, or unacceptably short reads. The 2023 ONT comparison provides expected performance for ligation, tagmentation, and amplification libraries. If a ligation library produces low output, check DNA quality and consider whether the sample requires tagmentation. If an amplification library produces high noise, consider whether the sample can provide sufficient DNA for ligation.
When to Consult Platform-Specific Resources
Consult platform-specific documentation and training resources when planning a new project or troubleshooting failed runs. The Galaxy Training Network provides accessible workflow training and analysis tutorials. The nf-core documentation describes community pipeline standards. The EMBL-EBI Training resources provide bioinformatics learning pathways and data-resource training. The Carpentries lessons provide foundational computing and data skills that support independent troubleshooting.
A Practical Decision Framework for Assembly Strategy Selection
The Coverage-Adjusted Cost Matrix
A practical decision framework must move beyond general recommendations and provide a structured method for comparing long-read-only and hybrid assembly options. The coverage-adjusted cost matrix is a record-keeping tool that translates sequencing output requirements into direct cost comparisons. This framework requires researchers to calculate the minimum coverage needed for each assembly strategy, then convert those coverage requirements into flow cell or chip counts and library preparation costs.
Start by determining the genome size in megabases. For a bacterial genome of 5 megabases, long-read-only assembly with ligation libraries typically requires 30 to 50-fold coverage. The 2023 ONT library comparison reported that ligation produced 33.62 gigabases of output, which is sufficient for hundreds of bacterial genomes at 50-fold coverage. However, the practical output per flow cell depends on the number of samples multiplexed and the expected read length distribution. For a single bacterial genome, one flow cell with a ligation library provides ample coverage, and the cost is dominated by the library preparation kit and flow cell instead of the sequencing output.
For hybrid assembly, the cost model must include both long-read and short-read components. Long-read coverage of 30 to 50-fold is still required for the initial assembly, and short-read coverage of 50 to 100-fold is needed for polishing. The short-read component adds a second library preparation and a second sequencing run. The cost difference between long-read-only and hybrid assembly for a single bacterial genome is therefore the cost of the short-read library and sequencing run. This difference may be small relative to the total project cost, making hybrid assembly the safer choice when accuracy is critical.
For eukaryotic genomes, the cost calculation changes substantially. A 1-gigabase genome requires 30 to 60 gigabases of long-read data for 30 to 60-fold coverage. This may require multiple flow cells or chips, depending on the platform. The short-read component for polishing adds another 50 to 100 gigabases of data. The total cost of hybrid assembly for a large genome can be two to three times the cost of long-read-only assembly. The decision framework must therefore weigh the cost of additional long-read sequencing against the cost of short-read polishing and the value of improved base-level accuracy.
The Decision Tree for Strategy Selection
The decision tree below provides a structured path for selecting between long-read-only and hybrid assembly. Each node in the tree corresponds to a specific question about the project, and the answers direct the researcher to the appropriate strategy.
Node 1: What is the target genome size?
If the genome is below 15 megabases, proceed to Node 2. If the genome is above 15 megabases, proceed to Node 3.
Node 2: Does the sample contain plasmids or mobile genetic elements?
If yes, proceed to Node 4. If no, long-read-only assembly with ligation libraries is the recommended strategy. The 2023 ONT comparison demonstrated that ligation libraries produce high output, low noise, and long reads suitable for complete bacterial assemblies. However, even for simple genomes, replicon verification is essential because the study detected plasmids assembled from leaked reads.
Node 3: Does the genome contain high repetitive content?
If yes, hybrid assembly is recommended. Short reads help resolve repetitive regions and correct homopolymer errors that persist in long-read assemblies. If no, long-read-only assembly may be sufficient, but base-level accuracy should be verified before proceeding to downstream analysis.
Node 4: Is plasmid content critical for the research question?
If yes, hybrid assembly is recommended. Short reads provide an independent check on plasmid content because short-read coverage of plasmid sequences should match the expected copy number. The 2023 ONT comparison warned that users should not accept assembly results at face value without careful replicon verification. If plasmid content is not critical, long-read-only assembly with ligation libraries may be sufficient, but the assembly should still be checked for circular contigs and coverage consistency.
Node 5: What is the DNA quantity and quality?
If high-molecular-weight DNA is available, ligation libraries are preferred. If DNA is limited or degraded, tagmentation libraries may be necessary. The 2023 ONT comparison found that tagmentation produced intermediate output at 11.72 gigabases with 87.3 percent mappability. Amplification libraries should be avoided for de novo assembly because they produced the smallest output, the shortest reads, and the highest noise with only 22.7 percent mappability.
The Decision Record Template
A standardized decision record ensures that the strategy selection process is documented and reproducible. The record should capture the following fields for each project:
Project identification fields: Sample identifier, species, source, collection date, and research question.
Genome assessment fields: Estimated genome size, estimated repetitive content, expected plasmid content, and ploidy.
Sample quality fields: DNA extraction method, DNA quantity in nanograms, DNA integrity measured by fragment analysis, and the presence of contaminants.
Coverage calculation fields: Target long-read coverage, target short-read coverage if hybrid assembly is considered, and the calculated output required in gigabases.
Cost comparison fields: Cost of long-read library preparation, cost of long-read sequencing, cost of short-read library preparation if applicable, cost of short-read sequencing if applicable, and total cost for each strategy.
Strategy selection field: The chosen strategy and the rationale for the choice.
Quality threshold fields: The minimum contig N50, minimum BUSCO completeness, and maximum error rate that the assembly must meet.
This record serves as the basis for the assembly quality assessment. If the assembly fails to meet the quality thresholds, the record provides the information needed to decide whether to add short-read data, change library preparation, or escalate to a specialist.
The Two-Pass Assembly Assessment Method
The two-pass assembly assessment method is a troubleshooting approach that separates the evaluation of assembly contiguity from the evaluation of base-level accuracy. This method is particularly useful for deciding whether a long-read-only assembly requires short-read polishing.
Pass 1: Contiguity assessment. Run the long-read assembly and evaluate contiguity metrics including contig N50, number of contigs, and total assembly size. For bacterial genomes, check whether the chromosome and each plasmid assembled as circular molecules. For eukaryotic genomes, check whether the assembly contains the expected number of chromosomes or chromosome-scale scaffolds. If contiguity is acceptable, proceed to Pass 2. If contiguity is poor, the problem is likely in the long-read data or the assembly parameters, and adding short reads will not fix the underlying fragmentation.
Pass 2: Base-level accuracy assessment. Map the long reads back to the assembly and estimate base-level accuracy. For ONT data, pay particular attention to homopolymer regions, which are known error hotspots. If base-level accuracy is below the threshold required for downstream analysis, short-read polishing is recommended. The 2026 soil microbiome review noted that the R10.4.1 flow cell chemistry has narrowed but not eliminated the accuracy gap with Illumina, indicating that even current ONT chemistry benefits from short-read correction for high-accuracy applications.
The two-pass method prevents a common failure pattern: adding short-read data to an assembly that is fragmented due to insufficient long-read coverage or poor library preparation. Short reads cannot resolve assembly gaps caused by missing long-read coverage across repetitive regions. The method also prevents the opposite failure pattern: accepting a contiguous assembly without verifying base-level accuracy, which can lead to errors in variant calling and gene annotation.
The Replicon Verification Protocol
The replicon verification protocol is a specific troubleshooting method for bacterial assemblies that addresses the warning from the 2023 ONT comparison about accepting assembly results at face value. The protocol consists of four checks that should be applied to every circular contig in a bacterial assembly.
Check 1: Coverage consistency. Calculate the read coverage across each circular contig. The chromosome should have uniform coverage, and plasmids present at high copy number should show higher coverage than the chromosome. Coverage that varies dramatically across a contig may indicate a misassembly or a chimeric sequence.
Check 2: Read mapping validation. Map the raw reads back to each circular contig and inspect the mapping quality. Reads that map with high identity across the full length of the contig support the assembly. Reads that map with poor identity or that span the contig boundaries may indicate errors.
Check 3: Short-read cross-validation. When short-read data is available, map the short reads to each circular contig. Short-read coverage should be consistent with the expected copy number of each replicon. Discrepancies between long-read and short-read coverage may indicate read leakage or assembly errors.
Check 4: Negative control verification. For multiplexed runs, verify that negative control samples contain no reads from other samples. The 2023 ONT comparison reported approximately 1.5 percent read leakage in barcoded libraries. The study detected plasmids assembled from leaked reads, demonstrating that this artifact can create false replicons in bacterial assemblies.
The Escalation Decision Matrix
The escalation decision matrix provides clear criteria for when to add short-read data, change library preparation, or seek specialist expertise. The matrix is based on the quality thresholds defined in the decision record.
Escalate to short-read addition when: The long-read assembly meets contiguity thresholds but fails base-level accuracy thresholds. The assembly contains unresolved repetitive regions that short reads could resolve. Plasmid content cannot be verified with confidence. The 2023 ONT comparison demonstrated that even high-quality ligation libraries benefit from careful replicon verification, so short-read addition is a reasonable response to verification failures.
Escalate to library preparation change when: The current library strategy produces insufficient output, excessive noise, or unacceptably short reads. The 2023 ONT comparison provides expected performance for ligation, tagmentation, and amplification libraries. If a ligation library produces low output, check DNA quality and consider whether the sample requires tagmentation. If an amplification library produces high noise, consider whether the sample can provide sufficient DNA for ligation.
Escalate to specialist expertise when: The assembly fails quality thresholds after two attempts. The genome contains unexpected structural features that cannot be resolved. The assembly produces conflicting results across different assemblers. The project requires regulatory submission, clinical reporting, or public health action. These contexts demand higher validation standards than research projects.
The Cost-Benefit Worksheet
The cost-benefit worksheet is a practical tool for comparing long-read-only and hybrid assembly options before committing to a sequencing strategy. The worksheet should be completed for each project and stored with the decision record.
Section 1: Sequencing cost estimates. List the cost of long-read library preparation, long-read sequencing, short-read library preparation if applicable, and short-read sequencing if applicable. Include the cost of any repeated runs that may be needed if the initial assembly fails quality checks.
Section 2: Computational cost estimates. List the cost of computational resources for assembly, polishing, and quality assessment. Hybrid assembly requires additional computational steps for short-read polishing, which increases the computational cost.
Section 3: Labor cost estimates. Estimate the hands-on time for library preparation, sequencing, and analysis for each strategy. Hybrid assembly requires additional library preparation and analysis steps.
Section 4: Risk assessment. Estimate the probability that each strategy will produce an assembly that meets the quality thresholds defined in the decision record. Consider the genome complexity, DNA quality, and the known performance of each library strategy from the 2023 ONT comparison.
Section 5: Total cost comparison. Sum the costs for each strategy and compare. The strategy with the lower total cost is not always the better choice. The worksheet should also consider the value of improved accuracy, the cost of repeated runs, and the consequences of assembly failure for the research question.
The cost-benefit worksheet is particularly important for metagenome projects. The 2026 soil microbiome review found that shotgun metagenomics reveals systematic biases in both short and long-read assembly that depend on population diversity within the sample. For high-diversity samples, the cost of deep short-read sequencing may exceed the cost of additional long-read sequencing, and the optimal strategy may be long-read-only with careful interpretation of the biases. The worksheet provides the structure needed to make this comparison explicit.
Frequently Asked Questions
What is the main advantage of long-read-only assembly over hybrid assembly?
Long-read-only assembly requires only one sequencing platform, reducing library preparation cost and workflow complexity. Long reads span repetitive elements and structural variants that short reads cannot resolve, producing highly contiguous assemblies. For small bacterial genomes with low repetitive content, long-read-only assembly with ligation libraries produces complete circular chromosomes and plasmids in a single workflow. The 2023 ONT library comparison demonstrated that ligation libraries produce high output, low noise, and long reads suitable for complete bacterial assemblies.
When should I choose hybrid assembly instead of long-read-only assembly?
Choose hybrid assembly when the genome contains complex repetitive regions, when base-level accuracy is critical for downstream analysis, or when plasmid content must be verified with independent evidence. Hybrid assembly adds short-read data that corrects homopolymer errors and resolves repetitive regions. For metagenomes with high population diversity, hybrid assembly improves the recovery of rare species and reduces assembly bias. The 2026 soil microbiome review found that shotgun metagenomics reveals systematic biases in both short and long-read assembly that depend on population diversity, suggesting that hybrid approaches may mitigate these biases.
How does library preparation choice affect long-read assembly quality?
Library preparation choice significantly affects output, read length, and noise. The 2023 ONT comparison found that ligation libraries produced the largest output at 33.62 gigabases, the highest mappability at 92.9 percent, and the lowest artifactual tandem content at 0.9 percent. Amplification libraries produced the smallest output, the shortest reads, and the highest noise with only 22.7 percent mappability. Tagmentation libraries performed intermediately. For de novo assembly projects, ligation libraries are the preferred choice when DNA quantity and quality permit.
What is read leakage and why does it matter for assembly?
Read leakage occurs when reads from one barcoded sample are incorrectly assigned to another sample during demultiplexing. The 2023 ONT comparison reported approximately 1.5 percent read leakage in barcoded libraries. Read leakage can create false plasmids or merge distinct replicons in bacterial assemblies. The study detected plasmids assembled from leaked reads, demonstrating that replicon verification is essential. Use negative controls and verify that control samples contain no reads from other samples.
Can long-read-only assembly produce accurate bacterial genomes?
Yes, long-read-only assembly can produce accurate bacterial genomes when using ligation libraries and appropriate assembly tools. The 2023 ONT comparison found that ligation libraries produced high-quality assemblies with low artifactual content. However, the study warned that users should not accept assembly results at face value without careful replicon verification. Base-level accuracy may still benefit from short-read polishing, particularly in homopolymer regions.
How does genome size affect the choice between long-read-only and hybrid assembly?
Genome size affects the cost and feasibility of each approach. Small bacterial genomes require relatively low sequencing output, making long-read-only assembly cost-effective. Large eukaryotic genomes require high coverage and may benefit from hybrid assembly to correct base-level errors. The cost of short-read sequencing for a large genome can be substantial, so the decision should compare the cost of additional long-read sequencing against the cost of short-read polishing.
What quality metrics should I report for a genome assembly?
Report contig N50, number of contigs, total assembly size, completeness based on BUSCO, and base-level accuracy estimated by mapping reads back to the assembly. For bacterial genomes, report whether the chromosome and each plasmid assembled as circular molecules. For metagenomes, report the recovery of expected species and the absence of chimeric contigs. Include the assembly method, sequencing platforms, coverage, and software versions in the submission to public databases.
How do I verify plasmid content in a bacterial assembly?
Verify plasmid content by checking that each circular contig has consistent read coverage and by mapping short reads to the assembly when available. Plasmids present at high copy number should show higher coverage than the chromosome. The 2023 ONT comparison detected plasmids assembled from leaked reads, so independent verification is essential. Compare the assembled plasmid profile with the expected plasmid content based on the species and any prior characterization.
Related Bioinformatics Guides
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
- Long-Read Genome Assembly and Polishing Strategies
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- A comparison of Oxford nanopore library strategies for bacterial genomics.. BMC genomics, 2023.
- Choosing Between Short-Read 16S, Full-Length ONT 16S, and Long-Read Shotgun Metagenomics for Soil Microbiome Studies: A Critical Review of the Benchmarking Evidence.. 2026.
- A comprehensive evaluation of long-read de novo transcriptome assembly.. 2026.
- Nanopore Sequencing in Veterinary Pathogen Detection: A Review of Technologies and Applications.. 2026.
- Genetic Characterization of Carbapenem-Resistant <,i>,Acinetobacter<,/i>, spp. Isolated from Diseased Companion Animals in Japan.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.