PacBio vs. Oxford Nanopore: A Side-by-Side Comparison for Genome Assembly Projects

By Dr. Zubair Khalid, DVM, MS, PhD ·

PacBio vs. Oxford Nanopore: A Side-by-Side Comparison for Genome Assembly Projects

Key Takeaways

  • PacBio HiFi excels in accuracy for reference-grade assemblies, yielding reads typically 10-25 kb with Q20+ accuracy, minimizing downstream polishing needs and reducing bioinformatics workload. This high accuracy is achieved through circular consensus sequencing, making it ideal for projects demanding minimal residual errors.
  • Oxford Nanopore (ONT) offers superior read length flexibility, with capabilities for ultralong reads (>100 kb), crucial for resolving highly repetitive genomic regions and complex structural variants. While per-read accuracy is lower and varies with base-calling models, ONT's scalability and lower entry cost make it attractive for large genomes and rapid turnaround projects.
  • The choice between platforms hinges on a trade-off between read length and accuracy, directly impacting assembly complexity and cost. PacBio's high accuracy simplifies assembly, whereas ONT's longer reads necessitate more rigorous error correction, often involving short-read polishing or advanced self-polishing algorithms.
  • Library preparation requirements differ significantly, with PacBio HiFi demanding substantial DNA input and shearing, while ONT requires less DNA and can process native, unsheared molecules. This distinction is critical for projects with limited sample material or degraded DNA.
  • Hybrid assembly strategies, combining long reads (PacBio or ONT) with Illumina short reads, are increasingly employed to leverage the strengths of both technologies. This approach is particularly effective for bacterial genomes and complex plant/animal genomes requiring chromosome-scale scaffolding, often incorporating Hi-C data for enhanced contiguity.
  • Total project cost must encompass sequencing, library preparation, and substantial bioinformatics resources, as ONT's lower per-base cost can be offset by increased compute time for error correction and assembly tool optimization. Conversely, PacBio's higher per-base cost may be justified by reduced downstream analysis effort and higher initial assembly quality.

Researchers planning a genome assembly project face a practical decision between the two dominant long-read sequencing platforms: Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT). Both platforms generate reads far longer than Illumina short reads, which simplifies assembly and resolves repetitive regions, but they differ substantially in read length profiles, base-calling accuracy, throughput, cost structure, and downstream bioinformatics requirements. This article provides a structured comparison for biology students, researchers, laboratory professionals, and life-science practitioners who need to select a platform for a specific assembly project. The decision framework considers genome size, complexity, budget, and downstream analytical needs, with reference to published assembly projects that used one or both platforms.

Platform Fundamentals and Data Output Differences

PacBio and ONT both produce long reads, but the underlying chemistry and data characteristics differ in ways that affect assembly strategy. PacBio systems use a zero-mode waveguide detector to observe DNA polymerase activity in real time during synthesis. The HiFi mode produces highly accurate circular consensus reads by sequencing the same molecule multiple times. ONT systems pass DNA through a protein nanopore embedded in a membrane and measure changes in ionic current as the molecule translocates, enabling direct detection of native DNA without synthesis.

The read length profiles differ meaningfully. PacBio HiFi reads typically range from 10 to 25 kilobases with high per-read accuracy. ONT can produce ultralong reads exceeding 100 kilobases, and in some projects reads of several hundred kilobases have been reported, but per-read accuracy is lower and varies with the base-calling model and flow cell generation. A study of the cynomolgus macaque major histocompatibility complex region used both ultralong ONT reads and PacBio HiFi reads to assemble a 5.2 megabase haplotype, demonstrating that the two platforms can be complementary for complex genomic regions 9.

For genome assembly, the key tradeoff is between read length and accuracy. Longer reads bridge more repetitive elements and structural variants, but lower accuracy requires polishing steps to correct errors. PacBio HiFi reads have high accuracy that reduces the need for extensive polishing, while ONT reads may require additional correction using short-read data or self-polishing approaches. The choice affects sequencing cost and the bioinformatics workload after sequencing.

At a Glance: Platform Comparison for Assembly Projects

The following table summarizes the main considerations for choosing between PacBio and ONT for genome assembly. Values represent typical ranges reported in the literature and should be confirmed with current vendor specifications for a specific project.

FeaturePacBio HiFiOxford NanoporePractical Implication
Typical read length10 to 25 kb10 to 100+ kb, ultralong possibleONT bridges larger repeats, PacBio provides consistent moderate length
Per-read accuracyHigh (Q20 or better typical)Lower, varies with base-calling modelPacBio requires less polishing, ONT may need short-read correction
Throughput per runModerate to high depending on instrumentVariable, depends on flow cell and pore countONT scales with flow cell number, PacBio scales with SMRT cell number
Instrument costHigh capital investmentLower entry cost, portable optionsONT suits smaller labs or field projects
Per-base costHigher per GbLower per Gb at high throughputONT can be economical for large genomes
Assembly complexitySimpler due to high accuracyMore complex due to error correction needsPacBio reduces bioinformatics burden
Best use caseReference-grade assemblies, complex genomesUltralong reads, structural variation, rapid turnaroundMatch platform to project goals

A second comparison table addresses the practical workflow differences that affect project planning.

Workflow StepPacBio HiFiOxford NanoporeDecision Point
Library preparationRequires substantial DNA input, shearing to target sizeLower DNA input, no shearing for native readsCheck DNA yield and quality before choosing
Sequencing run timeHours to days depending on instrumentReal-time data output, minutes to daysONT allows early stopping once coverage is sufficient
Base callingOn-instrument or cloud-basedReal-time or post-run with various modelsONT base-calling model affects accuracy
PolishingMinimal neededOften required, can use short readsBudget for additional compute time
Assembly toolsMany tools support HiFi directlyTools vary in handling error profilesVerify assembler compatibility before starting

Genome Size and Complexity Considerations

Genome size is the first filter in platform selection. Bacterial genomes of 4 to 5 megabases are now routinely assembled to completion with either platform. A study of nine drug-resistant Mycobacterium tuberculosis isolates compared assembly tools using Illumina, ONT PromethION, and PacBio data, finding that hybrid assemblies with long reads produced more complete genomes than short-read assembly alone 7. The M. tuberculosis genome is 4.4 megabases, and the study demonstrated that both ONT and PacBio long reads improved gene detection compared to Illumina-only assembly.

For larger genomes, the cost per gigabase becomes a dominant factor. Plant genomes frequently exceed 1 gigabase and can reach several gigabases. A study of the tetraploid forage crop sainfoin assembled a 2.36 gigabase genome using PacBio HiFi, ONT, Illumina short reads, and Hi-C data, resolving all 28 pseudochromosomes with high contiguity and gene completeness 10. This project used both long-read platforms in combination, suggesting that for very large and complex genomes, a single platform may not be optimal.

Genome complexity matters beyond size. Repetitive content, polyploidy, heterozygosity, and GC bias all influence assembly difficulty. The Ranunculus study compared Illumina short-read, ONT long-read, PacBio HiFi long-read, and hybrid strategies for a plant with a 2.69 gigabase genome, ultimately favoring a PacBio-based assembly polished with filtered short reads and scaffolded with Hi-C data 8. The authors noted that cost remains prohibitive for plant species with large, complex genomes, which is a practical consideration for any project budget.

For highly repetitive regions such as the major histocompatibility complex in macaques, ultralong ONT reads provided the ability to span duplicated segments, while PacBio HiFi reads contributed high accuracy for gene annotation 9. This project illustrates that the two platforms can be used together to overcome limitations of either alone.

Accuracy Profiles and Their Effect on Assembly Quality

Base-calling accuracy determines the error profile of the final assembly and the amount of polishing required. PacBio HiFi reads achieve high consensus accuracy because each molecule is sequenced multiple times to form a circular consensus sequence. The resulting reads have error rates low enough that assembly tools can produce high-quality contigs without additional short-read polishing.

ONT reads have a different error profile. The raw signal is translated into base calls using models that have improved over time, but errors remain more frequent than in PacBio HiFi reads. Insertions and deletions are more common than substitutions, which affects how assembly tools handle the data. Some projects use ONT reads for initial assembly and then polish with Illumina short reads to correct residual errors. The Mycobacterium tuberculosis study found that Illumina and ONT hybrid assemblies produced the highest number of SNPs, which may reflect either true variant detection or residual sequencing errors 7. This finding underscores the importance of validating variants in ONT-based assemblies.

The choice of assembly tool interacts with platform accuracy. The same M. tuberculosis study compared Unicycler, RagOut, and RagTag and found that Unicycler-based assemblies had significantly higher genome completeness at approximately 98.7 percent compared to the other tools 7. RagOut produced the fewest contigs and the longest genome size, which the authors selected for downstream analysis. This result demonstrates that platform choice alone does not determine assembly quality, the assembler must match the data type.

For projects that require reference-grade assemblies, PacBio HiFi data reduce the polishing burden. The sainfoin genome project used PacBio HiFi as the primary long-read data source and achieved reference-grade LTR assembly index scores across all haplotypes 10. The Ranunculus project similarly favored PacBio-based assembly for the nuclear genome, polishing three times with filtered short reads 8. These examples suggest that when the goal is a high-quality reference genome, PacBio HiFi is often the preferred primary platform.

Throughput and Scalability for Different Project Sizes

Throughput determines how many samples can be processed in a given time and how much sequencing is needed to reach target coverage. PacBio instruments have defined SMRT cell capacities, and throughput depends on the instrument model. ONT flow cells can be run individually or in parallel, and the real-time data output allows researchers to monitor coverage and stop runs once sufficient data have been collected.

For small genomes such as bacteria, a single flow cell or SMRT cell may provide enough coverage for multiple samples. The Vreelandella study used hybrid sequencing combining PacBio long reads and Illumina short reads to generate circular and complete genomes for two bacterial strains, one of 4.5 megabases and one of 3.6 megabases 11. This approach is common for bacterial projects because the cost per genome is modest and the resulting assemblies are complete.

For large genomes, throughput becomes a cost driver. A 2.69 gigabase plant genome requires substantial sequencing coverage, and the cost of PacBio HiFi at that scale may be prohibitive for some laboratories 8. ONT offers a lower per-base cost at high throughput, which can make it attractive for large genomes, but the additional bioinformatics work for error correction must be factored into the total project cost.

The real-time nature of ONT sequencing provides a practical advantage for projects with uncertain coverage requirements. Researchers can assess data quality and coverage during the run and decide whether to continue or stop. PacBio runs are typically completed in a fixed time window, and stopping early may waste the run if coverage is insufficient.

Cost Analysis and Budget Planning

Cost comparison between PacBio and ONT requires consideration of instrument purchase or access, consumables, library preparation, sequencing, and downstream bioinformatics. The capital cost of PacBio instruments is generally higher than ONT entry-level devices, but the per-run cost depends on throughput and the specific instrument model.

ONT offers a lower entry point with portable devices that can be used in field settings. This accessibility has made ONT attractive for projects with limited budgets or for laboratories that do not have access to a core sequencing facility. However, the lower per-base cost of ONT at high throughput must be balanced against the need for additional sequencing to compensate for lower accuracy and the compute time required for polishing.

PacBio HiFi has a higher per-base cost but produces reads that require less downstream processing. For projects where researcher time is the limiting factor, the higher sequencing cost may be offset by reduced bioinformatics effort. The Ranunculus study noted that cost can still be prohibitive for plant species with large, complex genomes, which suggests that budget constraints may push some projects toward ONT or hybrid approaches 8.

Library preparation costs also differ. PacBio HiFi requires shearing DNA to a target size and constructing circular consensus sequencing libraries. ONT library preparation is simpler and requires less DNA, which is an advantage for samples with limited yield. The sainfoin project used multiple data types including PacBio HiFi, ONT, Illumina, and Hi-C, which increased the total cost but produced a haplotype-resolved reference genome 10. Projects with limited budgets may need to choose a single platform and accept the associated tradeoffs.

Bioinformatics Workflows and Tool Compatibility

The bioinformatics pipeline for genome assembly differs between platforms in terms of base calling, assembly, polishing, and quality assessment. Both platforms have established workflows, and training resources are available through official channels. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover long-read assembly, and the nf-core Documentation describes community pipeline standards for reproducible analysis. These resources can help laboratories implement robust workflows without developing everything from scratch.

Base calling is the first step in the bioinformatics pipeline. PacBio instruments perform base calling during the run, producing reads in standard formats. ONT base calling can be performed in real time or after the run using various models, and the choice of base-calling model affects read accuracy. Researchers should document the base-calling model used because it affects downstream analysis and reproducibility.

Assembly tools have different levels of support for each platform. Some tools are designed specifically for PacBio HiFi data, while others handle ONT error profiles. The Mycobacterium tuberculosis study compared Unicycler, RagOut, and RagTag and found significant differences in assembly completeness and contiguity 7. This finding emphasizes the need to test multiple assemblers and select the one that performs best for the specific data type and genome characteristics.

Polishing is an additional step for ONT assemblies. Short-read polishing with Illumina data can correct residual errors, but this adds cost and complexity. Some projects use self-polishing approaches with the long-read data alone, but the effectiveness depends on the error profile and coverage. The Ranunculus project polished a PacBio-based assembly three times with filtered short reads, demonstrating that even high-accuracy long reads benefit from short-read polishing in complex genomes 8.

Quality assessment is essential after assembly. Completeness can be assessed using BUSCO, which checks for the presence of conserved single-copy genes. The Ranunculus assembly achieved 94.1 percent complete BUSCO genes 8, and the sainfoin assembly achieved high gene completeness across all haplotypes 10. These metrics provide a standardized way to compare assembly quality across projects and platforms.

Hybrid Assembly Strategies and When to Use Them

Hybrid assembly combines long reads from PacBio or ONT with short reads from Illumina to leverage the strengths of both data types. The long reads provide contiguity by spanning repeats, while the short reads provide high accuracy for error correction. This approach is common for bacterial genomes and is increasingly used for larger genomes.

The Vreelandella study used hybrid sequencing combining PacBio long reads and Illumina short reads to generate circular and complete genomes for two bacterial strains 11. The resulting genomes were complete and circular, demonstrating that hybrid assembly can produce reference-quality results for bacteria. The Mycobacterium tuberculosis study also used hybrid assemblies and found that long-read hybrid assemblies produced more genes than short-read assembly alone 7.

For larger genomes, hybrid assembly can be combined with Hi-C data for chromosome-scale scaffolding. The sainfoin project used PacBio HiFi, ONT, Illumina short reads, and Hi-C data to assemble a haplotype-resolved, chromosome-scale genome 10. The Ranunculus project used PacBio-based assembly polished with short reads and scaffolded with Hi-C data to produce pseudochromosomes 8. These projects demonstrate that hybrid strategies are valuable for complex genomes where a single platform may not suffice.

The decision to use hybrid assembly depends on the project goals and budget. For bacterial genomes, the additional cost of Illumina sequencing is modest and the benefit in assembly quality is clear. For large plant genomes, the cost of multiple data types can be substantial, and researchers must weigh the benefit of a reference-grade assembly against the budget available.

Practical Implementation Steps for Platform Selection

Selecting a platform for a genome assembly project requires a structured assessment of project goals, sample characteristics, and available resources. The following steps provide a practical framework for this decision.

First, define the assembly goals. A reference-grade genome for a new species requires higher accuracy and contiguity than a draft assembly for comparative genomics. The Ranunculus project aimed for a high-quality nuclear genome with pseudochromosomes, which required PacBio HiFi data and Hi-C scaffolding 8. A project focused on variant detection in a known species may require less assembly quality and could use a lower-cost platform.

Second, assess the genome size and complexity. Bacterial genomes of 4 to 5 megabases can be assembled with either platform, as demonstrated by the M. tuberculosis and Vreelandella studies 7 11. Larger genomes with high repeat content or polyploidy require more sequencing and may benefit from hybrid approaches 10.

Third, evaluate the DNA yield and quality. PacBio HiFi requires substantial DNA input and shearing to a target size. ONT requires less DNA and does not require shearing for native reads. Samples with limited DNA yield may be better suited to ONT.

Fourth, consider the bioinformatics capacity. PacBio HiFi data require less polishing and are compatible with a wide range of assembly tools. ONT data may require additional error correction and more careful tool selection. The M. tuberculosis study found that assembly tool choice significantly affected assembly quality 7, so laboratories should have capacity to test multiple tools.

Fifth, calculate the total cost including sequencing, bioinformatics, and researcher time. The lower per-base cost of ONT may be offset by additional compute time for polishing. The higher per-base cost of PacBio HiFi may be justified by reduced downstream effort and higher assembly quality.

Records and Measurements for Assembly Projects

Documentation is essential for reproducible genome assembly projects. Researchers should record the platform, instrument model, base-calling model, coverage, assembly tool and version, polishing steps, and quality metrics. This information allows others to reproduce the analysis and assess the reliability of the assembly.

Coverage is a critical measurement that affects assembly quality. Low coverage can result in fragmented assemblies, while excessive coverage increases cost without proportional benefit. The optimal coverage depends on the genome size, complexity, and platform accuracy. Researchers should monitor coverage during ONT runs and stop once sufficient data have been collected.

Quality metrics should be reported for every assembly. BUSCO completeness provides a standardized measure of gene content, and the Ranunculus and sainfoin projects both reported this metric 8 10. Contig N50 and assembly size provide measures of contiguity, and the M. tuberculosis study reported assembly sizes and contig counts for each tool 7. These metrics allow comparison across projects and platforms.

The NCBI Data Resources provide official descriptions of databases and search systems for depositing and accessing genome assemblies. Researchers should plan to deposit their assemblies in public databases to enable reuse and verification. The EMBL-EBI Training resources provide learning pathways for data submission and analysis.

Common Failure Patterns and How to Avoid Them

Several failure patterns recur in genome assembly projects, and understanding them can help researchers avoid costly mistakes.

Insufficient coverage is a common cause of fragmented assemblies. Researchers may underestimate the coverage needed for a complex genome or stop an ONT run too early. The real-time data output of ONT allows monitoring, but researchers must know the target coverage before starting. For PacBio runs, the fixed run time means that insufficient coverage cannot be corrected without an additional run.

Poor DNA quality is another frequent problem. PacBio HiFi requires high-molecular-weight DNA, and shearing or degradation during extraction can reduce read length. ONT is more tolerant of lower DNA quality but still benefits from high-molecular-weight DNA for ultralong reads. The sainfoin project used multiple data types, which required careful DNA extraction and quality assessment 10.

Inappropriate assembly tool selection can produce poor results even with good data. The M. tuberculosis study found significant differences in assembly completeness and contiguity among Unicycler, RagOut, and RagTag 7. Researchers should test multiple tools and select the one that performs best for their specific data type and genome.

Residual errors in ONT assemblies can lead to false variant calls. The M. tuberculosis study found that Illumina and ONT hybrid assemblies produced the highest number of SNPs 7, which may include false positives from sequencing errors. Researchers should validate variants using short-read data or orthogonal methods.

Limitations and Interpretation Boundaries

Both platforms have limitations that affect the interpretation of assembly results. PacBio HiFi reads are shorter than ONT ultralong reads, which can limit the ability to span very large repeats or structural variants. ONT reads have lower accuracy, which can complicate variant detection and gene annotation.

The choice of assembly tool can affect the final assembly more than the choice of platform. The M. tuberculosis study found that assembly completeness ranged from approximately 98.6 to 98.7 percent depending on the tool, and assembly size varied significantly 7. Researchers should not assume that a single tool works equally well for all data types.

Cost remains a barrier for large and complex genomes. The Ranunculus study noted that cost can be prohibitive for plant species with large genomes 8, and the sainfoin project used multiple data types that increased the total cost 10. Researchers should budget for the full pipeline, including bioinformatics, beyond the sequencing cost.

Safety and Regulatory Context for Sequencing Projects

Genome sequencing projects involving human or nonhuman primate samples may be subject to ethical and regulatory requirements. The cynomolgus macaque MHC study used nonhuman primate samples and required appropriate ethical approval 9. Researchers should confirm that their project has the necessary approvals before starting.

Data sharing and deposition may be subject to institutional or funder policies. The NCBI Data Resources provide official guidance on data submission and access. Researchers should plan for data deposition early in the project to avoid delays at the publication stage.

For projects involving pathogens such as Mycobacterium tuberculosis, biosafety considerations apply to sample handling and sequencing 7. Researchers should follow institutional biosafety guidelines and ensure that all personnel are trained in appropriate procedures.

Professional Escalation Criteria

Researchers should escalate to a supervisor, bioinformatics specialist, or sequencing facility manager when certain conditions arise. If the assembly quality metrics fall below the thresholds required for the project goals, such as low BUSCO completeness or excessive fragmentation, professional guidance may be needed to adjust the workflow.

If the sequencing run produces insufficient coverage or poor data quality, the sequencing facility should be consulted to determine whether the run can be repeated or whether the library preparation needs to be revised. If the assembly tools produce conflicting results, a bioinformatics specialist should be consulted to evaluate the options.

If the project involves human or nonhuman primate samples and the ethical or regulatory status is unclear, the institutional review board or ethics committee should be consulted before proceeding. If the project involves pathogens, the biosafety officer should be consulted to ensure compliance with institutional guidelines.

A Practical Decision Framework for Platform Selection Based on Project Constraints

Selecting between PacBio and Oxford Nanopore for a genome assembly project requires a structured evaluation that goes beyond comparing instrument specifications. The published projects cited in this article demonstrate that successful assemblies depend on matching platform capabilities to specific project constraints, including sample availability, timeline, bioinformatics capacity, and the intended use of the final assembly. This section provides a decision framework that researchers can apply before committing resources to either platform.

Step 1: Define the Assembly Endpoint and Quality Threshold

The first decision point is the required quality of the final assembly. A reference-grade genome intended for gene annotation, comparative genomics, or population studies demands higher contiguity and base-level accuracy than a draft assembly used for preliminary exploration. The Ranunculus project aimed for a haploid genome sequence of 2.69 Gbp with 94.1 percent complete BUSCO genes and 35,482 annotated genes, which required PacBio-based assembly polished three times with filtered short reads and scaffolded with Hi-C data 8. The sainfoin project achieved reference-grade LTR assembly index scores across all haplotypes using PacBio HiFi as the primary long-read platform 10. These examples establish that reference-grade assemblies are achievable with PacBio HiFi, but they also show that additional data types and polishing steps are often necessary.

For projects that do not require reference-grade quality, ONT may be sufficient. The Mycobacterium tuberculosis study demonstrated that ONT PromethION data could produce assemblies with approximately 98.7 percent genome completeness when using Unicycler 7. This level of completeness may be adequate for variant detection or comparative genomics in bacterial species, where the genome is small and the assembly problem is less complex than for large plant or animal genomes.

Researchers should write down the specific quality metrics their project requires before selecting a platform. These metrics include BUSCO completeness, contig N50, assembly size relative to the expected genome size, and the number of contigs. The M. tuberculosis study reported assembly sizes ranging from 4,377,642 bp with Unicycler to 4,418,574 bp with RagOut, compared to the H37Rv reference size of 4,411,532 bp 7. This variation of approximately 40,000 bp between assemblers shows that the choice of assembly tool can affect the final assembly size as much as the choice of sequencing platform.

Step 2: Assess Sample Quantity and DNA Quality

DNA yield and quality are often the most limiting factors in platform selection. PacBio HiFi library preparation requires substantial DNA input and includes a shearing step to fragment DNA to the target size. The shearing process is necessary because HiFi sequencing works best with reads in the 10 to 25 kb range, and molecules that are too long may not circularize efficiently. Samples with limited DNA yield, such as those from small organisms, archival specimens, or single cells, may not provide enough material for PacBio library construction.

ONT library preparation requires less DNA and does not require shearing for native reads. This makes ONT more suitable for projects with limited sample material. The real-time nature of ONT sequencing also allows researchers to assess data quality during the run and decide whether to continue or stop, which is an advantage when sample material is scarce and a failed run cannot be easily repeated.

DNA quality affects both platforms, but in different ways. High-molecular-weight DNA is essential for ultralong ONT reads, which are valuable for spanning large repeats and structural variants. The cynomolgus macaque MHC study used ultralong ONT reads to resolve an extended cluster of six Mafa-AG genes containing a recent duplication with a highly similar 48.5 kb block of sequence 9. This level of resolution required DNA that was not fragmented during extraction. For PacBio HiFi, DNA quality affects read length and accuracy, but the shearing step normalizes fragment size, which can partially compensate for degraded DNA.

Researchers should assess DNA quantity and quality before selecting a platform. A simple gel electrophoresis or fragment analyzer run can reveal whether high-molecular-weight DNA is present. If the DNA is degraded or the yield is low, ONT may be the only viable option. If the DNA is high quality and abundant, both platforms are feasible, and the decision shifts to cost and bioinformatics considerations.

Step 3: Evaluate Timeline and Throughput Requirements

The project timeline influences platform selection in two ways: the time required for sequencing and the time required for downstream analysis. ONT sequencing produces data in real time, with reads available for analysis as soon as they are generated. This allows researchers to monitor coverage and stop the run once sufficient data have been collected. For a small bacterial genome, this could mean completing the sequencing phase in a matter of hours. The Vreelandella study used hybrid sequencing with PacBio and Illumina data to generate circular and complete genomes for two bacterial strains 11, but a project with urgent timeline requirements might prefer ONT for its faster turnaround.

PacBio sequencing runs have a fixed duration that depends on the instrument and the sequencing mode. The run time is typically measured in hours to days, and the run cannot be stopped early without losing the remaining sequencing capacity. This makes PacBio less flexible for projects with uncertain coverage requirements or tight deadlines.

The downstream analysis timeline also differs between platforms. PacBio HiFi reads have high accuracy, which reduces the need for polishing and simplifies the assembly workflow. ONT reads may require additional error correction using short-read data or self-polishing approaches, which adds compute time and complexity. The M. tuberculosis study found that assembly tool choice significantly affected assembly quality, with Unicycler producing the highest genome completeness at approximately 98.7 percent 7. Testing multiple assemblers and selecting the best one adds time to the project, and this time should be factored into the overall timeline.

Step 4: Calculate Total Cost Including Bioinformatics

The total cost of a genome assembly project includes sequencing, library preparation, bioinformatics, and researcher time. A simple comparison of per-base sequencing costs is insufficient because the downstream analysis costs differ substantially between platforms.

PacBio HiFi has a higher per-base sequencing cost but requires less bioinformatics effort. The high accuracy of HiFi reads means that assembly tools can produce high-quality contigs without extensive polishing. The Ranunculus project polished a PacBio-based assembly three times with filtered short reads 8, which added bioinformatics cost, but this was for a complex 2.69 Gbp plant genome. For smaller or less complex genomes, the polishing requirement may be minimal.

ONT has a lower per-base sequencing cost but may require additional sequencing to compensate for lower accuracy and more bioinformatics work for error correction. The M. tuberculosis study found that Illumina and ONT hybrid assemblies produced the highest number of SNPs 7, which may include false positives from sequencing errors. Validating these variants requires additional analysis time and potentially additional sequencing.

The cost of bioinformatics training and infrastructure should also be considered. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover long-read assembly, and the nf-core Documentation describes community pipeline standards for reproducible analysis. The Bioconductor Project offers official package and workflow documentation for genomic analysis. These resources can reduce the time required to implement a robust bioinformatics pipeline, but they still require researcher time to learn and apply.

Step 5: Match Platform to Genome Complexity

Genome complexity is a critical factor that interacts with platform choice. Repetitive content, polyploidy, heterozygosity, and GC bias all influence assembly difficulty, and the optimal platform depends on which of these factors is most challenging for the target genome.

For genomes with large repeats or structural variants, ultralong ONT reads provide an advantage because they can span these regions in a single read. The cynomolgus macaque MHC study used ultralong ONT reads to resolve a recent duplication with a highly similar 48.5 kb block of sequence 9. This duplication would have been difficult to resolve with shorter reads, regardless of their accuracy.

For genomes with high heterozygosity or polyploidy, PacBio HiFi reads provide an advantage because their high accuracy allows haplotype-resolved assembly. The sainfoin project assembled a haplotype-resolved, chromosome-scale genome for a tetraploid species using PacBio HiFi, ONT, Illumina, and Hi-C data 10. The resulting assembly resolved all 28 pseudochromosomes, corresponding to four haplotypes of the seven base chromosomes. This level of resolution would be difficult to achieve with lower-accuracy reads.

For genomes with extreme GC bias, both platforms have strengths and weaknesses. PacBio HiFi reads are generated by polymerase synthesis, which can be affected by GC content. ONT reads are generated by direct translocation through a nanopore, which is less affected by GC content but has a different error profile. Researchers should consult the published literature for their specific organism or closely related species to understand which platform performs better.

Step 6: Consider Hybrid Approaches for Complex Genomes

For complex genomes, a hybrid approach that combines multiple data types may be necessary to achieve the desired assembly quality. The sainfoin project used PacBio HiFi, ONT, Illumina short reads, and Hi-C data to assemble a haplotype-resolved, chromosome-scale genome 10. The Ranunculus project used PacBio-based assembly polished with filtered short reads and scaffolded with Hi-C data 8. These projects demonstrate that hybrid approaches are valuable for complex genomes where a single platform may not suffice.

The decision to use a hybrid approach depends on the project goals and budget. For bacterial genomes, the additional cost of Illumina sequencing is modest and the benefit in assembly quality is clear. The Vreelandella study used hybrid sequencing combining PacBio long reads and Illumina short reads to generate circular and complete genomes 11. For large plant genomes, the cost of multiple data types can be substantial, and researchers must weigh the benefit of a reference-grade assembly against the budget available.

Step 7: Document the Decision and Record Key Parameters

Once a platform is selected, researchers should document the decision and record the key parameters that influenced it. This documentation is essential for reproducibility and for justifying the choice to funders, collaborators, or reviewers. The documentation should include the genome size and complexity, the required quality metrics, the DNA yield and quality, the project timeline, the total cost estimate, and the bioinformatics capacity.

The NCBI Data Resources provide official descriptions of databases and search systems for depositing and accessing genome assemblies. Researchers should plan to deposit their assemblies in public databases to enable reuse and verification. The EMBL-EBI Training resources provide learning pathways for data submission and analysis, and the The Carpentries Lessons offer foundational computing and data skills that are useful for managing sequencing data.

Common Failure Patterns in Platform Selection

Several failure patterns recur when researchers select a sequencing platform without a structured decision framework. Understanding these patterns can help researchers avoid costly mistakes.

The first failure pattern is selecting a platform based on per-base cost alone without considering the total project cost. A lower per-base cost may be offset by higher bioinformatics costs, additional sequencing for error correction, or the need for a second sequencing run to achieve sufficient coverage. The M. tuberculosis study found that assembly tool choice significantly affected assembly quality 7, and the time spent testing multiple tools adds to the project cost.

The second failure pattern is underestimating the importance of DNA quality. PacBio HiFi requires high-molecular-weight DNA, and shearing or degradation during extraction can reduce read length and accuracy. ONT is more tolerant of lower DNA quality but still benefits from high-molecular-weight DNA for ultralong reads. The cynomolgus macaque MHC study required ultralong ONT reads to resolve a complex duplication 9, which would not have been possible with degraded DNA.

The third failure pattern is assuming that a single assembly tool works equally well for all data types. The M. tuberculosis study found that Unicycler-based assemblies had significantly higher genome completeness at approximately 98.7 percent compared to RagOut and RagTag 7. Researchers should test multiple tools and select the one that performs best for their specific data type and genome characteristics.

The fourth failure pattern is not budgeting for the full pipeline, including bioinformatics. The Ranunculus study noted that cost can be prohibitive for plant species with large, complex genomes 8, and the sainfoin project used multiple data types that increased the total cost 10. Researchers should budget for the full pipeline, including bioinformatics, beyond the sequencing cost.

Professional Escalation Criteria for Platform Selection

Researchers should escalate to a supervisor, bioinformatics specialist, or sequencing facility manager when certain conditions arise during platform selection. If the DNA yield or quality is insufficient for the preferred platform, a specialist should be consulted to determine whether the extraction protocol can be optimized or whether an alternative platform should be used.

If the project timeline is too tight for the preferred platform, a sequencing facility manager should be consulted to determine whether faster turnaround is possible or whether an alternative platform should be used. If the bioinformatics capacity is insufficient for the preferred platform, a bioinformatics specialist should be consulted to determine whether additional training or support is needed.

If the project involves human or nonhuman primate samples and the ethical or regulatory status is unclear, the institutional review board or ethics committee should be consulted before proceeding. The cynomolgus macaque MHC study used nonhuman primate samples and required appropriate ethical approval 9. If the project involves pathogens, the biosafety officer should be consulted to ensure compliance with institutional guidelines.

Frequently Asked Questions

What is the main difference between PacBio and Oxford Nanopore reads?

PacBio HiFi reads are shorter, typically 10 to 25 kilobases, but have high per-read accuracy. Oxford Nanopore reads can be much longer, often exceeding 100 kilobases, but have lower per-read accuracy that varies with the base-calling model. The choice affects assembly strategy, polishing requirements, and downstream bioinformatics workload.

Which platform is better for bacterial genome assembly?

Both platforms can produce complete bacterial genomes. The Mycobacterium tuberculosis study found that hybrid assemblies with long reads produced more complete genomes than short-read assembly alone 7. The Vreelandella study used PacBio long reads combined with Illumina short reads to generate circular and complete genomes 11. The choice depends on budget, available instruments, and bioinformatics capacity.

Do I need Illumina short reads if I use PacBio HiFi?

PacBio HiFi reads have high accuracy and may not require short-read polishing for many projects. However, the Ranunculus project polished a PacBio-based assembly three times with filtered short reads 8, suggesting that short-read polishing can improve assembly quality in complex genomes. The decision depends on the genome complexity and the required assembly quality.

How much sequencing coverage do I need for a genome assembly?

The optimal coverage depends on the genome size, complexity, and platform accuracy. Bacterial genomes can be assembled with lower coverage than large plant genomes. Researchers should monitor coverage during ONT runs and stop once sufficient data have been collected. The M. tuberculosis study used data from nine isolates and compared assembly tools to determine the best approach 7.

Can I use both PacBio and Oxford Nanopore in the same project?

Yes, several projects have used both platforms in combination. The cynomolgus macaque MHC study used ultralong ONT reads and PacBio HiFi reads to assemble a complex genomic region 9. The sainfoin project used PacBio HiFi, ONT, Illumina, and Hi-C data 10. Combining platforms can overcome the limitations of either alone but increases cost and complexity.

What assembly tools work best with PacBio HiFi data?

Many assembly tools support PacBio HiFi data, but the best choice depends on the genome and project goals. The M. tuberculosis study compared Unicycler, RagOut, and RagTag and found significant differences in assembly quality 7. Researchers should test multiple tools and select the one that performs best for their specific data.

How do I assess the quality of a genome assembly?

BUSCO completeness provides a standardized measure of gene content, and the Ranunculus and sainfoin projects both reported this metric 8 10. Contig N50 and assembly size provide measures of contiguity. The M. tuberculosis study reported assembly sizes and contig counts for each tool 7.

What is the cost difference between PacBio and Oxford Nanopore?

PacBio instruments have higher capital costs, and the per-base cost is generally higher than ONT. ONT offers a lower entry point and lower per-base cost at high throughput, but the additional bioinformatics work for error correction must be factored into the total project cost. The Ranunculus study noted that cost can be prohibitive for large plant genomes 8.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.