The Cost and Throughput Trade-Offs of PacBio and Nanopore: A Budget-Focused Guide for Genome Centers

By Dr. Zubair Khalid, DVM, MS, PhD ·

The Cost and Throughput Trade-Offs of PacBio and Nanopore: A Budget-Focused Guide for Genome Centers

Key Takeaways

  • Capital Investment vs. Scalability: PacBio Sequel systems demand higher upfront capital, while Oxford Nanopore's PromethION offers a lower entry cost with modular flow cells enabling incremental scaling, making ONT more accessible for centers with variable demand or limited initial funding.
  • Per-Base Cost Dynamics: PacBio's per-GB cost decreases significantly with higher throughput, whereas ONT's cost advantage is more pronounced in hybrid assembly settings, particularly for bacterial workflows, due to lower consumables cost per isolate.
  • Accuracy vs. Read Length: PacBio HiFi excels in achieving the highest consensus accuracy, crucial for high-fidelity assemblies, while ONT generally provides the longest reads, offering superior contiguity for resolving complex genomic structures.
  • Workload-Specific Optimization: For small centers (<50 genomes/year) or those with bacterial-heavy workloads and existing Illumina infrastructure, ONT often presents a more cost-effective total cost of ownership due to lower capital and flexible consumable purchasing.
  • Total Cost of Ownership (TCO) is Paramount: For large production centers (>200 genomes/year), TCO, encompassing sequencing, library prep, computing, and bioinformatics labor, is the critical metric; benchmarking with representative samples and considering pipeline automation for hybrid assembly strategies is essential.

Genome center managers face a capital planning problem when choosing between Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) long-read sequencing platforms. The decision requires balancing per-base cost, instrument throughput, scalability, and assembly quality within fixed equipment and service budgets. This guide provides a total-cost-of-ownership framework for comparing these platforms across different operational scales, drawing on published benchmarking studies and official bioinformatics training resources.

Long-read sequencing has transformed genomics by overcoming limitations of short-read approaches, particularly the ability to resolve repeat sequences and large genomic rearrangements reliably [<a href="#ref-1">1</a>]. Both PacBio and ONT have pioneered competitive long-read platforms, with recent work focused on improving sequencing throughput and per-base accuracy [<a href="#ref-2">2</a>]. The cost structures of these two technologies differ substantially, and the optimal choice depends on your center's specific workflow, sample volume, and quality requirements.

At a Glance: Platform Cost and Throughput Comparison

The table below summarizes key differences between PacBio and ONT platforms based on published comparisons and operational considerations. These figures represent general patterns from benchmarking studies, not guaranteed pricing, as sequencing costs continue to decline with technological improvements [<a href="#ref-3">3</a>].

FactorPacBio Sequel SystemsOxford Nanopore PromethIONPractical Consideration
Capital costHigher instrument price, proprietary consumablesLower instrument entry cost, scalable flow cell purchasesONT offers lower barrier to entry for small centers
Per-base costHigher per-GB cost at low throughput, improves with scaleLower consumables cost per isolate in hybrid assembly settingsCost advantage depends on volume and workflow
Read lengthLong reads, with HiFi offering high accuracyLongest reads available, hundreds of kilobasesONT typically produces longer reads for contiguity
Consensus accuracyHighest consensus accuracy in direct comparisonsLower raw accuracy, improves with polishingPacBio HiFi excels for high-accuracy assemblies
Throughput scalabilityFixed instrument capacity, requires multiple instrumentsModular flow cells allow incremental scalingONT scales with demand more flexibly
Hybrid assembly fitWorks well with Illumina polishingMore cost-effective ONT-Illumina hybrid approachCost savings favor ONT for bacterial workflows

Published comparisons of long-read technologies applied to the same sample demonstrate that all three major long-read approaches produce highly contiguous and complete genome assemblies, but the cost associated with each method differs significantly [<a href="#ref-3">3</a>]. No single technology outperformed others in all metrics examined, meaning the ideal choice depends on the question under investigation [<a href="#ref-2">2</a>].

Understanding the Cost Components of Long-Read Sequencing

Capital Equipment Expenditure

The initial instrument purchase represents the most visible cost difference between PacBio and ONT platforms. PacBio Sequel systems require substantial capital investment, with the instrument serving as the central sequencing engine. ONT PromethION instruments have a lower entry price, and the modular flow cell design allows centers to purchase consumables incrementally based on demand.

For genome centers operating under fixed budgets, this capital cost difference affects cash flow planning. A center that anticipates steady, high-volume sequencing demand may justify the higher PacBio capital expenditure through lower per-run costs at scale. A center with variable demand or limited initial funding may prefer the lower ONT entry cost, accepting higher per-GB costs at lower volumes.

Consumables and Flow Cell Economics

PacBio sequencing requires proprietary SMRT cells and polymerase kits, with costs that scale with the number of runs performed. ONT uses flow cells that can be purchased individually, allowing centers to match consumable spending to actual demand. This difference matters for centers that experience seasonal variation in sequencing requests or that are building a user base gradually.

The published comparison of hybrid assembly approaches found that combining ONT and Illumina reads fully resolved most bacterial genomes at a lower consumables cost per isolate in the study setting [<a href="#ref-4">4</a>]. This finding suggests that for bacterial genome workflows, ONT consumables can be more economical than PacBio alternatives, particularly when hybrid assembly with short-read data is part of the standard workflow.

Per-GB Cost Dynamics

Per-gigabase sequencing costs follow different curves for the two platforms. PacBio costs per GB decrease as throughput increases because fixed run costs are distributed across more data. ONT costs per GB also decrease with volume, but the modular flow cell design means centers can avoid paying for unused capacity.

The GigaScience comparison noted that at the time of sequencing, the cost associated with each long-read method was significantly different, but continuous improvements have resulted in greater accuracy, increased throughput, and reduced costs [<a href="#ref-3">3</a>]. This rapid cost evolution means that budget models should include provisions for periodic reassessment instead of assuming static pricing.

Throughput Benchmarks and Scalability Planning

Instrument Throughput Capacity

PacBio Sequel II systems generate substantial throughput per run, with HiFi reads providing high accuracy for downstream assembly. ONT PromethION instruments offer comparable throughput potential, with the ability to run multiple flow cells simultaneously for increased capacity.

The G3 comparison of PacBio and ONT protocols found that ONT and PacBio CLR produced the longest reads, and genome contiguity was highest when assembling these datasets [<a href="#ref-2">2</a>]. This finding has direct budget implications: if your center prioritizes maximum contiguity for complex genomes, the read length advantages of ONT or PacBio CLR may justify their cost profiles.

Matching Throughput to Demand

Genome centers should model their expected monthly sequencing demand before committing to a platform. A center serving multiple research groups with diverse projects may benefit from ONT's flexible scaling, adding flow cells as demand grows. A center focused on high-throughput production of reference-quality genomes may achieve better per-GB economics with PacBio's fixed instrument capacity.

The hybrid assembly study emphasized that relative automation of the assembly process is crucial for high-throughput complete bacterial genome reconstruction, avoiding multiple bespoke filtering and data manipulation steps [<a href="#ref-4">4</a>]. When evaluating throughput, consider the bioinformatics pipeline required to convert raw reads into finished assemblies in addition to sequencing speed.

Scalability Over Time

Both platforms offer upgrade paths, but they differ in structure. PacBio upgrades typically involve instrument replacement or significant hardware modifications. ONT upgrades may involve new flow cell chemistries or software improvements that work with existing instruments.

Budget planning should account for technology obsolescence. The rapid evolution of sequencing technologies means that instruments purchased today may be superseded within a few years [<a href="#ref-3">3</a>]. Centers should establish equipment replacement funds and depreciation schedules that reflect this reality.

Total Cost of Ownership Model for Different Scales

Small Center Model: Under 50 Genomes Per Year

For a center processing fewer than 50 bacterial or small eukaryotic genomes annually, the ONT platform typically offers the most favorable total cost of ownership. The lower capital investment reduces financial risk, and the modular consumable structure prevents wasted spending during low-demand periods.

The hybrid assembly comparison found that ONT-Illumina hybrid approaches were more cost-effective for many users [<a href="#ref-2">2</a>]. For small centers that already have access to Illumina sequencing, adding ONT long-read capability provides a cost-effective path to complete genome assemblies without the higher capital commitment of PacBio.

Budget allocation for small centers should prioritize:

  • Instrument purchase or access arrangement
  • Flow cell inventory matched to quarterly demand
  • Computing resources for basecalling and assembly
  • Staff training on platform-specific workflows

Medium Center Model: 50 to 200 Genomes Per Year

Centers processing 50 to 200 genomes annually face a more complex decision. At this volume, per-GB cost differences become significant, and the choice between platforms may depend on the mix of project types.

If the center handles diverse projects including complex eukaryotic genomes, the higher consensus accuracy of PacBio HiFi may justify its cost premium [<a href="#ref-2">2</a>]. If the center primarily handles bacterial genomes and can implement hybrid assembly workflows, ONT's lower consumables cost may be more attractive [<a href="#ref-4">4</a>].

Medium centers should consider maintaining access to both platforms, either through in-house instruments or partnerships with core facilities. This dual-access strategy provides flexibility to match platform choice to project requirements.

Large Center Model: Over 200 Genomes Per Year

Large production-oriented centers processing hundreds of genomes annually should evaluate platforms based on total cost per finished assembly, not per-GB sequencing cost. This metric includes sequencing consumables, library preparation, computing resources, and bioinformatics staff time.

The GigaScience comparison emphasized that selecting the most cost-effective sequencing technology requires benchmarking different approaches applied to the same sample [<a href="#ref-3">3</a>]. Large centers should conduct their own benchmarking studies using representative samples from their typical workflows before committing to platform expansion.

At this scale, the automation and reproducibility advantages of established pipelines become critical. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-5">5</a>]. Centers should evaluate whether their chosen platform integrates with existing pipeline infrastructure.

Library Preparation Costs and Workflow Considerations

PacBio Library Preparation

PacBio HiFi library preparation requires specific DNA quality and quantity inputs. The SMRTbell library preparation process involves DNA damage repair, end repair, and adapter ligation steps. These procedures require specialized reagents and quality control checks to ensure optimal sequencing performance.

The G3 comparison noted that PacBio Sequel II HiFi assemblies had the highest consensus accuracy, even after accounting for differences in sequencing throughput [<a href="#ref-2">2</a>]. This accuracy advantage comes with library preparation requirements that may increase per-sample costs, particularly for challenging samples with limited DNA quantity.

ONT Library Preparation

ONT offers multiple library preparation protocols, including Rapid Sequencing and Ligation Sequencing kits. The Rapid Sequencing protocol requires minimal hands-on time and is suitable for samples with adequate DNA quantity. Ligation Sequencing provides higher throughput and is recommended for larger projects.

The G3 comparison found that ONT Rapid Sequencing libraries had the fewest chimeric reads and provided superior quantification of E. coli plasmids versus ligation-based libraries [<a href="#ref-2">2</a>]. This finding suggests that library preparation choice affects data quality characteristics in addition to cost.

DNA Quantity and Quality Requirements

The GigaScience comparison reported on DNA material requirements as part of the technology assessment [<a href="#ref-3">3</a>]. Different platforms and protocols require different amounts of high-molecular-weight DNA, and centers must consider whether their sample types can meet these requirements.

For samples with limited DNA quantity, the choice of library preparation protocol may be constrained. Centers should document DNA yield and quality metrics for each sample type and use this data to inform platform selection.

Assembly Quality Considerations and Their Budget Impact

Consensus Accuracy and Polishing Requirements

Assembly quality directly affects the total cost of a genome project. Lower-accuracy raw reads require additional polishing steps, which consume computing resources and staff time. The G3 comparison found that PacBio Sequel II assemblies had the highest consensus accuracy, potentially reducing downstream polishing requirements [<a href="#ref-2">2</a>].

The hybrid assembly study compared long-read-only assembly with Flye followed by short-read polishing with Pilon, finding that hybrid assembly with either PacBio or ONT reads was superior to the long-read assembly and polishing approach with respect to accuracy and completeness [<a href="#ref-4">4</a>]. This finding has budget implications: investing in hybrid assembly workflows may reduce total project costs by improving first-pass assembly quality.

Contiguity and Read Length

Genome contiguity, measured by metrics such as N50, affects the usefulness of an assembly for downstream analysis. The G3 comparison found that ONT and PacBio CLR produced the longest reads, and genome contiguity was highest when assembling these datasets [<a href="#ref-2">2</a>].

For projects where contiguity is the primary goal, such as finishing complete bacterial chromosomes or assembling complex eukaryotic genomes, the read length advantages of ONT may justify its cost profile. Centers should document contiguity requirements for each project type and match platform choice accordingly.

Completeness and Gene Content

Assembly completeness, often assessed through benchmarking universal single-copy orthologs, determines whether an assembly captures the full gene content of an organism. The GigaScience comparison reported that all three long-read technologies produced highly contiguous and complete genome assemblies of the test plant genome [<a href="#ref-3">3</a>].

Budget planning should include completeness assessment as a quality control step. Incomplete assemblies may require additional sequencing or assembly iterations, increasing project costs. Centers should establish minimum completeness thresholds and budget for potential re-sequencing.

Hybrid Assembly Strategies for Cost Optimization

Combining Long-Read and Short-Read Data

Hybrid assembly approaches combine long-read data for contiguity with short-read data for accuracy. The hybrid assembly study found that combining ONT and Illumina reads fully resolved most bacterial genomes without additional manual steps, at a lower consumables cost per isolate in the study setting [<a href="#ref-4">4</a>].

For centers with access to Illumina sequencing, hybrid assembly can reduce the amount of long-read data required for complete genome reconstruction. This strategy can lower total sequencing costs while maintaining assembly quality.

Pipeline Automation for Hybrid Assembly

The hybrid assembly study emphasized that relative automation of the assembly process is crucial for high-throughput complete bacterial genome reconstruction [<a href="#ref-4">4</a>]. Automated pipelines reduce staff time and minimize the risk of manual errors.

The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-5">5</a>]. Centers should evaluate whether existing pipelines support their chosen hybrid assembly strategy or whether custom pipeline development is required.

Cost Comparison of Assembly Strategies

The G3 comparison suggested that an ONT-Illumina hybrid approach would be more cost-effective for many users [<a href="#ref-2">2</a>]. This finding reflects the lower consumables cost of ONT sequencing combined with the accuracy benefits of Illumina short reads.

Centers should model the total cost of different assembly strategies, including:

  • Long-read-only assembly with polishing
  • Hybrid assembly with long-read and short-read data
  • HiFi assembly with PacBio high-accuracy reads

Each strategy has different cost and quality characteristics, and the optimal choice depends on project requirements.

Bioinformatics Infrastructure and Staffing Costs

Computing Resources for Basecalling and Assembly

Both PacBio and ONT platforms generate data that require substantial computing resources for processing. ONT basecalling is computationally intensive, particularly for high-accuracy basecalling models. PacBio HiFi data also requires significant computing resources for assembly.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help staff develop the skills needed for long-read data processing [<a href="#ref-6">6</a>]. Centers should budget for staff training as part of platform adoption.

Pipeline Development and Maintenance

Genome assembly pipelines require ongoing development and maintenance to incorporate new tools and best practices. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-5">5</a>]. Centers should evaluate whether to adopt community pipelines or develop custom workflows.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-7">7</a>]. These resources can support pipeline development and quality control efforts.

Staff Training Requirements

The Carpentries Lessons provide foundational computing, data, shell, Git, and programming training context [<a href="#ref-8">8</a>]. Staff working with long-read sequencing data need proficiency in command-line tools, scripting, and version control.

The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training [<a href="#ref-9">9</a>]. Centers should budget for ongoing staff development to maintain expertise in rapidly evolving analysis methods.

Quality Control Metrics and Their Cost Implications

Sequencing Run Quality Metrics

Both platforms generate quality metrics that should be monitored for each sequencing run. These metrics include read length distributions, per-base quality scores, and throughput per flow cell or SMRT cell.

The G3 comparison reported on sequencing throughput and per-base accuracy for different protocols [<a href="#ref-2">2</a>]. Centers should establish quality thresholds and document performance for each platform and protocol.

Assembly Quality Assessment

Assembly quality should be assessed using standardized metrics, including contiguity, base accuracy, and completeness. The GigaScience comparison reported on these metrics for assemblies generated with different technologies [<a href="#ref-3">3</a>].

The NCBI Data Resources provide official descriptions of NCBI databases, search systems, sequence resources, and analysis services [<a href="#ref-10">10</a>]. Centers should use these resources for assembly validation and submission.

Quality Control Costs

Quality control adds to the total cost of genome projects. Centers should budget for:

  • Sequencing run QC metrics monitoring
  • Assembly quality assessment
  • Sample-level QC for DNA quantity and quality
  • Bioinformatics pipeline validation

These costs are often overlooked in budget planning but are essential for producing reliable results.

Common Failure Patterns and Cost Overruns

Underestimating Computing Requirements

Genome centers frequently underestimate the computing resources required for long-read assembly. Basecalling, assembly, and polishing steps can require substantial CPU and memory resources, particularly for eukaryotic genomes.

Centers should benchmark their computing requirements using representative datasets before committing to a platform. The Galaxy Training Network provides accessible workflow training that can help staff understand computational requirements [<a href="#ref-6">6</a>].

Insufficient DNA Quality

Long-read sequencing requires high-molecular-weight DNA, and poor DNA quality leads to reduced throughput and shorter reads. The GigaScience comparison reported on DNA material requirements as part of the technology assessment [<a href="#ref-3">3</a>].

Centers should implement DNA quality assessment protocols before library preparation to avoid wasted sequencing runs. This quality control step adds cost but prevents more expensive failures.

Inadequate Staff Training

Long-read sequencing and assembly require specialized skills that differ from short-read workflows. The Carpentries Lessons provide foundational computing training that can support staff development [<a href="#ref-8">8</a>].

Centers should budget for staff training as an ongoing expense, not a one-time cost. The rapid evolution of sequencing technologies means that skills need regular updating.

Ignoring Pipeline Maintenance

Assembly pipelines require ongoing maintenance to incorporate new tools and fix bugs. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-5">5</a>].

Centers should allocate staff time for pipeline maintenance and version control. Neglecting this maintenance leads to reproducibility problems and wasted sequencing runs.

Limitations and Interpretation Boundaries

Platform Comparisons Are Time-Sensitive

Published platform comparisons reflect the technology state at the time of the study. The GigaScience comparison noted that continuous improvements have resulted in greater accuracy, increased throughput, and reduced costs [<a href="#ref-3">3</a>].

Budget models should include provisions for periodic reassessment of platform choices. A platform that is cost-effective today may not be optimal in two years.

Cost Data Vary by Setting

Published cost comparisons reflect specific laboratory settings and may not generalize to all centers. The hybrid assembly study noted that the cost advantage of ONT was specific to the study setting [<a href="#ref-4">4</a>].

Centers should conduct their own cost modeling using local pricing and workflow characteristics. Published comparisons provide useful benchmarks but should not replace site-specific analysis.

Quality Metrics Trade Off Against Each Other

No single technology outperforms others in all metrics examined [<a href="#ref-2">2</a>]. Centers must prioritize quality metrics based on project requirements, accepting trade-offs between contiguity, accuracy, and cost.

The ideal choice of long-read technology may depend on several factors including the question or hypothesis under examination [<a href="#ref-2">2</a>]. Centers should document decision criteria for each project type.

Professional Escalation Criteria

Genome center managers should escalate platform decisions to institutional leadership when:

  • Projected sequencing demand exceeds current instrument capacity by more than 50 percent
  • Multiple research groups report assembly quality issues that trace to platform limitations
  • Budget forecasts show consumables costs exceeding allocated funds by more than 20 percent
  • Staff training requirements exceed available time and resources
  • Institutional strategic plans require capabilities not available with current platforms

Escalation should include a written analysis of platform options, cost projections, and quality metrics from representative benchmarking studies.

A Practical Decision Framework for Platform Selection Based on Workload Profiles

Genome center managers often struggle to translate published cost comparisons into actionable procurement decisions because benchmarking studies reflect specific laboratory settings and time points. The GigaScience comparison of long-read technologies applied to the same plant genome noted that continuous improvements have resulted in greater accuracy, increased throughput, and reduced costs [<a href="#ref-3">3</a>]. This rapid evolution means that a decision framework based on workload profiles, instead of static price lists, provides more durable guidance for budget planning.

Defining Workload Profiles for Your Center

Before comparing platform costs, document the actual mix of projects your center handles. The G3 comparison of PacBio and ONT protocols across bacteria and fruit flies found that no single technology outperformed others in all metrics examined [<a href="#ref-2">2</a>]. This finding implies that the optimal platform depends on the specific characteristics of your sequencing requests.

Create a workload profile table that categorizes each project type by four variables:

Project TypeGenome SizeTarget CoverageQuality RequirementMonthly Volume
Bacterial isolates3 to 12 Mb50x to 100xComplete chromosome10 to 50
Small eukaryotic genomes50 to 500 Mb30x to 60xChromosome-scale scaffold2 to 10
Large eukaryotic genomes1 to 30 Gb20x to 40xHigh contiguity0.5 to 3
Targeted ampliconsVariable100x to 1000xHigh accuracy5 to 50

The hybrid assembly study of 20 bacterial isolates from the Enterobacteriaceae family demonstrated that complete genome reconstruction is relevant for precise understanding of antimicrobial resistance epidemiology [<a href="#ref-4">4</a>]. Centers serving clinical or public health users should expect bacterial projects to dominate their workload and plan platform selection accordingly.

Matching Platform Strengths to Workload Profiles

The G3 comparison reported that PacBio Sequel II assemblies had the highest consensus accuracy, even after accounting for differences in sequencing throughput [<a href="#ref-2">2</a>]. This accuracy advantage makes PacBio HiFi the preferred choice when your workload includes projects requiring high base-level accuracy, such as variant confirmation or clinical reporting.

The same comparison found that ONT and PacBio CLR produced the longest reads, and genome contiguity was highest when assembling these datasets [<a href="#ref-2">2</a>]. If your workload emphasizes maximum contiguity for complex genomes with large repeat structures, the read length advantages of ONT or PacBio CLR may justify their cost profiles.

The hybrid assembly study found that combining ONT and Illumina reads fully resolved most bacterial genomes without additional manual steps, at a lower consumables cost per isolate in the study setting [<a href="#ref-4">4</a>]. For centers with bacterial-heavy workloads and existing Illumina access, ONT provides a cost-effective path to complete genome assemblies.

Building a Decision Matrix for Platform Selection

Construct a decision matrix that scores each platform against your workload profile. Use a simple scoring system from 1 to 5 for each criterion, weighted by importance to your center.

CriterionWeightPacBio ScoreONT ScoreWeighted PacBioWeighted ONT
Consensus accuracy0.25531.250.75
Read length for contiguity0.20350.601.00
Consumables cost per isolate0.20340.600.80
Capital cost flexibility0.15250.300.75
Hybrid assembly compatibility0.10450.400.50
Automation and pipeline maturity0.10430.400.30
Total1.003.554.10

This example shows an ONT advantage for a bacterial-focused center with existing Illumina access. Adjust the weights and scores based on your specific workload profile and local pricing. The GigaScience comparison emphasized that selecting the most cost-effective sequencing technology requires benchmarking different approaches applied to the same sample [<a href="#ref-3">3</a>]. Your decision matrix should be validated with local benchmarking data.

Implementing a Pilot Benchmarking Protocol

Before committing to a platform purchase or service contract, run a pilot benchmarking study using representative samples from your actual workload. The GigaScience comparison of three long-read technologies applied to the same plant genome provides a model for this approach [<a href="#ref-3">3</a>]. The study generated sequencing data using Pacific Biosciences, Oxford Nanopore Technologies, and BGI technologies for the same sample, then compared assemblies for contiguity, base accuracy, and completeness, as well as sequencing costs and DNA material requirements.

Design your pilot study with these steps:

  1. Select three to five samples that represent your most common project types
  2. Prepare DNA from each sample using your standard extraction protocols
  3. Sequence each sample on both platforms using recommended protocols
  4. Assemble the data using the same assembler versions and parameters
  5. Compare assemblies using standardized metrics for contiguity, accuracy, and completeness
  6. Document actual consumables usage, staff time, and computing resources for each platform

The G3 comparison used whole-genome sequencing data produced by three PacBio protocols and two ONT protocols to compare assemblies of Escherichia coli and Drosophila ananassae [<a href="#ref-2">2</a>]. This multi-protocol approach revealed that library preparation choice affects data quality characteristics in addition to cost. Your pilot should include the specific library preparation protocols you plan to use in production.

Establishing a Record System for Cost and Quality Tracking

A durable record system that links sequencing costs to assembly quality outcomes provides the evidence base for ongoing platform decisions. The hybrid assembly study emphasized that relative automation of the assembly process is crucial for high-throughput complete bacterial genome reconstruction, avoiding multiple bespoke filtering and data manipulation steps [<a href="#ref-4">4</a>]. Your record system should capture both sequencing and downstream analysis costs.

Create a spreadsheet or database with these fields for each sequencing project:

FieldExample ValuePurpose
Project identifierBAC-2024-001Unique tracking
Sample typeBacterial isolateWorkload categorization
PlatformONT PromethIONCost attribution
Library protocolLigation SequencingProtocol comparison
Flow cell or SMRT cell count2Consumables tracking
Total raw bases generated8.5 GbThroughput measurement
Total Q20 bases7.2 GbQuality assessment
Staff hours for library prep3.5Labor cost tracking
Staff hours for assembly6.0Bioinformatics cost tracking
Computing hours48 CPU-hoursInfrastructure cost tracking
Assembly N504.2 MbContiguity metric
Assembly completeness99.1 percentCompleteness metric
Consensus accuracyQ40Accuracy metric
Total consumables cost420 USDDirect cost tracking
Total labor cost285 USDIndirect cost tracking
Cost per finished assembly705 USDTotal cost metric

The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that can support assembly validation and submission [<a href="#ref-10">10</a>]. Use these resources to verify assembly quality before recording final metrics in your tracking system.

Calculating Cost per Finished Assembly

The most meaningful cost metric for budget planning is cost per finished assembly, not per-GB sequencing cost. This metric includes sequencing consumables, library preparation, computing resources, and bioinformatics staff time. The GigaScience comparison reported on sequencing costs and DNA material requirements as part of the technology assessment [<a href="#ref-3">3</a>], but these published figures should be supplemented with your own cost tracking.

Calculate cost per finished assembly using this formula:

Total project cost equals consumables cost plus labor cost plus computing cost plus quality control cost.

Consumables cost includes library preparation kits, flow cells or SMRT cells, and any additional reagents. Labor cost includes staff time for DNA quality assessment, library preparation, sequencing operation, and bioinformatics analysis. Computing cost includes CPU and memory resources for basecalling, assembly, and polishing. Quality control cost includes sample QC, sequencing run QC, and assembly validation.

The hybrid assembly study found that combining ONT and Illumina reads fully resolved most bacterial genomes without additional manual steps, at a lower consumables cost per isolate in the study setting [<a href="#ref-4">4</a>]. This finding suggests that hybrid assembly can reduce total project costs by improving first-pass assembly quality and reducing the need for additional sequencing or manual finishing steps.

Troubleshooting Cost Overruns with Root Cause Analysis

When cost per finished assembly exceeds your budget threshold, use a structured root cause analysis to identify the source of the overrun. Common failure patterns include:

Insufficient DNA quality leading to wasted sequencing runs. The GigaScience comparison reported on DNA material requirements as part of the technology assessment [<a href="#ref-3">3</a>]. Poor DNA quality reduces throughput and read length, increasing the amount of sequencing required to achieve target coverage. Implement DNA quality assessment protocols before library preparation to avoid wasted runs.

Library preparation failures requiring repeated preparation. The G3 comparison found that ONT Rapid Sequencing libraries had the fewest chimeric reads in addition to superior quantification of E. coli plasmids versus ligation-based libraries [<a href="#ref-2">2</a>]. Different library protocols have different failure rates and cost profiles. Track library preparation success rates by protocol and sample type to identify problem areas.

Underestimating computing requirements for assembly. Basecalling, assembly, and polishing steps can require substantial CPU and memory resources, particularly for eukaryotic genomes. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help staff understand computational requirements [<a href="#ref-6">6</a>]. Benchmark your computing needs using representative datasets before committing to infrastructure investments.

Inadequate staff training leading to inefficient workflows. The Carpentries Lessons provide foundational computing, data, shell, Git, and programming training context [<a href="#ref-8">8</a>]. Staff working with long-read sequencing data need proficiency in command-line tools, scripting, and version control. Budget for staff training as an ongoing expense, not a one-time cost.

Establishing Performance Thresholds and Escalation Criteria

Define performance thresholds for each platform and workflow based on your pilot benchmarking data and published comparisons. The G3 comparison reported on sequencing throughput and per-base accuracy for different protocols [<a href="#ref-2">2</a>]. Use these benchmarks to establish minimum acceptable performance for production runs.

Set escalation criteria that trigger a formal review of platform choice or workflow configuration:

MetricThresholdAction
Cost per finished assemblyExceeds budget by 20 percentReview workflow for inefficiencies
Assembly N50Below target by 30 percentEvaluate read length and coverage
Assembly completenessBelow 95 percent for bacteriaAssess coverage and assembly parameters
Library preparation success rateBelow 80 percentReview DNA quality and protocol
Sequencing run failure rateAbove 15 percentInvestigate instrument or reagent issues
Staff hours per projectExceeds estimate by 50 percentReview training and automation

The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-5">5</a>]. Adopting standardized pipelines can reduce variability in assembly quality and staff time, making cost tracking more predictable.

Integrating Platform Decisions with Bioinformatics Training

The choice between PacBio and ONT affects also sequencing costs but also the bioinformatics skills your staff need. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training [<a href="#ref-9">9</a>]. Staff working with ONT data need proficiency in basecalling, which is computationally intensive and requires familiarity with different basecalling models. Staff working with PacBio HiFi data need skills in circular consensus sequencing analysis and HiFi-specific assembly tools.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-7">7</a>]. These resources can support pipeline development and quality control efforts for both platforms. Budget for staff training as part of platform adoption, and include training costs in your total cost of ownership model.

The Galaxy Training Network provides accessible workflow training that can help staff develop the skills needed for long-read data processing [<a href="#ref-6">6</a>]. Centers should evaluate whether their chosen platform integrates with existing pipeline infrastructure and whether staff have the skills to operate and maintain those pipelines.

Conducting Annual Platform Reviews

The rapid evolution of sequencing technologies means that platform decisions require periodic reassessment. The GigaScience comparison noted that continuous improvements have resulted in greater accuracy, increased throughput, and reduced costs [<a href="#ref-3">3</a>]. Schedule an annual platform review that includes:

  1. Updated pricing from both manufacturers
  2. New protocol and chemistry releases
  3. Published benchmarking studies from the past year
  4. Your center's cost and quality tracking data
  5. Changes in workload profile or project mix
  6. Staff feedback on workflow efficiency

The G3 comparison suggested that an ONT-Illumina hybrid approach would be more cost-effective for many users [<a href="#ref-2">2</a>]. This recommendation may change as PacBio introduces new chemistries or ONT improves raw read accuracy. Your annual review should evaluate whether your platform choice remains optimal given current technology and pricing.

Documenting Decision Criteria for Institutional Buy-In

Genome center managers often need to justify platform decisions to institutional leadership. Document your decision criteria, workload profile, pilot benchmarking results, and cost tracking data in a written analysis. The hybrid assembly study emphasized that relative automation of the assembly process is crucial for high-throughput workflows [<a href="#ref-4">4</a>]. Include automation and staff efficiency considerations in your documentation.

The NCBI Data Resources provide official descriptions of databases and analysis services that can support assembly validation and submission [<a href="#ref-10">10</a>]. Reference these resources in your documentation to demonstrate that your quality assessment procedures follow established standards.

Practical Implementation Steps

Implement this decision framework with these steps:

  1. Create a workload profile table for your center
  2. Score each platform against your workload using the decision matrix
  3. Run a pilot benchmarking study with representative samples
  4. Establish a cost and quality tracking spreadsheet or database
  5. Calculate cost per finished assembly for each platform
  6. Set performance thresholds and escalation criteria
  7. Schedule an annual platform review
  8. Document decision criteria for institutional stakeholders

The GigaScience comparison proposed updating their technology comparison regularly with reports on significant iterations of the sequencing technologies [<a href="#ref-3">3</a>]. Your center should adopt a similar approach, treating platform selection as an ongoing process instead of a one-time decision.

Common Failure Patterns in Platform Selection

Genome centers commonly make several mistakes when selecting between PacBio and ONT platforms:

Basing decisions on per-GB cost alone. Per-GB cost does not capture the full cost of producing a finished assembly. The hybrid assembly study found that combining ONT and Illumina reads fully resolved most bacterial genomes without additional manual steps [<a href="#ref-4">4</a>]. The cost per finished assembly is the metric that matters for budget planning.

Ignoring library preparation costs and failure rates. Library preparation costs vary significantly between protocols and platforms. The G3 comparison found that ONT Rapid Sequencing libraries had the fewest chimeric reads versus ligation-based libraries [<a href="#ref-2">2</a>]. Track library preparation success rates to identify hidden costs.

Underestimating bioinformatics staff time. Assembly, polishing, and quality assessment require substantial staff time. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-5">5</a>]. Adopting standardized pipelines can reduce staff time and improve reproducibility.

Failing to account for technology obsolescence. Sequencing technologies evolve rapidly, and instruments purchased today may be superseded within a few years [<a href="#ref-3">3</a>]. Establish equipment replacement funds and depreciation schedules that reflect this reality.

Neglecting to validate published comparisons with local data. Published cost comparisons reflect specific laboratory settings and may not generalize to all centers [<a href="#ref-4">4</a>]. Conduct your own benchmarking studies using representative samples from your typical workflows before committing to platform expansion.

Frequently Asked Questions

How do PacBio and ONT per-GB costs compare for bacterial genome projects?

Published comparisons indicate that ONT consumables can be more cost-effective for bacterial genome projects, particularly when combined with Illumina short reads in hybrid assembly workflows [<a href="#ref-4">4</a>]. The G3 comparison suggested that an ONT-Illumina hybrid approach would be more cost-effective for many users [<a href="#ref-2">2</a>]. However, exact cost differences depend on local pricing, sample volume, and workflow efficiency. Centers should model their specific costs instead of relying on published averages.

What is the minimum throughput needed to justify PacBio instrument purchase?

There is no universal threshold for justifying PacBio instrument purchase. The decision depends on the center's projected sequencing demand, available capital, and alternative access options such as core facilities or service providers. Centers should model total cost of ownership over a five-year period, including instrument depreciation, consumables, staffing, and computing resources. The GigaScience comparison noted that continuous improvements have resulted in reduced costs [<a href="#ref-3">3</a>], so current pricing should be obtained from the manufacturer.

Does ONT require more polishing than PacBio for comparable assembly accuracy?

Published comparisons found that PacBio Sequel II assemblies had the highest consensus accuracy, even after accounting for differences in sequencing throughput [<a href="#ref-2">2</a>]. This suggests that ONT assemblies may require more polishing to achieve comparable accuracy. However, hybrid assembly with Illumina short reads can improve accuracy for both platforms [<a href="#ref-4">4</a>]. The optimal polishing strategy depends on the assembly quality requirements for each project.

How should a genome center budget for computing infrastructure?

Computing infrastructure costs should be modeled based on the center's projected data volume and analysis requirements. Basecalling, assembly, and polishing steps require substantial CPU and memory resources. The Galaxy Training Network provides accessible workflow training that can help staff understand computational requirements [<a href="#ref-6">6</a>]. Centers should benchmark their computing needs using representative datasets before committing to infrastructure investments.

What are the hidden costs in long-read genome assembly projects?

Hidden costs include library preparation failures, insufficient DNA quality leading to wasted runs, staff training time, pipeline development and maintenance, and quality control assessments. The GigaScience comparison reported on DNA material requirements as part of the technology assessment [<a href="#ref-3">3</a>]. Centers should build contingency budgets for these costs, typically 10 to 20 percent of the total project budget.

Can a center switch between PacBio and ONT platforms without major workflow changes?

Switching platforms requires changes to library preparation protocols, sequencing workflows, and bioinformatics pipelines. The hybrid assembly study emphasized that relative automation of the assembly process is crucial for high-throughput workflows [<a href="#ref-4">4</a>]. Centers should evaluate the cost of workflow changes, including staff retraining and pipeline redevelopment, before switching platforms.

How does read length affect assembly cost?

Longer reads improve genome contiguity, potentially reducing the amount of sequencing required for complete assemblies. The G3 comparison found that ONT and PacBio CLR produced the longest reads, and genome contiguity was highest when assembling these datasets [<a href="#ref-2">2</a>]. However, longer reads may come with higher per-base costs or lower accuracy. Centers should evaluate whether contiguity improvements justify the cost premium for their specific projects.

What quality metrics should be tracked for budget planning?

Centers should track sequencing run quality metrics, including read length distributions, per-base quality scores, and throughput per flow cell or SMRT cell. Assembly quality metrics include contiguity, base accuracy, and completeness. The NCBI Data Resources provide official descriptions of databases and analysis services that can support quality assessment [<a href="#ref-10">10</a>]. These metrics should be linked to cost data to identify the most cost-effective workflows.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Next-Generation Sequencing and Emerging Technologies.](https://pubmed.ncbi.nlm.nih.gov/38692283). Seminars in thrombosis and hemostasis, 2024. [2] [Comparison of long-read sequencing technologies in interrogating bacteria and fly genomes.](https://pubmed.ncbi.nlm.nih.gov/33768248). G3 (Bethesda, Md.), 2021. [3] [Comparison of long-read methods for sequencing and assembly of a plant genome.](https://pubmed.ncbi.nlm.nih.gov/33347571). GigaScience, 2020. [4] [Comparison of long-read sequencing technologies in the hybrid assembly of complex bacterial genomes.](https://pubmed.ncbi.nlm.nih.gov/31483244). Microbial genomics, 2019. [5] [nf-core Documentation](https://nf-co.re/docs). nf-core. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [10] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.