# Troubleshooting Low Data Yield in Long-Read Sequencing: Common Causes and Fixes for PacBio and Nanopore

Low data yield in long-read sequencing is a practical problem that wastes flow cells, consumes budget, and delays project completion. This article gives biology students, researchers, and laboratory professionals a systematic method for diagnosing why a PacBio or Oxford Nanopore run produced less data than expected and for correcting the underlying causes before the next run. The approach covers library preparation, sequencing run parameters, and basecalling, with attention to the distinct failure modes of each platform.

## At a Glance

The table below summarizes the most common causes of low yield, the platform most often affected, the primary symptom, and the first corrective action to take.

| Cause Category | Platform Most Affected | Primary Symptom | First Corrective Action |
|---|---|---|---|
| Insufficient or degraded high molecular weight DNA input | Both PacBio and Nanopore | Low number of reads, short read lengths, low total bases | Quantify DNA with fluorometry, check integrity by pulsed-field gel electrophoresis or equivalent, re-extract if degraded |
| Suboptimal adapter ligation or motor protein loading | Nanopore | Low pore occupancy, many unproductive pores, low read count | Verify adapter concentration, check ligation reaction time and temperature, confirm motor protein is added |
| Flow cell quality or age | Both PacBio and Nanopore | Low number of active pores or zero-mode waveguides, high failure rate | Check flow cell QC metrics before loading, use fresh flow cells, run a control library to assess pore health |
| Basecalling parameters or software version | Both PacBio and Nanopore | Low yield after basecalling despite adequate raw signal | Update basecaller, check quality score thresholds, confirm adapter trimming settings |
| Library concentration overestimation | Both PacBio and Nanopore | Fewer reads than expected, low loading density | Use fluorometric quantification instead of spectrophotometry, verify with qPCR if available |
| PCR amplification bias or overcycling | PacBio | Uneven coverage, low yield for GC-rich regions | Reduce PCR cycles, use PCR-free library prep when input allows |

## Understanding Long-Read Sequencing Yield

Yield in long-read sequencing means the total number of usable bases produced from a run. For PacBio, yield is measured in gigabases of HiFi reads or continuous long reads. For Nanopore, yield is measured in gigabases of basecalled sequence passing quality thresholds. Yield depends on three linked factors: the number of productive sequencing units, the read length distribution, and the accuracy of basecalling.

The number of productive sequencing units is set by the flow cell. A PacBio SMRT cell contains zero-mode waveguides that each hold one polymerase molecule. A Nanopore flow cell contains pores that each pass one DNA strand. If the library does not load efficiently, many of these units remain dark and produce no data. If the library loads but the polymerase or pore fails quickly, the reads are short and total yield drops.

Read length distribution matters because long-read sequencing is purchased per flow cell, not per base. A run that produces 100,000 reads averaging 10 kilobases yields 1 gigabase. The same run producing reads averaging 20 kilobases yields 2 gigabases. Read length is governed by DNA fragment length in the library, by the processivity of the polymerase or the stability of the pore, and by the sequencing run duration.

Basecalling accuracy affects usable yield. Raw signal that cannot be converted to high-confidence sequence is discarded or down-weighted. For PacBio HiFi, circular consensus sequencing generates multiple passes of the same molecule, and the basecaller combines them into one accurate read. For Nanopore, the basecaller converts ionic current changes into nucleotide sequence, and quality scores determine which reads pass filtering.

## Library Preparation as the Primary Yield Determinant

Library preparation is where most yield problems originate. The library is the physical material loaded onto the flow cell, and its quality sets the ceiling for the entire run. A perfect run cannot rescue a poor library.

### High Molecular Weight DNA Input Quantity and Quality

Long-read sequencing requires high molecular weight DNA of adequate purity and integrity. Plant material is a useful example because leaves contain carbohydrates and secondary metabolites that contaminate DNA and impair downstream applications. Many extraction protocols require large amounts of starting material and still produce fragmented DNA, which makes sequencing suboptimal in read length and data yield. A protocol developed for plant high molecular weight DNA extraction from only 0.1 grams of starting material, completed in about 2.5 hours, successfully produced usable DNA from four plant families including Orchidaceae, Poaceae, Brassicaceae, and Asteraceae. For recalcitrant species, one additional purification step delivered a clean sample suitable for Oxford Nanopore PromethION sequencing, with or without a short fragment depletion kit.

The practical lesson is that input material quantity is not the only variable. Extraction method, purification steps, and fragment length assessment all determine whether the library will produce long reads. Researchers should measure DNA quantity with a fluorometric method instead of spectrophotometry, because spectrophotometry overestimates DNA concentration when contaminants absorb at 260 nanometers. DNA integrity should be checked by pulsed-field gel electrophoresis or an automated fragment analyzer system that resolves molecules above 50 kilobases.

For phage DNA, the challenge is different. Phage propagation can be difficult when a sensitive bacterial host is unavailable, and DNA extraction from induced bacterial cultures can yield very low amounts of genomic DNA. Many studies use tagmentation for amplification-free quantitative sequencing, but this technique loses phage genome ends and creates coverage bias. PCR-free sequencing is often recommended for unbiased phage genome characterization, yet sequencing very low DNA quantities without PCR is challenging, and library kit manufacturers guarantee results only with relatively high DNA inputs. A study testing phage genomic DNA with very low starting material found that high quality sequencing was achievable with DNA inputs 1000-fold lower than manufacturer recommendations, using both Illumina short-read and Nanopore long-read technologies.

The practical implication is that low input DNA does not automatically mean low yield. The library preparation method must match the input quantity. Researchers should not assume that falling below manufacturer recommendations guarantees failure, but they should also not assume that low input will produce the same yield as standard input. The correct approach is to test the library preparation method with the actual input quantity before committing an expensive flow cell.

### DNA Purity and Contaminant Effects

Contaminants in the DNA sample reduce yield by interfering with enzymatic steps in library preparation and with the sequencing chemistry itself. Common contaminants include polysaccharides, polyphenols, proteins, and residual extraction reagents such as ethanol or guanidine salts.

For plant samples, carbohydrates and secondary metabolites are the main concern. These compounds can coprecipitate with DNA during extraction and inhibit ligases and polymerases. The purification step described in the plant extraction protocol addresses this by adding a clean-up that removes the inhibitory compounds. For animal tissues, residual protein and lipid can similarly reduce enzyme activity.

The practical test for purity is the absorbance ratio at 260/280 nanometers, which should be near 1.8 for pure DNA, and at 260/230 nanometers, which should be above 2.0. Values below these thresholds indicate contamination that may reduce yield. However, absorbance ratios do not detect all inhibitors, and a clean absorbance profile does not guarantee that the library preparation enzymes will work. A more reliable test is to run a small-scale test library preparation with a fraction of the sample and measure the library yield before committing the full sample.

### Fragment Length Distribution

Read length in long-read sequencing is limited by the fragment length in the library. If the DNA is sheared to 10 kilobases during extraction, the sequencing run cannot produce 50 kilobase reads. Fragment length distribution should be measured after extraction and again after library preparation.

For PacBio HiFi, circular consensus sequencing reads the same molecule multiple times. The polymerase can process a circular template, and the read length is determined by the insert size and the number of passes. If the insert is short, the HiFi read is short. If the insert is long but the polymerase falls off early, the read is short.

For Nanopore, the read length is determined by the DNA fragment length and by the stability of the pore. A fragment that is 100 kilobases long can produce a 100 kilobase read if the pore remains active. However, if the DNA has nicks or damage, the pore may stall or the strand may break, producing shorter reads.

The practical management decision is to assess fragment length at two points: after DNA extraction and after library preparation. If the fragment length drops during library preparation, the protocol is causing shearing. Common causes include pipetting with narrow bore tips, vortexing, repeated freeze thaw cycles, and prolonged incubation at elevated temperatures. The fix is to use wide bore tips, mix by gentle inversion, and minimize handling steps.

### Adapter Ligation and Motor Protein Loading for Nanopore

Nanopore sequencing requires that each DNA fragment has adapters ligated to both ends. One adapter contains the motor protein that unwinds the double stranded DNA and feeds one strand through the pore. If adapter ligation is inefficient, many fragments lack the motor protein and cannot be sequenced. If the motor protein is not properly loaded, the pore may capture the DNA but fail to process it, producing no signal or a signal that cannot be basecalled.

The nCATS method for targeted nanopore sequencing uses Cas9 to cleave chromosomal DNA at specific sites and then ligates adapters to the cut ends. This approach requires about 3 micrograms of genomic DNA and can target many loci in a single reaction. The method achieved median sequencing coverage of 675 fold on a MinION flow cell and 34 fold on the smaller Flongle flow cell. The success of this method depends on efficient Cas9 cleavage and adapter ligation, which are the same enzymatic steps that determine yield in whole genome nanopore sequencing.

The practical checks for adapter ligation problems include measuring the library concentration after ligation, checking the fragment length distribution after ligation to confirm that adapters are attached, and running a small test on a single flow cell before committing multiple flow cells. If the library concentration drops substantially after ligation, the ligation reaction is failing. If the fragment length distribution shifts to shorter sizes, the DNA is being degraded during the ligation step.

## Sequencing Run Parameters That Control Yield

Run parameters determine how much data the flow cell can produce once the library is loaded. These parameters include loading concentration, run duration, temperature, and software settings.

### Loading Concentration and Pore Occupancy

Loading concentration is the amount of library added to the flow cell. If too little library is loaded, few pores capture a DNA molecule and yield is low. If too much library is loaded, multiple DNA molecules may compete for the same pore, or the pore may be blocked by an excess of adapters or contaminants.

For Nanopore, the optimal loading concentration depends on the flow cell type and the library preparation method. The nCATS study loaded libraries onto MinION and Flongle flow cells and achieved high coverage, demonstrating that the loading concentration must be adjusted for the flow cell format. A Flongle has far fewer pores than a MinION, so the same library concentration that works for a MinION may overload a Flongle.

For PacBio, the loading concentration determines how many zero-mode waveguides contain a polymerase bound to a template. If the concentration is too low, many waveguides are empty. If too high, multiple polymerases may bind in the same waveguide, and the signal becomes confounded.

The practical approach is to follow the manufacturer recommended loading concentration for the specific flow cell and library type, then adjust based on the observed pore occupancy in the first minutes of the run. Most sequencing software reports pore occupancy in real time. If occupancy is below the expected range, the next run should use a higher loading concentration. If occupancy is above the expected range but yield is still low, the problem is elsewhere.

### Run Duration and Read Length

Run duration affects yield in two ways. Longer runs allow more time for the polymerase or pore to produce data, but they also allow more time for the enzyme to fail. The optimal run duration balances these factors.

For PacBio HiFi, the polymerase processivity determines how long a single molecule can be read. The circular consensus sequencing approach reads the same molecule multiple times, and the number of passes determines the accuracy of the HiFi read. A longer run duration allows more passes but also allows the polymerase to fall off the template. The basecaller determines when to stop reading a molecule based on the quality of the signal.

For Nanopore, the pore can theoretically read a very long DNA fragment in a single pass. The run duration is limited by the stability of the pore and by the depletion of the motor protein. Most Nanopore runs are set for a fixed duration, typically 24 to 72 hours, and the yield increases with duration until the pores fail.

The practical decision is to monitor the run in real time and stop the run when the rate of new data production drops below a useful threshold. Continuing a run that is producing no new data wastes time and does not increase yield. Most sequencing software provides a plot of cumulative yield over time, and the curve should plateau when the run is complete.

### Temperature and Environmental Stability

Temperature affects enzyme activity and pore stability. Both PacBio and Nanopore instruments control temperature internally, but external temperature fluctuations can affect the instrument performance. The sequencing instrument should be placed in a room with stable temperature, away from direct sunlight, heating vents, and air conditioning drafts.

For Nanopore, temperature affects the rate of DNA translocation through the pore. Higher temperatures increase translocation speed, which can reduce read accuracy. Lower temperatures slow translocation, which can increase accuracy but reduce throughput. The instrument software controls temperature, and the user should not override the default settings unless there is a specific reason.

For PacBio, temperature affects polymerase activity and the stability of the zero-mode waveguide optics. The instrument maintains a constant temperature, and the user should ensure that the instrument is not placed near a heat source that could cause temperature drift.

### Software Settings and Run Configuration

The sequencing software controls the run configuration, including the expected read length, the quality thresholds, and the basecalling parameters. Incorrect software settings can reduce yield even when the library and flow cell are perfect.

For Nanopore, the software settings include the sequencing kit, the flow cell type, the expected read length, and the basecalling model. If the software is configured for the wrong kit or flow cell, the basecaller may apply incorrect parameters and produce low quality calls. If the expected read length is set too short, the software may stop reading a fragment prematurely.

For PacBio, the software settings include the sequencing mode, the expected insert size, and the polymerase binding kit. If the software is configured for the wrong insert size, the circular consensus sequencing may not generate enough passes for accurate HiFi reads.

The practical approach is to verify all software settings before starting the run. The settings should match the library preparation kit and the flow cell type. If the software has been updated, the settings may have changed, and the user should review the release notes for the new version.

## Basecalling and Post-Run Processing

Basecalling converts the raw signal from the sequencer into nucleotide sequence. The basecalling parameters determine how much of the raw signal becomes usable data. Poor basecalling can reduce yield even when the raw signal is excellent.

### Basecaller Version and Model Selection

Basecaller software improves over time, and newer versions generally produce more accurate calls. For Nanopore, the basecaller uses a neural network model that is trained on specific sequencing conditions. The model must match the kit and flow cell used for the run. If the model is outdated or mismatched, the basecaller may produce low quality calls that fail quality filters.

The practical approach is to use the latest basecaller version and the model that matches the sequencing kit. The basecaller software documentation specifies which model to use for each kit and flow cell combination. If the basecaller produces a high proportion of reads below the quality threshold, the model may be wrong.

For PacBio, the basecaller for HiFi reads is integrated into the instrument software. The basecaller uses the number of passes and the quality of each pass to generate a consensus read. The user can set the minimum number of passes and the minimum predicted accuracy. If these thresholds are set too high, many reads will be discarded. If set too low, the reads will have lower accuracy.

### Quality Score Thresholds and Read Filtering

Quality scores determine which reads are kept for downstream analysis. A quality score of Q20 means one error per 100 bases, and Q30 means one error per 1000 bases. The appropriate threshold depends on the application. For variant detection, higher quality is needed. For genome assembly, lower quality may be acceptable.

The practical decision is to set the quality threshold based on the downstream analysis. If the threshold is too high, many reads are discarded and yield drops. If too low, the reads contain errors that complicate analysis. The basecaller output includes a quality score for each read, and the user can plot the distribution of quality scores to choose an appropriate threshold.

For the diagnostic yield study using PacBio HiFi technology, the researchers generated hg38 aligned variants and de novo phased genome assemblies, then annotated, filtered, and curated variants using clinical standards. The study found new disease relevant findings in 16 of 96 probands, with 9 probands harboring pathogenic or likely pathogenic variants. Seven cases had variants that were only correctly interpreted in long-read data, including copy number variants, an inversion, a mobile element insertion, two low complexity repeat expansions, and a 1 base pair deletion. These variants were visible in short-read data in retrospect but were not called, were called with incorrect sizes or structures, or failed quality control and filtration. This finding shows that quality thresholds and filtering decisions directly affect the biological conclusions drawn from sequencing data.

### Adapter Trimming and Chimeric Read Removal

Adapter trimming removes the adapter sequences from the reads. If adapters are not trimmed, they appear as low complexity sequence at the ends of reads and can interfere with alignment and assembly. Chimeric reads are reads that contain sequence from two different DNA fragments, which can happen when two fragments ligate together during library preparation.

The practical approach is to use the adapter trimming and chimeric read removal tools provided by the sequencing software or by downstream analysis pipelines. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials that cover adapter trimming and quality filtering. The [nf-core Documentation](https://nf-co.re/docs) describes community pipeline standards for read processing, including adapter trimming and quality control steps.

## Platform-Specific Troubleshooting

PacBio and Nanopore have different failure modes, and the troubleshooting approach should be tailored to the platform.

### PacBio HiFi Yield Problems

PacBio HiFi sequencing uses circular consensus sequencing to generate accurate long reads. The polymerase reads the circular template multiple times, and the basecaller combines the passes into one consensus read. Yield depends on the number of polymerase molecules that bind a template, the processivity of the polymerase, and the number of passes per molecule.

Common PacBio yield problems include low polymerase binding, short polymerase processivity, and insufficient passes per molecule. Low polymerase binding can result from low library concentration, poor adapter ligation, or inhibitors in the library. Short processivity can result from DNA damage, suboptimal buffer conditions, or polymerase degradation. Insufficient passes can result from short run duration or early polymerase fall off.

The diagnostic yield study using PacBio HiFi technology on 96 short-read negative probands demonstrates the data quality that HiFi sequencing can achieve. The study generated hg38 aligned variants and de novo phased genome assemblies, and the long-read data allowed detection of structural variants that short-read data missed. The study found that long-read genome sequencing allowed substantial additional diagnostic yield beyond short-read sequencing, with 7 of 96 probands having variants only interpretable in long-read data. This finding underscores the value of troubleshooting yield problems, because the biological discoveries depend on sufficient data.

The practical troubleshooting steps for PacBio low yield are to check the polymerase binding rate, the read length distribution, and the number of passes per molecule. The instrument software reports these metrics. If the polymerase binding rate is low, increase the library concentration or check for inhibitors. If the read length distribution is shorter than expected, check the DNA fragment length and the polymerase processivity. If the number of passes is low, increase the run duration or check the polymerase activity.

### Nanopore Yield Problems

Nanopore sequencing uses protein pores embedded in a membrane. Each pore passes one DNA strand, and the ionic current changes are converted to sequence. Yield depends on the number of active pores, the loading efficiency, and the read length.

Common Nanopore yield problems include low pore occupancy, rapid pore failure, and short read length. Low pore occupancy can result from insufficient library loading, poor adapter ligation, or contaminants that block pores. Rapid pore failure can result from membrane instability, contaminants, or electrical issues. Short read length can result from DNA fragmentation, nicks in the DNA, or premature motor protein dissociation.

The nCATS study demonstrates the yield that can be achieved with targeted nanopore sequencing. The method achieved median sequencing coverage of 675 fold on a MinION flow cell and 34 fold on a Flongle flow cell, using about 3 micrograms of genomic DNA. The study applied nCATS to cell lines, a cell line derived xenograft, and normal and paired tumor normal primary human breast tissue. The method simultaneously assessed haplotype resolved single nucleotide variants, structural variations, and CpG methylation. This study shows that targeted enrichment can achieve high coverage even on small flow cells, which is relevant for troubleshooting yield because targeted methods can rescue projects that would otherwise require whole genome sequencing.

The practical troubleshooting steps for Nanopore low yield are to check the pore occupancy, the pore failure rate, and the read length distribution. The sequencing software reports these metrics in real time. If pore occupancy is low, increase the library concentration or check the adapter ligation. If pores fail rapidly, check the library purity and the flow cell quality. If read length is short, check the DNA fragment length and the motor protein loading.

## Common Failure Patterns and Their Fixes

The following failure patterns recur across laboratories and platforms. Recognizing the pattern quickly saves time and money.

### The Empty Flow Cell

The run completes with very few reads or no reads at all. The flow cell appears to have active pores or zero-mode waveguides, but no sequencing occurs. This pattern usually indicates that the library did not load. The fix is to check the library concentration after the final purification step, verify that the adapters are attached, and confirm that the loading buffer and protocol match the library type.

### The Short Read Run

The run produces many reads, but they are all short. The total yield is low because the read length is far below the expected range. This pattern usually indicates DNA fragmentation. The fix is to check the DNA integrity before library preparation, minimize handling steps, and use wide bore pipette tips. If the DNA was intact before library preparation but fragmented after, the library preparation protocol is causing shearing.

### The Rapid Pore Failure Run

The run starts with good pore occupancy, but the pores fail quickly and the yield plateaus early. This pattern usually indicates a contaminant in the library that damages the pores or the polymerase. The fix is to re-purify the library, check the absorbance ratios, and consider an additional clean-up step. For plant samples, the additional purification step described in the plant extraction protocol can remove recalcitrant contaminants.

### The Low Quality Run

The run produces many reads, but most fail the quality threshold. The total usable yield is low. This pattern usually indicates a basecalling problem or a sequencing chemistry problem. The fix is to check the basecaller version and model, verify that the software settings match the kit and flow cell, and review the quality score distribution. If the quality scores are uniformly low, the sequencing chemistry may be failing, and the flow cell or reagents may need replacement.

### The Uneven Coverage Run

The run produces good total yield, but coverage is highly uneven across the genome. Some regions have very high coverage and others have none. This pattern usually indicates amplification bias or library preparation bias. The fix is to reduce PCR cycles, use PCR-free library preparation when input allows, and check for GC bias. The phage sequencing study noted that tagmentation creates coverage bias and loses genome ends, which is relevant for projects that need uniform coverage.

## Records and Measurements for Yield Troubleshooting

Systematic troubleshooting requires records. The following measurements should be recorded for every sequencing run.

### Pre-Run Measurements

Record the DNA concentration measured by fluorometry, the absorbance ratios at 260/280 and 260/230, the fragment length distribution, and the library concentration after preparation. Record the library preparation kit and protocol version, the adapter lot number, and the enzyme lot numbers. Record the flow cell lot number and the flow cell quality control metrics reported by the manufacturer.

### During-Run Measurements

Record the pore occupancy or polymerase binding rate at the start of the run, at 1 hour, and at intervals throughout the run. Record the cumulative yield over time and the read length distribution. Record any instrument warnings or errors. For Nanopore, record the pore failure rate and the number of active pores over time. For PacBio, record the polymerase processivity and the number of passes per molecule.

### Post-Run Measurements

Record the total yield in gigabases, the read length N50, the mean read quality, and the number of reads passing the quality threshold. Record the basecaller version and model, the quality threshold used, and the adapter trimming settings. Record the downstream analysis results, including alignment rate, coverage, and variant calls.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of sequence databases and analysis services that can be used to store and analyze sequencing data. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) offers learning pathways for bioinformatics data resources and practical analysis education. The [Bioconductor Project](https://bioconductor.org/) provides official package and workflow documentation for reproducible genomic analysis. These resources support the analysis side of yield troubleshooting, because the yield problem may only become apparent during downstream analysis.

## Reproducibility and Workflow Standards

Yield troubleshooting is part of a broader effort to make sequencing reproducible. Reproducible sequencing requires consistent protocols, documented parameters, and version controlled analysis.

The [nf-core Documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflows. These standards include version control, containerization, and automated testing. Applying these standards to sequencing analysis ensures that the same data produces the same results regardless of who runs the analysis.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials that emphasize reproducibility. The tutorials cover quality control, alignment, variant calling, and other analysis steps. Following these tutorials ensures that the analysis steps are documented and repeatable.

The [Carpentries Lessons](https://carpentries.org/lessons) provide foundational computing, data, shell, Git, and programming training. These skills are essential for reproducible analysis, because they enable researchers to automate analysis steps, track changes, and share workflows.

The practical approach to reproducibility is to document every parameter that affects yield. This documentation should include the DNA extraction protocol, the library preparation protocol, the sequencing run parameters, and the basecalling parameters. The documentation should be stored with the sequencing data so that any researcher can reproduce the run.

## Limitations of Yield Troubleshooting

Yield troubleshooting has limits. Some yield problems cannot be fixed by adjusting parameters, and some require repeating the library preparation or replacing the flow cell.

### Input Material Limits

The input DNA quantity sets a hard limit on yield. If the sample has only 100 nanograms of DNA, the library preparation cannot produce the same yield as a sample with 1 microgram. The phage sequencing study showed that high quality sequencing is achievable with inputs 1000-fold lower than manufacturer recommendations, but the yield will be lower than with standard input. Researchers should set realistic yield expectations based on the input quantity.

### Flow Cell Quality Limits

The flow cell quality sets another hard limit. A flow cell with few active pores cannot produce the same yield as a flow cell with many active pores. The manufacturer provides quality control metrics for each flow cell, and researchers should check these metrics before loading. If the flow cell quality is poor, the yield will be low regardless of the library quality.

### Biological Sample Limits

The biological sample itself can limit yield. Some samples contain inhibitors that are difficult to remove, and some samples have DNA that is naturally fragmented. The plant extraction protocol addressed recalcitrant species with an additional purification step, but some samples may require multiple purification steps or alternative extraction methods.

### Basecalling Limits

Basecalling accuracy limits the usable yield. Even with perfect raw signal, the basecaller can only produce reads as accurate as the model allows. The basecaller model is trained on specific sequencing conditions, and if the run conditions differ from the training conditions, the basecaller may produce lower quality calls.

## Professional Escalation Criteria

Some yield problems require escalation to the instrument manufacturer or to a sequencing core facility. The following criteria indicate that escalation is appropriate.

### Repeated Failure With the Same Protocol

If the same protocol fails twice with the same yield problem, the protocol itself may be flawed. Escalate to the kit manufacturer or to a sequencing core facility for protocol optimization. The manufacturer may have updated protocols or troubleshooting guides that address the specific failure mode.

### Flow Cell Quality Below Manufacturer Specification

If the flow cell quality metrics are below the manufacturer specification, the flow cell may be defective. Escalate to the manufacturer for replacement. The manufacturer typically provides a warranty for flow cells that fail quality control.

### Instrument Malfunction

If the instrument reports errors or produces inconsistent results across runs, the instrument may be malfunctioning. Escalate to the instrument manufacturer for service. The manufacturer can run diagnostic tests and repair or replace faulty components.

### Unusual Contamination Patterns

If the sequencing data shows contamination from another organism, the contamination may come from the reagents, the laboratory environment, or the sample itself. Escalate to the laboratory safety officer and to the reagent manufacturer. Contamination can be investigated by sequencing a negative control and comparing the results.

### Clinical or Diagnostic Applications

If the sequencing is for clinical or diagnostic purposes, yield problems have regulatory implications. The diagnostic yield study using PacBio HiFi technology on 96 probands with rare diseases demonstrates the clinical value of long-read sequencing. The study found that long-read genome sequencing allowed substantial additional diagnostic yield beyond short-read sequencing, with variants including copy number variants, an inversion, a mobile element insertion, two low complexity repeat expansions, and a 1 base pair deletion. For clinical applications, yield problems must be escalated to the laboratory director and documented according to regulatory requirements.

## Safety and Regulatory Context

Sequencing laboratories operate under safety and regulatory requirements that affect yield troubleshooting.

### Laboratory Safety

Library preparation involves hazardous chemicals, including chaotropic salts, organic solvents, and enzymes. Researchers should follow the laboratory safety manual and use appropriate personal protective equipment. The [Carpentries Lessons](https://carpentries.org/lessons) include foundational computing training but do not cover laboratory safety. Laboratory safety training should be provided by the institution.

### Data Management

Sequencing data are large and require careful management. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of sequence databases and analysis services. Researchers should plan for data storage, backup, and sharing before starting a sequencing project. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) offers learning pathways for data management and analysis.

### Clinical Regulations

If the sequencing is for clinical or diagnostic purposes, the laboratory must comply with applicable regulations. The diagnostic yield study using PacBio HiFi technology on 96 probands with rare diseases was conducted under clinical standards, with variants annotated, filtered, and curated using clinical standards. The study found that long-read genome sequencing allowed detection of variants that were not called in short-read data, were represented by calls with incorrect sizes or structures, or failed quality control and filtration. For clinical applications, yield problems must be documented and reported according to regulatory requirements.

## Frequently Asked Questions

### What is the most common cause of low yield in long-read sequencing?

The most common cause is a library preparation problem, specifically insufficient or degraded high molecular weight DNA input. The DNA must be of adequate purity and integrity for the library preparation enzymes to work efficiently. Plant material often contains carbohydrates and secondary metabolites that impact DNA purity, and many extraction protocols lead to substantial DNA fragmentation. The fix is to measure DNA quantity with fluorometry, check integrity by pulsed-field gel electrophoresis, and re-extract if the DNA is degraded.

### How much DNA is needed for long-read sequencing?

The required input depends on the platform and the library preparation method. Manufacturers provide recommended inputs, but the phage sequencing study showed that high quality sequencing is achievable with inputs 1000-fold lower than manufacturer recommendations. The nCATS targeted nanopore method required about 3 micrograms of genomic DNA. The practical approach is to test the library preparation method with the actual input quantity before committing an expensive flow cell.

### Why are my Nanopore reads shorter than expected?

Short reads usually indicate DNA fragmentation. The DNA fragment length in the library sets the ceiling for read length. Check the DNA integrity before library preparation and after library preparation. If the DNA was intact before library preparation but fragmented after, the library preparation protocol is causing shearing. Use wide bore pipette tips, minimize handling steps, and avoid vortexing.

### Why did my PacBio run produce few reads despite good polymerase binding?

Low read count despite good polymerase binding can result from short polymerase processivity or insufficient passes per molecule. Check the read length distribution and the number of passes per molecule. If the reads are short, the polymerase is falling off the template early. This can result from DNA damage, suboptimal buffer conditions, or polymerase degradation.

### How do I know if my flow cell is defective?

Check the flow cell quality control metrics reported by the manufacturer before loading. If the metrics are below specification, the flow cell may be defective. During the run, monitor the pore occupancy or polymerase binding rate. If the occupancy is low from the start, the flow cell may be defective or the library may not have loaded.

### What basecalling settings should I use for Nanopore data?

Use the latest basecaller version and the model that matches the sequencing kit and flow cell. The basecaller software documentation specifies which model to use for each kit and flow cell combination. Set the quality threshold based on the downstream analysis. For variant detection, higher quality is needed. For genome assembly, lower quality may be acceptable.

### Can I sequence very low input DNA without PCR amplification?

Yes, but the yield will be lower than with standard input. The phage sequencing study showed that high quality sequencing is achievable with inputs 1000-fold lower than manufacturer recommendations, using both Illumina short-read and Nanopore long-read technologies. PCR-free sequencing is recommended for unbiased characterization, but it requires careful library preparation and realistic yield expectations.

### When should I escalate a yield problem to the manufacturer?

Escalate when the same protocol fails twice with the same yield problem, when the flow cell quality is below manufacturer specification, when the instrument malfunctions, when unusual contamination patterns appear, or when the sequencing is for clinical or diagnostic purposes. The manufacturer can provide updated protocols, troubleshooting guides, and diagnostic tests.

## Related Bioinformatics Guides

- [Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data](/knowledge/bioinformatics/long-read-metagenome-assembly-overcoming-challenges-with-nanopore-and-pacbio-data)
- [How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore](/knowledge/bioinformatics/how-to-choose-a-long-read-sequencing-platform-pacbio-vs-oxford-nanopore)
- [Long-Read Sequencing Technologies: PacBio and Oxford Nanopore](/knowledge/bioinformatics/long-read-sequencing-technologies-pacbio-and-oxford-nanopore)
- [Long-Read Sequencing Cost and Market: What to Expect](/knowledge/bioinformatics/long-read-sequencing-cost-and-market-what-to-expect)
- [Long-Read Sequencing for Isoform Quantification: Challenges and Solutions](/knowledge/bioinformatics/long-read-sequencing-for-isoform-quantification-challenges-and-solutions)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Long-read genome sequencing and variant reanalysis increase diagnostic yield in neurodevelopmental disorders.](https://pubmed.ncbi.nlm.nih.gov/39299904). Genome research, 2024.
- [Long-read genome sequencing and variant reanalysis increase diagnostic yield in neurodevelopmental disorders.](https://pubmed.ncbi.nlm.nih.gov/38585854). medRxiv : the preprint server for health sciences, 2024.
- [Targeted nanopore sequencing with Cas9-guided adapter ligation.](https://pubmed.ncbi.nlm.nih.gov/32042167). Nature biotechnology, 2020.
- [Short-read and Long-read PCR-Free Sequencing of Bacteriophages Using Ultra-Low Starting DNA Input.](https://pubmed.ncbi.nlm.nih.gov/40329981). Journal of biomolecular techniques : JBT, 2025.
- [Low-Input High-Molecular-Weight DNA Extraction for Long-Read Sequencing From Plants of Diverse Families.](https://pubmed.ncbi.nlm.nih.gov/35665166). Frontiers in plant science, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.