Understanding Quality Scores in Long-Read Sequencing: How to Interpret Phred Scores from PacBio and Nanopore

By Dr. Zubair Khalid, DVM, MS, PhD ·

Understanding Quality Scores in Long-Read Sequencing: How to Interpret Phred Scores from PacBio and Nanopore

Key Takeaways

  • Phred scores represent the logarithm of the inverse of the estimated probability of an incorrect base call (Q = -10 × log10(P)), with Q20 indicating a 1 in 100 error probability (99% accuracy) and Q30 indicating a 1 in 1,000 error probability (99.9% accuracy).
  • PacBio HiFi reads, generated via circular consensus sequencing (CCS), offer high per-base accuracy (often exceeding Q20) due to multiple passes over the same molecule, making them generally suitable for variant calling and assembly with Q20 or higher filtering.
  • Oxford Nanopore's systematic errors, particularly in homopolymers and methylation contexts, mean that per-base quality scores may overestimate true accuracy; higher quality thresholds and context-aware analysis are crucial for base-level decisions.
  • Applying universal quality filtering thresholds across PacBio and Nanopore platforms is inappropriate due to their distinct error profiles and consensus mechanisms; filtering decisions must account for these platform-specific characteristics.
  • For variant detection, especially single nucleotide variants (SNVs), PacBio HiFi data's high consensus accuracy facilitates straightforward analysis, while Nanopore data requires specialized callers and validation due to systematic error modes like homopolymer inaccuracies.
  • Structural variant detection relies more on read alignment patterns spanning breakpoints than on individual base accuracy, though filtering low-quality reads remains important to mitigate spurious alignment artifacts.

Quality scores in long-read sequencing are probability statements about base call accuracy, but their practical meaning differs substantially between PacBio and Oxford Nanopore platforms. A Phred score of Q20 does not guarantee the same downstream performance on both systems because error profiles, consensus mechanisms, and per-read context vary. This article explains the statistical foundation of Phred scores, how PacBio and Nanopore generate and report them, and how to use them correctly for read filtering, assembly, and variant detection. The guidance is written for biology students, researchers, laboratory professionals, and life-science practitioners who need to make defensible decisions about sequence data quality.

At a Glance: Quality Score Interpretation Across Platforms

The table below summarizes the key differences in how quality scores should be interpreted between the two major long-read platforms. These distinctions matter for practical decisions about filtering thresholds and downstream analysis.

PlatformError ProfileQuality Score MeaningPractical Filtering Guidance
PacBio HiFiRandom errors, low systematic biasPer-base accuracy estimate from circular consensus sequencingQ20 or higher generally suitable for variant calling and assembly
PacBio CLR (legacy)Higher error rate, some systematic componentsPer-base accuracy estimate from single-pass readsRequires consensus correction or hybrid approaches for most applications
Oxford NanoporeSystematic errors in homopolymers and methylation contextsPer-base accuracy estimate from raw signal-to-base conversionHigher quality thresholds needed for base-level decisions, consider context-specific errors

The central practical point is that a quality score from Nanopore and a quality score from PacBio are both Phred-scaled probabilities, but they are attached to reads with fundamentally different error structures. Filtering decisions should account for these differences instead of applying a universal threshold.

The Statistical Meaning of Phred Scores

Phred scores originated in the context of Sanger sequencing and were later adopted by high-throughput platforms. The score is defined as Q = -10 × log10(P), where P is the estimated probability that the base call is incorrect. A Q20 score corresponds to an error probability of 1 in 100, meaning the base is expected to be correct 99 percent of the time. A Q30 score corresponds to an error probability of 1 in 1,000, or 99.9 percent expected accuracy.

The relationship between Phred score and error probability is logarithmic, so each increase of 10 in the Phred score represents a tenfold reduction in error probability. This scaling means that the difference between Q10 and Q20 is much larger than the difference between Q20 and Q30 in terms of absolute error reduction. For long-read applications, understanding this logarithmic relationship is essential because quality thresholds are often set at values that seem numerically close but represent very different error rates.

Quality scores are estimates produced by the base calling software, not direct measurements of accuracy. The base caller uses models trained on known sequences and signal characteristics to assign probabilities. These estimates are generally well calibrated on average, but they can be systematically off in specific sequence contexts. For example, homopolymer regions in Nanopore data often have quality scores that overestimate the true accuracy because the base caller struggles to determine the exact number of repeated bases.

The practical implication is that quality scores should be treated as relative indicators of confidence instead of absolute guarantees. A read with a median quality of Q20 is generally more reliable than a read with a median quality of Q10, but the exact error rate for any particular base may differ from the predicted value. This is especially true for platforms with known systematic error modes.

Platform-Specific Quality Score Generation

PacBio Quality Scores

PacBio sequencing uses a single-molecule real-time approach where a polymerase incorporates fluorescently labeled nucleotides into a growing DNA strand. The original continuous long-read (CLR) mode produced reads with relatively high error rates, often in the range of 10 to 15 percent per base. These errors were largely random, which made them amenable to correction through consensus approaches.

The introduction of circular consensus sequencing (CCS) changed the quality landscape substantially. In CCS mode, the same DNA molecule is sequenced multiple times as the polymerase circles a template. The multiple passes are combined into a single high-accuracy read, and the quality score reflects the consensus of these passes. This approach produces what are now called HiFi reads, with per-base accuracy typically exceeding Q20 and often reaching Q30 or higher.

The quality score for a HiFi read is derived from the agreement among the individual subreads. If all subreads agree on a base, the quality score will be high. If there is disagreement, the quality score will be lower. This consensus-based approach means that HiFi quality scores are generally well calibrated because they are based on multiple independent observations of the same base.

For legacy CLR data, quality scores are less reliable because they are based on a single pass of the polymerase. The random error component is captured in the quality score, but systematic errors may not be fully reflected. Researchers working with CLR data should be cautious about relying on per-base quality scores for variant calling and should consider using consensus-based approaches instead.

Oxford Nanopore Quality Scores

Oxford Nanopore sequencing measures changes in electrical current as DNA passes through a protein nanopore. The base caller converts these current measurements into a sequence of bases, and the quality score reflects the confidence of this conversion. The quality score is assigned per base during the base calling process.

Nanopore quality scores have improved substantially with each generation of base calling software. Early versions produced reads with median quality scores around Q10 or lower, while current versions routinely produce reads with median quality scores above Q20. However, the error profile remains different from PacBio in important ways.

Nanopore errors are not uniformly random. The platform has known systematic issues with homopolymers, where the exact number of repeated bases is difficult to determine from the current signal. There are also context-dependent errors related to methylation and other base modifications. These systematic errors may not be fully captured in the per-base quality scores because the base caller may be confidently wrong in these contexts.

The practical consequence is that Nanopore quality scores are useful for relative comparisons between reads and for identifying low-quality regions, but they should not be interpreted as exact per-base accuracy probabilities in all sequence contexts. For applications that require base-level accuracy, such as identifying single nucleotide variants, additional validation steps may be necessary.

How Quality Scores Are Used in Read Filtering

Read filtering is one of the most common applications of quality scores in long-read sequencing workflows. The goal is to remove low-quality reads that could introduce errors into downstream analyses while retaining enough high-quality data for meaningful results.

The simplest filtering approach is to set a minimum average quality threshold for the entire read. Reads with an average quality below the threshold are discarded. This approach is straightforward to implement and can be effective for removing clearly poor reads. However, it does not account for variation in quality within a read, and a read with a good average quality may still contain problematic regions.

A more sophisticated approach is to filter based on the distribution of quality scores within the read. For example, a read might be retained if a certain percentage of bases exceed a quality threshold, or if there are no long stretches of low-quality bases. This approach is more computationally intensive but provides better control over the quality of the final dataset.

For PacBio HiFi data, the quality score is already a consensus of multiple passes, so filtering thresholds can be relatively lenient. A minimum quality of Q20 is commonly used and generally provides good results for assembly and variant calling. For Nanopore data, higher thresholds may be appropriate because the per-base quality scores are less reliable in systematic error contexts.

The choice of filtering threshold should also consider the downstream application. Assembly algorithms can often tolerate lower quality reads because they use overlapping information to correct errors. Variant calling, on the other hand, requires higher quality because a single base error can lead to a false variant call. Structural variant detection may be less sensitive to per-base quality because the analysis focuses on larger-scale patterns.

Quality Scores in Assembly Workflows

Long-Read-Only Assembly

Long-read-only assembly uses reads from a single platform to construct a genome. The quality of the assembly depends on both the read length and the read accuracy. Long reads provide the contiguity needed to span repetitive regions, but errors in the reads must be corrected during the assembly process.

For PacBio HiFi reads, the high per-base accuracy means that assembly algorithms can rely on the reads being largely correct. This allows for efficient assembly with less computational effort spent on error correction. The quality scores can be used to weight the confidence of individual bases during the assembly process.

For Nanopore reads, the lower per-base accuracy means that assembly algorithms must perform more aggressive error correction. This is typically done through a process of consensus building, where multiple reads are aligned to each other and the consensus sequence is used to correct errors. The quality scores can inform this process by indicating which bases are more reliable.

The choice between long-read-only and hybrid assembly depends on the goals of the project. Long-read-only assembly is simpler and requires less sequencing, but may produce assemblies with more errors. Hybrid assembly, which combines long reads with short reads, can produce higher quality assemblies but requires sequencing on two platforms.

Hybrid Assembly Approaches

Hybrid assembly combines long reads from PacBio or Nanopore with short reads from platforms like Illumina. The short reads provide high per-base accuracy and can be used to correct errors in the long reads. The long reads provide the contiguity needed to assemble complex genomes.

A comparative analysis of bacterial genome assembly using Illumina, Nanopore, and hybrid approaches demonstrated the value of combining platforms. The hybrid approach produced the highest contiguity assembly, with a dominant contig consistent with near-complete chromosomal representation. The hybrid assembly also improved annotation completeness and taxonomic resolution compared to single-platform approaches. This study, published in Pathogens, used SPAdes, Canu, Flye, and Unicycler assemblers and evaluated results with QUAST metrics.

In hybrid assembly workflows, quality scores from the long-read platform are used to guide the initial assembly, while the short-read data is used for polishing. The quality scores help the assembler decide which long reads to trust and how to weight their contribution to the consensus sequence.

The practical implication is that researchers with access to both platforms should consider hybrid approaches for applications that require high-quality assemblies. The additional cost of short-read sequencing is often justified by the improved accuracy of the final assembly.

Quality Scores in Variant Detection

Variant detection is one of the most demanding applications for quality scores because the analysis seeks to identify single base changes that may be present at low frequency. A single error in a read can be mistaken for a true variant, leading to false positive calls.

For PacBio HiFi data, the high per-base accuracy makes variant detection feasible with relatively straightforward approaches. The quality scores can be used to filter out low-confidence base calls, and the consensus nature of HiFi reads reduces the impact of random errors.

For Nanopore data, variant detection is more challenging because of the systematic error modes. Homopolymer errors can lead to false insertion or deletion calls, and context-dependent errors can affect specific positions. Researchers working with Nanopore data for variant detection should use variant callers that are specifically designed to handle the platform's error profile.

The quality scores from Nanopore base callers are useful for filtering, but they should not be the sole basis for variant calls. Additional validation, such as comparing results across multiple reads or using orthogonal methods, is often necessary to confirm true variants.

The impact of sample quality on variant detection is also important to consider. Studies of formalin-fixed paraffin-embedded tissue have shown that long-term storage can introduce sequencing artifacts that are difficult to distinguish from true variants. A study published in Clinical Chemistry found that storage of FFPET for even 2 to 4 years introduced sequencing artifacts totaling 191 false-positive variants compared with 154 true-positive variants. These artifacts included deamination single-nucleotide variants, nondeamination SNVs, and insertion or deletion variants, all of which were detected above 5 percent variant allele frequency. The quality scores from the sequencing platform may not fully capture these storage-related artifacts because they are systematic instead of random.

Quality Scores in Structural Variant Analysis

Structural variants, including large insertions, deletions, duplications, and rearrangements, are an important class of genomic variation. Long-read sequencing is particularly well suited for structural variant detection because the long reads can span breakpoints that are difficult to resolve with short reads.

Quality scores play a different role in structural variant detection than in single nucleotide variant detection. For structural variants, the key information is often the alignment pattern of the read instead of the accuracy of individual bases. A read that spans a breakpoint provides evidence for the structural variant regardless of the quality of individual bases within the read.

However, quality scores still matter for structural variant detection. Low-quality reads may produce spurious alignment patterns that are mistaken for structural variants. Filtering out low-quality reads can reduce the number of false positive calls.

The per-base quality scores are less critical for structural variant detection than for single nucleotide variant detection, but the overall read quality is still relevant. Researchers should consider both the average quality and the distribution of quality within reads when filtering for structural variant analysis.

Long-read sequencing is increasingly used in clinical applications such as cardiovascular genetics, where comprehensive assessment of both rare and common genetic variations is needed. A review published in Life discusses how next-generation sequencing has transformed cardiovascular genetics and highlights long-read sequencing among the new technologies likely to advance precision cardiology. The ability to detect structural variants is one of the advantages of long-read approaches in these settings.

Practical Workflow for Quality Score Assessment

Step 1: Examine Raw Quality Score Distributions

Before applying any filtering thresholds, examine the distribution of quality scores in your dataset. Most base calling software produces a quality score for each base, and these can be summarized at the read level. Plot the distribution of mean read quality to understand the overall quality of your sequencing run.

For PacBio HiFi data, the distribution should be centered at a relatively high quality, typically above Q20. If the distribution is shifted lower, this may indicate problems with the sequencing run or the library preparation.

For Nanopore data, the distribution may be broader and centered at a lower quality. This is expected given the platform's error profile. The distribution can also vary between runs depending on the flow cell generation, base calling software version, and library preparation method.

Step 2: Assess Per-Base Quality in Problematic Contexts

Quality scores can be misleading in specific sequence contexts. For Nanopore data, examine quality scores in homopolymer regions and in regions with known base modifications. For PacBio data, examine quality scores in GC-rich or GC-poor regions where polymerase behavior may differ.

This assessment can be done by aligning reads to a reference genome and comparing quality scores to known sequence features. If quality scores are systematically lower in certain contexts, this should inform your filtering decisions.

Step 3: Set Filtering Thresholds Based on Application

The appropriate filtering threshold depends on the downstream application. For assembly, lower quality reads can often be tolerated because the assembly process corrects errors through consensus. For variant calling, higher quality thresholds are needed to avoid false positive calls.

A reasonable starting point for PacBio HiFi data is a minimum mean quality of Q20. For Nanopore data, a minimum mean quality of Q15 to Q20 may be appropriate, depending on the application. These thresholds should be adjusted based on the quality distribution of your specific dataset.

Step 4: Validate Filtering Decisions

After filtering, validate that the retained reads produce the expected results. For assembly, check the assembly metrics such as contiguity and completeness. For variant calling, compare results to known variants or to results from an orthogonal method.

Validation is especially important when working with challenging samples. Studies of aged formalin-fixed paraffin-embedded specimens have shown that DNA quality declines with storage duration, and this decline affects sequencing metrics including read depth. A study published in the Journal of Clinical Laboratory Analysis examined thyroid cancer tissue samples preserved for up to 55 years and found that both DNA yield and DNA integrity number were negatively correlated with storage duration. Whole exome sequencing showed that storage duration was also negatively correlated with mean insert size and normalized read depth. The quality scores from the sequencing platform may not fully capture these sample-specific issues.

Step 5: Document Quality Score Decisions

Document the filtering thresholds and the rationale for those thresholds in your analysis protocol. This documentation is important for reproducibility and for interpreting results in the context of the data quality.

Reproducibility is a core principle of bioinformatics analysis. Training resources from organizations like the Galaxy Training Network and nf-core emphasize the importance of documenting analysis decisions and using standardized workflows. The European Bioinformatics Institute also provides training pathways that cover data-resource usage and practical analysis education.

Records and Measurements for Quality Assessment

Maintaining records of quality metrics is essential for tracking sequencing performance over time and for comparing results across runs. The following measurements should be recorded for each sequencing run:

Read length distribution, including the N50 value, which represents the read length at which half of the total sequence data is in reads of that length or longer. This metric is important for understanding the contiguity potential of the dataset.

Mean and median read quality, which provide an overall summary of the quality of the sequencing run. These values can be compared across runs to identify trends or problems.

Quality score distribution, which shows the proportion of bases at each quality level. This distribution is more informative than the mean or median because it reveals the shape of the quality profile.

Per-read quality variation, which identifies reads with unusual quality patterns. Reads with highly variable quality may be problematic even if the mean quality is acceptable.

For PacBio HiFi data, additional records should include the number of passes used for the circular consensus and the relationship between passes and quality. Higher numbers of passes generally produce higher quality reads but require more sequencing time.

For Nanopore data, records should include the base calling software version and the flow cell generation. These factors have a substantial impact on quality scores, and comparing quality across different versions can be misleading.

The table below provides a template for recording quality metrics across sequencing runs.

MetricPacBio HiFiOxford NanoporePurpose
N50 read lengthRecord in base pairsRecord in base pairsAssess contiguity potential
Mean read qualityTypically Q20 or higherVaries by base caller versionOverall run quality summary
Quality score distributionNarrow distribution at high valuesBroader distributionIdentify quality profile shape
Per-read quality variationLow variation expectedHigher variation expectedFlag problematic reads
Base caller versionRecord software and versionRecord software and versionEnable cross-run comparisons

Common Failure Patterns in Quality Score Interpretation

Overreliance on Mean Quality

A common mistake is to filter reads based solely on mean quality without examining the distribution of quality within reads. A read with a good mean quality may contain long stretches of low-quality bases that could introduce errors into downstream analyses. Conversely, a read with a lower mean quality may have high-quality regions that are perfectly usable.

The solution is to examine the distribution of quality within reads and to use filtering criteria that account for this distribution. For example, a read might be retained if it has no stretches of low-quality bases above a certain length, regardless of the mean quality.

Applying the Same Threshold to Different Platforms

Another common mistake is to apply the same quality threshold to PacBio and Nanopore data without considering the different error profiles. A Q20 threshold that works well for PacBio HiFi data may be too lenient for Nanopore data because of the systematic error modes.

The solution is to set platform-specific thresholds based on the error profile and the downstream application. This requires understanding the strengths and limitations of each platform.

Ignoring Context-Dependent Errors

Quality scores can be systematically wrong in specific sequence contexts. For Nanopore data, homopolymer regions are a well-known problem. The base caller may assign high quality scores to bases in homopolymers even when the exact number of repeats is uncertain.

The solution is to be aware of these context-dependent errors and to use additional validation for regions that are known to be problematic. This is especially important for variant calling, where a single base error can lead to a false positive call.

Confusing Read Quality with Consensus Quality

For PacBio HiFi data, the quality score reflects the consensus of multiple passes of the same molecule. This is different from the quality of a single-pass read. Researchers who are used to working with single-pass data may misinterpret HiFi quality scores.

The solution is to understand the consensus nature of HiFi reads and to recognize that the quality scores are generally more reliable than those from single-pass reads.

Neglecting Sample Quality Effects

The quality of the input DNA has a substantial impact on sequencing quality. Studies of formalin-fixed paraffin-embedded tissue have shown that long-term storage degrades DNA quality and introduces sequencing artifacts. These artifacts can be mistaken for true variants, and the quality scores from the sequencing platform may not fully capture them.

The solution is to assess input DNA quality before sequencing and to be cautious when interpreting results from samples with known quality issues. In some cases, pretreatment with a DNA repair enzyme may help restore sequencing quality. The study published in Clinical Chemistry found that storage-associated artifacts, including tumor mutation burden elevation, were largely restored by pretreatment with a DNA repair enzyme.

Limitations of Quality Scores

Quality scores are estimates, not measurements. They are produced by base calling software that uses models trained on known sequences. These models are generally well calibrated, but they can be systematically wrong in specific contexts.

The calibration of quality scores can vary between platforms, between base calling software versions, and even between sequencing runs. This means that a Q20 score from one run may not be exactly equivalent to a Q20 score from another run.

Quality scores also do not capture all sources of error. Systematic errors, such as those caused by base modifications or by the sequencing chemistry, may not be reflected in the quality scores. This is especially true for Nanopore data, where the error profile is complex and context-dependent.

For samples with known quality issues, such as aged formalin-fixed paraffin-embedded tissue, quality scores may be even less reliable. Studies have shown that DNA quality declines with storage duration, and this decline affects sequencing metrics. The study of thyroid cancer specimens preserved for up to 55 years found that while DNA yield and integrity declined with storage duration, the uniformity score, which reflects the evenness of sequence coverage across targeted regions, did not show a significant correlation with storage duration. This finding suggests that even degraded samples may be amenable to sequencing with increased read depth, but careful interpretation of results is warranted.

Researchers should therefore treat quality scores as one source of information about data quality, but not the only source. Other indicators, such as alignment patterns, coverage depth, and agreement between reads, should also be considered.

Safety and Regulatory Context for Clinical Applications

When long-read sequencing is used for clinical applications, quality scores take on additional significance. The accuracy of variant calls can have direct implications for patient care, and regulatory frameworks may require specific quality standards.

The use of next-generation sequencing in clinical settings is expanding, including for cardiovascular genetics. The review published in Life discusses how NGS has transformed the study of cardiovascular genetics, allowing researchers to move beyond single-gene analyses toward comprehensive assessments of both rare and common genetic variations. Long-read sequencing is among the new technologies highlighted as likely to further advance precision cardiology. However, the analytical and ethical challenges of clinical sequencing are substantial, and quality scores are a critical component of the analytical pipeline.

For clinical applications, quality scores should be interpreted in the context of validated workflows. The use of standardized pipelines and reference materials can help ensure that quality scores are meaningful and that variant calls are reliable. Resources such as the Bioconductor Project provide packages and workflows for reproducible genomic analysis, and the Galaxy Training Network offers accessible workflow training that emphasizes reproducibility.

Professional escalation criteria for clinical applications include situations where quality scores are unexpectedly low, where variant calls cannot be validated, or where sample quality is known to be compromised. In these situations, consultation with a clinical laboratory director or a bioinformatics specialist is appropriate.

Professional Escalation Criteria

There are situations where quality score issues indicate problems that require professional intervention. The following criteria should trigger escalation to a supervisor, a bioinformatics specialist, or a laboratory director:

Unexpectedly low quality scores across an entire sequencing run may indicate a problem with the sequencing instrument, the reagents, or the library preparation. This situation warrants investigation before proceeding with downstream analysis.

Quality scores that are systematically lower in specific sequence contexts may indicate a problem with the base calling model or with the sample. This situation may require consultation with the platform manufacturer or with a bioinformatics specialist.

Variant calls that cannot be validated by orthogonal methods may indicate a problem with the quality scores or with the variant calling approach. This situation warrants careful investigation before reporting results.

Samples with known quality issues, such as aged formalin-fixed paraffin-embedded tissue, may produce results that are difficult to interpret. Studies have shown that long-term storage can introduce sequencing artifacts that are difficult to distinguish from true variants. In these situations, consultation with a specialist is appropriate.

Results that are inconsistent with biological expectations may indicate a problem with the data quality or with the analysis approach. This situation warrants careful investigation before drawing conclusions.

A Decision Framework for Platform-Specific Quality Score Thresholds

Choosing a quality score threshold is not a one-time decision but a recurring judgment call that should be tied to the specific biological question, the error tolerance of the downstream analysis, and the cost of a wrong call. This section provides a structured framework for setting, testing, and revising quality thresholds for PacBio HiFi and Oxford Nanopore data. The framework is built around three questions: what error rate can your analysis tolerate, what is the cost of a false positive versus a false negative, and how much data can you afford to discard.

Step 1: Define the Error Tolerance of Your Downstream Analysis

Different analyses have different sensitivities to base-level errors. Before setting any threshold, write down the maximum per-base error rate that your analysis can tolerate without producing misleading results.

For variant calling, the tolerance is low. A single base error in a read can be mistaken for a true variant, especially when the variant is present at low frequency. The study of formalin-fixed paraffin-embedded tissue published in Clinical Chemistry demonstrated that storage-related artifacts produced 191 false-positive variants compared with 154 true-positive variants, all detected above 5 percent variant allele frequency. This finding shows that even modest error rates can overwhelm true signal when the biological variant is rare.

For assembly, the tolerance is higher. Assembly algorithms use overlapping reads to build a consensus, so random errors in individual reads are corrected during the process. A comparative analysis of bacterial genome assembly published in Pathogens showed that hybrid assembly with Unicycler produced the highest contiguity, yielding seven contigs and a dominant 4.55 Mb contig. The assembler was able to use lower quality long reads effectively because the short-read data provided the accuracy needed for correction.

For structural variant detection, the tolerance depends on the analysis method. If you are looking for large deletions or insertions, the alignment pattern matters more than individual base accuracy. A read with moderate quality can still provide strong evidence for a structural variant if it spans a breakpoint cleanly.

Step 2: Assign a Cost to Each Error Type

For each analysis, assign a relative cost to false positives and false negatives. This cost is not always symmetric.

In clinical variant screening, a false positive can lead to unnecessary follow-up testing, patient anxiety, or an incorrect treatment decision. A false negative can miss a clinically actionable variant. The review published in Life discusses how next-generation sequencing informs clinical practice for inherited cardiac disorders, where the cost of a wrong call is measured in patient outcomes.

In exploratory research, the cost structure may be different. A false positive variant call might waste time on validation experiments, while a false negative might mean a missed biological discovery. In this context, a more lenient threshold that retains more data may be appropriate, provided that validation is planned.

In assembly projects, the cost of a false positive is an assembly error that may propagate through downstream annotation and comparative analysis. The cost of a false negative is a gap in the assembly that may require additional sequencing to close.

Step 3: Set an Initial Threshold Based on Platform Error Profiles

Using the error tolerance and cost structure from the first two steps, set an initial threshold. The table below provides starting points based on platform and application.

ApplicationPacBio HiFi Initial ThresholdNanopore Initial ThresholdRationale
Variant calling (SNVs)Q30 mean read qualityQ20 mean read quality with context-aware filteringSNV calls require high per-base confidence
Variant calling (indels)Q20 mean read qualityQ15 mean read quality with homopolymer-aware callerIndel errors are common in Nanopore homopolymers
Assembly (long-read only)Q20 mean read qualityQ10 to Q15 mean read qualityAssembly corrects errors through consensus
Assembly (hybrid)Q10 mean read qualityQ10 mean read qualityShort reads provide the accuracy for correction
Structural variant detectionQ15 mean read qualityQ10 mean read qualityAlignment patterns matter more than base accuracy

These thresholds are starting points, not universal rules. The appropriate threshold for your dataset depends on the base calling software version, the library preparation method, and the sample quality.

Step 4: Test the Threshold on a Validation Set

Before applying the threshold to your full dataset, test it on a small validation set. This set should include samples with known variants or a reference genome with known sequence.

For variant calling, compare the variants called at your chosen threshold to the known variants. Calculate the sensitivity and specificity. If the sensitivity is too low, lower the threshold. If the specificity is too low, raise the threshold.

For assembly, assemble the validation set at your chosen threshold and compare the assembly to the reference. Check for misassemblies, base errors, and completeness. If the assembly has too many errors, raise the threshold or consider a hybrid approach.

For structural variant detection, check that known structural variants are detected and that no spurious variants are called. The validation set should include both positive and negative controls.

Step 5: Monitor Quality Metrics Across the Full Dataset

After applying the threshold to the full dataset, monitor the quality metrics to ensure that the filtering is working as expected. Track the number of reads retained, the mean quality of retained reads, and the distribution of quality scores.

If the threshold removes too much data, the coverage may be insufficient for the downstream analysis. If the threshold removes too little data, the error rate may be too high. Adjust the threshold based on the observed metrics.

For Nanopore data, pay particular attention to the quality scores in homopolymer regions. The base caller may assign high quality scores to bases in homopolymers even when the exact number of repeats is uncertain. If your analysis is sensitive to homopolymer errors, consider using a variant caller that is specifically designed to handle this error mode.

Step 6: Document the Decision and the Rationale

Record the threshold, the validation results, and the rationale for the final choice. This documentation is essential for reproducibility and for interpreting results in the context of data quality.

The Galaxy Training Network and nf-core provide guidance on documenting analysis decisions and using standardized workflows. The European Bioinformatics Institute offers training pathways that cover data-resource usage and practical analysis education. The Bioconductor Project provides packages for reproducible genomic analysis that can help you track filtering decisions.

Common Failure Patterns in Threshold Setting

One common failure is setting a threshold without considering the downstream analysis. A threshold that works for assembly may be too lenient for variant calling, and a threshold that works for variant calling may discard too much data for assembly.

Another common failure is applying the same threshold to all samples in a study without considering sample quality. Studies of aged formalin-fixed paraffin-embedded specimens published in the Journal of Clinical Laboratory Analysis found that DNA quality declines with storage duration, and this decline affects sequencing metrics including read depth. Samples with degraded DNA may require different thresholds than fresh samples.

A third common failure is ignoring the base calling software version. Quality scores from different versions of the same base caller are not directly comparable. If you change the base calling software version mid-project, you should re-evaluate your thresholds.

Records and Measurements for Threshold Decisions

For each project, record the following information:

The base calling software and version used for each platform. This information is essential for comparing quality scores across runs and for reproducing the analysis.

The initial threshold and the rationale for that threshold. This should include the error tolerance and cost structure from the first two steps.

The validation results, including sensitivity and specificity for variant calling, or assembly metrics such as contiguity and completeness.

The final threshold and any adjustments made during the analysis.

The number of reads retained and discarded at the final threshold, along with the mean quality of retained reads.

The table below provides a template for recording threshold decisions.

Decision PointPacBio HiFiOxford Nanopore
Base caller versionRecord software and versionRecord software and version
Initial thresholdRecord mean quality thresholdRecord mean quality threshold
Validation methodDescribe validation set and metricsDescribe validation set and metrics
Final thresholdRecord adjusted thresholdRecord adjusted threshold
Reads retainedRecord count and percentageRecord count and percentage
Mean quality of retained readsRecord valueRecord value

Professional Escalation Criteria for Threshold Decisions

Escalate to a bioinformatics specialist or laboratory director when the validation results do not meet the expected standards. This includes situations where sensitivity or specificity is below acceptable levels, where assembly metrics indicate poor quality, or where the threshold cannot be adjusted to achieve acceptable results.

Escalate when quality scores are systematically inconsistent with expected values for the platform and base caller version. This may indicate a problem with the sequencing run, the library preparation, or the sample.

Escalate when sample quality is known to be compromised and the threshold cannot compensate for the resulting artifacts. The study published in Clinical Chemistry found that pretreatment with a DNA repair enzyme largely restored storage-associated artifacts, including tumor mutation burden elevation. If you are working with aged formalin-fixed paraffin-embedded tissue and quality scores are poor, consider whether enzyme repair or other sample restoration methods are appropriate before adjusting thresholds.

Escalate when the cost of a wrong call is high and the quality scores do not provide sufficient confidence. This is especially important for clinical applications, where the accuracy of variant calls can have direct implications for patient care.

Frequently Asked Questions

What is the difference between a Phred score and a quality score?

A Phred score is a specific type of quality score that is defined as Q = -10 × log10(P), where P is the estimated probability that the base call is incorrect. The term quality score is more general and can refer to any measure of base call confidence. In practice, the terms are often used interchangeably in the context of sequencing data.

Why do PacBio and Nanopore quality scores need different interpretation?

PacBio and Nanopore platforms have different error profiles. PacBio HiFi reads are produced by consensus of multiple passes of the same molecule, which reduces random errors and makes quality scores more reliable. Nanopore reads are produced from a single pass of the molecule, and the error profile includes systematic components such as homopolymer errors. These differences mean that the same Phred score can have different practical implications on the two platforms.

What is a good quality score threshold for PacBio HiFi reads?

A minimum mean quality of Q20 is commonly used for PacBio HiFi reads and generally provides good results for assembly and variant calling. The consensus nature of HiFi reads means that quality scores are relatively reliable, and Q20 represents a good balance between data retention and quality.

What is a good quality score threshold for Nanopore reads?

The appropriate threshold for Nanopore reads depends on the application and on the base calling software version. A minimum mean quality of Q15 to Q20 is a reasonable starting point, but higher thresholds may be needed for applications that require base-level accuracy. The systematic error modes of Nanopore sequencing mean that quality scores should not be the sole basis for filtering decisions.

How do quality scores affect assembly quality?

Quality scores affect assembly quality by informing the error correction process. Higher quality reads require less correction and produce more accurate assemblies. Lower quality reads can still be used, but they require more aggressive error correction and may produce assemblies with more errors. Hybrid assembly approaches that combine long reads with short reads can improve assembly quality by using the short reads to correct errors in the long reads.

Can quality scores be used to detect sample quality problems?

Quality scores can provide some indication of sample quality problems, but they are not a reliable diagnostic tool. Samples with degraded DNA, such as aged formalin-fixed paraffin-embedded tissue, may produce lower quality scores, but the relationship is not consistent. Studies have shown that long-term storage can introduce sequencing artifacts that are not fully captured by quality scores.

How should quality scores be reported in publications?

Quality scores should be reported as part of the sequencing statistics for each dataset. This includes the mean and median read quality, the distribution of quality scores, and the filtering thresholds used in the analysis. Reporting these details is important for reproducibility and for allowing readers to assess the quality of the data.

What should I do if my quality scores are unexpectedly low?

If quality scores are unexpectedly low, first check the sequencing run metrics to identify any obvious problems. Then examine the quality score distribution to understand the pattern of low quality. If the problem persists, consult with a bioinformatics specialist or the platform manufacturer. Do not proceed with downstream analysis until the cause of the low quality scores is understood.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.