# Quality Score Recalibration in Polishing: How to Use Phred Scores to Improve Error Correction


## Key Takeaways

- Phred scores, a logarithmic representation of base call error probability (Q = -10 log10(P)), are fundamental to genome polishing tools like Racon and Medaka. Miscalibrated scores lead to overcorrection of true variants, undercorrection of residual errors, or introduction of new assembly mistakes.
- Raw quality scores from long-read sequencers (e.g., Oxford Nanopore, PacBio HiFi) are often miscalibrated due to variations in flow cell generation, base caller versions, run conditions, or sequence contexts like homopolymers. This necessitates empirical assessment and recalibration.
- Quality score recalibration adjusts reported Phred scores to align with empirically observed error rates, a process distinct from quality trimming which removes low-quality bases. Recalibration preserves read length while correcting score accuracy, crucial for polishing algorithms that weight read evidence.
- Tools like fastp and Qualimap facilitate quality score recalibration by comparing read alignments to a reference genome, estimating empirical error rates per quality score bin, and applying a transformation to correct the scores. This empirical error estimation is the cornerstone of accurate recalibration.
- Before polishing, assess calibration by aligning reads to a reference and calculating empirical error rates; recalibrate with tools like fastp if reported scores deviate significantly (e.g., by a factor of two) from observed error rates, especially when using Racon which relies heavily on input quality.
- Post-polishing validation using metrics like BUSCO, read mapping consistency, and direct comparison to a reference genome is essential to confirm accuracy improvements and detect any residual errors or new mistakes introduced during the process.

---

Base quality scores are the primary currency of error correction in genome assembly polishing. Polishing tools such as Medaka and Racon use per-base Phred scores to decide which reads support a consensus sequence and which mismatches represent true biological variation versus sequencing error. When quality scores are miscalibrated, polishing algorithms can overcorrect true variants, undercorrect residual errors, or introduce new mistakes into an otherwise accurate assembly. This article explains how Phred scores influence polishing behavior, why raw quality scores from long-read sequencers are often miscalibrated, and how recalibration with tools such as Qualimap or fastp can improve polishing accuracy. The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who generate or analyze genome assemblies and need practical decision criteria for when and how to recalibrate before polishing.

## The Role of Quality Scores in Assembly Polishing

### What Phred Scores Represent in Sequencing Data

A Phred score is a logarithmic transformation of the estimated probability that a base call is incorrect. A Phred score of Q20 corresponds to an error probability of 1 in 100, Q30 corresponds to 1 in 1000, and Q40 corresponds to 1 in 10,000. These scores are encoded in sequencing files such as FASTQ and SAM, and they accompany every base call produced by a sequencing instrument. The scores are not direct measurements of error. They are predictions generated by the base caller software that interprets the raw signal data from the sequencer.

The relationship between Phred score and error probability is defined as Q = -10 log10(P), where P is the probability that the base call is wrong. A base with Q30 has a 0.001 probability of being incorrect. This mathematical relationship is the foundation for downstream applications that weight evidence by quality, including variant calling, consensus generation, and assembly polishing.

### How Polishing Algorithms Consume Quality Scores

Polishing is the process of correcting errors in a draft assembly by aligning sequencing reads back to the assembly and using the read evidence to fix base-level mistakes. The two most widely used polishing tools for long-read assemblies are Racon and Medaka. Both tools operate on the principle that where multiple reads agree on a base that differs from the assembly, the reads are likely correct and the assembly is likely wrong. However, the tools differ substantially in how they weight read quality.

Racon performs polishing through a partial order alignment of reads against the assembly. It uses quality scores to weight the evidence contributed by each read at each position. A read with high quality scores at a given position contributes more evidence than a read with low quality scores. Racon does not perform any error correction on the reads themselves. It relies entirely on the quality scores provided in the input alignment.

Medaka uses a neural network model trained on specific sequencing platforms and base calling versions. The model learns the error profiles of the sequencing technology and uses this learned information alongside the quality scores to predict the most likely corrected base at each position. Medaka models are platform-specific. A model trained for Oxford Nanopore R10.4.1 flow cells with a particular base caller version will not perform optimally on data from a different platform or base caller version.

Both tools share a common vulnerability. If the quality scores in the input reads are systematically wrong, the polishing decision will be biased. Overly optimistic quality scores cause the polisher to trust erroneous reads. Overly pessimistic quality scores cause the polisher to ignore correct reads and retain assembly errors.

### Why Raw Quality Scores Are Often Miscalibrated

Sequencing instruments produce quality scores that are calibrated to the training data used to develop the base caller. When the actual sequencing conditions differ from the training conditions, the quality scores drift from their true values. Several factors contribute to miscalibration in long-read sequencing.

Oxford Nanopore sequencing produces quality scores that vary with flow cell generation, base caller version, pore type, and run conditions such as temperature and voltage. The same underlying DNA sequence can produce different quality score distributions across different runs. The base caller assigns quality scores based on the neural network confidence, but this confidence does not always match the empirical error rate observed when the reads are compared to a known reference.

PacBio HiFi reads have generally accurate quality scores because the circular consensus sequencing process generates multiple passes over the same molecule, and the quality scores are derived from the observed agreement across passes. However, systematic errors can still occur in homopolymer regions and other sequence contexts where the polymerase stutters or the signal interpretation is ambiguous.

The practical consequence is that raw quality scores from a sequencer are a starting point, not a final calibrated measurement. Before using these scores for polishing, it is prudent to assess whether they accurately reflect the empirical error rate of the reads.

## At a Glance

| Decision Point | Recommended Action | Rationale |
| --- | --- | --- |
| Before polishing a new assembly | Assess quality score calibration by comparing read error rates to reported Phred scores | Miscalibrated scores bias polishing decisions and can introduce new errors |
| When raw quality scores appear inflated | Recalibrate with fastp or similar tools before aligning reads for polishing | Recalibration aligns reported scores with empirical error rates |
| When using Medaka | Verify that the model matches the flow cell and base caller version | Medaka models are platform-specific and mismatched models reduce accuracy |
| When using Racon | Consider multiple rounds with quality-weighted alignments | Racon relies on input quality scores and benefits from calibrated inputs |
| After polishing | Evaluate assembly quality with independent metrics such as BUSCO and read mapping consistency | Polishing can introduce errors, independent validation is required |

## Core Principles of Quality Score Recalibration

### The Difference Between Trimming and Recalibration

Quality trimming and quality score recalibration are distinct operations that are often confused. Trimming removes low-quality bases from the ends of reads or from internal regions where quality falls below a threshold. Trimming reduces the amount of low-quality data that enters downstream analysis, but it does not change the quality scores of the bases that remain.

Recalibration changes the quality scores themselves. A recalibration step examines the relationship between reported quality scores and empirically observed errors, then adjusts the scores so that a reported Q30 actually corresponds to a 1 in 1000 error rate. Recalibration can increase or decrease individual quality scores depending on the direction of the miscalibration.

For polishing, recalibration is generally more important than aggressive trimming. Polishing algorithms need coverage across the entire assembly, including regions where quality is moderate. Aggressive trimming can remove useful evidence from the ends of reads, reducing coverage in repetitive or GC-rich regions. Recalibration preserves the read length while correcting the quality score values.

### Empirical Error Estimation

The foundation of recalibration is empirical error estimation. To know whether quality scores are calibrated, the reads must be compared to a trusted reference sequence. The reference can be a closely related genome, a high-quality assembly of the same organism, or a synthetic construct with known sequence. For each read, the alignment to the reference reveals the positions where the read differs from the reference. These differences are assumed to be sequencing errors, provided the reference is accurate and the organism matches.

The empirical error rate is then calculated for each reported quality score bin. For example, all bases with reported Q30 are grouped together, and the fraction of those bases that are errors is calculated. If the fraction is 0.001, the quality scores are well calibrated at Q30. If the fraction is 0.01, the quality scores are overconfident by a factor of ten, and a reported Q30 actually corresponds to Q20.

This empirical error estimation is the same principle used by the Genome Analysis Toolkit Base Quality Score Recalibration module for short-read data. For long-read data, the same principle applies, but the tools differ. fastp can perform quality score recalibration for both short and long reads, and Qualimap provides quality metrics that can inform recalibration decisions.

### Recalibration as a Transformation

Recalibration applies a transformation to the reported quality scores. The transformation can be as simple as a lookup table that maps each reported score to a corrected score, or it can be a more complex model that accounts for sequence context, position in read, and other covariates.

The lookup table approach is straightforward. For each reported quality score value, the empirical error rate is calculated, and the corrected Phred score is derived from that error rate. A reported Q30 with an empirical error rate of 0.01 would be corrected to Q20. A reported Q20 with an empirical error rate of 0.005 would be corrected to Q23.

The covariate approach accounts for systematic patterns in error rates. For example, errors may be more common at the ends of reads, in homopolymer regions, or after specific sequence motifs. A covariate-based recalibration model can adjust quality scores differently depending on these contextual factors. This approach is more accurate but requires more data to estimate the model parameters reliably.

## Practical Workflow for Recalibration Before Polishing

### Step 1: Assess the Current Calibration State

Before deciding whether recalibration is needed, assess the current state of quality score calibration. This assessment requires a reference sequence to compare against. If a high-quality reference genome exists for the organism being sequenced, use it. If not, a closely related species can serve as a proxy, though the results will include true biological divergence in addition to sequencing error.

Align a sample of reads to the reference using a long-read aligner such as minimap2. Extract the reported quality scores and the match or mismatch status for each aligned base. Calculate the empirical error rate for each quality score bin and compare to the expected error rate derived from the Phred formula.

The assessment can be performed with Qualimap, which provides detailed quality metrics for aligned reads, or with custom scripts that parse alignment files. The key output is a table showing reported quality score, observed error rate, and the difference between the two.

### Step 2: Decide Whether Recalibration Is Necessary

The decision to recalibrate depends on the magnitude of the discrepancy between reported and empirical error rates. Small discrepancies of a few percent may not materially affect polishing outcomes. Large discrepancies, where reported quality scores are off by a factor of two or more, are likely to bias polishing decisions.

A practical threshold is to recalibrate when the empirical error rate differs from the reported error rate by more than a factor of two at any quality score bin that contributes substantially to the data. For example, if the median read quality is Q20 and the empirical error rate at Q20 is 0.02 instead of 0.01, recalibration is warranted.

The decision also depends on the polishing tool. Medaka neural network models may partially compensate for miscalibrated quality scores because the model learns error patterns from training data. Racon has no such compensation mechanism and relies entirely on the input quality scores. If using Racon, recalibration is more critical.

### Step 3: Perform Recalibration

fastp is a widely used tool that can perform quality score recalibration. The tool was originally developed for short-read data but has been extended to support long reads. fastp can estimate the empirical error rates from an alignment to a reference and apply the correction to the reads.

The fastp recalibration workflow requires a reference sequence and an alignment of the reads to that reference. fastp analyzes the alignment, builds the error model, and outputs recalibrated reads with corrected quality scores. The recalibrated reads are then used as input for the polishing pipeline.

Alternative approaches include custom scripts that implement the lookup table transformation. For research groups with specific recalibration needs, a custom implementation may be preferable because it allows full control over the covariates included in the model.

### Step 4: Verify the Recalibration

After recalibration, verify that the corrected quality scores match the empirical error rates. Repeat the alignment and error rate calculation from Step 1 using the recalibrated reads. The reported quality scores should now closely match the empirical error rates.

Verification is an essential quality control step. Recalibration can introduce its own errors if the reference used for calibration is divergent from the sample or if the error model is poorly estimated. Verification catches these problems before the recalibrated reads are used for polishing.

### Step 5: Polish with Recalibrated Reads

With recalibrated reads, proceed with the polishing pipeline. Align the recalibrated reads to the draft assembly, then run the polishing tool. For Racon, use the recalibrated reads directly. For Medaka, use the recalibrated reads as input to the model, ensuring that the model matches the sequencing platform and base caller version.

After polishing, evaluate the assembly quality using independent metrics. BUSCO completeness scores, read mapping consistency, and comparison to a reference genome if available are all useful validation approaches. The polished assembly should show improved accuracy compared to the unpolished assembly, and the improvement should be attributable to the recalibration instead of to random variation.

## Tools for Quality Score Assessment and Recalibration

### Qualimap for Quality Assessment

Qualimap is a tool for evaluating the quality of aligned sequencing data. It provides a range of metrics including per-base quality distributions, error rates, and coverage statistics. For quality score calibration assessment, Qualimap can show the relationship between reported quality scores and observed mismatches in the alignment.

Qualimap is available through the [Bioconductor project](https://bioconductor.org/), which provides official package documentation and installation guidance for reproducible genomic analysis workflows. The Bioconductor ecosystem includes many tools for quality assessment and processing of sequencing data, and its documentation supports reproducible analysis practices.

The practical use of Qualimap for calibration assessment is to generate a report for a sample of reads aligned to a reference. The report shows whether the quality scores are consistent with the observed error rates. If the report indicates systematic discrepancies, recalibration is warranted.

### fastp for Recalibration

fastp is a tool that performs quality control, trimming, and recalibration of sequencing reads. It supports both short and long reads and can estimate error models from alignments to a reference. The recalibration output is a new FASTQ file with corrected quality scores.

The fastp documentation and usage examples are available through the [Galaxy Training Network](https://training.galaxyproject.org/), which provides accessible workflow training and analysis tutorials. The Galaxy platform allows users to run fastp and other tools through a web interface, which is useful for researchers who prefer not to work exclusively on the command line.

fastp is also integrated into community workflow pipelines. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline usage and configuration, and many nf-core pipelines include fastp as a quality control step. Understanding these pipeline standards helps researchers integrate recalibration into larger analysis workflows.

### Custom Recalibration Scripts

For research groups with specific needs, custom recalibration scripts offer the most control. A custom script can implement the lookup table transformation, include covariates specific to the sequencing platform, and produce detailed reports on the calibration process.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in computing, data analysis, shell scripting, and programming that is useful for developing custom bioinformatics scripts. The Carpentries curriculum emphasizes reproducible data analysis practices, which are essential for maintaining the integrity of recalibration workflows.

Custom scripts require careful validation. The error model must be estimated from sufficient data, and the transformation must be tested on independent data to ensure it generalizes. The verification step described in the workflow is essential when using custom scripts.

## Options and Tradeoffs in Recalibration

### Reference-Based Versus Reference-Free Approaches

Reference-based recalibration requires a trusted reference sequence to estimate empirical error rates. The advantage is accuracy. The error model is directly estimated from observed errors. The disadvantage is the requirement for a reference, which may not be available for non-model organisms or may be too divergent to provide reliable error estimates.

Reference-free approaches estimate error rates from the reads themselves, typically by examining read-to-read overlaps or by using the assembly as a proxy for the reference. These approaches are less accurate because they cannot distinguish sequencing errors from true biological variation. For polishing applications, reference-based recalibration is preferred when a suitable reference exists.

### Platform-Specific Recalibration

Different sequencing platforms have different error profiles, and recalibration should account for these differences. Oxford Nanopore reads have error rates that vary with base caller version and flow cell generation. PacBio HiFi reads have lower error rates but can still show systematic errors in specific sequence contexts.

Medaka models are trained on specific platform and base caller combinations, and the [EMBL-EBI training resources](https://www.ebi.ac.uk/training) provide learning pathways for understanding how platform-specific error profiles affect downstream analysis. The training materials cover data resource usage and practical analysis education, which helps researchers understand the context for platform-specific recalibration decisions.

### Recalibration Frequency

Recalibration is not a one-time operation. Sequencing conditions can change between runs, and base caller versions are updated regularly. A recalibration model estimated from one run may not apply to the next run, even on the same instrument.

The practical approach is to assess calibration for each new sequencing run or whenever the sequencing conditions change. The assessment is relatively inexpensive compared to the cost of a polishing error, and it provides confidence that the quality scores are reliable.

## Observations and Measurements for Calibration Assessment

### Metrics to Collect

The calibration assessment should collect several metrics for each quality score bin. The reported quality score is the starting point. The observed error rate is calculated from the alignment to the reference. The number of bases in each bin provides context for the reliability of the error rate estimate.

Additional metrics include the position in read, the sequence context around each base, and the strand. These covariates can reveal systematic patterns in error rates that a simple lookup table would miss. For example, if errors are concentrated at read ends, the recalibration model should account for position.

### Interpreting the Results

The interpretation of calibration assessment results depends on the magnitude of the discrepancies. Small discrepancies within the noise of the error rate estimate are not concerning. Large discrepancies that are consistent across multiple quality score bins indicate systematic miscalibration.

A useful visualization is a plot of reported quality score versus empirical error rate. A well-calibrated dataset produces points along the diagonal where reported and empirical values match. Miscalibrated datasets show systematic deviation from the diagonal, either above for overconfident scores or below for underconfident scores.

### Recording Calibration Information

Calibration information should be recorded as part of the analysis metadata. The record should include the tool versions, the reference used for calibration, the date of the assessment, and the observed discrepancies. This information is essential for reproducibility and for diagnosing problems if the polishing results are unexpected.

The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of sequence databases and analysis services that support the deposition and retrieval of sequencing data and associated metadata. Recording calibration information in a structured format facilitates data sharing and reanalysis by other researchers.

## Common Failure Patterns in Polishing Without Recalibration

### Overcorrection of True Variants

When quality scores are overconfident, polishing tools trust reads that contain errors. If the errors are consistent across multiple reads, the polisher may change a correct base in the assembly to an incorrect base. This overcorrection is particularly problematic in regions of biological variation, where the polisher may eliminate true variants.

Overcorrection is difficult to detect because the polished assembly looks clean and the quality scores are high. The errors are only revealed by comparison to an independent reference or by resequencing the region. Recalibration reduces the risk of overcorrection by ensuring that quality scores accurately reflect error probabilities.

### Undercorrection of Assembly Errors

When quality scores are underconfident, polishing tools ignore reads that contain correct information. The polisher retains assembly errors because the read evidence is not weighted sufficiently. Undercorrection is less damaging than overcorrection because the assembly is no worse than before polishing, but it wastes the opportunity to improve accuracy.

Undercorrection is more common when quality scores are systematically low, which can happen with older base caller versions or suboptimal sequencing conditions. Recalibration corrects the quality scores upward, allowing the polisher to use the read evidence effectively.

### Model Mismatch in Medaka

Medaka models are specific to sequencing platforms and base caller versions. Using a model that does not match the data can produce poor polishing results regardless of quality score calibration. The model mismatch can cause systematic errors that are difficult to diagnose.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on assembly polishing that include guidance on selecting appropriate Medaka models. Following these tutorials helps researchers avoid model mismatch and other common polishing errors.

### Inconsistent Results Across Polishing Rounds

Multiple rounds of polishing can produce inconsistent results, particularly when the quality scores are miscalibrated. The first round may correct some errors, but subsequent rounds may introduce new errors or reverse previous corrections. This instability is a sign that the polishing inputs are not reliable.

Recalibration before polishing reduces the risk of inconsistent results by providing stable quality scores across rounds. If instability persists after recalibration, the assembly itself may have structural errors that polishing cannot correct.

## Limitations of Quality Score Recalibration

### Recalibration Cannot Fix Systematic Sequencing Errors

Recalibration adjusts quality scores to match empirical error rates, but it cannot fix the underlying sequencing errors. If the sequencing process produces systematic errors in specific sequence contexts, recalibration will correctly report that those bases have low quality, but the errors remain in the reads. Polishing tools will still struggle in these regions because the read evidence is genuinely unreliable.

### Reference Divergence Limits Calibration Accuracy

The accuracy of reference-based recalibration depends on the divergence between the sample and the reference. If the reference is highly divergent, the observed mismatches include true biological differences in addition to sequencing errors. The error model will overestimate error rates, and the recalibrated quality scores will be too low.

For highly divergent references, the calibration should be interpreted with caution. The empirical error rates will be inflated, and the recalibrated quality scores will be conservative. Conservative quality scores are safer than overconfident scores for polishing, but they may reduce the effectiveness of polishing in regions where the reads are actually accurate.

### Recalibration Models Can Overfit

Recalibration models with many covariates can overfit the training data. The model may capture noise in the error rate estimates instead of true patterns. Overfit models perform poorly on new data and can introduce systematic biases in the recalibrated quality scores.

The risk of overfitting is reduced by using a simple model with few covariates and by validating the model on independent data. The verification step in the workflow is essential for detecting overfitting.

## Quality Controls and Validation After Polishing

### Independent Assembly Quality Metrics

After polishing, the assembly should be evaluated with metrics that are independent of the polishing process. BUSCO completeness scores measure the presence of conserved single-copy genes, providing a proxy for assembly completeness. Read mapping consistency measures whether the reads align cleanly to the polished assembly, indicating that the assembly is consistent with the read data.

The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) provide learning pathways for genome assembly and quality assessment that cover these metrics in detail. The training materials help researchers understand the strengths and limitations of different quality metrics.

### Comparison to Reference Genomes

When a reference genome is available, the polished assembly can be compared directly to the reference. The comparison reveals the number and types of remaining errors, including single nucleotide errors, insertions, deletions, and structural differences. This comparison is the most direct validation of polishing accuracy.

The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and associated annotation data that support these comparisons. The NCBI databases are the primary repository for genomic data and provide the infrastructure for comparative analysis.

### Read Back-Mapping Consistency

Mapping the original reads back to the polished assembly and examining the consistency of the alignments can reveal polishing errors. If the polished assembly contains an error, reads that were previously consistent with the assembly may now show mismatches at the error position. Conversely, if the polishing corrected an error, reads should now align more cleanly.

This validation approach is particularly useful for detecting overcorrection, where the polisher changed a correct base to an incorrect one. The read evidence at the changed position will show the discrepancy.

## Professional Escalation Criteria

### When to Seek Expert Assistance

Several situations warrant escalation to a bioinformatics specialist or the tool developers. If polishing produces inconsistent results across multiple rounds despite recalibration, the problem may be in the assembly structure instead of the quality scores. If the recalibration assessment shows unexpected patterns that cannot be explained by known sequencing conditions, the problem may be in the sequencing process itself.

The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline usage and troubleshooting. The nf-core community provides support for pipeline-related issues, and the documentation includes guidance on diagnosing common problems.

### When to Reconsider the Sequencing Strategy

If recalibration reveals severe quality score miscalibration that cannot be corrected, the sequencing strategy may need to be reconsidered. The sequencing conditions may be suboptimal, the base caller version may be inappropriate, or the sequencing platform may not be suitable for the application.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on sequencing data quality assessment that help researchers diagnose sequencing problems. The tutorials cover the relationship between sequencing conditions and data quality, providing context for deciding whether to resequence or proceed with the available data.

### Documentation for Reproducibility

All recalibration and polishing steps should be documented for reproducibility. The documentation should include the tool versions, parameters, reference sequences, and quality metrics at each step. This documentation is essential for troubleshooting and for publishing the analysis in a reproducible form.

The [Carpentries lessons](https://carpentries.org/lessons) provide training in reproducible data analysis practices, including version control, documentation, and workflow management. These practices are essential for maintaining the integrity of complex bioinformatics analyses.

## Decision Framework for Choosing Between Recalibration and Alternative Polishing Strategies

### When Recalibration Is the Correct Choice

Recalibration is the appropriate strategy when the draft assembly is structurally sound and the dominant source of residual error is base-level miscalls that polishing can correct. The decision to recalibrate should be based on three observable conditions in your data. First, the read alignment to a trusted reference shows a systematic discrepancy between reported Phred scores and empirical error rates. Second, the assembly itself shows good contiguity metrics, such as reasonable N50 values and complete BUSCO gene content, indicating that the scaffolding and contig layout are reliable. Third, the polishing errors you observe are distributed as single nucleotide substitutions and small indels instead of large structural rearrangements.

Under these conditions, recalibration directly addresses the root cause of poor polishing performance. The quality scores are the input that Racon weights and that Medaka uses alongside its neural network predictions. When those inputs are systematically wrong, no amount of parameter tuning in the polishing tool will fully compensate. Recalibration corrects the input data itself, which is a more fundamental fix than adjusting polishing parameters.

### When to Skip Recalibration and Polish Directly

Recalibration is not always necessary or beneficial. If your calibration assessment shows that the reported quality scores already match the empirical error rates within a factor of two at the quality score bins that dominate your data, the additional recalibration step adds computational cost without improving polishing accuracy. The assessment from Step 1 of the workflow provides the evidence for this decision.

Direct polishing without recalibration is also appropriate when you are using Medaka with a model that exactly matches your sequencing platform and base caller version. Medaka models are trained on data with specific error profiles, and the neural network has learned to interpret the quality score distributions from that platform. If the model matches and the quality scores are reasonably calibrated, the model can compensate for minor calibration discrepancies. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) provide learning pathways that explain how platform-specific error profiles affect model performance and help researchers understand when direct polishing is appropriate.

A third situation where skipping recalibration is reasonable is when you have a highly divergent reference for calibration. If the only available reference is from a different species or a very distant strain, the empirical error rates will be inflated by true biological divergence. The recalibrated quality scores will be systematically too low, and polishing will undercorrect assembly errors. In this case, direct polishing with the raw quality scores may produce better results than polishing with poorly calibrated scores.

### When to Reassemble Instead of Polish

Recalibration and polishing cannot fix structural assembly errors. If the draft assembly has misjoins, collapsed repeats, or incorrect contig ordering, polishing will only correct base-level errors within the flawed structure. The polished assembly will still have the structural problems, and the polishing process may even make the assembly worse by introducing errors in regions where the read evidence conflicts with the incorrect assembly structure.

Signs that reassembly is needed instead of polishing include highly fragmented assemblies with many short contigs, BUSCO scores that show substantial missing or fragmented genes, and read alignment patterns that show reads spanning contig boundaries in inconsistent ways. If your assembly shows these signs, the appropriate action is to revisit the assembly parameters or assembler choice instead of to invest effort in recalibration and polishing.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on assembly quality assessment that help researchers distinguish between base-level errors that polishing can fix and structural errors that require reassembly. These tutorials cover the relationship between assembly metrics and underlying assembly problems.

### When to Change the Sequencing Strategy

Severe quality score miscalibration that persists across multiple runs and cannot be corrected by recalibration may indicate a problem with the sequencing strategy itself. If the empirical error rates are dramatically higher than the reported quality scores across all quality bins, the base caller may be inappropriate for the sequencing conditions, or the sequencing chemistry may be suboptimal.

In this situation, the options include updating the base caller version, adjusting sequencing conditions such as flow cell loading or run duration, or switching to a different sequencing platform. The decision to change the sequencing strategy should be based on the calibration assessment results and the cost of resequencing compared to the cost of proceeding with unreliable data.

The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide access to sequencing data from diverse platforms and conditions, which can help researchers understand the quality score distributions that are achievable with different sequencing approaches. Comparing your calibration assessment to published data from similar platforms provides context for deciding whether your quality scores are within the expected range.

### Decision Matrix for Polishing Strategy Selection

| Condition | Recommended Strategy | Rationale |
| --- | --- | --- |
| Quality scores match empirical error rates within factor of two | Polish directly without recalibration | Recalibration adds cost without improving accuracy |
| Quality scores systematically overconfident by factor of two or more | Recalibrate before polishing | Overconfident scores cause overcorrection of true variants |
| Quality scores systematically underconfident by factor of two or more | Recalibrate before polishing | Underconfident scores cause undercorrection of assembly errors |
| Highly divergent reference only available for calibration | Polish directly with raw quality scores | Recalibration with divergent reference produces overly conservative scores |
| Assembly shows structural errors such as misjoins or collapsed repeats | Reassemble instead of polishing | Polishing cannot fix structural errors and may worsen them |
| Severe miscalibration persists across multiple runs | Reconsider sequencing strategy | Base caller or sequencing conditions may be inappropriate |
| Medaka model exactly matches platform and base caller | Polish directly with Medaka | Model compensates for minor calibration discrepancies |

### Implementing the Decision Framework in Practice

The decision framework should be applied at two points in the analysis workflow. The first decision point is after the initial assembly and before any polishing. At this point, run the calibration assessment from Step 1 of the workflow. The assessment results determine whether to recalibrate, polish directly, or reassemble.

The second decision point is after the first round of polishing. Evaluate the polished assembly using independent metrics such as BUSCO completeness and read back-mapping consistency. If the polished assembly shows improvement, proceed with additional polishing rounds or finalize the assembly. If the polished assembly shows no improvement or has degraded, revisit the decision framework to determine whether recalibration was insufficient, the assembly structure is flawed, or the sequencing data quality is the limiting factor.

This two-point application of the decision framework prevents wasted computational effort on polishing runs that cannot succeed and provides a structured approach to troubleshooting when polishing results are unexpected. The framework is particularly valuable for research groups processing multiple assemblies, where consistent decision criteria improve reproducibility across samples.

### Recording the Decision Rationale

Each polishing decision should be recorded with the evidence that supported it. The record should include the calibration assessment results, the assembly quality metrics, the polishing tool and parameters used, and the rationale for the chosen strategy. This documentation serves two purposes. It provides the metadata needed for reproducible analysis, and it creates a reference for future decisions on similar data.

The [Carpentries lessons](https://carpentries.org/lessons) provide training in reproducible data analysis practices, including documentation standards and workflow management. Applying these practices to polishing decisions ensures that the decision framework is applied consistently and that the rationale is available for review by collaborators or reviewers.

The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline usage and configuration that include structured logging and reporting. Integrating the decision rationale into the pipeline logs ensures that the information is captured automatically as part of the analysis workflow instead of requiring separate manual documentation.

## Frequently Asked Questions

### What is the difference between quality trimming and quality score recalibration?

Quality trimming removes low-quality bases from reads, reducing the amount of unreliable data that enters downstream analysis. Quality score recalibration changes the quality score values themselves so that reported scores accurately reflect empirical error rates. Trimming reduces data volume while recalibration preserves data volume but corrects the quality score values. For polishing, recalibration is generally more important because polishing algorithms need coverage across the entire assembly, and aggressive trimming can remove useful evidence.

### How do I know if my quality scores need recalibration?

Assess the calibration by aligning a sample of reads to a trusted reference and comparing the reported quality scores to the empirically observed error rates. If the reported scores are systematically higher than the empirical error rates justify, the scores are overconfident and recalibration is warranted. If the reported scores are systematically lower, the scores are underconfident and recalibration can improve polishing effectiveness. A discrepancy of more than a factor of two at quality score bins that contribute substantially to the data indicates a need for recalibration.

### Can I use fastp for long-read quality score recalibration?

fastp supports both short and long reads and can perform quality score recalibration for both data types. The tool estimates empirical error rates from an alignment to a reference and applies a correction to the quality scores. The recalibrated reads are output in FASTQ format and can be used as input for polishing tools such as Racon and Medaka.

### Does Medaka require recalibrated quality scores?

Medaka uses neural network models that are trained on specific sequencing platforms and base caller versions. The models learn error patterns from training data and may partially compensate for miscalibrated quality scores. However, the models cannot fully compensate for severe miscalibration, and recalibrated quality scores generally improve Medaka polishing accuracy. The most important factor for Medaka is using a model that matches the sequencing platform and base caller version.

### What reference should I use for recalibration?

Use a high-quality reference genome of the same species if one is available. If no same-species reference exists, a closely related species can serve as a proxy, but the calibration will be less accurate because the observed mismatches include true biological divergence. For highly divergent references, the recalibrated quality scores will be conservative, which is safer than overconfident scores but may reduce polishing effectiveness.

### How often should I recalibrate quality scores?

Assess calibration for each new sequencing run or whenever sequencing conditions change. Base caller versions are updated regularly, and flow cell generations change, both of which can affect quality score calibration. The assessment is relatively inexpensive and provides confidence that the quality scores are reliable for polishing.

### What are the signs that polishing introduced new errors?

Signs of polishing-induced errors include inconsistent results across multiple polishing rounds, read back-mapping inconsistencies at positions that were changed during polishing, and discrepancies between the polished assembly and an independent reference. If the polished assembly shows lower BUSCO completeness or higher read mapping inconsistency than the unpolished assembly, the polishing likely introduced errors.

### When should I escalate polishing problems to an expert?

Escalate when polishing produces inconsistent results despite recalibration, when the recalibration assessment shows unexpected patterns, or when the polished assembly fails independent quality checks. The nf-core community and the Galaxy Training Network provide support resources for diagnosing polishing problems. If the sequencing conditions are suspected to be the cause, reconsider the sequencing strategy before proceeding with additional analysis.

## Related Bioinformatics Guides

- [FASTQ File Format: Decoding Phred Quality Scores and Quality Control Workflows](/knowledge/bioinformatics/fastq-file-format-phred-scores-qc)
- [RNA-Seq Batch Effect Detection and Correction](/knowledge/bioinformatics/rna-seq-batch-effect-detection-and-correction)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Digital Pathology Scanners: A Buyer's Guide for Clinical and Research Use](/knowledge/bioinformatics/digital-pathology-scanners-a-buyer-s-guide-for-clinical-and-research-use)
- [Persistent Identifiers for Research Data: A Guide to Selection and Use](/knowledge/bioinformatics/persistent-identifiers-for-research-data-a-guide-to-selection-and-use)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [AI-driven cardiovascular risk prediction in patients with diabetes: bridging algorithmic innovation to equitable clinical application.](https://doi.org/10.3389/fmed.2026.1831220). 2026.
- [Designing an Indicator-Driven, Value-Based Architecture for Pneumonia Prevention in Japan: A Formative Policy Viewpoint on Adult Vaccination and Oral Care.](https://doi.org/10.2196/86912). 2026.
- [A Feasibility Study of Splintage by 3D Scanning and Printing: Process and Evaluation of Current 3D Printing Material.](https://doi.org/10.3390/ma19061146). 2026.
- [Fouling Control of Ion-Selective Electrodes (ISEs) in Aquatic and Aquacultural Environments: A Comprehensive Review.](https://doi.org/10.3390/s25247515). 2025.
- [Protocol for untargeted lipidomics of human serum using LC-TIMS-PASEF.](https://doi.org/10.1016/j.xpro.2026.104607). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.