# Troubleshooting Basecalling Failures in Oxford Nanopore Sequencing: Common Artifacts and Solutions


## Key Takeaways

- Basecalling failures in Oxford Nanopore sequencing manifest as low read identity, high unclassified reads, systematic substitutions, or pipeline crashes, stemming from signal acquisition, pore/chemistry degradation, model mismatch, or downstream artifacts.
- A systematic diagnostic approach is crucial, beginning with raw signal integrity checks (pore occupancy, read length distribution, yield) before evaluating basecalling model compatibility with flow cell and chemistry versions.
- Demultiplexing performance is sensitive to basecalling accuracy; high unclassified read fractions may necessitate re-basecalling with higher accuracy models or employing alignment-based demultiplexers like MysteryMaster.
- Methylation-induced errors at specific motifs can be mitigated by re-basecalling with modified-base-aware models or through raw signal reanalysis, preserving raw data is key for future corrections.
- Pipeline crashes are often due to insufficient computational resources (GPU memory, disk space) or incorrect configuration, requiring validation against documentation like nf-core standards.
- Maintaining detailed run metadata and a failure signature log is essential for identifying recurring patterns and enabling efficient, cross-run troubleshooting and knowledge building.

---

Basecalling failures in Oxford Nanopore sequencing typically present as low read identity, high unclassified read fractions, systematic base substitutions, or outright pipeline crashes. These failures usually trace to one of four root causes: signal acquisition problems, pore or chemistry degradation, model mismatch between the sequencing chemistry and the basecaller, or downstream analysis artifacts that are mistaken for basecalling errors. This article provides a systematic diagnostic approach for researchers and laboratory professionals who need to distinguish between these causes and apply targeted corrections.

The diagnostic framework presented here follows a decision-tree logic. Start with the raw signal data, then move to the basecalling model, then to demultiplexing, and finally to downstream alignment and variant calling. Each stage has distinct failure signatures that point to specific corrective actions. The guidance draws on published datasets, open-source tooling documentation, and community workflow standards.

## At a Glance

The table below summarizes the most common basecalling failure modes, their typical signatures, likely root causes, and first-line corrective actions. Use this table as a triage tool before diving into detailed troubleshooting sections.

| Failure Signature | Observed Symptom | Likely Root Cause | First-Line Correction |
| --- | --- | --- | --- |
| Low read identity across all reads | Alignment identity below expected threshold for the chemistry and model combination | Model mismatch, degraded flow cell, or poor library preparation | Confirm the basecalling model matches the flow cell version and chemistry kit, then check pore occupancy and read length distribution |
| High fraction of unclassified reads after demultiplexing | Many reads lack a barcode assignment or map to multiple barcodes | Barcode sequence errors, adapter contamination, or basecalling errors in barcode regions | Re-run demultiplexing with a higher accuracy basecall model or use an alternative demultiplexer that performs alignment-based barcode assignment |
| Systematic base substitutions at specific motifs | Recurring errors at methylated sites or homopolymer regions | Methylation-induced basecalling errors or model limitations on repetitive sequence | Re-basecall with a model that accounts for modified bases or use raw signal reanalysis |
| Basecalling pipeline crash or out-of-memory error | Process terminates during basecalling or consumes excessive resources | Incorrect configuration, insufficient computational resources, or corrupted input files | Validate input file integrity, check workflow configuration against the pipeline documentation, and allocate resources according to the pipeline requirements |
| High error rate in specific read subsets | Errors concentrated in reads from one barcode or one sample | Sample-specific issues such as degraded input DNA or contamination | Check per-sample quality metrics and library preparation records for the affected sample |

## Understanding the Basecalling Pipeline Architecture

Basecalling converts raw electrical current signals from the nanopore into nucleotide sequences. This conversion happens through neural network models that learn the relationship between ionic current disruptions and the DNA or RNA sequence passing through the pore. The basecaller output feeds directly into downstream steps including demultiplexing, alignment, and variant calling. Errors introduced at the basecalling stage propagate through the entire analysis, so identifying whether a problem originates in basecalling or in a later step is the first diagnostic priority.

### Signal Acquisition and Raw Data Integrity

The raw signal data, often stored in POD5 or FAST5 format, contains the measured ion current values over time. These signals retain information beyond the nucleotide sequence, including evidence for nucleotide modifications such as methylation. A published dataset of nanopore raw signals from bacterial isolates demonstrates that raw ion current data can support reanalysis for modified bases, resistance genes, and strain differentiation. This dataset, generated on R10.4.1 flow cells with V14 chemistry, produced roughly 16 gigabases per flow cell across 79 bacterial isolates from six species. The availability of raw signal data enables researchers to re-basecall with newer models as they become available, which is particularly useful for correcting methylation-based calling errors that affect SNP profiling and core genome multilocus sequence typing.

Before troubleshooting basecalling failures, verify that the raw signal files are complete and uncorrupted. Check the file count against the sequencing run output, confirm that the flow cell produced the expected yield, and inspect the pore occupancy plots from the sequencing software. If the raw signal files are incomplete or corrupted, no basecalling model will produce usable results.

### Basecalling Models and Chemistry Compatibility

Oxford Nanopore basecallers use different models optimized for specific flow cell versions, chemistry kits, and accuracy tiers. The choice between Fast, High Accuracy (HAC), and Super Accuracy (SUP) models involves a tradeoff between computational cost and read accuracy. A comparative study of demultiplexing tools across three datasets found that the choice of basecalling model affects downstream demultiplexing performance. The study compared Oxford Nanopore's Dorado and Guppy tools against a third-party demultiplexer called MysteryMaster across Fast, HAC, and SUP basecalled data. The results showed that demultiplexing performance varied by basecalling model, with the third-party tool performing slightly better on Fast basecalled data while performing similarly to Dorado on HAC and SUP data.

The practical implication is that basecalling model selection is not neutral. If you plan to demultiplex barcoded samples, the basecalling model affects how many reads receive correct barcode assignments. Using a lower accuracy model such as Fast may increase the unclassified read fraction, while using SUP increases computational time. The optimal choice depends on your accuracy requirements and available compute resources.

## Diagnostic Workflow for Basecalling Failures

The following workflow provides a structured approach to identifying the root cause of basecalling failures. Work through these steps in order, recording observations at each stage.

### Step 1: Verify Raw Signal Quality

Start by examining the raw signal data before any basecalling occurs. Check the following metrics from the sequencing run output:

- Pore occupancy over time, which indicates how many pores were actively sequencing
- Read length distribution, which reveals whether reads are being truncated prematurely
- Yield per flow cell, which should match expected values for the chemistry and flow cell version
- Channel activity, which identifies whether specific channels failed

If pore occupancy dropped early in the run, the flow cell may have degraded or the library may have contained inhibitors. If read lengths are shorter than expected, the library preparation may have fragmented the DNA or the pore may have become blocked. These signal-level issues produce poor basecalling results regardless of the model used.

### Step 2: Confirm Model and Chemistry Match

Verify that the basecalling model matches the flow cell version and chemistry kit used for the sequencing run. Using a model designed for a different chemistry produces systematic errors that appear as low identity across all reads. Check the basecaller documentation for the correct model names and versions.

The raw signal dataset described earlier used R10.4.1 flow cells with V14 chemistry, which represents a current standard configuration. If your run used different flow cells or chemistry, confirm that the basecalling model corresponds to that specific combination. Model mismatches are among the most common causes of poor basecalling quality and are also the easiest to fix by re-running basecalling with the correct model.

### Step 3: Assess Basecalling Accuracy Metrics

After basecalling, examine the output quality metrics. Most basecallers report per-read quality scores, and downstream alignment tools provide identity statistics. Compare these metrics against expected values for your chemistry and model combination. If the observed accuracy falls substantially below expectations, proceed to the next diagnostic steps.

### Step 4: Evaluate Demultiplexing Performance

For multiplexed runs, check the fraction of reads assigned to each barcode and the fraction left unclassified. A high unclassified fraction may indicate basecalling errors in the barcode regions, adapter contamination, or barcode sequence diversity that the demultiplexer cannot resolve. The MysteryMaster study provides evidence that alternative demultiplexing approaches can recover reads that standard tools leave unclassified. The study demonstrated that a sequential approach using Dorado followed by MysteryMaster produced the best overall performance, with the third-party tool achieving a false positive rate of 0.41 percent with default settings.

If demultiplexing performance is poor, consider re-basecalling with a higher accuracy model or using an alignment-based demultiplexer that can tolerate errors in barcode sequences.

### Step 5: Examine Downstream Alignment and Variant Calling

If basecalling and demultiplexing appear satisfactory but downstream analysis still shows problems, examine the alignment and variant calling steps. Poor alignment may result from incorrect reference sequences, inappropriate alignment parameters, or structural variants that the aligner cannot handle. Long-read sequencing excels at detecting structural variants and telomere analysis, but these applications require specialized pipelines. The TARPON pipeline for telomere analysis demonstrates how application-specific processing steps, including quality filtering and boundary detection, are necessary to obtain reliable results from nanopore data.

## Common Failure Patterns and Their Corrections

### Methylation-Induced Basecalling Errors

Modified bases such as 5-methylcytosine alter the ionic current signal in ways that can confuse basecallers. These modifications produce systematic errors at specific sequence motifs, which can affect SNP calling and genotyping accuracy. The raw signal dataset publication emphasizes that re-basecalling with models that account for modified bases can mitigate methylation-based calling errors, enhancing the reliability of SNP profiling and core genome multilocus sequence typing analyses.

If you observe systematic base substitutions at CpG motifs or other methylation-prone sequences, consider the following actions:

- Re-basecall using a model that detects modified bases
- Compare basecalls from modified-base-aware and standard models to identify affected sites
- Use raw signal reanalysis to confirm methylation status at error-prone positions

The ability to re-basecall from raw signals is a key advantage of nanopore sequencing. Raw signal data retains information that can support future model improvements, so archiving raw signals instead of only basecalled reads enables retrospective correction of systematic errors.

### Barcode Misassignment and Unclassified Reads

Demultiplexing failures produce two distinct problems: reads assigned to the wrong sample and reads left unclassified. Both problems waste data and can compromise downstream analysis. The MysteryMaster study provides evidence that demultiplexing performance depends on both the basecalling model and the demultiplexing algorithm. The study compared three tools across three basecalling models and found that the optimal tool choice depended on the basecalling model used.

For barcode-related failures, consider these corrections:

- Re-basecall with a higher accuracy model to reduce errors in barcode regions
- Use a demultiplexer that performs alignment-based barcode assignment instead of exact matching
- Apply a sequential approach that combines multiple demultiplexers to maximize read recovery
- Check barcode sequences for compatibility with the demultiplexer's expected format

The false positive rate of 0.41 percent reported for MysteryMaster with default settings indicates that alignment-based demultiplexing can achieve high specificity, but the optimal configuration depends on your data characteristics.

### Pore Blockage and Read Truncation

Pore blockage occurs when DNA or RNA molecules become stuck in the pore, preventing further sequencing. This produces truncated reads and reduced yield. Pore blockage can result from library preparation artifacts, contaminants, or secondary structures in the nucleic acid. While pore blockage is primarily a sequencing problem instead of a basecalling problem, it manifests as poor basecalling quality because the basecaller receives incomplete signal data.

If pore blockage is the suspected cause, examine the pore occupancy plot and the read length distribution. A high proportion of short reads or a rapid decline in active pores suggests blockage. Corrective actions include optimizing library preparation to remove contaminants, adjusting the sequencing run parameters, and using fresh flow cells for critical samples.

### Computational Resource Failures

Basecalling with high accuracy models requires substantial computational resources. The SUP model in particular demands significant GPU memory and processing time. Pipeline crashes during basecalling often result from insufficient resources or incorrect configuration. The [nf-core documentation](https://nf-co.re/docs) provides standards for reproducible pipeline configuration, and following these standards can prevent many configuration-related failures.

If basecalling crashes or terminates unexpectedly, check the following:

- GPU memory availability and driver compatibility
- Disk space for intermediate and output files
- Configuration file syntax and parameter values
- Input file format and integrity

The [Galaxy Training Network](https://training.galaxyproject.org/) and [Bioconductor](https://bioconductor.org/) provide training materials for reproducible genomic analysis that include guidance on resource allocation and pipeline configuration. These resources can help laboratory professionals develop the computational skills needed to diagnose and fix basecalling pipeline failures.

## Demultiplexing Strategies for Problematic Reads

Demultiplexing assigns reads to samples based on barcode sequences. When basecalling errors affect barcode regions, standard demultiplexers may fail to assign reads correctly. The MysteryMaster study provides a detailed comparison of demultiplexing approaches and offers practical guidance for improving read classification.

### Comparing Demultiplexing Tools

The study compared Oxford Nanopore's Dorado and Guppy demultiplexers against MysteryMaster, a demultiplexer that uses the Cola aligner for sequence alignment. Across three datasets of 37 diverse samples with established ground truth, the study found that MysteryMaster identified a similar or greater percentage of reads compared to the standard tools across Fast, HAC, and SUP basecalling models. The performance difference was most notable for Fast basecalled data, where MysteryMaster performed slightly better than the other tools.

The practical implication is that the choice of demultiplexer matters, particularly for lower accuracy basecalling models. If you use Fast basecalling for rapid turnaround, consider using an alignment-based demultiplexer to maximize read recovery. For HAC and SUP data, the standard tools perform comparably to the third-party option.

### Sequential Demultiplexing Approach

The study concluded that the sequential application of Dorado followed by MysteryMaster produced the best overall performance. This approach uses Dorado for initial demultiplexing and then applies MysteryMaster to the unclassified reads to recover additional assignments. The sequential approach maximizes read recovery while maintaining low false positive rates.

Implementing a sequential demultiplexing workflow requires attention to the output formats and the handling of reads that remain unclassified after both tools. Document the fraction of reads assigned at each stage to track the improvement from the sequential approach.

## Application-Specific Basecalling Considerations

Different applications impose different requirements on basecalling accuracy and downstream analysis. The following sections address application-specific considerations for common nanopore sequencing use cases.

### Bacterial Genotyping and Methylation Analysis

Bacterial whole-genome sequencing with nanopore technology supports genotyping, resistance gene detection, and methylation analysis. The raw signal dataset described earlier was generated specifically to support these applications, with raw signals from 79 isolates across six bacterial species. The dataset includes triplicate samples from three different laboratories, enabling assessment of inter-laboratory reproducibility.

For bacterial genotyping applications, basecalling accuracy directly affects SNP detection and core genome multilocus sequence typing. Methylation-induced basecalling errors can produce false SNPs at modified sites, compromising strain differentiation. The dataset publication emphasizes that re-basecalling with future models can mitigate these errors, highlighting the value of retaining raw signal data.

Practical recommendations for bacterial genotyping:

- Retain raw signal data to enable re-basecalling with improved models
- Use modified-base-aware basecalling models when methylation analysis is a goal
- Validate SNP calls against known reference sequences or orthogonal methods
- Include replicate samples to assess technical variability

### Telomere Analysis

Telomere analysis requires nucleotide-level resolution of chromosome arm-specific telomere length. The TARPON pipeline demonstrates a best-practices approach for this application, using Nextflow for reproducible workflow execution. The pipeline isolates telomeric repeat-containing reads, assigns strand specificity, and applies quality filtering to remove low-quality or subtelomeric reads.

Basecalling errors in telomeric repeats are particularly problematic because these regions consist of repetitive sequence that can confuse basecallers. The TARPON pipeline addresses this by applying quality filtering after telomere boundary detection, ensuring that only full-length telomeres are included in the analysis.

For telomere analysis applications, consider the following:

- Use the highest accuracy basecalling model feasible for your computational resources
- Apply application-specific quality filtering to remove reads that do not meet telomere analysis requirements
- Validate results against known telomere length standards or orthogonal methods
- Document the pipeline parameters used for each analysis to ensure reproducibility

The TARPON pipeline can be executed via command line or integrated into ONT's EPI2ME agent, providing options for users with different computational backgrounds. The container-based architecture eliminates dependency conflicts, which is particularly valuable for clinical or diagnostic contexts where reproducibility is critical.

### Direct RNA Sequencing

Direct RNA sequencing (DRS) presents unique basecalling challenges because RNA molecules differ from DNA in their chemical properties and may contain modified bases such as pseudouridine. A published protocol for pseudouridine detection using nanopore DRS describes the use of unmodified transcriptome controls to enable differential analysis of RNA modifications.

The protocol covers RNA extraction, poly(A) selection, DRS library preparation, and knockdowns of writer proteins as critical controls. For basecalling troubleshooting in DRS applications, the key considerations include:

- Using basecalling models optimized for RNA instead of DNA
- Including unmodified transcriptome controls to distinguish modification signals from basecalling errors
- Applying modification-specific analysis tools that use raw signal data
- Validating pseudouridine detection against orthogonal methods

The protocol emphasizes the importance of controls for distinguishing true modifications from technical artifacts. Without unmodified controls, basecalling errors can be mistaken for RNA modifications, producing false biological conclusions.

### Metabarcoding and Environmental Applications

Nanopore sequencing supports metabarcoding applications such as pollen DNA analysis for ecological research. A study of honeybee foraging behavior used nanopore metabarcoding with the trnL chloroplast region to assess how formic acid treatment influences foraging preferences. The study detected significant differences in foraging composition between treated and control hives, demonstrating the utility of nanopore sequencing for ecological monitoring.

For metabarcoding applications, basecalling accuracy affects species identification because errors in the amplified marker region can lead to incorrect taxonomic assignments. The study used real-time nanopore sequencing, which requires basecalling during the sequencing run instead of after completion. Real-time basecalling imposes constraints on model selection because computational resources must keep pace with sequencing throughput.

Practical recommendations for metabarcoding:

- Use reference databases appropriate for the amplified marker region
- Apply taxonomic assignment methods that tolerate basecalling errors
- Include positive and negative controls to assess contamination and amplification artifacts
- Consider the tradeoff between basecalling accuracy and real-time analysis requirements

## Records and Measurements for Basecalling Troubleshooting

Systematic troubleshooting requires systematic record keeping. Maintain the following records for each sequencing run to support diagnosis of basecalling failures:

### Run-Level Records

- Flow cell lot number and version
- Chemistry kit version and lot number
- Library preparation protocol and any deviations
- Basecalling model name and version
- Basecaller software version and configuration parameters
- Demultiplexer software version and parameters
- Computational resources used for basecalling

### Quality Metrics

- Total yield in gigabases
- Read N50 and read length distribution
- Pore occupancy over time
- Per-read quality score distribution
- Alignment identity statistics
- Unclassified read fraction after demultiplexing
- Per-barcode read counts and quality metrics

### Troubleshooting Log

- Date and time of the failure
- Observed symptoms and error messages
- Steps taken to diagnose the problem
- Corrective actions applied
- Results after correction

These records enable comparison across runs and support identification of systematic issues. For example, if multiple runs using the same chemistry kit show similar basecalling errors, the problem may be kit-specific instead of run-specific.

## Common Failure Patterns in Basecalling Pipelines

The following failure patterns recur across nanopore sequencing projects. Recognizing these patterns speeds up diagnosis and correction.

### Pattern 1: Uniform Quality Degradation

When all reads show lower than expected quality, the cause is typically a model mismatch, degraded flow cell, or poor library preparation. Check the model and chemistry compatibility first, then examine the flow cell metrics. If the flow cell produced low pore occupancy or short reads, the problem is upstream of basecalling.

### Pattern 2: Barcode-Specific Failures

When quality problems concentrate in reads from specific barcodes, the cause may be sample-specific degradation, contamination, or barcode sequence issues. Check the library preparation records for the affected samples and examine the barcode sequences for compatibility with the demultiplexer.

### Pattern 3: Motif-Specific Errors

When errors occur at specific sequence motifs, such as methylated sites or homopolymers, the cause is likely a model limitation. Re-basecalling with a modified-base-aware model or a higher accuracy model may resolve the issue. Raw signal reanalysis can confirm whether the errors correspond to true modifications.

### Pattern 4: Pipeline Crashes

When the basecalling pipeline crashes, the cause is typically computational resource exhaustion or configuration errors. Check GPU memory, disk space, and configuration file syntax. The [nf-core documentation](https://nf-co.re/docs) provides standards for pipeline configuration that can prevent many of these failures.

### Pattern 5: Demultiplexing Failures

When a high fraction of reads remain unclassified after demultiplexing, the cause may be basecalling errors in barcode regions or limitations of the demultiplexing algorithm. Consider re-basecalling with a higher accuracy model or using an alignment-based demultiplexer.

## Limitations of Basecalling Troubleshooting

Basecalling troubleshooting has inherent limitations that should be acknowledged. The following limitations affect the diagnostic process:

### Model Availability

Basecalling models are developed for specific chemistry and flow cell combinations. If you use a non-standard configuration, appropriate models may not be available. Check the basecaller documentation for supported configurations before starting a sequencing run.

### Computational Constraints

Higher accuracy basecalling models require more computational resources. Laboratories without access to high-performance GPUs may be limited to lower accuracy models, which can affect downstream analysis quality. The [Galaxy Training Network](https://training.galaxyproject.org/) provides training on using shared computational infrastructure for genomic analysis, which can help laboratories access the resources needed for high accuracy basecalling.

### Reference Dependence

Alignment-based quality assessment requires an appropriate reference sequence. For organisms without a reference genome or for metagenomic samples, assessing basecalling quality is more challenging. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides sequence resources that can support reference-based quality assessment for many organisms.

### Error Correction Limitations

Basecalling errors can sometimes be corrected through downstream analysis, but not all errors are correctable. Errors in repetitive regions or regions with high modification density may persist through error correction. Understanding the limitations of error correction helps set realistic expectations for basecalling quality.

## Professional Escalation Criteria

Some basecalling problems require escalation to specialized support or collaboration with bioinformatics experts. Escalate when:

- Basecalling failures persist across multiple runs with different flow cells and chemistry kits
- Systematic errors appear at specific motifs and re-basecalling with different models does not resolve them
- Demultiplexing failures produce unclassified read fractions that exceed acceptable thresholds for your application
- Pipeline crashes cannot be resolved through configuration changes
- You need to implement application-specific pipelines such as TARPON and require guidance on parameter selection

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides learning pathways for bioinformatics data analysis that can help laboratory professionals develop the skills needed to address complex basecalling problems. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational computing skills that support reproducible data analysis.

## Building a Basecalling Troubleshooting Record System and Decision Framework

Systematic troubleshooting of basecalling failures requires more than isolated corrective actions. Without a structured record system and decision framework, laboratories repeat the same diagnostic steps across runs, fail to recognize recurring patterns, and lose time on problems that previous runs already solved. This section provides a practical record system, a tiered decision framework, and a comparison of diagnostic approaches that complement the workflow described earlier.

### The Run Metadata Register

The first component of a functional troubleshooting system is a run metadata register that captures every variable that can influence basecalling outcomes. This register differs from the run-level records described earlier because it organizes information in a queryable format that supports cross-run comparison. Create a spreadsheet or database table with one row per sequencing run and columns for each of the following categories.

Flow cell and chemistry variables include the flow cell lot number, flow cell version, chemistry kit version, and kit lot number. These four fields identify the physical consumables and enable detection of lot-specific defects. If two runs using the same flow cell lot show similar failure signatures, the problem may trace to manufacturing variability instead of laboratory procedure.

Library preparation variables include the input nucleic acid type, extraction method, library preparation kit version, input quantity, and any deviations from the standard protocol. Record the person who performed the preparation and the date. Human variability in library preparation is a documented source of sequencing quality differences, and the personnel field supports identification of technique-related patterns.

Basecalling variables include the basecaller software name and version, the model name and version, the accuracy tier, the compute hardware used, and the basecalling parameters. Record whether basecalling ran in real time during sequencing or as a post-run step. Real-time basecalling imposes different resource constraints than post-run basecalling, and the distinction matters when diagnosing pipeline crashes.

Downstream analysis variables include the demultiplexer name and version, the alignment tool and version, the reference sequence version, and the variant caller or analysis pipeline used. These fields support diagnosis of problems that appear to be basecalling failures but actually originate in downstream steps.

The register should include a free-text notes column for observations that do not fit structured fields. Examples include unusual pore occupancy plots, unexpected yield patterns, or environmental conditions such as laboratory temperature fluctuations. These unstructured observations often provide the clue that structured fields miss.

### The Failure Signature Log

The second component is a failure signature log that records each basecalling problem in a standardized format. This log enables pattern recognition across runs and supports the development of laboratory-specific troubleshooting knowledge. Each entry should capture the following elements.

The observed symptom describes what the user sees, such as low read identity, high unclassified fraction, systematic substitutions, or pipeline crash. Use the language of the sequencing software output and downstream analysis reports instead of interpretive language. For example, record "alignment identity 82 percent" instead of "poor quality reads."

The failure context records the run metadata register fields for the affected run, including flow cell lot, chemistry version, basecalling model, and demultiplexer. This context enables queries such as "show all failures with this flow cell lot" or "list all runs using this basecalling model version."

The diagnostic steps taken records each action performed during troubleshooting, the date and time of each action, and the person who performed it. This field prevents duplicate diagnostic work and provides a basis for evaluating which diagnostic steps produce useful information.

The root cause determination records the final diagnosis, even if the cause remains unknown. For unresolved cases, record the leading hypotheses and the evidence for and against each hypothesis. This documentation supports future diagnosis when new information becomes available.

The corrective action and outcome records what was changed and whether the change resolved the problem. Include quantitative before and after metrics where available. For example, record the unclassified read fraction before and after switching demultiplexers.

The recurrence flag indicates whether the same failure signature has appeared in previous runs. When a signature recurs, the log supports rapid diagnosis by surfacing previously successful corrective actions.

### The Tiered Decision Framework

The decision framework presented here organizes troubleshooting into three tiers based on the scope of the problem and the resources required to address it. This tiered structure prevents over-investment in simple problems and under-investment in complex ones.

Tier one addresses single-run failures with obvious causes. These include model chemistry mismatches, configuration errors, and computational resource exhaustion. The diagnostic workflow described earlier in this article resolves most tier one problems within hours. The corrective actions are model changes, configuration corrections, and resource reallocation. Document the resolution in the failure signature log and proceed with the analysis.

Tier two addresses problems that recur across runs or that persist after tier one corrections. These include systematic motif-specific errors, persistent demultiplexing failures, and quality degradation that correlates with specific library preparation batches. Tier two problems require cross-run comparison using the run metadata register and failure signature log. The corrective actions may involve changing library preparation protocols, switching basecalling models, or implementing alternative demultiplexing strategies. The MysteryMaster study provides evidence that alternative demultiplexing approaches can recover reads that standard tools leave unclassified, and the study's finding that sequential application of Dorado followed by MysteryMaster produced the best overall performance supports a tier two corrective action for persistent demultiplexing failures.

Tier three addresses problems that resist resolution through standard corrective actions and require specialized expertise or external support. These include novel failure signatures, problems that persist across multiple flow cell lots and chemistry kit versions, and issues that require custom pipeline development. Tier three problems warrant escalation to bioinformatics support, instrument manufacturer technical support, or collaboration with specialized analysis groups. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program and [Galaxy Training Network](https://training.galaxyproject.org/) provide educational pathways that can help laboratory personnel develop the skills to address tier three problems internally, while the [nf-core documentation](https://nf-co.re/docs) supports implementation of community-standard pipelines that may resolve persistent configuration issues.

### Quantitative Thresholds for Decision Points

The tiered framework requires quantitative thresholds to guide decision-making. These thresholds should be established locally based on the laboratory's historical performance data and the requirements of the specific application. The following threshold categories provide a starting point for developing laboratory-specific criteria.

Yield thresholds compare observed yield against expected yield for the flow cell and chemistry combination. The raw signal dataset publication reports an average of 16 gigabases per flow cell for R10.4.1 flow cells with V14 chemistry across 79 bacterial isolates. If a run produces substantially less than the expected yield, investigate signal acquisition problems before basecalling issues. Establish a local threshold, such as 70 percent of expected yield, below which the run triggers a tier one investigation.

Read identity thresholds compare alignment identity against expected values for the basecalling model and chemistry combination. Establish separate thresholds for Fast, HAC, and SUP models because their expected accuracy differs substantially. The MysteryMaster study demonstrates that basecalling model choice affects downstream demultiplexing performance, which implies that identity thresholds should be model-specific.

Unclassified read fraction thresholds guide demultiplexing investigations. Establish a baseline unclassified fraction from historical runs using the same barcoding kit and demultiplexing approach. When the unclassified fraction exceeds the baseline by a defined margin, such as 50 percent above the historical mean, trigger a demultiplexing investigation. The MysteryMaster study reports a false positive rate of 0.41 percent with default settings, which provides a reference point for evaluating demultiplexer specificity.

Recurrence thresholds trigger tier escalation. Define a problem as recurring when the same failure signature appears in a specified number of runs within a defined time window, such as three occurrences within ten runs. Recurrence triggers a tier two investigation that includes cross-run comparison using the run metadata register.

### The Diagnostic Decision Matrix

The decision matrix below organizes common failure signatures against diagnostic actions and corrective options. This matrix complements the failure signature log by providing a structured reference for selecting diagnostic steps.

| Failure Signature | Primary Diagnostic Action | Secondary Diagnostic Action | Corrective Options |
| --- | --- | --- | --- |
| Low identity across all reads | Verify model chemistry match | Check pore occupancy and read length distribution | Re-basecall with correct model, investigate flow cell quality |
| High unclassified fraction | Examine per-barcode read counts | Compare demultiplexer performance across models | Re-basecall with higher accuracy model, use alignment-based demultiplexer |
| Motif-specific substitutions | Identify error positions and sequence context | Compare modified-base-aware and standard model outputs | Re-basecall with modified-base-aware model, use raw signal reanalysis |
| Pipeline crash | Check resource availability | Validate configuration against pipeline documentation | Reallocate resources, correct configuration, verify input integrity |
| Barcode-specific quality problems | Review library preparation records for affected samples | Examine barcode sequences for compatibility | Repeat library preparation, switch barcoding kit, use alternative demultiplexer |
| Read truncation | Examine read length distribution | Check pore occupancy over time | Optimize library preparation, adjust run parameters, use fresh flow cell |

The matrix supports rapid triage by matching the observed signature to the appropriate diagnostic actions. The corrective options column provides starting points for resolution, with the understanding that the optimal correction depends on the specific context recorded in the failure signature log.

### Implementing the Record System

Implementation of this record system requires commitment to consistent data entry and regular review. Assign responsibility for maintaining the run metadata register and failure signature log to a specific laboratory role, such as the sequencing facility manager or a designated bioinformatics lead. Establish a review schedule, such as monthly, to examine the logs for patterns and update the decision thresholds based on accumulated data.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in spreadsheet organization, data management, and reproducible workflows that support consistent record keeping. The [Bioconductor](https://bioconductor.org/) project provides tools for reproducible genomic analysis that can support automated quality metric collection and reporting. The [Galaxy Training Network](https://training.galaxyproject.org/) offers tutorials on workflow construction that can help automate the collection of quality metrics into structured formats.

### Common Implementation Failures

Laboratories commonly encounter several implementation failures when establishing a basecalling troubleshooting record system. Recognizing these failures helps prevent them.

The first implementation failure is incomplete metadata capture. Laboratories often record the basecalling model but omit the model version, or record the flow cell version but omit the lot number. These omissions prevent cross-run comparison and pattern recognition. Address this failure by using a structured form with required fields and validation rules that prevent submission of incomplete records.

The second implementation failure is inconsistent terminology. Different laboratory members may use different terms for the same failure signature, such as "low quality" versus "poor identity" versus "bad reads." This inconsistency prevents effective querying of the failure signature log. Address this failure by establishing a controlled vocabulary for failure signatures and training laboratory members on its use.

The third implementation failure is the absence of baseline data. Laboratories attempt to set decision thresholds without historical performance data, resulting in thresholds that are either too strict or too lenient. Address this failure by collecting quality metrics for a defined period, such as ten runs, before setting thresholds. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides sequence resources that can support reference-based quality assessment for establishing baseline identity thresholds.

The fourth implementation failure is the failure to review the logs. Laboratories maintain the records but never examine them for patterns, defeating the purpose of the system. Address this failure by scheduling regular review meetings and assigning responsibility for presenting patterns and recommendations based on the log data.

### Comparison of Diagnostic Approaches

The record system and decision framework described here complement the diagnostic workflow presented earlier in this article. The workflow provides a linear sequence of diagnostic steps for individual failures, while the record system provides the infrastructure for pattern recognition across runs. The decision framework provides escalation criteria that determine when individual troubleshooting transitions to systematic investigation.

The [nf-core documentation](https://nf-co.re/docs) describes community standards for reproducible pipeline configuration that support consistent basecalling and downstream analysis. Adopting these standards reduces configuration variability across runs, which simplifies the interpretation of the run metadata register. When every run uses the same pipeline configuration, differences in outcomes more clearly trace to differences in flow cells, chemistry, libraries, or samples.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides learning pathways that cover data management and reproducible analysis practices. These skills support the implementation of the record system described here. The [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible tutorials on quality assessment and workflow construction that can help laboratory personnel develop the skills needed to maintain and interpret the records.

### Records as the Foundation for Troubleshooting

The record system described in this section transforms basecalling troubleshooting from a reactive process into a systematic practice. The run metadata register captures the variables that influence basecalling outcomes. The failure signature log records problems in a standardized format that supports pattern recognition. The tiered decision framework guides resource allocation based on problem scope. The quantitative thresholds provide objective criteria for decision-making. The diagnostic decision matrix supports rapid triage of common failure signatures.

Laboratories that implement this system accumulate institutional knowledge that reduces troubleshooting time with each successive run. When a failure signature recurs, the log provides previously successful corrective actions. When a new failure signature appears, the metadata register supports identification of variables that changed. This institutional knowledge is particularly valuable in facilities with high personnel turnover, where the record system preserves knowledge that would otherwise leave with departing staff.

The raw signal dataset publication emphasizes the value of retaining raw signal data for re-basecalling with future models. The record system described here extends this principle to the metadata and quality metrics that support troubleshooting. Just as raw signals enable retrospective re-analysis, comprehensive records enable retrospective diagnosis of problems that were not recognized at the time of the run.

## Frequently Asked Questions

### What is the most common cause of poor basecalling quality in nanopore sequencing?

The most common cause is a mismatch between the basecalling model and the flow cell or chemistry version. Using a model designed for a different chemistry produces systematic errors across all reads. Check the basecaller documentation for the correct model for your specific flow cell and chemistry kit combination. The raw signal dataset publication demonstrates that R10.4.1 flow cells with V14 chemistry represent a current standard configuration, and models should be selected to match this combination.

### How do I know if my basecalling failures are caused by methylation?

Methylation-induced basecalling errors appear as systematic base substitutions at specific sequence motifs, particularly CpG sites in eukaryotic genomes. If you observe recurring errors at the same positions across multiple reads, methylation may be the cause. Re-basecalling with a modified-base-aware model can confirm whether the errors correspond to true modifications. The raw signal dataset publication emphasizes that raw signal data retains information about nucleotide modifications, enabling reanalysis with improved models.

### Why do I have so many unclassified reads after demultiplexing?

High unclassified read fractions typically result from basecalling errors in barcode regions or limitations of the demultiplexing algorithm. The MysteryMaster study found that demultiplexing performance varies by basecalling model, with lower accuracy models producing more unclassified reads. Consider re-basecalling with a higher accuracy model or using an alignment-based demultiplexer that can tolerate errors in barcode sequences.

### Should I use Fast, HAC, or SUP basecalling models?

The choice depends on your accuracy requirements and computational resources. Fast basecalling is quick but produces lower accuracy reads, which can affect demultiplexing and variant calling. SUP produces the highest accuracy but requires substantially more computational resources. The MysteryMaster study provides evidence that demultiplexing performance differs across these models, with the performance gap between tools being most notable for Fast basecalled data.

### Can I re-basecall my data with a different model?

Yes, if you retained the raw signal data. Raw signal files contain the information needed for re-basecalling with different models. The raw signal dataset publication emphasizes that re-basecalling with future models can mitigate methylation-based calling errors and improve genotyping accuracy. Archiving raw signal data is recommended best practice for nanopore sequencing projects.

### What should I do if my basecalling pipeline crashes?

Check your computational resources first, including GPU memory, disk space, and CPU availability. Then verify your configuration file syntax and parameter values against the pipeline documentation. The [nf-core documentation](https://nf-co.re/docs) provides standards for reproducible pipeline configuration that can prevent many common failures. If the crash persists, check the input file integrity and format.

### How do I assess basecalling quality for organisms without a reference genome?

For organisms without a reference genome, use alternative quality metrics such as per-read quality scores, read length distributions, and consistency across overlapping reads. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides sequence resources that may support reference-based assessment for some organisms. For metagenomic samples, consider using taxonomic assignment consistency as a quality indicator.

### When should I escalate basecalling problems to specialized support?

Escalate when problems persist across multiple runs with different flow cells and chemistry kits, when systematic errors cannot be resolved through model changes, or when you need to implement application-specific pipelines that require specialized expertise. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program and [Galaxy Training Network](https://training.galaxyproject.org/) provide educational resources that can help you develop the skills to address complex problems independently.

## Related Bioinformatics Guides

- [Oxford Nanopore Sequencing: From Sample to Base Calls](/knowledge/bioinformatics/oxford-nanopore-sequencing-from-sample-to-base-calls)
- [How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore](/knowledge/bioinformatics/how-to-choose-a-long-read-sequencing-platform-pacbio-vs-oxford-nanopore)
- [Long-Read Sequencing for Isoform Quantification: Challenges and Solutions](/knowledge/bioinformatics/long-read-sequencing-for-isoform-quantification-challenges-and-solutions)
- [Genomic Data Processing: From Raw Sequencing to Analysis-Ready Files](/knowledge/bioinformatics/genomic-data-processing-from-raw-sequencing-to-analysis-ready-files)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [A whole-genome sequencing dataset of nanopore raw signals for bacterial genotyping and methylation analysis.](https://doi.org/10.1038/s41597-025-06319-4). 2025.
- [MysteryMaster: scraping the bottom of the barrel of barcoded Oxford nanopore reads.](https://doi.org/10.1186/s12859-025-06266-2). 2025.
- [TARPON-A Telomere Analysis and Research Pipeline Optimized for Nanopore.](https://doi.org/10.1371/journal.pcbi.1013915). 2026.
- [Protocol for differential analysis of pseudouridine modifications using nanopore DRS and unmodified transcriptome control.](https://doi.org/10.1016/j.xpro.2025.103948). 2025.
- [Effect of formic acid treatment on Apis mellifera foraging behavior using nanopore metabarcoding technologies.](https://doi.org/10.1371/journal.pone.0343810). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.