# Raw Signal Processing in Oxford Nanopore: From Squiggles to Basecalls - A Technical Overview


## Key Takeaways

- Oxford Nanopore sequencing generates raw ionic current data ("squiggles") as DNA/RNA translocates through a protein pore, with current fluctuations reflecting k-mer composition within the pore.
- Basecalling computationally converts these raw signals into nucleotide sequences, a process revolutionized by deep learning models (e.g., recurrent neural networks, transformers) that directly map signal to sequence, bypassing error-prone segmentation.
- Signal processing involves normalization (e.g., z-score) to account for pore-specific and run-specific variations, and conditioning to mitigate drift, ensuring comparability across reads.
- Modern basecalling models are trained on large datasets of known sequences, and their accuracy is influenced by training data diversity, including coverage of modified bases and varying signal-to-noise ratios.
- Raw signal data retains value beyond basecalling, enabling direct detection of base modifications (e.g., methylation) and facilitating signal-level demultiplexing and adaptive sampling for targeted sequencing.
- File formats like FAST5 and the more efficient SLOW5 are critical for raw signal storage, impacting processing speed and accessibility for advanced signal-level analyses.

---

Oxford Nanopore sequencing measures changes in ionic current as DNA, RNA, or protein molecules pass through a protein nanopore embedded in an electrically resistant membrane. The raw electrical signal, commonly called a squiggle, is a time series of current measurements that must be converted into nucleotide sequence through a computational process called basecalling. This article explains how raw signals become basecalls, what decisions researchers make at each stage, and how signal processing choices affect downstream data quality. The intended reader is a biology student, researcher, or laboratory professional who has generated or plans to generate Nanopore data and needs to understand the signal-to-basecall pipeline well enough to make informed analysis choices.

## The Physical Basis of Nanopore Signal Generation

A nanopore sequencing device applies a voltage across a membrane that contains a single protein pore, with an electrolyte solution filling both sides. When no molecule occupies the pore, a steady baseline ionic current flows through the pore. When a polymer such as DNA enters the pore, it partially blocks the flow of ions, reducing the measured current. The magnitude of the current reduction depends on which nucleotides occupy the narrowest constriction of the pore at any given moment.

For DNA sequencing, a motor protein controls the rate of translocation so that the DNA moves through the pore in a controlled, stepwise fashion. As the DNA advances, the combination of nucleotides in the pore sensing region changes, producing a characteristic current level for each k-mer, which is a short sequence of consecutive nucleotides. The resulting current trace is a series of plateaus at different current levels, with transitions between plateaus corresponding to nucleotide movement through the pore. This raw current trace is the fundamental data type produced by the instrument.

The relationship between nucleotide sequence and current level is not a simple one-to-one mapping. A given position in the trace reflects the influence of multiple consecutive nucleotides, typically five to ten depending on the pore chemistry. This means the same nucleotide at the same position can produce different current readings depending on its neighbors. The basecaller must therefore solve an inverse problem: given a current trace, infer the most likely underlying sequence that could have produced it.

The same physical principle applies to direct RNA sequencing, where RNA molecules pass through the pore without reverse transcription, and to emerging protein sequencing approaches. Recent work has demonstrated that a commercial nanopore sensor array can read intact protein strands by using the ClpX unfoldase to ratchet proteins through a CsgG nanopore, achieving sensitivity to single amino acids on synthetic protein strands hundreds of amino acids in length. This work also showed that raw nanopore signals can be simulated a priori based on residue volume and charge, which enhances interpretation of raw signal data. While protein sequencing is not yet a routine laboratory method, the signal processing principles are shared across analyte types.

## The Basecalling Problem in Formal Terms

Basecalling is the computational process of converting the raw current signal into a sequence of nucleotides. The input is a time series of current measurements, typically sampled at thousands of readings per second. The output is a DNA or RNA sequence with an associated quality score for each base.

The core difficulty is that the signal is noisy and the mapping from sequence to signal is complex. Thermal noise, electronic noise, and stochastic variation in pore occupancy all contribute to measurement uncertainty. The translocation speed is not perfectly uniform, so the same sequence can produce signals of slightly different durations in different reads. The k-mer composition of the pore affects the current level, and the relationship between k-mer and current is not known analytically for most pore chemistries.

Early basecallers used hidden Markov models to segment the signal into discrete events and then classify each event as a particular k-mer. This approach required a segmentation step that identified the boundaries between successive translocation steps. The segmentation step was error-prone because the signal does not always contain clear boundaries between events, especially in regions of low current difference between adjacent k-mers.

Deep learning approaches changed this landscape. Chiron was the first deep learning model to achieve end-to-end basecalling, directly translating raw signal to DNA sequence without the error-prone segmentation step. Trained with only a small set of 4,000 reads, Chiron provided state-of-the-art basecalling accuracy even on previously unseen species, and achieved basecalling speeds of more than 2,000 bases per second using desktop computer graphics processing units. This demonstrated that neural networks could learn the complex mapping from signal to sequence directly from data.

Modern basecallers use recurrent neural networks, convolutional neural networks, or transformer architectures to process the raw signal. These models are trained on large datasets of reads with known reference sequences. The training process adjusts the network weights to minimize the difference between predicted and actual sequences. Once trained, the model can process new signals without reference information.

## Signal Acquisition and File Formats

The raw signal is recorded by the sequencing device and stored in a file format that preserves the full resolution of the current measurements. The original format, called FAST5, uses the Hierarchical Data Format 5 (HDF5) container. Each FAST5 file typically contains the raw signal for a single read, along with metadata about the sequencing run, the pore, the channel, and the basecalling status.

FAST5 files are large and can be slow to process because each read is stored in a separate file. The SLOW5 format was developed to address these limitations. SLOW5 stores nanopore signal data in a more efficient binary format that enables faster compression and decompression, and it supports random access to individual reads. The slow5tools toolkit provides lossless data conversion between FAST5 and SLOW5, along with tools for interacting with SLOW5 files. Slow5tools uses multi-threading, multi-processing, and other engineering strategies to achieve fast data conversion and manipulation, including live FAST5-to-SLOW5 conversion during sequencing.

The choice of file format affects the analysis workflow. Researchers who need to process large volumes of raw signal data, such as those developing new basecallers or signal analysis algorithms, may benefit from converting to SLOW5 for faster processing. Researchers who are using standard analysis pipelines may never need to interact with raw signal files directly, because most pipelines handle basecalling automatically.

The raw signal data has value beyond basecalling. Signal-level analysis can detect base modifications, measure poly(A) tail lengths, and identify barcodes directly from the current trace. Tools such as SquiggleKit simplify file handling, data extraction, visualization, and signal processing for researchers who want to work with raw signal data directly. SquiggleKit is cross-platform and freely available, with all tools designed to operate in Python 2.7 or later with minimal additional libraries.

### At a Glance: Signal Processing Stages and Key Decisions

| Pipeline Stage | Primary Input | Key Decisions | Common Output |
| --- | --- | --- | --- |
| Signal acquisition | Ionic current measurements | FAST5 versus SLOW5 format, storage capacity | Raw signal files with per-read metadata |
| Basecalling | Raw current trace | Model selection, accuracy versus speed tradeoff, GPU availability | Nucleotide sequence with per-base quality scores |
| Signal-level analysis | Raw current trace | Modification detection, demultiplexing, adaptive sampling | Base modification calls, barcode assignments, enriched read sets |
| Quality assessment | Basecalled reads | Quality thresholds, read length filters, error rate estimation | Filtered read sets for downstream analysis |
| Downstream analysis | Filtered reads | Alignment, assembly, variant calling approach | Aligned reads, contigs, variant calls |

## Segmentation and Event Detection

Before the advent of end-to-end deep learning basecallers, the first step in basecalling was segmentation, also called event detection. Segmentation divides the continuous current trace into discrete events, where each event corresponds to a period during which a particular k-mer occupies the pore sensing region.

The segmentation algorithm identifies transitions in the current level. A typical approach uses a sliding window to compute a statistic such as the variance or the difference between adjacent windows. When the statistic exceeds a threshold, the algorithm marks a transition point. The current levels between transitions are then averaged to produce a single measurement for each event.

Segmentation is complicated by the fact that transitions between k-mers are not instantaneous. The current changes gradually as the DNA moves through the pore, and the exact boundary between events is ambiguous. Short events, where the DNA moves quickly through the pore, may be missed by the segmentation algorithm. Long events, where the DNA pauses, may be split into multiple events incorrectly.

The quality of segmentation directly affects basecalling accuracy. If the segmentation misses an event, the basecaller will not see the corresponding k-mer. If the segmentation splits one event into two, the basecaller will see a spurious extra k-mer. Both errors propagate through the rest of the basecalling pipeline.

End-to-end deep learning basecallers avoid explicit segmentation by treating the signal as a continuous input and letting the neural network learn to map signal regions to sequence positions. This approach eliminates the error-prone segmentation step and has become the standard for modern basecalling. However, segmentation remains useful for certain signal analysis tasks, such as detecting base modifications or analyzing the kinetics of translocation.

## Normalization and Signal Conditioning

Raw current measurements are affected by factors unrelated to the sequence, including the specific pore, the temperature, the salt concentration, and the electronics of the individual channel. These factors cause the baseline current and the scale of current changes to vary between pores and between runs. Normalization is the process of transforming the raw signal so that it is comparable across pores and runs.

A common normalization approach is to estimate the mean and standard deviation of the signal for a read and then transform each measurement to have zero mean and unit standard deviation. This z-score normalization removes the overall offset and scale differences while preserving the relative differences between k-mer levels.

More sophisticated normalization approaches account for drift over time. The baseline current can drift during a run as the pore ages or as the salt concentration changes. Some basecallers use a moving average or a median filter to estimate the local baseline and subtract it from the signal. Others use a calibration step at the start of the run to establish the expected current levels for known k-mers.

The choice of normalization method affects basecalling accuracy, particularly for reads with low signal-to-noise ratio. Over-normalization can remove real signal variation, while under-normalization leaves systematic differences that confuse the basecaller. Most modern basecallers include normalization as an internal step, so the user does not need to perform it separately. However, researchers developing custom signal analysis methods need to understand normalization to interpret their results correctly.

## Neural Network Architectures for Basecalling

The core of a modern basecaller is a neural network that maps the normalized signal to a sequence of nucleotides. Different architectures make different tradeoffs between accuracy, speed, and computational cost.

Recurrent neural networks, particularly long short-term memory (LSTM) networks, were the first deep learning architecture applied to basecalling. LSTMs process the signal sequentially, maintaining a hidden state that summarizes the information seen so far. This architecture is well suited to time series data because it can capture long-range dependencies in the signal. However, LSTMs are slow to train and run because they process the signal one time step at a time.

Convolutional neural networks (CNNs) process the signal in parallel, applying filters that detect local patterns. CNNs are faster than LSTMs because they can process many time steps simultaneously. However, CNNs have a limited receptive field, meaning they can only see a limited window of the signal at any given time. This limitation can be addressed by stacking multiple convolutional layers or by using dilated convolutions that increase the receptive field.

Transformer architectures, which use attention mechanisms to weigh the importance of different signal regions, have become popular for many sequence processing tasks. Transformers can capture long-range dependencies and are highly parallelizable. However, they require large amounts of training data and memory, and they may not be as well suited to streaming basecalling, where the signal is processed as it is generated.

Hybrid architectures combine multiple approaches. For example, a basecaller might use a convolutional front end to extract local features from the signal, followed by a recurrent or transformer layer to capture long-range dependencies, followed by a classification layer that outputs the sequence.

The choice of architecture affects accuracy and the computational resources required. Some basecallers are designed to run on graphics processing units (GPUs) for speed, while others can run on central processing units (CPUs) with acceptable performance. The basecalling speed determines whether real-time basecalling is possible, which is important for applications that require immediate feedback, such as adaptive sampling.

## Basecalling Models and Training Data

The accuracy of a basecaller depends on the quality and diversity of its training data. Training data consists of reads with known reference sequences, typically obtained by sequencing a well-characterized genome or amplicon. The basecaller learns to map the raw signal to the known sequence by adjusting its weights to minimize the prediction error.

Training data must cover the diversity of signals that the basecaller will encounter in practice. This includes different k-mer compositions, different GC contents, different read lengths, and different levels of signal noise. A basecaller trained only on high-quality reads from a single species may perform poorly on noisy reads from a different species.

The choice of training data also affects the basecaller's ability to handle base modifications. Modified bases such as 5-methylcytosine produce different current levels than unmodified bases. A basecaller trained only on unmodified DNA will misread modified bases. Some basecallers are trained to detect modifications directly from the signal, while others require a separate modification calling step.

Basecalling models are updated as new pore chemistries and sequencing kits are released. Each new chemistry may have different signal characteristics, requiring retraining or fine-tuning of the basecaller. Researchers should use the basecalling model recommended for their specific pore and kit combination, and they should be aware that changing the model can affect downstream analysis results.

The choice of basecalling model is a key decision in the analysis workflow. High-accuracy models produce better basecalls but are slower and require more computational resources. Fast models produce lower-accuracy basecalls but can keep up with real-time sequencing. Some workflows use a fast model for real-time quality assessment and then re-basecall with a high-accuracy model after the run is complete.

## Quality Scores and Their Interpretation

Basecallers assign a quality score to each base in the called sequence. The quality score is a Phred-scaled probability that the base is correct. A quality score of Q10 corresponds to 90% accuracy, Q20 to 99% accuracy, and Q30 to 99.9% accuracy.

Quality scores are useful for filtering reads and for downstream analysis. Reads with low average quality can be removed from the dataset, and individual bases with low quality can be masked or ignored. However, quality scores are only as good as the model that produces them. If the basecaller is overconfident, the quality scores will be too high, and if it is underconfident, they will be too low.

The relationship between quality scores and actual accuracy can be assessed by comparing basecalls to a known reference. This is typically done by aligning the basecalled reads to a reference genome and counting mismatches. The observed error rate can then be compared to the predicted error rate from the quality scores.

Quality scores are also affected by the basecalling model and the signal quality. Reads with low signal-to-noise ratio, such as those from degraded samples or from pores with unstable baselines, will have lower quality scores. The distribution of quality scores across a run can be used to monitor the health of the sequencing run and to identify problematic pores or channels.

For many applications, the average read quality is more important than the quality of individual bases. A read with an average quality of Q20 may be perfectly adequate for assembly, even if some individual bases have lower quality. The appropriate quality threshold depends on the application. For variant calling, higher quality thresholds are typically used, while for assembly, lower quality reads can be included because the assembler can correct errors through consensus.

## The Role of Raw Signal in Base Modification Detection

Base modifications, such as methylation, produce characteristic changes in the raw signal. These changes can be detected by comparing the observed signal to the expected signal for the unmodified sequence. The detection can be performed during basecalling or as a separate step after basecalling.

During basecalling, some models are trained to output modification probabilities alongside the nucleotide sequence. These models learn to recognize the signal patterns associated with modified bases. The accuracy of modification detection depends on the training data, which must include reads with known modification status.

After basecalling, modification detection can be performed by re-analyzing the raw signal at specific positions. The signal at a position is compared to the expected signal for the called base, and a deviation indicates a possible modification. This approach requires access to the raw signal, which is one reason why preserving raw signal data is important.

The ability to detect modifications from raw signal is a major advantage of Nanopore sequencing over short-read sequencing. Short-read sequencing typically requires a separate bisulfite conversion step to detect methylation, which damages the DNA and introduces biases. Nanopore sequencing can detect modifications directly from the native molecule, preserving the original sequence and modification status.

Direct RNA sequencing takes this a step further by detecting modifications on RNA molecules. The WarpDemuX approach for direct RNA sequencing demonstrates that raw signal processing can be used for adapter-barcoding and demultiplexing, and that integrating signal processing into sequencing control software enables real-time enrichment of target molecules through barcode-specific adaptive sampling. This work showed that raw signal processing can identify systematic differences in transcript abundance and poly(A) tail lengths during infection.

## Demultiplexing and Adaptive Sampling at the Signal Level

Demultiplexing is the process of assigning reads to their samples of origin based on barcode sequences. In standard workflows, demultiplexing is performed after basecalling by identifying barcode sequences in the called reads. However, demultiplexing can also be performed at the signal level, before basecalling.

Signal-level demultiplexing has several advantages. It can be faster than basecalling followed by demultiplexing, because the signal-level barcode detection is simpler than full basecalling. It can also be more accurate for barcodes that are difficult to basecall, such as those with low sequence complexity. Signal-level demultiplexing is particularly useful for direct RNA sequencing, where the barcodes are attached to the RNA molecules and may be affected by RNA secondary structure.

The WarpDemuX approach for direct RNA sequencing uses fast processing of the raw nanopore signal and a light-weight machine-learning algorithm to achieve ultra-fast and highly accurate adapter-barcoding and demultiplexing. This approach was demonstrated with SQK-RNA002 and SQK-RNA004 chemistries, and it enabled rapid phenotypic profiling of different SARS-CoV-2 viruses through multiplexed sequencing of longitudinal samples on a single flowcell.

Adaptive sampling is a technique that uses real-time basecalling or signal analysis to decide whether to accept or reject a molecule as it is being sequenced. If the initial portion of a read indicates that the molecule is not of interest, the device can reverse the voltage and eject the molecule from the pore, freeing the pore for the next molecule. This enriches the sequencing run for molecules of interest.

Adaptive sampling requires fast basecalling or signal analysis, because the decision must be made before the molecule has fully translocated through the pore. The speed of the basecaller is therefore a critical factor. Signal-level analysis can be faster than full basecalling, enabling more efficient adaptive sampling. The integration of WarpDemuX into sequencing control software enabled real-time enrichment of target molecules through barcode-specific adaptive sampling, demonstrating the potential of signal-level processing for this application.

## Practical Workflow for Signal Processing

A typical Nanopore sequencing analysis workflow involves several stages, from raw signal acquisition to final sequence analysis. Understanding each stage helps researchers make informed decisions about their analysis.

The first stage is signal acquisition. The sequencing device records the raw current signal and stores it in FAST5 or SLOW5 format. The choice of format affects the speed and efficiency of subsequent processing. Researchers who plan to work with raw signal data should consider converting to SLOW5 for faster processing.

The second stage is basecalling. The raw signal is converted to nucleotide sequence using a basecaller. The choice of basecaller and model depends on the application. High-accuracy models are appropriate for applications that require precise sequence information, such as variant calling. Fast models are appropriate for real-time applications, such as adaptive sampling or quality monitoring.

The third stage is quality assessment. The basecalled reads are examined for quality, length, and other characteristics. This assessment can be performed during the run for real-time monitoring or after the run for final quality control. Poor-quality reads can be filtered out at this stage.

The fourth stage is downstream analysis. The basecalled reads are aligned to a reference genome, assembled into contigs, or analyzed for specific features such as base modifications. The choice of downstream analysis depends on the research question.

The fifth stage is interpretation and reporting. The results of the analysis are interpreted in the context of the research question, and the methods and parameters used are reported. Reproducibility requires careful documentation of all analysis steps, including the versions of all software and the parameters used.

### Practical Implementation Steps for Signal Processing

| Step | Action | Purpose | Verification |
| --- | --- | --- | --- |
| 1 | Select file format for raw signal storage | Balance processing speed against storage requirements | Confirm SLOW5 conversion preserves signal fidelity |
| 2 | Choose basecalling model matched to pore and kit chemistry | Ensure compatibility between signal characteristics and model expectations | Check manufacturer recommendations for the specific chemistry |
| 3 | Run basecalling with selected model | Convert raw signal to nucleotide sequence | Monitor basecalling speed and quality score distribution |
| 4 | Assess read quality and length distributions | Identify problematic pores, channels, or library issues | Compare quality metrics against expected values for the chemistry |
| 5 | Perform signal-level analysis if needed | Detect modifications, demultiplex, or enable adaptive sampling | Validate results against known controls or reference data |
| 6 | Document all software versions and parameters | Enable reproducibility and troubleshooting | Record versions, model names, and parameter values in laboratory records |

## Choosing Between Basecalling Options

The choice of basecalling software and model is one of the most important decisions in the Nanopore analysis workflow. Several factors should be considered.

Accuracy is the primary consideration for most applications. Higher accuracy reduces the need for downstream error correction and improves the reliability of variant calls. However, higher accuracy typically requires more computational resources and longer processing times.

Speed is important for real-time applications and for large datasets. Real-time basecalling enables adaptive sampling and provides immediate feedback on run quality. For large datasets, the time required for basecalling can be a bottleneck, and faster basecallers can significantly reduce the overall analysis time.

Computational resources are a practical constraint. Some basecallers require GPUs for acceptable performance, while others can run on CPUs. The availability of GPUs in the laboratory or on a computing cluster will influence the choice of basecaller.

The specific pore and kit chemistry must be matched to the basecalling model. Using an incompatible model will produce poor results. The sequencing device manufacturer provides recommended models for each chemistry, and these recommendations should be followed unless there is a specific reason to deviate.

The choice between a single basecalling pass and multiple passes depends on the application. Some workflows use a fast basecaller for initial quality assessment and then re-basecall with a high-accuracy model for final analysis. This two-pass approach combines the speed of the fast model with the accuracy of the high-accuracy model.

## Records and Measurements for Quality Control

Maintaining records of sequencing and basecalling parameters is essential for quality control and reproducibility. The following measurements should be recorded for each sequencing run.

The pore occupancy, which is the fraction of pores that are actively sequencing, indicates the health of the flowcell. Low pore occupancy may indicate a problem with the flowcell or the library preparation.

The read length distribution, including the N50 read length, indicates the quality of the library and the sequencing conditions. Short reads may indicate DNA degradation or problems with the library preparation.

The basecalling quality scores, including the distribution of quality scores across reads, indicate the overall accuracy of the basecalls. A shift in the quality score distribution during a run may indicate pore deterioration or other problems.

The basecalling speed, measured in bases per second, indicates whether the basecaller is keeping up with the sequencing rate. If the basecaller falls behind, the sequencing data will accumulate and may need to be basecalled after the run.

The yield, measured in total bases or total reads, indicates the overall output of the run. The yield depends on the number of active pores, the read length, and the sequencing duration.

These measurements should be recorded in a laboratory notebook or electronic record system. The records should include the software versions and parameters used for basecalling, because changing the basecaller or model can affect the results.

## Common Failure Patterns in Signal Processing

Several common problems can arise during signal processing and basecalling. Recognizing these problems early can save time and resources.

Low pore occupancy is a common problem. If fewer pores than expected are actively sequencing, the yield will be low. This can be caused by a faulty flowcell, poor library preparation, or problems with the sequencing device. The pore occupancy should be monitored during the run, and if it drops significantly, the run may need to be restarted.

High error rates in basecalls can be caused by several factors. The basecalling model may be incompatible with the pore or kit chemistry. The signal quality may be poor due to noise or drift. The library may contain contaminants that interfere with sequencing. The error rate should be assessed by aligning a subset of reads to a known reference.

Basecalling speed that falls behind the sequencing rate can cause data to accumulate. This can be addressed by using a faster basecaller, using a GPU, or processing the data after the run. If real-time basecalling is required, the basecaller must be fast enough to keep up.

File format issues can cause problems with data processing. FAST5 files can be corrupted or incomplete, and SLOW5 conversion can fail if the files are not properly formatted. The integrity of the raw signal files should be checked before processing.

Signal drift can cause the baseline current to change over time, affecting the accuracy of basecalling. This can be addressed by using a basecaller that accounts for drift or by normalizing the signal before basecalling.

## Limitations of Signal Processing Approaches

Signal processing approaches have inherent limitations that researchers should understand.

The accuracy of basecalling is limited by the information content of the signal. The current level depends on the k-mer composition of the pore, but the relationship is not one-to-one. Different k-mers can produce similar current levels, making them difficult to distinguish. This is a fundamental limitation that cannot be fully overcome by better algorithms.

The signal-to-noise ratio varies between pores and between reads. Some pores produce noisier signals than others, and some reads have lower signal-to-noise ratios due to the sequence composition or the translocation speed. The basecaller must be robust to this variation, but there is a limit to how much noise can be tolerated.

The training data for basecallers may not cover all possible sequences and conditions. A basecaller trained on a particular set of species and conditions may perform poorly on novel sequences or unusual conditions. Researchers should validate the basecaller performance on their specific samples.

The computational resources required for basecalling can be substantial. High-accuracy basecalling of large datasets can take hours or days, even with GPUs. This can be a bottleneck for large-scale projects.

The interpretation of raw signal data requires specialized expertise. Tools such as SquiggleKit and slow5tools provide access to raw signal data, but using these tools effectively requires an understanding of the signal characteristics and the analysis methods.

## Reproducibility and Documentation

Reproducibility is a core principle of scientific research, and it is particularly important for computational analyses. The following practices support reproducibility in Nanopore signal processing.

Document all software versions. The basecaller, the model, and all analysis tools should be recorded with their version numbers. Software updates can change results, so the exact versions used should be documented.

Document all parameters. The basecalling parameters, such as the model, the quality threshold, and the chunk size, should be recorded. The parameters for downstream analysis should also be recorded.

Document the data processing steps. The order of operations, from raw signal to final results, should be documented. This includes any filtering, normalization, or transformation steps.

Use version control for analysis scripts. Version control systems such as Git track changes to scripts and enable reproducibility of the analysis. Training in version control is available through resources such as The Carpentries lessons, which provide foundational computing and data skills.

Use workflow management tools. Workflow managers such as those provided by nf-core enable reproducible and portable analysis pipelines. The nf-core documentation provides standards for pipeline usage and configuration, supporting reproducible workflow context.

Use containerization. Containers package software and its dependencies, ensuring that the analysis runs in the same environment regardless of the host system. This is particularly important for complex analyses with many dependencies.

## Training and Learning Resources

Researchers new to Nanopore signal processing can benefit from structured training. Several resources provide accessible introductions to bioinformatics and genomic analysis.

The Galaxy Training Network provides accessible workflow training and analysis tutorials. These tutorials cover a range of topics, from basic sequence analysis to advanced genomic workflows, and they emphasize reproducibility. The Galaxy platform provides a web-based interface that lowers the barrier to entry for researchers without programming experience.

The EMBL-EBI Training program provides bioinformatics learning pathways and data-resource training. These courses cover the use of major biological databases and analysis tools, and they provide practical analysis education. The training is designed for researchers at various levels, from beginners to advanced users.

The Carpentries lessons provide foundational computing, data, shell, Git, and programming training. These lessons are designed for researchers who want to develop the computational skills needed for modern biological research. The lessons are hands-on and practical, and they are taught by certified instructors.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation. Bioconductor packages are widely used for genomic analysis, and the documentation provides guidance on using these packages effectively.

The nf-core documentation provides community pipeline standards, usage, configuration, and reproducible workflow context. nf-core pipelines are widely used for genomic analysis, and the documentation provides guidance on using and configuring these pipelines.

The NCBI Data Resources provide official descriptions of NCBI databases, search systems, sequence resources, and analysis services. These resources are essential for accessing reference sequences and for depositing and retrieving sequence data.

## Professional Escalation Criteria

Researchers should escalate signal processing problems to more experienced colleagues or to the sequencing device manufacturer when certain conditions are met.

If the basecalling accuracy is consistently below the expected level for the pore and kit chemistry, and if changing the basecalling model does not improve the results, the problem may be with the library preparation or the sequencing device. This should be escalated to the manufacturer's technical support.

If the pore occupancy drops significantly during a run, and if the drop cannot be explained by normal pore aging, the flowcell may be faulty. This should be escalated to the manufacturer for replacement.

If the raw signal files are corrupted or incomplete, and if the data loss is significant, the run may need to be repeated. The manufacturer should be consulted to determine the cause of the file corruption.

If the computational resources required for basecalling exceed the available capacity, and if the analysis cannot be completed in a reasonable time, the researcher should consult with a bioinformatics specialist or a computing facility for guidance on optimizing the analysis.

If the interpretation of the results is uncertain, and if the uncertainty could affect the conclusions of the study, the researcher should consult with a statistical geneticist or a bioinformatics expert for guidance.

## Frequently Asked Questions

### What is a squiggle in Nanopore sequencing?

A squiggle is the raw electrical signal produced by a nanopore sequencing device. It is a time series of current measurements that changes as a DNA, RNA, or protein molecule passes through the pore. The current level at any moment depends on which nucleotides or amino acids occupy the pore sensing region. The squiggle is the fundamental data type produced by the instrument, and it must be converted into sequence through basecalling.

### Why is basecalling necessary for Nanopore data?

Basecalling is necessary because the raw signal is not directly interpretable as sequence. The current level depends on the k-mer composition of the pore, but the relationship is complex and noisy. Basecalling uses computational models, typically neural networks, to infer the most likely sequence that could have produced the observed signal. Without basecalling, the raw signal cannot be used for most downstream analyses.

### What is the difference between FAST5 and SLOW5 file formats?

FAST5 is the original file format for Nanopore raw signal data, using the Hierarchical Data Format 5 container. Each FAST5 file typically contains the raw signal for a single read. SLOW5 is a newer format that stores the same data more efficiently, enabling faster compression and decompression and supporting random access to individual reads. The slow5tools toolkit provides lossless conversion between the two formats.

### How does a neural network basecaller work?

A neural network basecaller processes the raw signal through multiple layers of computation. The network learns to recognize patterns in the signal that correspond to particular sequences. The training process adjusts the network weights to minimize the difference between predicted and actual sequences. Once trained, the network can process new signals without reference information. Different architectures, such as recurrent, convolutional, or transformer networks, make different tradeoffs between accuracy and speed.

### What is adaptive sampling in Nanopore sequencing?

Adaptive sampling is a technique that uses real-time analysis of the signal to decide whether to accept or reject a molecule as it is being sequenced. If the initial portion of a read indicates that the molecule is not of interest, the device can eject the molecule from the pore, freeing the pore for the next molecule. This enriches the sequencing run for molecules of interest. Adaptive sampling requires fast basecalling or signal analysis to make the decision before the molecule has fully translocated.

### How are base modifications detected from raw signal?

Base modifications produce characteristic changes in the raw signal. These changes can be detected by comparing the observed signal to the expected signal for the unmodified sequence. Detection can be performed during basecalling, using models trained to output modification probabilities, or as a separate step after basecalling. The accuracy of modification detection depends on the training data and the signal quality.

### What quality scores should I use for filtering Nanopore reads?

The appropriate quality threshold depends on the application. For assembly, lower quality reads can be included because the assembler can correct errors through consensus. For variant calling, higher quality thresholds are typically used. The quality score distribution should be examined for each run, and the threshold should be chosen based on the specific requirements of the analysis.

### How can I make my Nanopore analysis reproducible?

Reproducibility requires documentation of all analysis steps, including software versions, parameters, and data processing steps. Version control systems such as Git track changes to analysis scripts. Workflow management tools and containerization ensure that the analysis runs in the same environment regardless of the host system. Training in these practices is available through resources such as The Carpentries lessons and the Galaxy Training Network.

## Related Bioinformatics Guides

- [Genomic Data Processing: From Raw Sequencing to Analysis-Ready Files](/knowledge/bioinformatics/genomic-data-processing-from-raw-sequencing-to-analysis-ready-files)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight](/knowledge/bioinformatics/metabolomics-data-analysis-workflow-from-raw-data-to-biological-insight)
- [Oxford Nanopore Sequencing: From Sample to Base Calls](/knowledge/bioinformatics/oxford-nanopore-sequencing-from-sample-to-base-calls)
- [Metagenomics Data Analysis: From Raw Reads to Biological Insights](/knowledge/bioinformatics/metagenomics-data-analysis-from-raw-reads-to-biological-insights)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Multi-pass, single-molecule nanopore reading of long protein strands.](https://pubmed.ncbi.nlm.nih.gov/39261738). Nature, 2024.
- [Chiron: translating nanopore raw signal directly into nucleotide sequence using deep learning.](https://pubmed.ncbi.nlm.nih.gov/29648610). GigaScience, 2018.
- [Flexible and efficient handling of nanopore sequencing signal data with slow5tools.](https://pubmed.ncbi.nlm.nih.gov/37024927). Genome biology, 2023.
- [Demultiplexing and barcode-specific adaptive sampling for nanopore direct RNA sequencing.](https://pubmed.ncbi.nlm.nih.gov/40258808). Nature communications, 2025.
- [SquiggleKit: a toolkit for manipulating nanopore signal data.](https://pubmed.ncbi.nlm.nih.gov/31332428). Bioinformatics (Oxford, England), 2019.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.