# How to Tune Minimap2 Parameters for PacBio HiFi vs. Oxford Nanopore Reads: A Practical Guide


## Key Takeaways

- Minimap2 parameter selection for long-read mapping is critically dependent on platform-specific read characteristics; PacBio HiFi reads exhibit high per-base accuracy (>99.9%) with rare, isolated substitutions, whereas Oxford Nanopore (ONT) reads possess higher error rates, particularly in homopolymer regions, with a greater proportion of insertions and deletions.
- The k-mer size is a primary determinant of seeding sensitivity and specificity; smaller k-mers (e.g., 15 or lower for ONT) increase sensitivity by tolerating more errors but can lead to spurious seeds in repetitive regions, while larger k-mers (e.g., 19 for HiFi) offer better specificity when error rates are low.
- Scoring parameters, including match score, mismatch penalty, and gap penalties, must be tuned to balance error tolerance with alignment precision; for ONT data with higher indel rates, gap penalties may require adjustment, while for HiFi data, a higher mismatch penalty can improve precision in repetitive regions.
- Evaluating mapping rate, alignment identity distribution, and mapping quality distribution are essential diagnostic steps; low mapping rates often indicate insufficient seeding sensitivity (requiring smaller k-mers), while poor alignment quality despite high mapping rates suggests overly permissive scoring parameters.
- Comprehensive documentation of parameter choices, including the full command line, reference genome version, read data source, and observed mapping statistics, is crucial for reproducibility and troubleshooting, especially when comparing results across different sequencing platforms or batches.
- Parameter tuning is constrained by reference genome completeness and computational resources; a highly fragmented or erroneous reference will limit achievable mapping rates, and computationally intensive parameter sets may necessitate trade-offs for large-scale analyses.

---

Bioinformaticians frequently run minimap2 with default settings for all long-read mapping tasks, yet the optimal parameter set differs between PacBio HiFi reads and Oxford Nanopore Technologies (ONT) reads because the two platforms produce distinct error profiles, read length distributions, and base quality characteristics. This article provides a parameter-by-parameter tuning framework for minimap2, with recommended settings for HiFi, ONT, and ultra-long reads, including executable command-line examples. The guidance applies to researchers, laboratory professionals, and students who need reliable read mapping for genome assembly, structural variant detection, and downstream comparative analysis.

## Understanding Platform-Specific Read Characteristics Before Tuning

The decision to adjust minimap2 parameters begins with an accurate understanding of what each sequencing platform produces. PacBio HiFi reads are circular consensus sequences with per-base accuracy typically above 99.9 percent, meaning errors are rare and mostly isolated substitutions. ONT reads, depending on the flow cell version and basecalling model, carry higher error rates with a greater proportion of insertions and deletions, particularly in homopolymer regions. These differences matter because minimap2 uses a seed-and-extension algorithm where initial matches are identified by exact k-mer matches, and the sensitivity of this seeding step depends on both the k-mer size and the error tolerance of the reference index.

The read length distribution also influences parameter choice. HiFi reads commonly range from 10 to 25 kilobases, while ONT ultra-long reads can exceed 100 kilobases. Longer reads increase the chance that a single read spans repetitive elements or structural variation breakpoints, but they also increase the computational cost of alignment. The minimap2 presets for `map-hifi` and `map-ont` encode different assumptions about error rates and read lengths, and understanding these assumptions helps you decide when to modify individual parameters instead of accept the preset as fixed.

For researchers working with newly assembled genomes, the choice of mapping parameters directly affects the reported mapping rate and the confidence in downstream variant calls. Recent telomere-to-telomere assembly projects have reported mapping rates above 97 percent for both ONT ultra-long reads and PacBio HiFi reads when aligned against their respective finished assemblies, which demonstrates that high mapping rates are achievable when parameters match the data characteristics. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provides access to reference sequences and read archives that can be used to test parameter choices on real data before committing to a full analysis run.

## Core Principles of Minimap2 Parameter Selection

Minimap2 operates through a series of stages: indexing the reference, seeding candidate alignments with exact matches, chaining seeds into longer alignments, and finally computing base-level alignments with a scoring scheme. Each stage has parameters that can be tuned, but the most consequential parameters for long-read mapping are the k-mer size, the scoring matrix for matches and mismatches, gap penalties, and the threshold for reporting secondary alignments.

The k-mer size controls the sensitivity and specificity of the seeding step. Smaller k-mers increase sensitivity because they are more likely to match despite sequencing errors, but they also increase the number of spurious seeds in repetitive regions. Larger k-mers reduce false positives but may miss true alignments when errors disrupt the exact match. For HiFi reads with low error rates, larger k-mers work well because the reads are accurate enough to produce reliable exact matches. For ONT reads with higher error rates, smaller k-mers are often necessary to achieve adequate sensitivity.

The scoring parameters determine how the alignment algorithm balances matches against mismatches and gaps. The default scoring in minimap2 uses a match score of 2, a mismatch penalty of 4, and gap penalties that depend on the preset. Increasing the mismatch penalty makes the aligner more conservative, requiring stronger evidence for a mismatch, while decreasing it allows more mismatches to be tolerated. Gap penalties control the cost of insertions and deletions, which is particularly important for ONT data where indel errors are common.

Secondary alignment reporting controls how many alternative mappings are reported for each read. For reads that map to repetitive regions, the primary alignment may not be the biologically correct one, and reporting secondary alignments allows downstream tools to handle multi-mapping reads appropriately. However, reporting too many secondary alignments increases file size and may confuse downstream analysis if the tools do not account for multi-mapping reads.

## At a Glance: Recommended Parameter Sets by Platform

The following table summarizes recommended minimap2 parameter choices for common long-read mapping scenarios. These recommendations assume a standard reference genome and do not account for specialized cases such as mapping to a highly repetitive genome or using a custom scoring scheme for a specific downstream application.

| Parameter | PacBio HiFi (map-hifi preset) | ONT Standard (map-ont preset) | ONT Ultra-long (custom tuning) |
| --- | --- | --- | --- |
| k-mer size | 19 (default in preset) | 15 (default in preset) | 15 or lower for high-error data |
| Match score | 2 (default) | 2 (default) | 2 (default) |
| Mismatch penalty | 4 (default) | 4 (default) | 4 to 6 for stricter filtering |
| Gap open penalty | 4 (default) | 4 (default) | 4 to 8 for longer reads |
| Gap extension penalty | 2 (default) | 2 (default) | 2 to 4 for error-prone reads |
| Secondary alignments | Report up to 5 | Report up to 5 | Report up to 5 or more for repetitive genomes |
| Recommended command | `minimap2 -ax map-hifi ref.fa reads.fq.gz` | `minimap2 -ax map-ont ref.fa reads.fq.gz` | `minimap2 -ax map-ont -k15 -A2 -B4 -O4,2 ref.fa reads.fq.gz` |

The command-line examples in the table represent starting points, not fixed protocols. You should test these parameters on a subset of your data and compare the mapping statistics before processing the full dataset. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials on read mapping and quality assessment that can help you build the skills needed to evaluate parameter choices systematically.

## Practical Workflow for Parameter Tuning

### Step 1: Assess Your Read Data Characteristics

Before adjusting any parameters, you need to know the error profile and length distribution of your reads. Run `NanoPlot` or `fastqc` on a random sample of 10,000 to 50,000 reads to obtain summary statistics including read length N50, mean read quality, and the distribution of quality scores along reads. For ONT data, also check the basecalling model used, because newer basecalling models produce more accurate reads and may allow more aggressive parameter settings.

For PacBio HiFi data, verify that the reads have been properly processed through the circular consensus sequencing pipeline. HiFi reads should have quality scores above Q20, meaning an expected error rate below 1 percent. If your HiFi reads have lower quality, treat them as intermediate-quality data and use parameters closer to the ONT settings.

### Step 2: Run Minimap2 with the Platform Preset

Start with the preset that matches your platform. Use `-ax map-hifi` for PacBio HiFi reads and `-ax map-ont` for ONT reads. The `-a` flag outputs SAM format, and the `-x` flag selects the preset. Run the mapping on a subset of reads, typically 100,000 to 500,000 reads or roughly 1 to 2 gigabases of sequence, to keep the test run fast.

Capture the mapping statistics from the SAM file header or use `samtools stats` to obtain the mapping rate, the distribution of mapping qualities, and the number of reads that map to multiple locations. These statistics provide the baseline against which you will compare parameter adjustments.

### Step 3: Evaluate Mapping Rate and Quality

The mapping rate is the proportion of reads that produce at least one alignment. For HiFi reads against a high-quality reference, mapping rates above 95 percent are typical. For ONT reads, mapping rates above 90 percent are common, though the exact rate depends on the basecalling quality and the reference completeness. Recent telomere-to-telomere assembly projects have demonstrated that mapping rates above 97 percent are achievable for both platforms when the reference is complete and the parameters match the data. The [mandarin fish genome assembly](https://doi.org/10.1038/s41597-026-07113-6) achieved mapping rates above 97 percent for ONT ultra-long reads, PacBio HiFi reads, and Hi-C data, and the [pond smelt genome assembly](https://doi.org/10.1038/s41597-026-07078-6) used a combined strategy with PacBio HiFi and ONT ultra-long reads to produce a gap-free reference.

Beyond the raw mapping rate, examine the distribution of mapping qualities. Reads with mapping quality below 20 are less reliable and may indicate problems with the reference or the parameters. Also check the proportion of reads that map to multiple locations, as a high multi-mapping rate suggests that the k-mer size is too small or that the reference contains many repetitive regions.

### Step 4: Adjust Parameters Based on Observed Failures

If the mapping rate is lower than expected, the most likely cause is insufficient seeding sensitivity. Reduce the k-mer size from the preset default and test again. For ONT data, reducing the k-mer from 15 to 13 or even 11 can improve sensitivity, though it will increase the runtime and the number of spurious seeds. For HiFi data, reducing the k-mer below 19 is rarely necessary unless the reads are unusually error-prone.

If the mapping rate is acceptable but the alignment quality is poor, indicated by low mapping qualities or short aligned blocks, adjust the scoring parameters. Increasing the mismatch penalty makes the aligner more conservative and may improve the precision of alignments in repetitive regions. Increasing the gap open penalty reduces the number of gaps in the alignment, which can help when reads contain many small indels that are likely sequencing errors instead of biological variation.

### Step 5: Validate with Downstream Analysis

The ultimate test of parameter choices is whether the downstream analysis produces biologically meaningful results. If you are using the alignments for structural variant detection, run your variant caller on a small test region and compare the calls against known variants or against calls from a different aligner. If you are using the alignments for genome assembly, check the assembly continuity and completeness metrics.

The [Bioconductor project](https://bioconductor.org/) provides numerous R packages for analyzing alignment data, including tools for visualizing alignments and computing coverage statistics. These tools can help you identify systematic problems in your alignments that might not be apparent from summary statistics alone.

## Options and Tradeoffs in Parameter Selection

### K-mer Size: Sensitivity versus Specificity

The k-mer size is the most influential parameter for mapping sensitivity. Smaller k-mers increase the chance of finding seeds in error-prone reads, but they also increase the number of false seeds in repetitive regions. The optimal k-mer size depends on the read error rate and the repetitiveness of the reference genome.

For HiFi reads with error rates below 0.1 percent, a k-mer size of 19 or 21 provides good sensitivity with acceptable specificity. For ONT reads with error rates of 5 to 15 percent, a k-mer size of 15 is a reasonable starting point, but you may need to reduce it to 13 for particularly error-prone data. The tradeoff is computational cost: smaller k-mers produce more seeds, which increases the time required for chaining and alignment.

For ultra-long ONT reads, the read length itself provides additional mapping information because a single read can span multiple unique regions. In this case, a slightly larger k-mer may be acceptable because the long read provides multiple independent seeds that can be chained into a confident alignment.

### Scoring Parameters: Balancing Match Evidence against Error Tolerance

The scoring parameters control the alignment extension stage, where the algorithm decides how to align the sequence between seeds. The default scoring in minimap2 uses a match score of 2, a mismatch penalty of 4, and gap open and extension penalties of 4 and 2 respectively. These values were chosen to work well across a range of data types, but they may not be optimal for your specific data.

For HiFi data, the low error rate means that mismatches and gaps in the alignment are likely to be biologically real. You may want to increase the mismatch penalty to reduce false alignments in repetitive regions, but be careful not to make it so high that genuine mismatches are not tolerated. A mismatch penalty of 5 or 6 is a reasonable test range for HiFi data.

For ONT data, the higher error rate means that many mismatches and gaps are sequencing errors. You may want to decrease the mismatch penalty to tolerate these errors, but this will also increase the number of false alignments. A mismatch penalty of 3 or 4 is a reasonable test range for ONT data, with the exact value depending on the basecalling quality.

### Secondary Alignments: Handling Multi-Mapping Reads

The `--secondary` parameter controls how many secondary alignments are reported for each read. The default is to report up to 5 secondary alignments, which is appropriate for most applications. For reads that map to repetitive regions, the primary alignment may not be the biologically correct one, and reporting secondary alignments allows downstream tools to handle these reads appropriately.

If you are using the alignments for variant calling, you may want to reduce the number of secondary alignments to avoid false positive variant calls from multi-mapping reads. If you are using the alignments for genome assembly, you may want to increase the number of secondary alignments to ensure that all possible mappings are considered during the assembly process.

The [nf-core documentation](https://nf-co.re/docs) provides guidance on how different bioinformatics pipelines handle multi-mapping reads, which can help you decide on the appropriate secondary alignment setting for your specific workflow.

## Observations and Measurements for Parameter Evaluation

### Mapping Rate as a Primary Metric

The mapping rate is the most direct measure of whether your parameters are appropriate for your data. A low mapping rate indicates that the parameters are too stringent, either because the k-mer size is too large or the scoring parameters are too conservative. A high mapping rate with poor alignment quality indicates that the parameters are too permissive, allowing spurious alignments.

To measure the mapping rate, use `samtools stats` on the sorted BAM file and examine the "reads mapped" field. For a quick check, you can also count the number of reads with a mapping quality above 0 in the SAM file, since reads with mapping quality 0 are typically unmapped or multi-mapping.

### Alignment Identity and Coverage Distribution

Beyond the mapping rate, examine the distribution of alignment identities. The alignment identity is the proportion of matching bases in the alignment, and it should be consistent with the expected error rate of your platform. For HiFi reads, alignment identities above 99 percent are typical. For ONT reads, alignment identities of 90 to 95 percent are common, depending on the basecalling model.

The coverage distribution across the reference genome is another useful diagnostic. Uneven coverage may indicate that some regions are difficult to map, either because of repetitive content or because of reference errors. The Integrative Genomics Viewer or similar visualization tools can help you examine coverage in specific regions.

### Runtime and Memory Usage

Parameter choices affect computational cost. Smaller k-mers and more permissive scoring parameters increase the number of candidate alignments, which increases runtime and memory usage. If your parameter choices make the mapping step prohibitively slow, you may need to find a balance between sensitivity and computational efficiency.

The [EMBL-EBI training materials](https://www.ebi.ac.uk/training) provide guidance on optimizing bioinformatics workflows for computational efficiency, including strategies for parallelizing read mapping across multiple cores.

## Records and Documentation for Reproducible Mapping

### Recording Parameter Choices

Every parameter choice should be recorded in a way that allows another researcher to reproduce your analysis. The simplest approach is to include the full minimap2 command line in your analysis scripts or workflow files. If you use a workflow manager such as Nextflow, the [nf-core documentation](https://nf-co.re/docs) provides guidance on how to document parameter choices in a reproducible way.

For each mapping run, record the following information: the minimap2 version, the reference genome version and source, the read data version and source, the full command line including all parameters, the date of the run, and the computational environment. This information allows you to compare results across runs and to troubleshoot problems that may arise.

### Version Control for Parameters

Parameter choices should be treated as part of your analysis code and placed under version control. If you change parameters, record the reason for the change and the effect on the mapping statistics. This documentation is particularly important when you are developing a new analysis pipeline or when you are comparing results across different versions of the reference genome.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training on version control with Git, which is essential for managing analysis code and parameter choices in a reproducible way.

### Quality Control Reports

Generate a quality control report for each mapping run that includes the mapping rate, the distribution of mapping qualities, the alignment identity distribution, and the coverage statistics. This report serves as a record of the analysis and provides a baseline for comparison when you adjust parameters.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on generating quality control reports for sequencing data, including read mapping quality assessment.

## Common Failure Patterns and Troubleshooting

### Low Mapping Rate with ONT Reads

If your ONT reads produce a mapping rate below 90 percent, the most likely cause is that the k-mer size is too large for the error rate of your data. Reduce the k-mer size from 15 to 13 or 11 and test again. Also verify that your basecalling model is appropriate for the flow cell version and that the reads have not been trimmed too aggressively.

Another possible cause is that the reference genome is incomplete or contains errors. Check the mapping rate against a different reference or against the same reference with a different aligner to determine whether the problem is with the parameters or the reference.

### Low Mapping Rate with HiFi Reads

HiFi reads should map at rates above 95 percent against a high-quality reference. If the mapping rate is lower, check that the reads are truly HiFi and not a mix of HiFi and lower-quality reads. Also verify that the reference genome is appropriate for the species and that the reads have not been contaminated with adapter sequences.

If the mapping rate is still low, try reducing the k-mer size from 19 to 17 or 15. This is rarely necessary for HiFi data, but it can help if the reads have lower quality than expected.

### High Multi-Mapping Rate

A high proportion of reads mapping to multiple locations indicates that the reference contains many repetitive regions or that the k-mer size is too small. If the multi-mapping rate is above 10 percent, examine the repetitive content of the reference and consider whether the downstream analysis can handle multi-mapping reads appropriately.

For structural variant detection, multi-mapping reads are often excluded from the analysis because they cannot be confidently placed. For genome assembly, multi-mapping reads can be useful for resolving repetitive regions, but they require careful handling.

### Poor Alignment Quality Despite High Mapping Rate

If the mapping rate is high but the alignment quality is poor, indicated by low alignment identities or short aligned blocks, the scoring parameters may be too permissive. Increase the mismatch penalty and the gap open penalty to make the aligner more conservative. Test a range of values and compare the alignment statistics to find the optimal setting.

## Limitations of Parameter Tuning

### Reference Genome Completeness

No parameter tuning can compensate for a reference genome that is missing large regions or contains significant errors. The mapping rate and alignment quality are fundamentally limited by the reference quality. Recent telomere-to-telomere assembly projects have demonstrated that complete references produce higher mapping rates and more reliable alignments, but such references are not available for all species.

The [Zi goose genome assembly](https://doi.org/10.1038/s41597-026-07143-0) achieved a contig N50 of 54.1 Mb with 13 chromosomes assembled gap-free, and the [stonefly genome assembly](https://doi.org/10.1038/s41597-025-06430-6) anchored 99.76 percent of sequence into 11 pseudochromosomes. These examples show that reference completeness directly affects the confidence of downstream read mapping.

If your reference genome is incomplete, you may observe lower mapping rates and more reads mapping to multiple locations. In this case, consider whether a different reference or a reference-free approach would be more appropriate for your analysis.

### Platform-Specific Error Profiles

The parameter recommendations in this article are based on the typical error profiles of PacBio HiFi and ONT reads. However, both platforms are evolving rapidly, and newer sequencing chemistries and basecalling models may produce reads with different error characteristics. Always verify that your parameter choices are appropriate for your specific data by examining the mapping statistics and downstream results.

### Computational Constraints

Parameter choices that improve sensitivity often increase computational cost. If you are working with very large datasets or limited computational resources, you may need to accept a lower mapping rate or use a more stringent parameter set to keep the analysis tractable. The [EMBL-EBI training materials](https://www.ebi.ac.uk/training) provide guidance on optimizing bioinformatics workflows for computational efficiency.

## Welfare and Safety Context for Laboratory Practice

While parameter tuning is a computational task, it is part of a broader laboratory workflow that involves handling biological samples and sequencing reagents. Standard laboratory safety practices apply to the sample preparation and sequencing steps that produce the data used for mapping. These practices include proper handling of biological samples, appropriate use of personal protective equipment, and adherence to institutional biosafety guidelines.

For researchers working with agricultural species, the mapping parameters you choose can affect the reliability of genomic analyses that inform breeding decisions. The [mandarin fish genome assembly](https://doi.org/10.1038/s41597-026-07113-6) established a genomic resource for molecular breeding strategies in this aquaculture species, and the [pond smelt genome assembly](https://doi.org/10.1038/s41597-026-07078-6) provides a high-quality reference for molecular breeding and evolutionary analyses of cold-water fishes. Similarly, the [grain aphid genome assembly](https://doi.org/10.1038/s41597-026-07147-w) serves as a resource for pest management research, where accurate read mapping is essential for identifying genetic variants associated with host adaptation.

## Professional Escalation Criteria

If you encounter persistent problems with read mapping that you cannot resolve through parameter tuning, consider escalating the issue to a bioinformatics specialist or a core facility. The following situations warrant professional consultation:

- Mapping rates remain below 85 percent for HiFi data or below 80 percent for ONT data after extensive parameter tuning
- The reference genome is suspected to contain significant errors or missing regions
- The downstream analysis produces inconsistent results across different parameter sets
- You need to map reads to a highly repetitive genome or a genome with unusual characteristics
- You are developing a new analysis pipeline and need guidance on parameter optimization

Bioinformatics core facilities and specialized consultants can provide expertise in parameter optimization and troubleshooting. The [Galaxy Training Network](https://training.galaxyproject.org/) and [EMBL-EBI training materials](https://www.ebi.ac.uk/training) provide additional learning resources that can help you build the skills needed to resolve mapping problems independently.

## Building a Parameter Decision Log for Cross-Platform Comparison

When you manage mapping tasks across multiple sequencing platforms or multiple batches of reads from the same platform, the absence of a structured record system becomes a practical bottleneck. You may tune parameters successfully for one dataset, only to repeat the same trial-and-error process weeks later when a new batch arrives. A parameter decision log solves this problem by capturing also the final command line but also the evidence that led to each parameter choice, the observed mapping statistics, and the downstream validation results. This section provides a concrete framework for building such a log, with specific fields, comparison methods, and escalation triggers that fit into existing laboratory workflows.

### Why a Decision Log Matters for Multi-Platform Facilities

Facilities that process both PacBio HiFi and ONT reads face a recurring challenge: the same reference genome may be used for both platforms, but the optimal parameters differ. Without a decision log, each researcher independently discovers that `map-hifi` and `map-ont` presets produce different mapping rates, alignment identities, and multi-mapping proportions. The knowledge stays in individual notebooks or memory, and the next person repeats the exploration from scratch.

A decision log converts this tacit knowledge into an explicit, queryable record. It also supports cross-platform comparisons that are difficult to perform informally. For example, you may want to know whether a particular ONT basecalling model produces reads that map as reliably as HiFi reads from the same sample. A structured log lets you compare mapping statistics side by side, identify systematic differences, and decide whether parameter adjustments are needed for one platform or both.

The [Galaxy Training Network](https://training.galaxyproject.org/) emphasizes reproducible analysis workflows, and a decision log is a lightweight complement to workflow tools. It does not replace version control or workflow managers, but it captures the reasoning that those tools cannot record. The [nf-core documentation](https://nf-co.re/docs) similarly stresses the importance of documenting parameter choices, and a decision log provides a human-readable layer on top of machine-readable workflow definitions.

### Core Fields for Each Log Entry

Each entry in the parameter decision log should capture enough information to reproduce the mapping run and to understand why specific parameters were chosen. The following fields provide a practical minimum:

| Field | Purpose | Example Value |
| --- | --- | --- |
| Run identifier | Unique label for the mapping run | `2025-06-14_ont-r10-batch3` |
| Date | When the run was executed | `2025-06-14` |
| Operator | Person who performed the run | `J. Chen` |
| Minimap2 version | Exact software version | `2.28` |
| Reference genome | Version and source | `Siniperca scherzeri v1.0, NCBI` |
| Read data source | Platform, flow cell, basecalling model | `ONT R10.4.1, dorado v0.7` |
| Read subset size | Number of reads or bases used for testing | `200,000 reads` |
| Full command line | Complete minimap2 invocation | `minimap2 -ax map-ont -k13 -A2 -B4 -O4,2 ref.fa reads.fq.gz` |
| Mapping rate | Proportion of reads mapped | `94.2%` |
| Median mapping quality | Central tendency of MAPQ distribution | `42` |
| Multi-mapping proportion | Reads with more than one reported alignment | `6.8%` |
| Median alignment identity | Proportion of matching bases in alignments | `93.1%` |
| Runtime and peak memory | Computational cost | `18 min, 12 GB` |
| Downstream validation | Result of variant calling or assembly check | `F1 score 0.97 on chr1 test set` |
| Rationale for changes | Why parameters were adjusted from the previous run | `Reduced k-mer from 15 to 13 to improve sensitivity` |

The table above is a template, not a fixed schema. You may add fields for reference genome assembly quality metrics, such as contig N50 or BUSCO completeness, when those metrics are relevant to mapping performance. Recent genome assembly projects have shown that reference completeness directly affects mapping rates. The [mandarin fish genome assembly](https://doi.org/10.1038/s41597-026-07113-6) reported mapping rates above 97 percent for ONT ultra-long reads, PacBio HiFi reads, and Hi-C data against a telomere-to-telomere reference, and the [pond smelt genome assembly](https://doi.org/10.1038/s41597-026-07078-6) achieved a contig N50 of 20.23 Mb with all sequences anchored to pseudochromosomes. If your reference has lower continuity, you should record that context in the log because it explains why mapping rates may differ from published benchmarks.

### A Structured Comparison Method for Parameter Candidates

The decision log becomes most useful when you use it to compare parameter candidates systematically. A simple grid search over two or three parameters can generate many combinations, and recording each one in the log lets you identify trends instead of isolated results.

Start with the platform preset as the baseline. For PacBio HiFi reads, use `-ax map-hifi`. For ONT reads, use `-ax map-ont`. Record the mapping statistics for this baseline in the log. Then vary one parameter at a time while holding the others constant. The most informative parameters to vary first are the k-mer size and the mismatch penalty, because they have the largest effect on sensitivity and alignment precision.

For each candidate parameter set, run minimap2 on the same read subset and reference genome. Record the mapping rate, median mapping quality, multi-mapping proportion, and median alignment identity. Also record the runtime, because a parameter set that doubles the runtime may not be practical for full-scale analysis even if it improves sensitivity by a small margin.

After testing three to five candidates, compare the results in the log. Look for the parameter set that achieves the highest mapping rate without a substantial increase in multi-mapping proportion or a decrease in alignment identity. If two parameter sets produce similar mapping statistics, choose the one with the lower runtime or the one that is closer to the platform preset, because deviations from presets are harder to maintain across minimap2 versions.

The [EMBL-EBI training materials](https://www.ebi.ac.uk/training) provide guidance on designing systematic comparisons in bioinformatics, and the same principles apply to parameter tuning. The key is to change one variable at a time so that you can attribute any change in mapping statistics to a specific parameter adjustment.

### Recording Observations That Do Not Fit the Table

Some observations are qualitative and do not fit neatly into a table field. For example, you may notice that a particular parameter set produces alignments that are systematically truncated at the ends of reads, or that reads from a specific genomic region consistently fail to map. These observations are valuable for troubleshooting, and the log should have a free-text field for them.

When you record a qualitative observation, include enough context for someone else to understand it. Instead of writing "alignments look bad," write "alignments from reads spanning the rDNA repeat cluster on chromosome 2 are truncated at the array boundary, suggesting that the k-mer size is too large to seed within the repeat." This level of detail helps you or a colleague diagnose the problem without repeating the full analysis.

The [Carpentries lessons](https://carpentries.org/lessons) teach the value of clear documentation in computational work, and the same principle applies to a decision log. A log entry that records only the command line and the mapping rate is useful, but a log entry that also records why you chose a particular k-mer size and what you observed in the alignments is far more valuable for future troubleshooting.

### Using the Log for Cross-Platform Benchmarking

One of the most practical uses of a parameter decision log is benchmarking one platform against another on the same reference genome. This comparison is common in facilities that evaluate whether to switch from one sequencing platform to another, or that need to combine data from both platforms in a single analysis.

To perform this comparison, create a log entry for each platform using the appropriate preset. Then create additional entries where you tune parameters for each platform to achieve the best mapping statistics. Compare the best achievable mapping rate, alignment identity, and multi-mapping proportion between platforms. Also compare the runtime, because the computational cost of mapping can be a deciding factor for large projects.

Recent genome assembly projects provide useful benchmarks for this comparison. The [Zi goose genome assembly](https://doi.org/10.1038/s41597-026-07143-0) used both ONT long reads and PacBio HiFi reads to produce a telomere-to-telomere reference, and the [stonefly genome assembly](https://doi.org/10.1038/s41597-025-06430-6) used PacBio HiFi reads and Hi-C data. These projects demonstrate that both platforms can produce high-quality mapping results when parameters are matched to the data, but the optimal parameter sets differ.

When you record cross-platform comparisons in the log, include the reference genome version and the read data source for each entry. This context is essential because a comparison performed on one reference genome may not generalize to another. The [grain aphid genome assembly](https://doi.org/10.1038/s41597-026-07147-w) and the [stonefly genome assembly](https://doi.org/10.1038/s41597-025-06430-6) have different genome sizes and repetitive content, and mapping parameters that work well for one may be suboptimal for the other.

### Common Failure Patterns in Parameter Documentation

The most common failure in parameter documentation is recording the command line but not the rationale. A command line alone tells you what was run, but it does not tell you why those parameters were chosen or what alternatives were considered. When you revisit the log months later, you may not remember whether the k-mer size of 13 was chosen because the default of 15 produced a low mapping rate, or because a colleague recommended it based on a different dataset.

Another common failure is recording only the final parameter set without recording the intermediate candidates. If the final parameter set produces poor results in a downstream analysis, you have no record of what else was tried. Recording all candidates, including those that were rejected, gives you a complete picture of the parameter space you explored.

A third failure is inconsistent use of read subsets. If you test parameters on 100,000 reads one day and 500,000 reads the next, the mapping statistics are not directly comparable. The log should record the read subset size for each entry, and you should use the same subset size for all entries in a comparison series.

### Professional Escalation Criteria for Parameter Troubleshooting

The decision log also serves as evidence for professional escalation. If you have systematically tested parameter candidates and recorded the results, you can present this evidence to a bioinformatics specialist or core facility consultant. The log shows what you have tried, what the results were, and where the remaining problems lie.

Escalate to a specialist when any of the following conditions are met:

- Mapping rates remain below 85 percent for HiFi data or below 80 percent for ONT data after testing at least five parameter combinations
- The multi-mapping proportion exceeds 15 percent and does not decrease when the k-mer size is increased
- Alignment identities are consistently below 98 percent for HiFi data or below 85 percent for ONT data, suggesting a systematic problem with the reads or the reference
- The same parameter set produces very different mapping statistics across batches of reads from the same platform, suggesting a batch effect or a change in data quality
- You need to map reads to a reference genome with unusual characteristics, such as extreme repetitive content or a high level of heterozygosity

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) can help you access reference genomes and read archives for comparison, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials that can help you build the skills to resolve mapping problems independently. However, when the problem persists despite systematic parameter tuning, professional consultation is the appropriate next step.

### Integrating the Decision Log with Existing Workflow Tools

The decision log does not replace workflow managers or version control systems, but it complements them. Workflow managers such as Nextflow, documented in the [nf-core documentation](https://nf-co.re/docs), capture the command lines and parameter values in a machine-readable format. The decision log adds the human reasoning that workflow managers cannot record.

A practical integration strategy is to store the decision log as a spreadsheet or a plain-text table in the same repository as your analysis scripts. Each log entry can reference the workflow file or script that executed the mapping run. This approach gives you both the machine-readable record and the human-readable rationale.

For researchers who use R for downstream analysis, the [Bioconductor project](https://bioconductor.org/) provides packages for reading and visualizing alignment data. You can use these tools to generate the mapping statistics that go into the decision log, and you can also use them to create plots that compare parameter candidates visually. These plots can be attached to the log entry as supplementary evidence.

### A Practical Example of a Log Entry

To illustrate the framework, consider a concrete example. A researcher is mapping ONT ultra-long reads from a fish species to a newly assembled reference genome. The baseline run with `-ax map-ont` produces a mapping rate of 91.5 percent, a median mapping quality of 38, and a multi-mapping proportion of 7.2 percent. The researcher records this baseline in the log.

The researcher then tests a k-mer size of 13 while keeping all other parameters at the preset values. The mapping rate increases to 94.2 percent, the median mapping quality increases to 42, and the multi-mapping proportion increases slightly to 7.8 percent. The runtime increases from 14 minutes to 18 minutes. The researcher records this candidate in the log and notes that the improvement in mapping rate justifies the modest increase in runtime.

The researcher also tests a mismatch penalty of 6 while keeping the k-mer at 15. The mapping rate decreases to 89.7 percent, and the median alignment identity increases from 93.1 percent to 94.5 percent. The researcher records this candidate and notes that the higher mismatch penalty reduces sensitivity more than it improves precision, so it is not a good choice for this dataset.

After testing three candidates, the researcher selects the k-mer size of 13 with the preset scoring parameters. The decision log shows the complete comparison, the rationale for the final choice, and the downstream validation results. When a new batch of ONT reads arrives, the researcher can start with the recorded parameter set and only adjust if the mapping statistics differ from the baseline.

This example demonstrates the practical value of the decision log. It turns parameter tuning from an ad hoc process into a systematic, documented activity that produces reproducible results and supports informed decision making.

## Frequently Asked Questions

### What is the difference between the map-hifi and map-ont presets in minimap2?

The map-hifi preset is optimized for PacBio HiFi reads, which have high accuracy and low error rates. It uses a larger k-mer size of 19 and scoring parameters that assume most mismatches and gaps are biologically real. The map-ont preset is optimized for Oxford Nanopore reads, which have higher error rates with more insertions and deletions. It uses a smaller k-mer size of 15 and scoring parameters that tolerate more errors. The presets also differ in other parameters such as the minimum seed length and the chaining options.

### How do I know if my minimap2 parameters are appropriate for my data?

The most direct way to evaluate parameter appropriateness is to examine the mapping statistics, including the mapping rate, the distribution of mapping qualities, and the alignment identity distribution. Compare these statistics against the expected values for your platform and data quality. If the mapping rate is lower than expected, the parameters are likely too stringent. If the mapping rate is high but the alignment quality is poor, the parameters are likely too permissive. You should also validate the downstream analysis results to ensure that the alignments produce biologically meaningful conclusions.

### Can I use the same parameters for PacBio HiFi and ONT reads?

You can use the same parameters for both platforms, but the results will likely be suboptimal for at least one platform. HiFi reads have lower error rates and can tolerate larger k-mers and more stringent scoring parameters, while ONT reads require smaller k-mers and more permissive scoring to achieve adequate sensitivity. Using the platform-specific presets is the simplest way to ensure appropriate parameters, but you may need to adjust individual parameters based on your specific data characteristics.

### What is the best k-mer size for ONT ultra-long reads?

The optimal k-mer size for ONT ultra-long reads depends on the error rate of the basecalling model. For standard ONT data, a k-mer size of 15 is a reasonable starting point. For ultra-long reads with higher error rates, you may need to reduce the k-mer to 13 or 11 to achieve adequate sensitivity. The long read length provides additional mapping information through multiple independent seeds, so the sensitivity loss from a larger k-mer may be partially compensated by the ability to chain seeds across long distances.

### How do secondary alignments affect downstream analysis?

Secondary alignments represent alternative mappings for reads that match multiple locations in the reference genome. For variant calling, secondary alignments are typically ignored or filtered out because they cannot be confidently placed. For genome assembly, secondary alignments can be useful for resolving repetitive regions, but they require careful handling to avoid misassemblies. The appropriate number of secondary alignments depends on your downstream analysis and the repetitiveness of your reference genome.

### Should I adjust the scoring parameters for structural variant detection?

The scoring parameters can affect the sensitivity and precision of structural variant detection. More permissive scoring parameters may increase sensitivity for detecting variants in repetitive regions, but they also increase the number of false positive calls. More stringent scoring parameters reduce false positives but may miss genuine variants. The optimal scoring parameters depend on the specific variant caller you are using and the characteristics of your data. Test a range of parameter values and compare the variant calls against known variants or against calls from a different aligner.

### How do I document my minimap2 parameters for reproducibility?

Record the full minimap2 command line, including all parameters, in your analysis scripts or workflow files. Include the minimap2 version, the reference genome version and source, the read data version and source, and the date of the run. Place your parameter choices under version control so that changes are tracked and documented. The [nf-core documentation](https://nf-co.re/docs) provides guidance on documenting parameter choices in reproducible workflows.

### What should I do if my mapping rate is still low after parameter tuning?

If the mapping rate remains low after extensive parameter tuning, consider whether the reference genome is appropriate for your data. Check for reference errors or missing regions by examining the coverage distribution and the mapping quality of reads in specific genomic regions. You may also want to try a different aligner to determine whether the problem is specific to minimap2 or a general issue with your data and reference. If the problem persists, consult a bioinformatics specialist for guidance.

## Related Bioinformatics Guides

- [How to Choose a Long-Read Sequencing Platform: PacBio vs Oxford Nanopore](/knowledge/bioinformatics/how-to-choose-a-long-read-sequencing-platform-pacbio-vs-oxford-nanopore)
- [Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data](/knowledge/bioinformatics/long-read-metagenome-assembly-overcoming-challenges-with-nanopore-and-pacbio-data)
- [Long-Read Sequencing Technologies: PacBio and Oxford Nanopore](/knowledge/bioinformatics/long-read-sequencing-technologies-pacbio-and-oxford-nanopore)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Long-Read Sequencing Cost and Market: What to Expect](/knowledge/bioinformatics/long-read-sequencing-cost-and-market-what-to-expect)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Telomere-to-telomere gapless genome assembly of Siniperca scherzeri.](https://doi.org/10.1038/s41597-026-07113-6). 2026.
- [A telomere-to-telomere reference genome assembly of the Hypomesus nipponensis.](https://doi.org/10.1038/s41597-026-07078-6). 2026.
- [Telomere-to-telomere-level genome assembly and annotation of the Zi goose Anser cygnoides.](https://doi.org/10.1038/s41597-026-07143-0). 2026.
- [The chromosome-level genome assembly of Sitobion avenae Fabricius (Hemiptera: Aphididae).](https://doi.org/10.1038/s41597-026-07147-w). 2026.
- [Chromosome-level genome assembly of the stonefly Indonemoura scalprata (Plecoptera: Nemouridae).](https://doi.org/10.1038/s41597-025-06430-6). 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.