Minimap2 vs. NGMLR vs. GraphMap: A Benchmark of Long-Read Aligners for Structural Variant Detection

By Dr. Zubair Khalid, DVM, MS, PhD ·

Minimap2 vs. NGMLR vs. GraphMap: A Benchmark of Long-Read Aligners for Structural Variant Detection

Key Takeaways

  • Minimap2 is the current de facto standard for long-read structural variant (SV) detection workflows, offering the optimal balance of speed, accuracy, and broad community support for modern PacBio HiFi and Oxford Nanopore data, making it suitable for population-scale studies due to its computational efficiency.
  • NGMLR retains utility for specific, error-prone datasets or targeted validation of complex SV loci, particularly with legacy PacBio CLR data or when high gap tolerance is paramount, though its slower runtime and higher memory footprint limit its scalability for large cohorts.
  • GraphMap is largely obsolete for production pipelines, having been designed for early, high-error Oxford Nanopore data; it lacks the performance and feature set required for current sequencing chemistries and downstream SV caller compatibility.
  • Aligner choice critically impacts SV detection accuracy, breakpoint resolution, and variant type sensitivity, with downstream SV callers interpreting read placements; inconsistent alignment of reads at complex SVs and inversions is a primary driver of discordance between different detection strategies.
  • Reproducibility necessitates rigorous documentation and standardization, including reference genome versions, aligner parameters, software versions, and crucially, read order and sorting algorithms, as demonstrated by studies showing significant SV call variability due to input data ordering.

Researchers building structural variant (SV) detection pipelines for PacBio or Oxford Nanopore long-read data face a practical decision at the alignment step: which aligner produces the most accurate read placements for downstream SV calling? This article compares minimap2, NGMLR, and GraphMap across alignment accuracy, SV breakpoint resolution, runtime, and memory consumption, using evidence from published benchmarks and official bioinformatics training resources. The direct answer is that minimap2 currently offers the best balance of speed, accuracy, and community support for most SV detection workflows, while NGMLR retains value for specific error-prone datasets and GraphMap is largely obsolete for production pipelines. The comparison below gives concrete criteria for choosing among them, recording performance metrics, and escalating when alignment artifacts compromise SV calls.

Scope and Reader Context

This benchmark targets biology students, wet-lab researchers transitioning to computational analysis, laboratory professionals validating pipelines, and life-science practitioners who need reproducible long-read SV detection. The practical outcome is a decision framework for aligner selection based on data type, available compute, SV class priorities, and downstream caller compatibility. The comparison draws on peer-reviewed benchmarking studies, official training materials from EMBL-EBI and Galaxy, and community pipeline documentation from nf-core. It does not replace hands-on validation with your own data, but it provides the measurement approach and quality thresholds that make local benchmarking meaningful.

The three aligners differ fundamentally in design philosophy. GraphMap was developed for early Oxford Nanopore data with high error rates and has not kept pace with sequencing chemistry improvements. NGMLR was designed specifically for SV detection with long reads and uses a sparse seeding approach that tolerates large gaps. Minimap2 is a general-purpose aligner that handles both short and long reads, uses minimizer-based seeding, and has become the default in most modern pipelines. Understanding these design differences explains why performance varies across datasets and why no single aligner dominates every benchmark.

Structural Variant Detection with Long Reads

Structural variants include deletions, insertions, duplications, inversions, and translocations larger than 50 base pairs. Short-read sequencing detects these poorly because reads cannot span large repetitive or duplicated regions. Long reads from PacBio and Oxford Nanopore platforms produce fragments of 10 kilobases or more, which can span entire SV breakpoints and resolve complex rearrangements. The Human Genome Structural Variation Consortium analysis found that long-read assembly captured roughly 25,000 SVs per genome compared with about 11,000 from short-read whole-genome sequencing, with long reads uniquely accessing insertions and repeat expansions in segmental duplications and simple repeats [<a href="#ref-1">1</a>]. This sensitivity advantage depends entirely on alignment quality, because SV callers interpret read placements relative to a reference genome.

Alignment-based SV detection works by mapping each long read to the reference, then identifying reads that disagree with the reference structure. A deletion appears as a read with a large internal gap in its alignment. An insertion appears as a read with extra sequence that does not align contiguously. Inversions and translocations produce split alignments where different segments of one read map to different locations or orientations. The aligner must place these reads correctly despite sequencing errors, repetitive sequence, and genuine structural variation. If the aligner misplaces reads or fragments them incorrectly, the SV caller either misses variants or produces false positives.

The choice of aligner affects SV calls in ways that are measurable but not always predictable. A 2023 benchmark of read-based and assembly-based SV detection pipelines found that variant type, size, and breakpoint accuracy from read-based strategies were greatly affected by the aligner used [<a href="#ref-2">2</a>]. The same study reported that up to 80% of SVs could be detected by both strategies across different long-read datasets, but discordance between strategies was largely caused by complex SVs and inversions resulting from inconsistent alignment of reads and assemblies at those loci [<a href="#ref-2">2</a>]. This means aligner choice directly influences which SVs your pipeline finds and how precisely it defines their breakpoints.

Aligner Design and Error Profiles

Minimap2 Architecture

Minimap2 uses minimizer-based seeding, where the read and reference are reduced to a subset of representative k-mers, and candidate alignments are built from colinear minimizer matches. It then performs chaining to identify the best alignment path and base-level alignment for the final output. This design gives minimap2 exceptional speed because it avoids exhaustive seed comparison. It handles both PacBio and Oxford Nanopore data with preset parameter sets that account for platform-specific error profiles. The aligner also supports spliced alignment, which makes it useful beyond SV detection for transcriptome analysis.

Minimap2 produces SAM output with mapping quality scores and supplementary alignments for split reads. These features matter for SV callers because split-read evidence is essential for detecting inversions and translocations. The aligner also outputs the CIGAR string that records gaps and mismatches, which SV callers use to infer deletion and insertion sizes. Minimap2 has become the default aligner in nf-core community pipelines and Galaxy workflows because of its speed and maintained development [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].

NGMLR Design

NGMLR was developed specifically for structural variant detection with long reads. It uses a sparse seeding approach where seeds are selected based on their uniqueness in the reference, then extended with a dynamic programming step that tolerates large gaps. This design targets the problem of aligning reads that span SV breakpoints, where a read may have a long unaligned segment corresponding to a deletion or insertion. NGMLR was optimized for PacBio data and early Oxford Nanopore data with error rates around 10% to 15%.

The tradeoff is computational cost. NGMLR runs slower than minimap2 because the sparse seeding and gap-tolerant extension require more computation per read. Memory usage is also higher because the aligner keeps more intermediate data structures. For small datasets or targeted validation of specific loci, this cost is acceptable. For population-scale studies with hundreds of samples, the runtime difference becomes prohibitive.

GraphMap Architecture

GraphMap was one of the first aligners designed for Oxford Nanopore reads, using a seed-and-vote approach that builds candidate regions from k-mer matches, then refines alignments with a banded dynamic programming step. It was designed for the high error rates and systematic base-calling biases of early nanopore chemistry. GraphMap introduced a robust alignment model that could handle insertion and deletion errors common in nanopore data.

The limitation is that GraphMap has not been actively maintained as nanopore chemistry improved. Modern nanopore data with Q20 or better base quality does not need the aggressive error tolerance that GraphMap provides, and the aligner lacks features that modern SV callers expect, such as supplementary alignment output and compatibility with current SAM specifications. Most contemporary benchmarks and pipelines have dropped GraphMap in favor of minimap2 or NGMLR.

Benchmark Evidence from Published Studies

Read-Based versus Assembly-Based Detection

A 2023 benchmark in Briefings in Bioinformatics compared 20 read-based and eight assembly-based SV detection pipelines on six datasets from the HG002 human genome [<a href="#ref-2">2</a>]. The study found that up to 80% of SVs could be detected by both strategies across different long-read datasets, but the aligner used for read-based detection greatly affected variant type, size, and breakpoint accuracy [<a href="#ref-2">2</a>]. For high-confidence insertions and deletions at non-tandem repeat regions, 82% of assembly-based calls and 93% of read-based calls were captured by both reads and assemblies, accounting for around 4000 SVs [<a href="#ref-2">2</a>]. Discordance between strategies was largely caused by complex SVs and inversions, which resulted from inconsistent alignment of reads and assemblies at these loci [<a href="#ref-2">2</a>].

The clinical relevance finding is that read-based strategy recall reached 77% on 5X coverage data, while assembly-based strategy required 20X coverage to achieve similar performance [<a href="#ref-2">2</a>]. This supports the use of alignment-based SV detection for low-coverage studies where assembly is not feasible. It also means aligner choice has outsized importance at low coverage, because each read carries more weight in the SV call.

Systematic Comparison of Alignment and Assembly Methods

A 2024 Nature Communications study systematically compared 14 read alignment-based SV calling methods, four assembly-based methods, four upstream aligners, and seven assemblers [<a href="#ref-5">5</a>]. The study found that assembly-based tools excel in detecting large SVs, especially insertions, and exhibit robustness to evaluation parameter changes and coverage fluctuations [<a href="#ref-5">5</a>]. Alignment-based tools demonstrate superior genotyping accuracy at low sequencing coverage of 5X to 10X and excel in detecting complex SVs like translocations, inversions, and duplications [<a href="#ref-5">5</a>]. The study emphasized that no universally superior tool exists and provided guidelines across 31 criteria combinations to aid tool selection [<a href="#ref-5">5</a>].

This benchmark matters for aligner choice because the four upstream aligners included minimap2 and NGMLR, and the performance of downstream SV callers depended on which aligner produced the input. The study's conclusion that alignment-based methods are favored for computational efficiency and lower coverage requirements supports the practical use of minimap2 for most workflows [<a href="#ref-5">5</a>].

Read Order Sensitivity

A 2024 PeerJ study examined whether FASTQ read order affects SV calling from long-read data [<a href="#ref-6">6</a>]. Using PacBio data from 15 Caenorhabditis elegans strains and four Arabidopsis thaliana ecotypes, the study found that the order of input data affected the SVs predicted by each caller [<a href="#ref-6">6</a>]. The pbsv caller was highly sensitive to read order, especially at the highest depths where over 70% of SV calls generated from pairs of differently ordered FASTQ files were in disagreement [<a href="#ref-6">6</a>]. The SAMtools alignment sorting algorithm was identified as a source of variability following read order randomization [<a href="#ref-6">6</a>].

This finding has direct implications for aligner benchmarking. If you compare minimap2 and NGMLR on the same dataset, you must control for read order and sorting behavior. Differences in SV calls between aligners could be confounded by read order sensitivity instead of true alignment quality differences. The study suggests that researchers should standardize read order and document sorting parameters in their protocols [<a href="#ref-6">6</a>].

At a Glance

AlignerDesign PurposeBest Use CaseRuntime ProfileMemory ProfileSV Caller Compatibility
Minimap2General-purpose long and short read alignmentModern PacBio and Nanopore data, production pipelines, population studiesFast, suitable for high-throughput datasetsModerate, scales to whole-genome datasetsDefault in most SV callers including Sniffles2, cuteSV, and pbsv
NGMLRSV-specific alignment for error-prone long readsLegacy PacBio data, validation of complex loci, small targeted panelsSlow, impractical for large cohortsHigh, requires substantial RAMDesigned for Sniffles, compatible with other callers
GraphMapEarly Oxford Nanopore alignmentHistorical nanopore datasets, educational comparisonSlow, not maintained for modern dataModerateLimited, lacks modern SAM output features

The table summarizes the practical distinctions. Minimap2 is the default choice for new projects. NGMLR is a fallback for specific validation scenarios. GraphMap is not recommended for new pipelines.

Practical Workflow for Aligner Selection

Step 1: Define Data Characteristics

Record the sequencing platform, chemistry version, base-calling model, and estimated error rate for your dataset. PacBio HiFi data has error rates below 1% and benefits from minimap2 preset for HiFi reads. PacBio continuous long reads have error rates around 10% to 15% and may benefit from NGMLR. Oxford Nanopore data varies by chemistry and base-calling model, with recent versions approaching HiFi accuracy. The official EMBL-EBI training materials recommend understanding your data type before selecting analysis tools, because each platform has distinct error profiles that affect alignment parameters [<a href="#ref-7">7</a>].

Step 2: Select Candidate Aligners

For most projects, minimap2 is the primary candidate. If you are working with legacy PacBio data or need to validate SV calls at complex loci, include NGMLR as a secondary aligner. GraphMap should only be included if you are reproducing historical analyses or teaching alignment concepts. The nf-core documentation provides pipeline configurations that show minimap2 as the default aligner in community-maintained workflows, which reduces the burden of parameter optimization [<a href="#ref-4">4</a>].

Step 3: Run Controlled Comparison

Create a test dataset that includes known SVs if available, or use a well-characterized reference sample like HG002. Run each aligner with platform-appropriate presets. Record runtime and peak memory usage for each aligner using standard system monitoring tools. The Galaxy Training Network provides tutorials on running alignment workflows that include parameter selection and output validation steps [<a href="#ref-3">3</a>].

Step 4: Evaluate Alignment Statistics

Use SAMtools or equivalent tools to compute alignment statistics for each aligner output. Record the number of mapped reads, mapping quality distribution, and the proportion of reads with supplementary alignments. High mapping rates with low mapping quality suggest the aligner is placing reads but with low confidence. A high proportion of supplementary alignments may indicate genuine structural variation or alignment fragmentation, and you need to distinguish these cases.

Step 5: Run SV Calling and Compare

Run the same SV caller on each aligner output. Sniffles2, cuteSV, and pbsv are common choices, and each has documented compatibility with minimap2 output. Compare the number of SV calls, the size distribution, and the breakpoint coordinates. The 2023 benchmark found that aligner choice affects variant type, size, and breakpoint accuracy, so differences in these metrics between aligners are expected [<a href="#ref-2">2</a>]. Focus on high-confidence calls that are supported by multiple reads and are reproducible across aligners.

Step 6: Validate with Visualization

Use Ribbon or similar visualization tools to inspect alignments at SV loci where aligners disagree. Ribbon shows how alignments are positioned within both the reference and read contexts, giving an intuitive view that enables better understanding of structural variants and the read evidence supporting them [<a href="#ref-8">8</a>]. This step is essential for distinguishing true SVs from alignment artifacts. The tool was developed specifically for curating complex structural variant calls and determining whether each was well supported by long-read evidence [<a href="#ref-8">8</a>].

Step 7: Document and Standardize

Record the aligner version, parameter settings, reference genome version, and read order for every analysis. The read order sensitivity study found that input data order affected SV calls and that the SAMtools alignment sorting algorithm was a source of variability [<a href="#ref-6">6</a>]. Standardizing these parameters across samples and studies is necessary for reproducibility. The Carpentries lessons on version control and reproducible research provide foundational practices for documenting analysis workflows [<a href="#ref-9">9</a>].

Records and Measurements

Runtime and Memory Logging

For each aligner comparison, record the following metrics in a structured table:

MetricMinimap2NGMLRGraphMap
Input read countRecord actual countRecord actual countRecord actual count
Total bases alignedRecord actual basesRecord actual basesRecord actual bases
Wall clock timeRecord hours and minutesRecord hours and minutesRecord hours and minutes
Peak memoryRecord gigabytesRecord gigabytesRecord gigabytes
Mapped read percentageRecord percentageRecord percentageRecord percentage
Median mapping qualityRecord Phred scoreRecord Phred scoreRecord Phred score
Supplementary alignment percentageRecord percentageRecord percentageRecord percentage
SV calls from downstream callerRecord countRecord countRecord count
SV calls validated by visualizationRecord countRecord countRecord count

These records allow you to compare aligners on your specific data instead of relying only on published benchmarks. The Nature Communications study provided performance insights across 31 criteria combinations, but your data may differ in error rate, coverage, and variant composition [<a href="#ref-5">5</a>].

SV Call Concordance Tracking

When comparing aligners, track the concordance of SV calls between aligner outputs. For each SV call from the primary aligner, check whether the secondary aligner produces a call at the same locus with overlapping breakpoint coordinates. Record the number of calls unique to each aligner and the number shared. The 2023 benchmark found that discordance between read-based and assembly-based strategies was largely caused by complex SVs and inversions [<a href="#ref-2">2</a>]. Similar discordance between aligners indicates loci where alignment uncertainty is high and manual curation is needed.

Breakpoint Precision Measurement

For validated SV calls, measure breakpoint precision by comparing the reported breakpoint coordinates with the true breakpoints from a curated truth set or from assembly-based calls. The 2023 benchmark found that breakpoints detected by read-based strategy were greatly affected by aligners [<a href="#ref-2">2</a>]. Record the distribution of breakpoint offsets for each aligner. Smaller offsets indicate more precise alignment at SV boundaries.

Common Failure Patterns

Aligner-Specific Failure Modes

Minimap2 can fail on reads that span very large insertions because the minimizer-based seeding may not find enough anchors across the insertion. This produces fragmented alignments where the insertion is represented as multiple supplementary alignments instead of one continuous read placement. NGMLR handles large insertions better because of its gap-tolerant extension, but it can produce false alignments in repetitive regions where sparse seeding finds spurious anchors. GraphMap fails on modern high-quality data because its error model expects more errors than are present, leading to unnecessary alignment fragmentation.

Read Order and Sorting Artifacts

The read order sensitivity study found that SV callers can produce different results depending on the order of reads in the input FASTQ file, with pbsv showing over 70% disagreement between differently ordered files at high depth [<a href="#ref-6">6</a>]. This means that if you run the same aligner and caller twice with different read orders, you may get different SV calls. The SAMtools sorting algorithm was identified as a source of variability following read order randomization [<a href="#ref-6">6</a>]. To avoid this failure mode, standardize read order across samples and document the sorting procedure.

Coverage-Dependent Failures

At low coverage of 5X to 10X, alignment-based SV detection has superior genotyping accuracy compared with assembly-based methods [<a href="#ref-5">5</a>]. However, low coverage also means fewer reads support each SV call, increasing the chance that aligner errors produce false calls. At high coverage, read order sensitivity increases for some callers [<a href="#ref-6">6</a>]. The practical implication is that you should validate SV calls at coverage extremes with visualization tools and consider whether additional sequencing is needed.

Complex SV Misalignment

Complex SVs involving multiple breakpoints in the same region are difficult for all aligners. The 2023 benchmark found that discordance between read-based and assembly-based strategies was largely caused by complex SVs and inversions, which resulted from inconsistent alignment of reads and assemblies at these loci [<a href="#ref-2">2</a>]. When aligners disagree on complex loci, manual curation with Ribbon is necessary to determine whether the SV is real and which alignment is correct [<a href="#ref-8">8</a>].

Quality Controls and Reproducibility

Reference Genome Version Control

Record the exact reference genome version and any patches or alt-contig masking used for alignment. Different reference versions can change SV calls at repetitive and duplicated loci. The NCBI provides official reference genome resources and documentation on assembly versions [<a href="#ref-10">10</a>]. Using the same reference version across all samples in a study is necessary for comparability.

Parameter Documentation

Record all aligner parameters, including preset selections, minimum seed length, and mapping quality thresholds. The nf-core documentation emphasizes the importance of configuration management for reproducible workflows [<a href="#ref-4">4</a>]. The Galaxy Training Network provides tutorials that show how parameter choices affect alignment output and downstream analysis [<a href="#ref-3">3</a>].

Pipeline Version Pinning

Pin the versions of all software in your pipeline, including the aligner, SV caller, and sorting tools. The read order sensitivity study found that the SAMtools alignment sorting algorithm was a source of variability [<a href="#ref-6">6</a>]. Version changes in any of these tools can alter results. The Bioconductor project provides documentation on reproducible package management for R-based analysis workflows [<a href="#ref-11">11</a>].

Validation with Independent Methods

For high-confidence SV calls, validate with an independent method such as assembly-based detection or PCR amplification across the breakpoint. The 2023 benchmark found that integrating SVs from read and assembly is suggested for general-purpose detection because of inconsistently detected complex SVs and inversions [<a href="#ref-2">2</a>]. Assembly-based validation is optional for applications with limited resources, but it provides the strongest confirmation for clinically relevant calls [<a href="#ref-2">2</a>].

Limitations of Aligner Benchmarks

Dataset Specificity

Published benchmarks use specific datasets with particular error profiles, coverage levels, and variant compositions. The Nature Communications study found that no universally superior tool exists and provided guidelines across 31 criteria combinations [<a href="#ref-5">5</a>]. Your data may not match the benchmark conditions, so local validation is necessary.

Rapid Tool Development

Aligners and SV callers are under active development. A benchmark published in 2023 or 2024 may not reflect the current version of minimap2 or NGMLR. The 2024 Nature Communications study compared 14 read alignment-based SV calling methods and four assembly-based methods, but newer versions may have improved performance [<a href="#ref-5">5</a>]. Check the release notes for each tool and rerun benchmarks when major versions change.

Truth Set Limitations

Benchmark studies rely on truth sets that may not include all SV types or genomic contexts. The Human Genome Structural Variation Consortium analysis found that detection power and precision varied dramatically by genomic context and variant class, with 91.4% of deletions specifically discovered by long-read sequencing localizing to segmental duplications and simple repeats [<a href="#ref-1">1</a>]. Truth sets that do not cover these regions may overstate or understate aligner performance.

Computational Resource Constraints

Runtime and memory benchmarks depend on the compute environment. A comparison run on a high-memory server may show different relative performance than one run on a standard workstation. The 2024 Nature Communications study noted that assembly-based methods demand significantly more computational resources than alignment-based methods [<a href="#ref-5">5</a>]. Record your compute environment when benchmarking aligners so that runtime comparisons are interpretable.

Safety and Regulatory Context

Clinical Validation Requirements

If SV calls are used for clinical decision-making, the entire pipeline must be validated according to applicable regulations. The American Journal of Human Genetics study noted that virtually all genome sequencing efforts in national biobanks and medical genetic initiatives rely on short-read sequencing, which presents challenges for SV detection relative to long-read technologies [<a href="#ref-1">1</a>]. Long-read SV detection pipelines used in clinical contexts require validation against curated truth sets and documentation of performance characteristics.

Data Privacy and Security

Long-read sequencing data contains sensitive genomic information. Alignment and SV calling workflows must comply with data protection regulations applicable to your institution and jurisdiction. The NCBI provides documentation on data submission and access policies for genomic data [<a href="#ref-10">10</a>]. Ensure that your compute environment meets institutional security requirements before processing human data.

Reproducibility for Publication

Journals increasingly require that bioinformatics analyses be reproducible. The read order sensitivity study highlighted implications for the replication of SV studies and the development of consistent SV calling protocols [<a href="#ref-6">6</a>]. Documenting aligner versions, parameters, and read order is necessary for meeting reproducibility standards. The Carpentries lessons provide foundational training on reproducible research practices [<a href="#ref-9">9</a>].

Professional Escalation Criteria

When to Seek Expert Consultation

Escalate to a bioinformatics specialist or computational genomics expert when you observe any of the following:

  • SV call concordance between aligners drops below 70% for high-confidence calls
  • Breakpoint precision varies by more than 100 base pairs between aligners for the same SV
  • Runtime or memory usage exceeds your compute environment capacity for the primary aligner
  • SV calls at medically relevant genes cannot be validated by visualization or independent methods
  • Read order sensitivity produces different SV calls in replicate runs of the same data

When to Consider Assembly-Based Validation

The 2023 benchmark found that read-based strategy recall reached 77% on 5X coverage data, while assembly-based strategy required 20X coverage to achieve similar performance [<a href="#ref-2">2</a>]. If your SV calls are used for clinical reporting or if you need comprehensive detection of insertions, consider assembly-based validation. The 2024 Nature Communications study found that assembly-based tools excel in detecting large SVs, especially insertions, and exhibit robustness to coverage fluctuations [<a href="#ref-5">5</a>]. Assembly-based validation is optional for applications with limited resources [<a href="#ref-2">2</a>].

When to Update Your Pipeline

Update your aligner and SV caller versions when major releases include bug fixes or new features that affect SV detection. The 2024 Nature Communications study provided directions for further method development, indicating that the field is actively evolving [<a href="#ref-5">5</a>]. Monitor release notes and rerun benchmarks on a validation dataset before adopting new versions in production.

Decision Framework for Aligner Selection Across Study Scales

The practical question for most research groups is not which aligner performs best on a published benchmark, but which aligner should be used for a specific study design with defined sample counts, coverage targets, and computational budgets. Published benchmarks provide performance rankings, but they do not translate directly into a selection decision because study scale changes the relative importance of runtime, memory, and accuracy. A decision framework that maps study characteristics to aligner choice helps researchers avoid both over-investment in compute resources and under-powered validation.

Study Scale Categories and Aligner Fit

Small Targeted Panels

For studies analyzing fewer than 20 samples at targeted loci, runtime and memory constraints are rarely limiting. The dominant consideration is alignment accuracy at the specific regions of interest, particularly if those regions contain complex SVs or repetitive elements. NGMLR becomes a viable primary aligner in this context because its slower runtime is acceptable at small scale, and its gap-tolerant extension can produce more complete alignments across large insertions and complex breakpoints. The 2024 Nature Communications benchmark found that alignment-based tools demonstrate superior genotyping accuracy at low sequencing coverage of 5X to 10X [<a href="#ref-5">5</a>]. For targeted panels where each sample is sequenced at moderate depth, NGMLR provides a useful cross-check against minimap2 output, and the computational cost remains manageable.

The workflow for small panels should include running both minimap2 and NGMLR on the same samples, then comparing SV calls at the targeted loci. The 2023 benchmark found that variant type, size, and breakpoint detected by read-based strategy were greatly affected by aligners [<a href="#ref-2">2</a>]. When the two aligners disagree at a targeted locus, manual curation with Ribbon is warranted because the region is small enough that visual inspection of every discordant call is feasible [<a href="#ref-8">8</a>]. This dual-aligner approach converts the aligner selection problem into a validation opportunity instead of a one-time choice.

Medium Cohort Studies

For studies with 20 to 200 samples, runtime becomes a meaningful constraint but not the dominant one. The primary consideration shifts to consistency across samples. If the same aligner and parameters are used for all samples, systematic alignment biases will affect all samples equally, which is preferable for case-control comparisons. Minimap2 is the recommended primary aligner for this scale because its speed allows the full cohort to be processed within reasonable timeframes, and its output is compatible with the widest range of SV callers. The nf-core documentation shows minimap2 as the default aligner in community-maintained pipelines, which reduces the burden of parameter optimization and simplifies reproducibility across the cohort [<a href="#ref-4">4</a>].

The critical practice at this scale is to run a pilot comparison on a subset of samples before committing to the full cohort. Select five to ten samples that represent the range of coverage and quality in your dataset. Run both minimap2 and NGMLR on these samples, compare SV calls, and measure runtime and memory. If the SV call concordance between aligners is high, proceed with minimap2 for the full cohort. If concordance is low, investigate whether the discordance is concentrated at specific loci or variant types before deciding whether NGMLR is needed for the full cohort.

Population-Scale Studies

For studies with more than 200 samples, runtime and memory dominate the aligner selection decision. Minimap2 is the only viable choice among the three aligners for population-scale work. NGMLR runtime becomes prohibitive at this scale, and GraphMap is not maintained for modern data. The 2024 Nature Communications study noted that alignment-based methods are favored for their computational efficiency and lower coverage requirements [<a href="#ref-5">5</a>]. Minimap2 delivers this efficiency while maintaining compatibility with the SV callers used in population studies.

The practical implication is that population-scale studies must accept the alignment bias profile of minimap2 and design validation strategies around it. The read order sensitivity study found that SV callers can produce different results depending on the order of reads in the input FASTQ file, with pbsv showing over 70% disagreement between differently ordered files at high depth [<a href="#ref-6">6</a>]. At population scale, standardizing read order and sorting procedures across all samples is essential to prevent this source of variability from confounding downstream analyses. The SAMtools alignment sorting algorithm was identified as a source of variability following read order randomization [<a href="#ref-6">6</a>], so the sorting step must be documented and controlled.

Coverage-Based Selection Criteria

Low Coverage Studies at 5X to 10X

The 2024 Nature Communications benchmark found that alignment-based tools demonstrate superior genotyping accuracy at low sequencing coverage of 5X to 10X [<a href="#ref-5">5</a>]. This finding supports the use of alignment-based SV detection for studies where sequencing depth is limited by budget or sample availability. At this coverage level, each read carries more weight in the SV call, so alignment accuracy per read matters more than at higher coverage.

For low coverage studies, minimap2 is the recommended primary aligner because its mapping quality scores and supplementary alignment output provide the information SV callers need to make confident calls from limited read support. The 2023 benchmark found that read-based strategy recall reached 77% on 5X coverage data [<a href="#ref-2">2</a>]. This recall level is achievable with minimap2, but it requires careful parameter selection and validation of low-support calls. The Galaxy Training Network provides tutorials on running alignment workflows that include parameter selection and output validation steps [<a href="#ref-3">3</a>].

High Coverage Studies at 20X and Above

At high coverage, the read order sensitivity problem becomes more pronounced. The 2024 PeerJ study found that pbsv was highly sensitive to the order of the input data, especially at the highest depths where over 70% of SV calls generated from pairs of differently ordered FASTQ files were in disagreement [<a href="#ref-6">6</a>]. This finding has direct implications for aligner benchmarking at high coverage. If you compare minimap2 and NGMLR on the same high-coverage dataset, differences in SV calls between aligners could be confounded by read order sensitivity instead of true alignment quality differences.

For high coverage studies, the recommendation is to use minimap2 as the primary aligner and to implement strict read order standardization across all samples. The read order sensitivity study suggested that researchers should standardize read order and document sorting parameters in their protocols [<a href="#ref-6">6</a>]. This practice is more important at high coverage because the increased read count amplifies the effect of order-dependent variability in the sorting and calling steps.

Computational Resource Planning

Runtime Budget Calculation

Before selecting an aligner, calculate the expected runtime for your dataset using the read count and average read length. Minimap2 processes long reads at a rate that makes whole-genome alignment feasible on standard compute servers. NGMLR requires substantially more time per read because of its sparse seeding and gap-tolerant extension. The exact runtime difference depends on your dataset and compute environment, so the calculation should be based on a pilot run instead of published benchmarks.

The practical approach is to run each aligner on a subset of 100,000 to 500,000 reads, measure the wall clock time, and extrapolate to the full dataset. Record the input read count, total bases aligned, and wall clock time for each aligner in a structured format. This pilot measurement provides a realistic runtime estimate for your specific data and compute environment.

Memory Allocation Planning

Memory usage differs substantially between the three aligners. Minimap2 has moderate memory requirements that scale with reference genome size and read length. NGMLR requires more memory because it keeps additional intermediate data structures for gap-tolerant extension. GraphMap has moderate memory requirements but is not maintained for modern data.

For population-scale studies, memory planning must account for the peak memory usage of the aligner plus the downstream SV caller. The 2024 Nature Communications study noted that assembly-based methods demand significantly more computational resources than alignment-based methods [<a href="#ref-5">5</a>]. Within alignment-based methods, minimap2 provides the most favorable memory profile for large cohorts.

Validation Strategy by Study Scale

Single Sample Validation

For single sample analyses, such as clinical case reports or targeted investigations, the validation strategy should include running both minimap2 and NGMLR, comparing SV calls, and manually curating discordant loci with Ribbon. The 2023 benchmark found that discordance between read-based and assembly-based strategies was largely caused by complex SVs and inversions [<a href="#ref-2">2</a>]. Similar discordance between aligners at specific loci indicates regions where alignment uncertainty is high and manual curation is needed.

Cohort-Level Validation

For cohort studies, validation should include a subset of samples where both aligners are run and compared. The concordance rate between aligners across this subset provides a measure of alignment robustness for the study. If concordance is high, the study can proceed with a single aligner for the full cohort. If concordance is low, the discordant loci should be characterized to determine whether they represent genuine alignment challenges or systematic biases in one aligner.

Population-Level Validation

For population studies, validation is necessarily limited to a subset of samples because running multiple aligners on the full cohort is computationally prohibitive. The validation subset should be selected to represent the range of coverage, read quality, and variant composition in the full cohort. The 2024 Nature Communications study provided guidelines across 31 criteria combinations to aid tool selection [<a href="#ref-5">5</a>]. These guidelines can inform the selection of validation samples and the interpretation of validation results.

Documentation Requirements for Aligner Decisions

The decision framework requires documentation of the rationale for aligner selection, the pilot comparison results, and the validation strategy. The read order sensitivity study highlighted implications for the replication of SV studies and the development of consistent SV calling protocols [<a href="#ref-6">6</a>]. Documentation should include the aligner version, parameter settings, reference genome version, read order, and sorting procedure for every analysis.

The Carpentries lessons on version control and reproducible research provide foundational practices for documenting analysis workflows [<a href="#ref-9">9</a>]. The Bioconductor project provides documentation on reproducible package management for R-based analysis workflows [<a href="#ref-11">11</a>]. These resources support the documentation practices needed for reproducible aligner selection and SV detection.

Escalation Criteria for Aligner Selection

Escalate to a bioinformatics specialist when the pilot comparison reveals SV call concordance below 70% between aligners for high-confidence calls, when breakpoint precision varies by more than 100 base pairs between aligners for the same SV, or when runtime or memory usage exceeds your compute environment capacity for the primary aligner. The 2023 benchmark found that up to 80% of SVs could be detected by both read-based and assembly-based strategies across different long-read datasets [<a href="#ref-2">2</a>]. Concordance below this level between aligners suggests systematic alignment differences that require expert investigation.

Escalate when read order sensitivity produces different SV calls in replicate runs of the same data. The 2024 PeerJ study found that the order of input data affected the SVs predicted by each caller [<a href="#ref-6">6</a>]. If replicate runs with different read orders produce substantially different SV calls, the pipeline requires standardization before results can be interpreted.

Integration with Assembly-Based Validation

The decision framework should include criteria for when to integrate assembly-based SV detection. The 2023 benchmark found that integrating SVs from read and assembly is suggested for general-purpose detection because of inconsistently detected complex SVs and inversions [<a href="#ref-2">2</a>]. Assembly-based validation is optional for applications with limited resources [<a href="#ref-2">2</a>].

For studies where comprehensive detection of insertions is critical, assembly-based validation should be considered. The 2024 Nature Communications study found that assembly-based tools excel in detecting large SVs, especially insertions, and exhibit robustness to coverage fluctuations [<a href="#ref-5">5</a>]. The 2023 benchmark found that read-based strategy recall reached 77% on 5X coverage data, while assembly-based strategy required 20X coverage to achieve similar performance [<a href="#ref-2">2</a>]. These findings support the use of assembly-based validation for studies with sufficient coverage and computational resources.

Practical Implementation Checklist

The following checklist summarizes the decision framework for aligner selection:

  1. Record the sequencing platform, chemistry version, base-calling model, and estimated error rate for your dataset
  2. Determine the study scale as small targeted panel, medium cohort, or population-scale
  3. Calculate the runtime and memory budget for your compute environment
  4. Run a pilot comparison on a subset of reads with minimap2 and NGMLR
  5. Measure runtime, memory, mapping statistics, and SV call concordance for each aligner
  6. Select the primary aligner based on study scale, coverage, and computational budget
  7. Standardize read order and sorting procedures across all samples
  8. Document aligner version, parameters, reference genome version, and read order
  9. Validate SV calls with Ribbon visualization at discordant loci
  10. Escalate to a bioinformatics specialist if concordance falls below 70% or breakpoint precision varies by more than 100 base pairs

This checklist provides a structured approach to aligner selection that accounts for study-specific factors beyond the published benchmarks. The 2024 Nature Communications study emphasized that no universally superior tool exists and provided guidelines across 31 criteria combinations [<a href="#ref-5">5</a>]. The decision framework translates these guidelines into practical steps that researchers can implement with their own data.

Frequently Asked Questions

Which aligner should I use for PacBio HiFi data?

Minimap2 with the HiFi preset is the recommended choice for PacBio HiFi data. HiFi reads have error rates below 1%, and minimap2 is optimized for this accuracy level. The 2024 Nature Communications benchmark included minimap2 among the upstream aligners and found that alignment-based methods are favored for computational efficiency and lower coverage requirements [<a href="#ref-5">5</a>]. NGMLR was designed for error-prone long reads and does not provide advantages for HiFi data.

Which aligner should I use for Oxford Nanopore data?

Minimap2 with the Nanopore preset is the recommended choice for modern Oxford Nanopore data. Current nanopore chemistry and base-calling models produce reads with accuracy approaching HiFi, and minimap2 handles these error profiles efficiently. GraphMap was designed for early nanopore data with higher error rates and is not maintained for current chemistry. The EMBL-EBI training materials recommend understanding your data type before selecting analysis tools, and minimap2 is the default in most community pipelines [<a href="#ref-7">7</a>][<a href="#ref-4">4</a>].

Does aligner choice affect SV breakpoint accuracy?

Yes. The 2023 benchmark found that variant type, size, and breakpoint detected by read-based strategy were greatly affected by aligners [<a href="#ref-2">2</a>]. Different aligners place reads differently at SV boundaries, producing breakpoint coordinates that can vary by tens or hundreds of base pairs. For clinical reporting, breakpoint precision matters, and you should validate breakpoints with visualization tools like Ribbon [<a href="#ref-8">8</a>].

How much does runtime differ between minimap2 and NGMLR?

Minimap2 is substantially faster than NGMLR because of its minimizer-based seeding approach. NGMLR uses sparse seeding with gap-tolerant extension, which requires more computation per read. The exact runtime difference depends on your dataset size, error rate, and compute environment. Record runtime for your specific data using the measurement approach described in this article.

Can I use GraphMap for modern nanopore data?

GraphMap is not recommended for modern nanopore data. It was designed for early nanopore chemistry with high error rates and has not been actively maintained. Modern nanopore data does not need the aggressive error tolerance that GraphMap provides, and the aligner lacks features that current SV callers expect. Use minimap2 for new projects.

How does read order affect SV calling?

The 2024 PeerJ study found that the order of input data affected the SVs predicted by each caller, with pbsv showing over 70% disagreement between differently ordered FASTQ files at high depth [<a href="#ref-6">6</a>]. The SAMtools alignment sorting algorithm was identified as a source of variability following read order randomization [<a href="#ref-6">6</a>]. Standardize read order across samples and document sorting parameters to ensure reproducibility.

Should I use assembly-based SV detection instead of alignment-based?

Assembly-based detection excels in detecting large SVs, especially insertions, and exhibits robustness to coverage fluctuations [<a href="#ref-5">5</a>]. Alignment-based detection demonstrates superior genotyping accuracy at low sequencing coverage of 5X to 10X and excels in detecting complex SVs like translocations, inversions, and duplications [<a href="#ref-5">5</a>]. The 2023 benchmark found that read-based strategy recall reached 77% on 5X coverage data, while assembly-based strategy required 20X coverage to achieve similar performance [<a href="#ref-2">2</a>]. Integrating SVs from read and assembly is suggested for general-purpose detection [<a href="#ref-2">2</a>].

How do I validate SV calls when aligners disagree?

Use Ribbon or similar visualization tools to inspect alignments at SV loci where aligners disagree. Ribbon shows how alignments are positioned within both the reference and read contexts, giving an intuitive view that enables better understanding of structural variants and the read evidence supporting them [<a href="#ref-8">8</a>]. If visualization confirms the SV is supported by reads from both aligners, the call is likely real. If only one aligner supports the call, investigate whether the other aligner is misplacing reads at that locus.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Expectations and blind spots for structural variation detection from long-read assemblies and short-read genome sequencing technologies.](https://pubmed.ncbi.nlm.nih.gov/33789087). American journal of human genetics, 2021. [2] [Comparison and benchmark of structural variants detected from long read and long-read assembly.](https://pubmed.ncbi.nlm.nih.gov/37200087). Briefings in bioinformatics, 2023. [3] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [4] [nf-core Documentation](https://nf-co.re/docs). nf-core. [5] [Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data.](https://pubmed.ncbi.nlm.nih.gov/38503752). Nature communications, 2024. [6] [The impact of FASTQ and alignment read order on structural variant calling from long-read sequencing data.](https://pubmed.ncbi.nlm.nih.gov/38500526). PeerJ, 2024. [7] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [8] [Ribbon: intuitive visualization for complex genomic variation.](https://pubmed.ncbi.nlm.nih.gov/32766814). Bioinformatics (Oxford, England), 2021. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [10] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [11] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.