Closing Gaps in Genome Assemblies: A Comparative Guide to Tools and Strategies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Read length is paramount for gap closure: Long reads (e.g., from PacBio or Oxford Nanopore) are significantly more effective than short reads for spanning repetitive regions and bridging gaps, directly providing contiguous sequence information. Raw long reads are often preferred over error-corrected reads, as error correction can inadvertently remove reads crucial for resolving difficult, error-prone segments.
- Gap size dictates strategy and tool selection: Small gaps (<1 kb) may be addressable with short-read extension methods, while medium gaps (1-100 kb) necessitate long reads. Large gaps (>10 kb), particularly in centromeric or telomeric regions, require ultra-long reads or specialized tools like the quarTeT GapFiller module.
- Gap closure is an iterative process: A single pass with one tool is insufficient; optimal results are achieved through multi-round strategies, potentially involving different tools, read sets, or parameter adjustments to address remaining gaps. Validation of closed gaps through read alignment consistency and junction quality is critical to ensure accuracy.
- Understanding gap origin is crucial for tool selection: Gaps arising from assembly limitations, repetitive sequences, or structural variations require different approaches. Repeat-derived gaps, for instance, demand tools adept at using flanking unique sequences for disambiguation, while structural gaps may benefit from hybrid sequencing strategies.
- Reproducibility hinges on meticulous documentation: Maintaining detailed records of gap inventories (before and after), tool versions, command-line parameters, input read data characteristics, and validation results is essential for reproducible workflows and transparent reporting of assembly quality.
Genome assemblies produced by both short-read and long-read pipelines routinely contain gaps that interrupt contiguity and obscure functional elements. These gaps concentrate in repetitive regions, segmental duplications, centromeric arrays, and other sequences that challenge graph construction and consensus calling. For researchers working with plant, animal, fungal, or microbial genomes, the practical question is which gap-closing tool to apply given the available read data, the size distribution of the gaps, and the quality of the initial assembly. This article compares LR_Gapcloser, TGS-GapCloser, GapFiller, and related strategies, with explicit guidance on read type selection, gap size thresholds, parameter choices, and validation steps. The focus is on reproducible workflow decisions that a laboratory bioinformatics team can implement and document.
The Gap Problem in Modern Genome Assembly
Gaps in genome assemblies represent regions where the assembler could not determine the intervening sequence with sufficient confidence. They appear as runs of N characters in FASTA files and are typically classified by their origin. Assembly gaps arise when read coverage drops below the threshold needed to bridge two adjacent contigs. Repeat-derived gaps occur when identical or near-identical sequences confuse the overlap graph, causing the assembler to collapse or fragment the region. Structural gaps reflect genuine difficulty in resolving large insertions, deletions, or rearrangements that differ between haplotypes.
The consequences of unresolved gaps extend beyond cosmetic incompleteness. Genes embedded in or near gap regions may be missing from annotation sets. Regulatory elements, telomeres, and centromeres are frequently under-represented in draft assemblies. Comparative genomics analyses that depend on complete gene models or syntenic blocks lose power when gaps interrupt the alignment. For clinical or agricultural applications, a gap that hides a disease-associated variant or a trait-linked locus can invalidate downstream conclusions.
The scale of the problem is substantial. Many reference-quality assemblies still contain hundreds or thousands of gaps. The human CHM1 genome, used as a benchmark in gap-closing studies, had a contig N50 of only 143 kb before gap closure, which improved to 19 Mb after applying long-read-based gap filling, a 132-fold increase in contiguity. Large repeat-rich genomes such as wheat show similar patterns, with gap closure increasing contig N50 by approximately 40 percent. These improvements change the analytical utility of the assembly.
The emergence of telomere-to-telomere assembly projects has raised expectations for gap-free genomes. These projects aim to resolve every chromosome end-to-end, including the highly repetitive centromeric and subtelomeric regions that were previously considered intractable. The quarTeT toolkit, developed for this purpose, includes a GapFiller module designed to close unclosed gaps using additional ultra-long sequences. Its successful application to the Actinidia chinensis genome produced an assembly comparable in quality to a manually curated reference, demonstrating that automated gap closure can approach the quality of expert human curation.
For most research groups, the goal is not necessarily a complete telomere-to-telomere assembly. The realistic objective is to close as many gaps as possible with the data on hand, to prioritize which gaps matter for the biological questions being asked, and to document the limitations of the final assembly honestly. This requires a systematic approach to gap identification, tool selection, parameter tuning, and validation.
Core Principles of Gap Closing
Gap closing is a local assembly problem. The input is a set of reads that span or partially overlap the gap region, and the output is a consensus sequence that replaces the N run with high-confidence bases. The difficulty lies in the fact that the reads available for gap closure are often the same reads that failed to assemble the region in the first place. The gap exists because the assembler could not resolve the region, so the gap-closing tool must use different algorithmic strategies or additional data to succeed.
The first principle is that read length matters more than read depth for most gap-closing applications. Long reads from third-generation sequencing platforms can span entire gap regions in a single molecule, providing direct evidence for the intervening sequence. Short reads, by contrast, must be extended stepwise from the gap flanks, which is error-prone in repetitive regions. The LR_Gapcloser tool was specifically designed to exploit long reads for this purpose, using a tiling path approach that selects reads based on their alignment to the gap flanks and then constructs a consensus from the selected reads.
The second principle is that raw reads often outperform error-corrected reads for gap closure. This counterintuitive finding, reported in the LR_Gapcloser study, reflects the fact that error correction can remove the very reads that span difficult regions. If a read contains errors in a repetitive segment, the error-correction step may discard it as low quality, even though it provides the only link across the gap. Gap-closing tools that work with raw reads preserve this information and can fill more gaps as a result.
The third principle is that gap size determines the appropriate strategy. Small gaps of a few hundred base pairs can often be closed with short-read extension methods. Medium gaps of several kilobases require long reads or linked reads. Large gaps of tens of kilobases or more, particularly those in centromeric or telomeric regions, may require ultra-long reads or targeted sequencing approaches. The quarTeT toolkit's GapFiller module, for example, relies on additional ultra-long sequences to fill gaps, reflecting the reality that standard long-read libraries may not provide sufficient span for the most difficult regions.
The fourth principle is that gap closure is iterative. A single pass with one tool will close some gaps but not others. Re-running the tool with different parameters, using a different read set, or applying a complementary tool can close additional gaps. The optimal workflow treats gap closure as a multi-round process, with each round addressing the gaps that remain after the previous round.
At a Glance: Tool Comparison for Gap Closing
The following table summarizes the key characteristics of commonly used gap-closing tools and the situations in which each is most appropriate. The comparisons are based on published performance data and the documented design goals of each tool.
| Tool | Input Read Type | Gap Size Range | Key Strength | Primary Limitation | Best Use Case |
|---|---|---|---|---|---|
| LR_Gapcloser | Long reads (raw or error-corrected) | 1 kb to 100 kb | Fast, memory-efficient, uses raw reads effectively | Requires long-read data, not designed for short-read input | Closing gaps in draft assemblies when long-read data is available |
| TGS-GapCloser | Long reads (third-generation sequencing) | 1 kb to 100 kb | Optimized for TGS data, handles repetitive regions | May require parameter tuning for different genomes | Gap closure in large, repeat-rich genomes |
| GapFiller (quarTeT module) | Ultra-long sequences | 10 kb to megabase scale | Designed for telomere-to-telomere assembly | Requires ultra-long read data, web-based tool | Closing large gaps in T2T assembly projects |
| GapFiller (short-read based) | Short reads (paired-end or mate-pair) | 100 bp to 5 kb | Works with standard Illumina data | Limited by read length in repetitive regions | Closing small gaps when only short-read data is available |
| Blackbird | Synthetic long reads plus low-coverage long reads | 50 bp to 10 kb | Hybrid approach, reduces long-read coverage requirement | Requires synthetic long-read technology | Structural variant detection and gap closing with limited long-read coverage |
The choice of tool depends on the data available, beyond on the tool's theoretical capabilities. A research group with only Illumina data cannot use LR_Gapcloser effectively, regardless of how well it performs with long reads. Conversely, a group with high-coverage long-read data should not limit itself to short-read gap-closing methods. The decision matrix in the table reflects these practical constraints.
Read Type Selection and Data Requirements
The type of sequencing data available is the single most important factor in gap-closing strategy. Each read type has distinct strengths and weaknesses for gap closure, and the choice of tool must align with the data.
Long Reads from Third-Generation Sequencing
Long reads from platforms such as Pacific Biosciences and Oxford Nanopore provide the most direct path to gap closure. A single long read can span a gap of several kilobases, providing contiguous sequence information that short reads cannot offer. The LR_Gapcloser study demonstrated that long reads can close gaps in assemblies produced by different methods, including de novo assemblies, repeat-derived gaps, and real gaps from reference genomes.
The key advantage of long reads is their ability to bridge repetitive regions. Short reads that fall entirely within a repeat cannot be uniquely placed, but a long read that extends from a unique flanking region through the repeat into the unique region on the other side provides unambiguous placement. This is why long-read-based gap closure succeeds where short-read methods fail.
The practical requirement is coverage. Gap closure needs reads that span the gap, which means the read length must exceed the gap size plus the flanking alignment length. For a 10 kb gap, reads of at least 15 to 20 kb are needed to provide comfortable flanking alignments. Coverage of 10 to 30 times is typically sufficient for gap closure, though higher coverage improves the accuracy of the consensus sequence.
Raw reads are preferred over error-corrected reads for gap closure. The LR_Gapcloser study found that using raw reads filled more gaps than using error-corrected reads. This is because error correction can discard reads that span difficult regions, particularly if the errors are concentrated in repetitive sequence. The error rate of raw long reads is acceptable for gap closure because the consensus step averages out random errors.
Short Reads from Next-Generation Sequencing
Short reads from platforms such as Illumina can close small gaps, but their utility decreases rapidly as gap size increases. Paired-end reads with insert sizes of 300 to 500 base pairs can bridge gaps of a few hundred base pairs. Mate-pair libraries with insert sizes of 2 to 10 kb can bridge larger gaps, but the preparation is more complex and the data quality is more variable.
The fundamental limitation of short reads is that they cannot span large repetitive regions. If a gap contains a repeat that is longer than the read length, the reads will not provide unique placement information. This is why short-read gap closure works well for small gaps in unique sequence but fails for large gaps or gaps in repetitive regions.
For research groups with only short-read data, the realistic goal is to close gaps of less than 5 kb. Tools such as GapFiller, which use paired-end and mate-pair information to extend contigs into gap regions, can achieve this. The quality of the closure depends on the insert size distribution of the library and the depth of coverage in the flanking regions.
Synthetic Long Reads and Hybrid Approaches
Synthetic long-read technologies, such as those that barcode short reads from the same long DNA molecule, offer a middle ground between short-read cost and long-read span. The Blackbird tool uses synthetic long reads together with low-coverage long reads to improve structural variant detection and assembly. Its sliding window approach uses barcode information to assemble small segments accurately, then uses long reads for gap closing and contig assembly.
The practical advantage of hybrid approaches is cost reduction. Blackbird achieves results comparable to state-of-the-art long-read tools using only 5 times coverage of long reads, compared to the 10 times coverage typically required by other tools. This can substantially reduce sequencing costs for large genomes.
The limitation is that synthetic long-read technologies are not available in all sequencing facilities, and the barcode information requires specialized library preparation. Research groups considering this approach should verify that their sequencing provider can produce the required data format.
Practical Workflow for Gap Closing
A systematic gap-closing workflow proceeds through several stages: gap identification, tool selection, parameter tuning, execution, validation, and iteration. Each stage has specific decisions that affect the final assembly quality.
Step 1: Identify and Characterize Gaps
The first step is to inventory the gaps in the assembly. This involves scanning the FASTA file for runs of N characters and recording their positions, lengths, and surrounding sequence context. Most assembly tools provide gap statistics in their output, but a dedicated analysis is more thorough.
For each gap, record the following information:
- Chromosome or contig identifier and position
- Gap length in base pairs
- GC content of the flanking 1 kb on each side
- Presence of known repeats in the flanking sequence
- Whether the gap is internal to a contig or at a contig end
This information guides tool selection. Gaps of less than 1 kb may be closable with short-read methods. Gaps of 1 to 10 kb require long reads. Gaps of more than 10 kb, particularly those in centromeric or telomeric regions, may require ultra-long reads or specialized tools.
Step 2: Select the Gap-Closing Tool
The tool selection depends on the read data available and the gap size distribution. The decision table in the At a Glance section provides a starting point, but the final choice should consider the specific characteristics of the assembly.
For long-read data, LR_Gapcloser is a strong default choice. It is fast, memory-efficient, and has been validated on human, wheat, and other complex genomes. Its tiling path approach selects reads that align to the gap flanks and constructs a consensus from the selected reads. The tool accepts both raw and error-corrected reads, with raw reads generally producing better results.
For ultra-long read data and telomere-to-telomere projects, the quarTeT GapFiller module is appropriate. It is designed to fill gaps using additional ultra-long sequences and can be combined with the other quarTeT modules for telomere and centromere identification.
For short-read data, GapFiller or similar tools that use paired-end and mate-pair information are the practical choice. These tools extend contigs into gap regions using the insert size distribution of the library to determine the expected distance between paired reads.
Step 3: Prepare the Input Data
The input data preparation depends on the tool. For LR_Gapcloser, the required inputs are the assembly FASTA file and the long-read data in FASTA or FASTQ format. The reads should be filtered to remove adapters and low-quality sequences, but error correction is optional and may reduce the number of gaps closed.
For the quarTeT GapFiller module, the inputs are the assembly and the ultra-long sequence data. The web-based interface accepts these inputs directly, and the tool runs on the quarTeT server.
For short-read gap fillers, the inputs are the assembly and the paired-end or mate-pair reads. The insert size distribution must be known or estimated, as the tool uses this information to determine the expected distance between paired reads.
Step 4: Run the Gap-Closing Tool
The execution parameters depend on the tool and the data. For LR_Gapcloser, the key parameters are the minimum read length, the minimum alignment identity, and the number of threads. The default parameters work well for most genomes, but adjusting the minimum read length can improve results for genomes with specific read length distributions.
For the quarTeT GapFiller module, the web interface provides a simple parameter set. The main decision is the minimum sequence length for gap filling, which should be set based on the gap size distribution.
For short-read gap fillers, the key parameters are the insert size, the insert size standard deviation, and the minimum overlap length. These parameters should be estimated from the library preparation information or from the alignment of reads to the assembly.
Step 5: Validate the Closed Gaps
Validation is the most important step in the workflow. A gap-closing tool can produce a sequence that is incorrect, particularly in repetitive regions. The validation process should check both the sequence accuracy and the placement accuracy.
The first validation is to re-align the reads used for gap closure to the new assembly. Reads that were previously unaligned or misaligned should now align to the closed gap region. The alignment should be checked for consistency, with no reads showing conflicting placements.
The second validation is to check the gap-flanking junctions. The closed sequence should connect smoothly to the flanking sequence, with no unexpected insertions or deletions. A multiple sequence alignment of the reads spanning the junction can reveal errors.
The third validation is to check for assembly artifacts. If the gap closure introduced a tandem duplication or a mis-assembly, the surrounding gene structure or syntenic relationships may be disrupted. Comparing the closed region to a related genome can reveal such artifacts.
Step 6: Iterate
Gap closure is rarely complete after a single pass. Some gaps will remain, either because no reads spanned them or because the reads were insufficient to produce a confident consensus. The remaining gaps should be re-analyzed to determine whether additional data or different parameters could close them.
The iteration strategy depends on the gap size distribution of the remaining gaps. If most remaining gaps are large, additional long-read sequencing may be needed. If most remaining gaps are small but in repetitive regions, a different tool or different parameters may help.
Records and Measurements for Gap Closure
Documenting the gap-closing process is essential for reproducibility and for honest reporting of assembly quality. The following records should be maintained for each gap-closing run.
Gap Inventory Before and After
The primary record is the gap inventory. Before gap closure, record the total number of gaps, the total number of N bases, the N50 gap size, and the distribution of gap sizes. After gap closure, record the same statistics. The difference between the two inventories shows the improvement achieved.
The LR_Gapcloser study provides a useful benchmark for expected improvements. In the human CHM1 genome, gap closure improved contig N50 from 143 kb to 19 Mb, a 132-fold increase. In the Triticum urartu genome, a large genome rich in repeats, contig N50 increased by 40 percent. These benchmarks help set realistic expectations for other genomes.
Tool Parameters and Versions
Record the exact tool version, the command line used, and all parameter values. This information is essential for reproducing the run and for troubleshooting if problems arise. The Bioconductor project and the Galaxy Training Network provide guidance on reproducible bioinformatics workflows, and the nf-core documentation describes community standards for pipeline configuration and usage.
Read Data Used
Record the read data used for gap closure, including the sequencing platform, the read length distribution, the coverage, and whether raw or error-corrected reads were used. This information is important for interpreting the results and for planning additional sequencing if needed.
Validation Results
Record the results of the validation steps, including the read alignment statistics, the junction quality, and any assembly artifacts detected. This information supports the quality claims made in the final assembly report.
Common Failure Patterns in Gap Closing
Understanding why gap closure fails is as important as understanding why it succeeds. The following failure patterns are common across tools and genomes.
Insufficient Read Span
The most common cause of gap-closure failure is that no read spans the gap. If the longest read in the dataset is shorter than the gap plus the required flanking alignment, the gap cannot be closed with the available data. This is a data limitation, not a tool limitation. The solution is to generate longer reads or to use a different sequencing platform.
Repetitive Sequence Confusion
Gaps in repetitive regions are difficult to close because the reads do not provide unique placement information. A read that aligns to a repeat in the gap region could come from any copy of the repeat in the genome. The gap-closing tool must use the flanking unique sequence to disambiguate the placement, but this is not always possible.
The centromere provides an extreme example. Centromeric regions are composed of highly repetitive alpha-satellite arrays that can extend for megabases. The recent characterization of 2,110 human centromeres from diverse individuals identified 226 centromere haplotypes and 1,870 alpha-satellite higher-order repeat variants, illustrating the complexity of these regions. Standard gap-closing tools are not designed for such regions, and specialized approaches are needed.
Error Correction Artifacts
Using error-corrected reads can reduce the number of gaps closed, as the LR_Gapcloser study demonstrated. The error-correction process can discard reads that span difficult regions, particularly if the errors are concentrated in repetitive sequence. If a gap-closing run with error-corrected reads produces poor results, re-running with raw reads may improve the outcome.
Parameter Mismatch
Gap-closing tools have parameters that must match the data. If the minimum read length is set too high, reads that could span a gap are excluded. If the minimum alignment identity is set too high, reads with sequencing errors are rejected. If the insert size is set incorrectly for short-read gap fillers, the expected distance between paired reads is wrong, and the extension process fails.
The solution is to check the parameter values against the actual data characteristics. The read length distribution, the error rate, and the insert size distribution should be measured from the data, not assumed from the library preparation protocol.
Assembly Errors in Flanking Regions
Gap closure assumes that the flanking sequence is correct. If the assembly has errors in the flanking regions, the gap-closing tool may not be able to align reads correctly, or it may produce a consensus that is inconsistent with the true sequence. This is particularly problematic in regions with segmental duplications or other complex structures.
The solution is to validate the flanking regions before attempting gap closure. If the flanking sequence is suspect, the gap closure will inherit the errors.
Quality Controls and Assembly Validation
Gap closure should be subject to the same quality controls as the initial assembly. The following checks are recommended for every gap-closing run.
Read Alignment Consistency
After gap closure, re-align the reads to the new assembly. The reads that were used for gap closure should align to the closed gap region with high identity and consistent placement. Reads that show conflicting placements or low identity indicate potential errors in the closed sequence.
Junction Quality
The junctions between the closed sequence and the flanking sequence should be checked for accuracy. A multiple sequence alignment of the reads spanning each junction can reveal errors. If the reads consistently support a different junction sequence than the one produced by the gap-closing tool, the closed sequence should be corrected.
Gene Structure Preservation
If the assembly includes annotated genes, check that the gap closure did not disrupt gene structures. A gap that interrupts a gene model should be closed with a sequence that restores the correct reading frame and exon structure. If the closed sequence introduces a frameshift or a premature stop codon, the closure is likely incorrect.
Syntenic Comparison
If a related genome is available, compare the closed region to the syntenic region in the related genome. Conservation of gene order and sequence identity supports the accuracy of the closure. Disruptions in synteny may indicate mis-assembly.
BUSCO and Other Completeness Metrics
Completeness metrics such as BUSCO provide a global assessment of assembly quality. Comparing BUSCO scores before and after gap closure can reveal whether the closure improved the completeness of the assembly. However, BUSCO scores are not sensitive to local errors, so they should be used in combination with the local validation checks described above.
Limitations of Gap-Closing Tools
Gap-closing tools have inherent limitations that should be acknowledged in any assembly report. These limitations do not negate the value of gap closure, but they define the boundaries of what can be achieved.
Sequence Accuracy in Repetitive Regions
The consensus sequence produced by gap-closing tools is only as accurate as the reads that support it. In repetitive regions, the reads may be misaligned, and the consensus may contain errors. The error rate in closed gaps is typically higher than in the rest of the assembly, particularly for long-read-based closure where the raw read error rate is higher.
Inability to Close All Gaps
No gap-closing tool can close every gap. Some gaps are simply too large, too repetitive, or too poorly covered to be resolved with the available data. The goal of gap closure is to close as many gaps as possible, not to achieve a gap-free assembly. The remaining gaps should be reported honestly, with an explanation of why they could not be closed.
Computational Resource Requirements
Some gap-closing tools require substantial computational resources. The LR_Gapcloser study emphasized the tool's speed and memory efficiency compared to other approaches, but even efficient tools require significant resources for large genomes. Research groups with limited computational infrastructure should plan accordingly.
Dependence on Input Data Quality
The quality of the gap-closing output depends on the quality of the input data. Low-coverage read data, short reads, or reads with high error rates will produce poorer results than high-coverage, long, accurate reads. The input data quality should be assessed before gap closure, and the limitations should be documented.
Safety and Reproducibility Context
Gap closure is a computational process, but it has implications for the safety and reproducibility of downstream analyses. The following considerations are relevant for research groups and laboratory professionals.
Reproducibility Standards
Reproducible gap closure requires documented workflows, versioned tools, and recorded parameters. The nf-core documentation describes community standards for pipeline configuration and usage, and the Galaxy Training Network provides accessible workflow training and analysis tutorials. The Carpentries lessons offer foundational computing and data skills that support reproducible analysis. Research groups should adopt these standards to ensure that their gap-closing results can be reproduced by others.
Data Management
The read data used for gap closure should be archived and made available with the assembly. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support data sharing and archival. The EMBL-EBI Training resources provide guidance on bioinformatics learning pathways and data-resource training. Depositing the assembly and the supporting read data in a public repository ensures that the gap-closing results can be verified and reused.
Version Control
The assembly, the gap-closing scripts, and the parameter files should be under version control. This allows the research group to track changes, revert to previous versions, and document the exact state of the analysis at each stage. The Carpentries lessons provide training in Git and other version-control tools.
Professional Escalation Criteria
When gap closure produces unexpected results, or when the remaining gaps are in biologically important regions, professional escalation may be appropriate. The following criteria indicate when to seek additional expertise:
- The gap-closing tool produces a sequence that conflicts with experimental data, such as PCR products or optical maps
- The remaining gaps are in regions known to be biologically important, such as disease-associated loci or trait-linked genes
- The gap-closing results are inconsistent between different tools or different parameter sets
- The assembly is intended for clinical or regulatory use, where accuracy standards are higher
In these cases, consultation with a bioinformatics specialist, a genome assembly expert, or the tool developers may be necessary.
A Decision Framework for Matching Gap-Closing Tools to Assembly Context
Selecting a gap-closing tool from a table of features is only the first step. The harder problem is matching the tool to the specific assembly context, which includes the assembly graph structure, the gap size distribution, the repeat content of the flanking regions, and the intended downstream use of the finished assembly. This section provides a practical decision framework that researchers can apply before committing computational resources to a gap-closing run.
Step 1: Classify the Assembly by Gap Origin
The first decision point is to classify the gaps in the assembly by their likely origin. This classification determines which tool is most likely to succeed. Assembly gaps, where coverage simply dropped below the threshold needed to bridge two contigs, respond well to any tool with sufficient read span. Repeat-derived gaps, where identical or near-identical sequences confused the graph construction, require tools that can use flanking unique sequence for placement. Structural gaps, which reflect genuine difficulty in resolving large insertions, deletions, or rearrangements, may require hybrid approaches or specialized tools.
To classify gaps, examine the flanking sequence context. Gaps flanked by unique sequence with moderate GC content are likely assembly gaps. Gaps flanked by low-complexity sequence, tandem repeats, or known repeat families are likely repeat-derived gaps. Gaps that appear at the ends of contigs in syntenic comparisons with related genomes are likely structural gaps. The classification can be recorded in the gap inventory table described in the Records and Measurements section.
Step 2: Assess the Read Span Capability
The second decision point is to assess whether the available read data can physically span the gaps. This is a simple calculation. For each gap, compare the gap length to the read length distribution of the available data. A read must span the entire gap plus provide sufficient flanking alignment on both sides. For a 1 kb gap, reads of at least 3 to 5 kb are needed. For a 10 kb gap, reads of at least 15 to 20 kb are needed. For a 100 kb gap, reads of at least 150 kb are needed, which requires ultra-long sequencing.
This assessment should be done before running any tool. If the read length distribution cannot span the gap, no tool will close it, and the researcher should either generate longer reads or accept the gap as unresolved. The LR_Gapcloser study demonstrated that the tool uses raw reads to fill more gaps than error-corrected reads, but even raw reads cannot span a gap longer than the read length.
Step 3: Evaluate the Repeat Content of Flanking Regions
The third decision point is to evaluate the repeat content of the flanking regions. This is critical because it determines whether the gap-closing tool can uniquely place the reads that span the gap. If the flanking regions contain repeats that are longer than the read length, the reads will not provide unique placement information, and the gap closure will be unreliable.
The centromere provides the most extreme example. Centromeric regions are composed of highly repetitive alpha-satellite arrays that can extend for megabases. The recent characterization of 2,110 human centromeres from diverse individuals identified 226 centromere haplotypes and 1,870 alpha-satellite higher-order repeat variants, illustrating the complexity of these regions. Standard gap-closing tools are not designed for such regions, and specialized approaches are needed. The quarTeT toolkit includes a CentroMiner module specifically for identifying centromeric regions, which can be used to flag gaps that require specialized treatment.
For gaps in less extreme repeat contexts, the decision is whether the flanking unique sequence is sufficient for read placement. A practical rule is to require at least 500 bp of unique sequence on each side of the gap for long-read-based closure. If the unique flanking sequence is shorter than this, the gap closure may produce a sequence that is placed incorrectly.
Step 4: Match the Tool to the Gap Size Distribution
The fourth decision point is to match the tool to the gap size distribution. The At a Glance table provides a starting point, but the final choice should consider the specific characteristics of the assembly.
For assemblies with mostly small gaps of less than 5 kb, short-read-based gap fillers such as GapFiller may be sufficient, particularly if the flanking regions are unique. For assemblies with gaps of 1 to 100 kb, LR_Gapcloser is a strong default choice for long-read data. For assemblies with gaps larger than 100 kb, particularly those in centromeric or telomeric regions, the quarTeT GapFiller module or specialized approaches are needed.
The gap size distribution should be plotted before selecting the tool. If the distribution is bimodal, with a cluster of small gaps and a cluster of large gaps, a two-tool strategy may be appropriate. Run the short-read tool first to close the small gaps, then run the long-read tool to close the remaining larger gaps.
Step 5: Consider the Downstream Use of the Assembly
The fifth decision point is to consider the downstream use of the assembly. This is often overlooked but is critical for determining the acceptable error rate in the closed gaps.
For assemblies intended for gene annotation, the closed gaps must preserve reading frames and exon structures. A gap closure that introduces a frameshift or a premature stop codon is worse than an unresolved gap, because it creates a false gene model. For assemblies intended for comparative genomics, the closed gaps must maintain syntenic relationships with related genomes. For assemblies intended for clinical or regulatory use, the accuracy standards are higher, and professional escalation may be appropriate if the gap closure produces unexpected results.
The downstream use also determines the validation stringency. An assembly for a model organism with extensive experimental resources can tolerate a higher error rate than an assembly for a clinical application. The validation steps described in the Quality Controls section should be adjusted accordingly.
Step 6: Plan the Iteration Strategy
The sixth decision point is to plan the iteration strategy before running the first tool. Gap closure is rarely complete after a single pass. Some gaps will remain, either because no reads spanned them or because the reads were insufficient to produce a confident consensus.
The iteration strategy should specify which tool to run first, which parameters to use, and which validation checks to perform after each round. The strategy should also specify the criteria for stopping. A common stopping criterion is to stop when the number of gaps closed in a round falls below a threshold, such as 5 percent of the remaining gaps. Another criterion is to stop when the remaining gaps are all larger than the read length distribution, indicating that additional rounds with the same data will not help.
The LR_Gapcloser study provides a useful benchmark for iteration planning. The tool closed a higher number of gaps faster and with a lower error rate than existing tools when tested on de novo assembled gaps, repeat-derived gaps, and real gaps. The contig N50 of the human CHM1 genome improved from 143 kb to 19 Mb, a 132-fold increase, after gap closure. This suggests that a single round of long-read-based gap closure can produce substantial improvements, but multiple rounds may be needed for the most difficult regions.
Step 7: Document the Decision Rationale
The seventh step is to document the decision rationale for each gap-closing run. This documentation should include the gap classification, the read span assessment, the repeat content evaluation, the tool selection, the parameter choices, and the validation results. This documentation serves two purposes. First, it supports reproducibility, allowing other researchers to understand why specific decisions were made. Second, it supports troubleshooting, allowing the research group to identify which decisions led to poor results and which decisions should be changed in a subsequent round.
The documentation should follow the reproducibility standards described in the Safety and Reproducibility Context section. The nf-core documentation describes community standards for pipeline configuration and usage, and the Galaxy Training Network provides accessible workflow training and analysis tutorials. The Carpentries lessons offer foundational computing and data skills that support reproducible analysis.
Applying the Framework to Common Scenarios
The framework can be applied to several common scenarios that research groups encounter.
Scenario 1: Short-read assembly with long-read data available. The gap classification will likely show a mix of assembly gaps and repeat-derived gaps. The read span assessment will show that the long reads can span most gaps. The repeat content evaluation will identify gaps that require specialized treatment. The tool selection should favor LR_Gapcloser, which was designed for this exact scenario. The downstream use will determine the validation stringency.
Scenario 2: Long-read assembly with gaps in centromeric regions. The gap classification will show that the remaining gaps are concentrated in centromeric regions. The read span assessment will show that standard long reads cannot span the largest gaps. The repeat content evaluation will show that the flanking regions are highly repetitive. The tool selection should favor the quarTeT GapFiller module, which is designed for telomere-to-telomere assembly, combined with the CentroMiner module to identify the centromeric regions. The downstream use will determine whether the remaining gaps are acceptable.
Scenario 3: Low-coverage long-read data with limited budget. The gap classification will show that the gaps are distributed across the genome. The read span assessment will show that the read length distribution can span most gaps, but the coverage is low. The repeat content evaluation will identify gaps that require higher coverage for confident closure. The tool selection should consider hybrid approaches such as Blackbird, which uses synthetic long reads together with low-coverage long reads to improve gap closing and contig assembly. The downstream use will determine whether the lower accuracy of low-coverage closure is acceptable.
Common Mistakes in Tool Selection
Several common mistakes lead to poor gap-closing results. The first is selecting a tool based on popularity instead of data compatibility. A tool that performs well with high-coverage long-read data will not perform well with low-coverage short-read data. The second mistake is failing to assess the read span capability before running the tool. This wastes computational resources on gaps that cannot be closed with the available data. The third mistake is ignoring the repeat content of the flanking regions. A gap in a repetitive region will not be closed reliably by a tool that cannot place reads uniquely. The fourth mistake is using error-corrected reads when raw reads would produce better results. The LR_Gapcloser study found that raw reads filled more gaps than error-corrected reads.
The fifth mistake is failing to validate the closed gaps. A gap-closing tool can produce a sequence that is incorrect, particularly in repetitive regions. The validation steps described in the Quality Controls section should be performed for every closed gap. The sixth mistake is stopping after a single round of gap closure. Multiple rounds with different tools or parameters can close additional gaps.
When to Escalate to Specialized Approaches
The decision framework includes criteria for when to escalate to specialized approaches. These criteria are based on the gap characteristics and the downstream use of the assembly.
Escalate when the remaining gaps are in centromeric or telomeric regions. These regions require specialized tools such as the quarTeT toolkit, which includes modules for telomere and centromere identification. The recent characterization of 2,110 human centromeres from diverse individuals demonstrated that specialized bioinformatic tools are needed to assemble and characterize these regions. Standard gap-closing tools are not designed for such regions.
Escalate when the gap closure produces results that conflict with experimental data. If PCR products, optical maps, or other experimental evidence contradict the closed sequence, the gap closure is likely incorrect, and manual curation or additional sequencing may be needed.
Escalate when the assembly is intended for clinical or regulatory use. The accuracy standards for such assemblies are higher, and the gap-closing results should be reviewed by a bioinformatics specialist or a genome assembly expert.
Escalate when the remaining gaps are in biologically important regions, such as disease-associated loci or trait-linked genes. The cost of an incorrect gap closure in such regions is high, and specialized approaches may be warranted.
Frequently Asked Questions
What is the difference between assembly gaps and sequencing gaps?
Assembly gaps are regions of N characters in the assembled sequence where the assembler could not determine the intervening bases. Sequencing gaps are regions where no sequencing data was generated, either because the library preparation failed to cover the region or because the sequencing platform could not read through it. Assembly gaps can sometimes be closed with additional analysis of existing data, while sequencing gaps require additional sequencing.
How many gaps can I expect to close with long-read gap closure?
The number of gaps closed depends on the read length, the coverage, and the gap size distribution. In the LR_Gapcloser study, the tool closed a higher number of gaps than existing tools when tested on de novo assembled gaps, repeat-derived gaps, and real gaps. The human CHM1 genome improved from a contig N50 of 143 kb to 19 Mb, and the Triticum urartu genome improved by 40 percent. These results suggest that most gaps of 1 to 100 kb can be closed when adequate long-read data is available.
Should I use raw reads or error-corrected reads for gap closure?
Raw reads generally produce better gap-closure results than error-corrected reads. The LR_Gapcloser study found that using raw reads filled more gaps than using error-corrected reads. Error correction can discard reads that span difficult regions, particularly if the errors are concentrated in repetitive sequence. If you have both raw and error-corrected reads, try both and compare the results.
What is the minimum read length needed for gap closure?
The minimum read length depends on the gap size. A read must span the entire gap plus provide sufficient flanking alignment on both sides. For a 1 kb gap, reads of at least 3 to 5 kb are needed. For a 10 kb gap, reads of at least 15 to 20 kb are needed. The read length distribution of your dataset determines which gaps can be closed.
Can I close gaps in a genome assembled from short reads using long-read data?
Yes. Long-read data can be used to close gaps in assemblies produced by any method, including short-read assemblies. The LR_Gapcloser study demonstrated that the tool is applicable to gaps in assemblies by different approaches and from large and complex genomes. The long reads provide the span needed to bridge the gaps that short reads could not resolve.
How do I know if a closed gap is correct?
Validation is essential. Re-align the reads used for gap closure to the new assembly and check for consistent placement and high identity. Check the junctions between the closed sequence and the flanking sequence. Compare the closed region to a related genome if one is available. If the closed sequence disrupts gene structures or syntenic relationships, it is likely incorrect.
What should I do if gaps remain after gap closure?
Analyze the remaining gaps to determine why they could not be closed. If the gaps are larger than the longest reads, additional long-read sequencing with longer reads may be needed. If the gaps are in repetitive regions, specialized tools or manual curation may be required. If the gaps are in biologically important regions, consider targeted sequencing or PCR-based closure.
Are there tools that can close gaps without additional sequencing?
Some gaps can be closed with existing data by using different tools or parameters. The quarTeT toolkit includes a GapFiller module that uses additional ultra-long sequences, but if such sequences are not available, the module cannot be used. The Blackbird tool uses synthetic long reads and low-coverage long reads, which may be available from existing datasets. However, gaps that lack spanning reads cannot be closed without additional sequencing.
Related Bioinformatics Guides
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Evaluating Genome Assembly Quality: Metrics and Tools
- Metagenome Co-Assembly: Strategies for Multi-Sample Data
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
- Proteomics Analysis Tools: A Comparative Guide for Functional Interpretation
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- quarTeT: a telomere-to-telomere toolkit for gap-free genome assembly and centromeric repeat identification.. Horticulture research, 2023.
- LR_Gapcloser: a tiling path-based gap closer that uses long reads to complete genome assembly.. GigaScience, 2019.
- A global view of human centromere variation and evolution.. Nature, 2026.
- A global view of human centromere variation and evolution.. bioRxiv : the preprint server for biology, 2025.
- Blackbird: structural variant detection using synthetic and low-coverage long-reads.. bioRxiv : the preprint server for biology, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.