Why Did My Assembly Collapse Repeats? Troubleshooting Misjoins and Underrepresented Repeat Regions
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Genome assembly collapse occurs when repetitive regions are represented as a single copy, leading to underestimated gene copy numbers and missing paralogs. This is often caused by read lengths shorter than repeat units, insufficient sequencing depth, or high heterozygosity confusing overlap graphs.
- Diagnostic signatures include abnormal read depth at repeat loci, where collapsed regions exhibit higher coverage than flanking unique sequences. Analyzing the ratio of observed depth to median genome depth provides a quantitative estimate of collapsed copies.
- Assembler parameter choices significantly influence repeat resolution; adjusting coverage cutoffs, increasing overlap identity thresholds, or switching to haplotype-aware assemblers can mitigate collapse. Overly aggressive repeat masking prior to assembly can also lead to misjoins and underrepresentation.
- Long-read sequencing platforms (e.g., Oxford Nanopore, Pacific Biosciences) are generally superior for resolving complex repeat structures due to their ability to span longer repeat units, whereas short-read platforms may struggle with repeats exceeding read length.
- Hybrid assembly approaches, combining short and long reads, or utilizing scaffolding data like linked reads or optical maps, offer enhanced contiguity and repeat resolution by leveraging complementary information.
- Inspecting the assembly graph is crucial for understanding collapse, revealing nodes with high depth and multiple connections that indicate dispersed repeats, or linear structures representing collapsed tandem arrays.
Genome assembly collapse occurs when repetitive regions that exist in multiple copies within a genome are represented as a single copy in the final assembly, or when reads originating from distinct repeat copies are incorrectly merged into one consensus sequence. This problem manifests as underestimated gene copy numbers, missing paralogs, truncated tandem arrays, and apparent absence of transposable element families that are known to be present from other evidence. The root causes are typically insufficient read length relative to repeat unit size, inadequate sequencing depth across repetitive loci, high heterozygosity that confuses read overlap graphs, or assembler parameter choices that favor conservative merging. This article provides a diagnostic framework for identifying why repeats collapsed in your assembly, practical steps for quantifying the extent of collapse using read depth analysis, and criteria for selecting alternative assembly strategies that preserve repeat copies.
Scope and Reader Context
This troubleshooting guide is written for researchers who have produced a draft genome assembly and suspect that repetitive regions are underrepresented. The intended reader is a biology student, research scientist, laboratory professional, or life-science practitioner who has basic familiarity with sequencing platforms and has encountered a specific problem: the assembly shows fewer copies of a gene family than expected from PCR, Southern blot, or comparative genomics evidence, or a repeat family that should be present is entirely missing. The guidance assumes you have access to raw sequencing reads, the assembly file, and basic command-line skills. The diagnostic steps described here use read depth analysis, assembly graph inspection, and re-assembly with alternative tools. The scope covers both short-read and long-read assembly scenarios, with emphasis on practical decisions you can make with the data you already have.
The outcome you should expect from working through this material is a clear diagnosis of whether your assembly collapsed repeats, identification of the most likely cause given your sequencing strategy, and a concrete plan for either rescuing the collapsed regions through targeted analysis or generating a new assembly with parameters and tools that preserve repeat copies. The article does not provide a single universal solution because the correct approach depends on the repeat architecture of your organism and the sequencing data available. Instead, it gives you a decision framework grounded in assembly theory and practical bioinformatics practice.
At a Glance
The table below summarizes the most common causes of repeat collapse, the diagnostic signatures you can observe in your data, and the first-line corrective action for each scenario.
| Cause of Collapse | Diagnostic Signature | First-Line Corrective Action |
|---|---|---|
| Read length shorter than repeat unit | Reads cannot span the full repeat, depth at repeat loci is roughly half the flanking unique depth | Switch to long-read sequencing or use an assembler that exploits linked reads or optical maps |
| Insufficient sequencing depth | Overall low coverage with high variance, repeat copies have near-zero depth while unique regions have expected depth | Increase sequencing depth, especially for libraries with high molecular weight DNA |
| High heterozygosity between repeat copies | Repeat copies differ by SNPs or indels, assembler merges them into a single mosaic consensus | Use an assembler with haplotype-aware mode or increase the overlap identity threshold |
| Overly aggressive repeat masking | RepeatMasker or similar masking removed reads before assembly | Re-assemble without masking or use a masking strategy that preserves reads for graph construction |
| Assembler parameter bias | Assembly graph shows low connectivity at repeat nodes, unitig depth is bimodal | Adjust coverage cutoff parameters or switch to a different assembler family |
Understanding Repeat Collapse in Genome Assembly
What Collapse Means at the Sequence Level
When an assembler collapses repeats, it produces a single contig or scaffold where the true genome has multiple similar or identical copies. The collapsed region may be a consensus of all copies, a mosaic of segments from different copies, or a copy that is identical to one true copy with the others missing entirely. The biological consequence is that any analysis relying on copy number, such as gene family counting, transposable element annotation, or telomere length estimation, will produce incorrect results. For example, if a genome contains a tandem array of 20 ribosomal RNA genes and the assembler collapses them into 3 copies, the assembly will suggest a much smaller rDNA array than the true genome. Similarly, if a transposable element family has 500 copies dispersed across the genome and the assembler merges them into 10 consensus loci, the assembly will severely underestimate the TE content. The NCBI Data Resources provide reference genomes and annotation databases that can help you compare your assembly against closely related species to identify suspicious reductions in repeat content.
The impact of collapse extends beyond simple copy number errors. When repeats are collapsed, the flanking unique sequence may also be affected because the assembler must choose which genomic context to attach to the collapsed node. This can lead to misjoins where sequence from two different genomic locations is concatenated into a single contig. The resulting chimeric contigs can produce false gene fusions, incorrect gene order, and misleading synteny analyses. Understanding the distinction between collapse and misassembly is important because the corrective actions differ. Collapse requires strategies that provide more information about copy number, while misassembly requires strategies that improve the graph structure and resolve ambiguous junctions.
Why Assemblers Merge Repeat Copies
Assemblers work by finding overlaps between reads and building a graph that represents the genome. When two reads come from different copies of a repeat that are identical or nearly identical, the assembler cannot distinguish them and treats them as coming from the same genomic location. This is the fundamental reason for collapse. The problem is exacerbated when the repeat copies are longer than the reads, because no single read can span from unique sequence on one side of the repeat to unique sequence on the other side. In that case, the assembler has no information to resolve the copy number and will typically merge all copies into one path through the graph. The Galaxy Training Network offers accessible tutorials on genome assembly that explain the relationship between read length, overlap detection, and repeat resolution in practical terms.
The graph-based nature of modern assemblers means that repeat resolution depends on the connectivity of the assembly graph. In a de Bruijn graph, repeats appear as nodes with high k-mer coverage. The assembler must decide whether a high-coverage node represents a single genomic region that was sequenced deeply or multiple genomic regions that share the same sequence. This decision is made based on the graph topology and the coverage distribution. If the graph shows a simple linear path through the high-coverage node, the assembler will collapse the repeats. If the graph shows multiple incoming and outgoing edges that connect to different flanking sequences, the assembler may be able to resolve the copy number. The quality of this resolution depends on the read length and the sequencing depth.
The Role of Read Length and Sequencing Platform
Read length is the single most important factor determining whether repeats can be resolved. If the repeat unit is shorter than the read length, a single read can span the entire repeat and the flanking unique sequence, allowing the assembler to place each copy correctly. If the repeat unit is longer than the read length, the assembler must rely on the graph structure to resolve copy number, which is error-prone. Short-read platforms such as Illumina produce reads of 150 to 300 base pairs, which are sufficient for resolving short tandem repeats and small gene families but inadequate for large segmental duplications or long transposable elements. Long-read platforms such as Oxford Nanopore and Pacific Biosciences produce reads of 10 to 100 kilobases, which can span most repeat units found in eukaryotic genomes. The EMBL-EBI Training portal provides learning pathways on sequencing technologies and their applications to genome assembly that can help you understand the tradeoffs between platforms.
The relationship between read length and repeat resolution is not simply a matter of read length exceeding repeat length. The read must also contain sufficient unique sequence on at least one side of the repeat to anchor it to a specific genomic location. If the read spans the entire repeat but the flanking sequence is also repetitive, the read may still be ambiguous. This is why telomeric and subtelomeric regions are particularly difficult to assemble. The TARPON pipeline for telomere analysis demonstrates the importance of specialized approaches for repetitive regions that are refractory to standard assembly methods. The pipeline uses capture probes and sliding-window analysis to identify full-length telomeres, highlighting the need for targeted strategies when standard assembly fails to represent repetitive regions accurately.
Core Principles of Repeat Resolution
Depth of Coverage as a Copy Number Signal
The most direct way to detect collapsed repeats is to examine sequencing depth across the assembly. In a correctly assembled genome, the depth of coverage should be approximately uniform across all regions, with some variation due to GC bias and other sequencing artifacts. If a region is collapsed, the depth at that region will be higher than the flanking unique regions, because reads from all copies of the repeat map to the single collapsed copy. The ratio of observed depth to expected depth gives an estimate of the number of collapsed copies. For example, if the genome-wide average depth is 30x and a particular contig region has 90x depth, that region likely represents three collapsed copies. This depth-based approach is the standard first diagnostic for repeat collapse and requires only the assembly and the aligned reads.
The depth signal is most reliable when the sequencing coverage is uniform and the repeat copies are identical or nearly identical. When the repeat copies have diverged, reads from different copies may not map to the collapsed consensus with equal efficiency. Reads with more mismatches may be clipped or mapped with lower quality, reducing the observed depth at the collapsed region. This can lead to an underestimate of the copy number. You should therefore interpret depth ratios as minimum estimates when the repeat copies are known to be divergent. The Bioconductor project hosts packages for analyzing genomic variation and copy number that can help you distinguish between true biological variation and assembly artifacts.
Heterozygosity and Haplotype Divergence
When the copies of a repeat are not identical but differ by single nucleotide polymorphisms or small indels, the assembler faces a choice. It can either keep the copies separate, producing multiple contigs that represent the different haplotypes, or it can merge them into a single mosaic sequence. The decision depends on the assembler's overlap threshold and the degree of divergence between copies. Highly divergent copies are more likely to be kept separate, while nearly identical copies are more likely to be merged. In diploid or polyploid organisms, the two alleles at a locus may also be collapsed into a single haploid representation, which is a form of collapse that affects all regions, beyond repeats. The Bioconductor project hosts packages for analyzing genomic variation and copy number that can help you distinguish between true biological variation and assembly artifacts.
Heterozygosity creates a particular challenge for repeat resolution because the assembler must distinguish between allelic variation and paralogous variation. Allelic variation occurs between the two copies of a chromosome in a diploid organism, while paralogous variation occurs between different loci that share a common ancestor. Both types of variation can cause the assembler to either merge or separate sequences, depending on the degree of divergence. In highly heterozygous organisms, the assembler may produce a mosaic assembly that combines alleles from both haplotypes, leading to an overestimate of the number of repeat copies. In contrast, in organisms with low heterozygosity, the assembler may collapse paralogs that differ by only a few variants. The appropriate assembly strategy depends on the heterozygosity level of your organism and the biological question you are addressing.
Graph Topology and Repeat Boundaries
Assembly graphs contain nodes that represent sequence segments and edges that represent overlaps or adjacencies. Repeats appear in the graph as nodes with high depth or as structures where multiple paths converge and diverge. A repeat that is longer than the reads creates a node that is traversed by all reads from all copies, resulting in a node with depth equal to the sum of depths of all copies. The graph may show a "bubble" structure where the repeat node is flanked by unique sequence on both sides, and the number of paths through the bubble indicates the copy number. Inspecting the assembly graph is a powerful way to understand why collapse occurred and whether the reads contain enough information to resolve the repeats. The nf-core Documentation describes community pipelines for genome assembly that include graph-based quality assessment steps.
The graph structure also reveals the difference between tandem and dispersed repeats. Tandem repeats appear as a series of identical or similar nodes connected in a linear chain, while dispersed repeats appear as a single node with multiple connections to different flanking sequences. The resolution strategy differs for these two types. Tandem repeats can be resolved by estimating the number of copies from the depth of the repeat node and the length of the array. Dispersed repeats require information about the flanking sequence to place each copy in its correct genomic context. The TARPON pipeline for telomere analysis provides an example of how specialized graph analysis can resolve repetitive regions that are refractory to standard assembly methods.
Practical Workflow for Diagnosing Collapse
Step 1: Align Reads Back to the Assembly
The first step is to align your raw sequencing reads back to the assembly using a read aligner appropriate for your data type. For short reads, use a splice-aware or unspliced aligner depending on whether you are assembling a genome with introns. For long reads, use a long-read aligner that can handle high error rates. The alignment should be done with settings that report all secondary alignments, because reads from collapsed repeats may map to multiple locations. After alignment, compute the depth of coverage in sliding windows across each contig. The Galaxy Training Network provides tutorials on read alignment and depth calculation that can be adapted to this purpose.
The choice of alignment parameters is critical for accurate depth estimation. If you use stringent mapping criteria that discard multi-mapping reads, you will underestimate the depth at collapsed repeats because reads from different copies will be counted only once. If you use permissive mapping criteria that allow reads to map to multiple locations, you may overestimate the depth at unique regions because reads from repeats will map there as well. A balanced approach is to report all alignments with a minimum mapping quality threshold and then compute depth as the sum of all alignments at each position. This approach captures the total read signal at each locus while filtering out low-quality alignments.
Step 2: Identify Regions with Abnormal Depth
Once you have depth values for each window, identify regions where the depth is significantly higher than the genome-wide median. A common threshold is two times the median depth, but the appropriate threshold depends on the variance in your data. Regions with high depth are candidates for collapsed repeats. You should also look for regions with zero or near-zero depth, which may indicate that the assembly contains sequence that is not supported by reads, or that reads from a repeat family were excluded during assembly. The NCBI Data Resources provide tools for visualizing depth of coverage in the context of annotated genomes, which can help you interpret the depth profile.
The depth distribution across the assembly provides additional diagnostic information. In a correctly assembled genome, the depth distribution should be approximately normal with a single peak at the genome-wide median. If there are secondary peaks at higher depth, these indicate regions with higher copy number, which may be collapsed repeats or genuine amplifications. The presence of a long tail of high-depth regions suggests widespread collapse, while the presence of a single high-depth peak suggests a specific repeat family that was collapsed. You should examine the genomic context of the high-depth regions to determine whether they correspond to known repeat families or gene families.
Step 3: Compare Depth to Expected Copy Number
For each region with abnormal depth, estimate the number of collapsed copies by dividing the observed depth by the median genome depth. This gives you a copy number estimate that you can compare against external evidence. For example, if you know from quantitative PCR that a gene is present in 10 copies in the genome, and the depth analysis suggests 3 copies, you have strong evidence of collapse. If you have no external evidence, you can use the depth ratio as a provisional copy number estimate and validate it with targeted experiments or with a different assembly approach.
The comparison between depth-based copy number and external evidence should account for the limitations of both methods. Quantitative PCR can underestimate copy number if the primers do not amplify all copies equally, and it can overestimate copy number if there are pseudogenes or partial copies that are not counted in the assembly. Southern blot can provide a more accurate estimate of copy number for tandem arrays, but it requires careful calibration and is not suitable for all repeat types. The EMBL-EBI Training portal provides resources on experimental validation of genome assemblies that can help you design appropriate validation experiments.
Step 4: Inspect the Assembly Graph
If you have access to the assembly graph from the assembler you used, inspect the graph around the collapsed regions. Look for nodes with high depth that are connected to multiple flanking nodes. The number of connections may indicate the number of distinct genomic contexts in which the repeat appears. If the graph shows a single node with many connections, the repeat is likely dispersed and the assembler collapsed all copies. If the graph shows a tandem array structure with multiple copies of the same node in a row, the assembler may have collapsed the array to a shorter length. The nf-core Documentation describes pipelines that produce assembly graphs as part of their output, and the Bioconductor project has packages for graph analysis.
Graph inspection requires familiarity with the specific graph format produced by your assembler. Some assemblers produce GFA (Graphical Fragment Assembly) files that can be visualized with dedicated tools, while others produce custom graph formats that require custom scripts to parse. The Galaxy Training Network provides tutorials on assembly graph visualization that can help you interpret the graph structure. The key features to look for are high-depth nodes, nodes with multiple connections, and structures that suggest tandem arrays or dispersed repeats.
Step 5: Test Alternative Assembly Parameters
If the depth analysis confirms collapse, the next step is to re-assemble with modified parameters. The specific parameters depend on the assembler you are using, but common adjustments include increasing the minimum overlap length, increasing the overlap identity threshold, disabling repeat masking, and adjusting the coverage cutoff. For short-read assemblers, increasing the k-mer size can help resolve repeats that are shorter than the k-mer. For long-read assemblers, adjusting the minimum read length filter can retain reads that span repeats. The EMBL-EBI Training portal offers courses on assembly parameter optimization that provide practical guidance.
The parameter space for genome assembly is large, and testing all combinations is not feasible. A practical approach is to start with the default parameters and then test one parameter at a time, measuring the effect on the depth profile and the assembly statistics. The nf-core Documentation describes community pipelines that include parameter sweeps and quality assessment steps that can automate this process. You should document the parameters used for each assembly attempt and the resulting quality metrics so that you can compare the results systematically.
Options and Tradeoffs in Assembly Strategies
Short-Read Assemblers
Short-read assemblers such as SPAdes, ABySS, and SOAPdenovo2 build de Bruijn graphs from k-mers. The k-mer size is a critical parameter that determines the minimum repeat length that can be resolved. If the k-mer is shorter than the repeat unit, the repeat appears as a single node in the graph and all copies are collapsed. Increasing the k-mer size can resolve longer repeats, but it also increases the memory requirement and may reduce the contiguity of the assembly if the sequencing depth is insufficient. Short-read assemblies are generally less effective at resolving repeats than long-read assemblies, but they are more cost-effective for large genomes. The Galaxy Training Network provides tutorials on short-read assembly that explain the relationship between k-mer size and repeat resolution.
The choice of k-mer size involves a tradeoff between repeat resolution and assembly continuity. Larger k-mers provide more specific matches that can distinguish between repeat copies, but they also require higher sequencing depth to achieve the same coverage of the k-mer space. If the sequencing depth is insufficient, larger k-mers will produce a fragmented assembly with many small contigs. Smaller k-mers produce a more contiguous assembly but cannot resolve repeats that are longer than the k-mer. A common strategy is to assemble with multiple k-mer sizes and then merge the results, but this approach can introduce new errors if the merging is not done carefully.
Long-Read Assemblers
Long-read assemblers such as Flye, Canu, and HiCanu use overlap-layout-consensus algorithms that can exploit the full length of long reads. These assemblers are better suited for resolving repeats because a single read can span the entire repeat unit and the flanking unique sequence. However, long-read assemblers have their own failure modes. They may collapse repeats that are longer than the reads, and they may produce chimeric contigs if the error rate is high. The choice of assembler and parameters depends on the read length distribution and the error profile of your sequencing platform. The nf-core Documentation describes long-read assembly pipelines that include quality control steps for detecting collapsed regions.
Long-read assemblers are particularly effective for resolving transposable elements and other dispersed repeats because the long reads can span the entire element and the flanking unique sequence. The study of transposable elements in parasitic wasps demonstrates the importance of accurate repeat representation for comparative genomics. The study found that TE abundance and diversity were highly variable across Braconidae species, and this variability would be obscured if the assemblies collapsed TE copies. Accurate repeat resolution is therefore essential for studies that compare repeat content across species.
Hybrid Assembly Approaches
Hybrid assembly combines short reads and long reads to leverage the advantages of both. The short reads provide high accuracy for base calling, while the long reads provide the contiguity needed to resolve repeats. Hybrid assemblers such as MaSuRCA and Unicycler use the short reads to build an initial assembly and then use the long reads to scaffold and fill gaps. Hybrid approaches are often the most effective for resolving repeats in genomes that have a mix of short and long repeat units. The tradeoff is increased computational cost and complexity. The EMBL-EBI Training portal provides resources on hybrid assembly strategies.
The success of hybrid assembly depends on the quality and quantity of both data types. The short reads must provide sufficient depth to correct errors in the long reads, and the long reads must provide sufficient coverage to span the repeats. If either data type is inadequate, the hybrid assembly may not resolve the repeats any better than a single-platform assembly. You should therefore assess the quality of both data types before embarking on a hybrid assembly project.
Scaffolding and Linkage Information
If you have access to additional data such as linked reads, Hi-C, or optical maps, you can use these to resolve repeats that are ambiguous in the sequence data alone. Linked reads from 10x Genomics or similar platforms provide long-range information by tagging reads that originate from the same DNA molecule. Hi-C data provides information about the spatial proximity of genomic regions, which can be used to order and orient contigs and to identify misjoins. Optical maps provide restriction enzyme cut site patterns that can be used to validate the assembly structure. These approaches are particularly useful for resolving large segmental duplications and other complex repeat architectures. The NCBI Data Resources provide access to databases of genomic variation and structural variation that can help you interpret the results of these analyses.
The integration of scaffolding data requires specialized tools and careful quality control. The scaffolding process can introduce new errors if the linkage information is misinterpreted, particularly in regions with complex repeat structure. You should validate the scaffolded assembly by checking the depth profile and the assembly graph after scaffolding. The nf-core Documentation describes pipelines that include scaffolding and validation steps that can help you assess the quality of the final assembly.
Observations and Measurements for Collapse Detection
Depth Ratio as a Quantitative Measure
The depth ratio, defined as the observed depth at a candidate collapsed region divided by the genome-wide median depth, is the primary quantitative measure for detecting collapse. A depth ratio of 2 suggests two collapsed copies, a ratio of 3 suggests three copies, and so on. However, the depth ratio can be affected by GC bias, which causes regions with extreme GC content to have lower depth, and by copy number variation that is real biological variation instead of assembly artifact. You should therefore interpret depth ratios in the context of the overall depth distribution and validate candidate collapsed regions with independent evidence.
The precision of the depth ratio estimate depends on the sequencing depth and the length of the collapsed region. For short regions, the depth estimate will have high variance because the number of reads mapping to the region is small. For long regions, the depth estimate will be more precise. You should calculate confidence intervals for the depth ratio and consider the length of the region when interpreting the results. The Bioconductor project provides packages for calculating depth statistics and confidence intervals.
Read Pair and Read Span Information
For short-read data, the insert size of the library provides additional information about repeat structure. If the insert size is larger than the repeat unit, read pairs can span the repeat and provide evidence for the number of copies. If the insert size is smaller than the repeat unit, read pairs will map within the repeat and provide no information about copy number. For long-read data, the read length itself provides the spanning information. You can measure the fraction of reads that span a candidate collapsed region and use this to assess whether the data contain sufficient information to resolve the repeat.
The analysis of read pairs and read spans requires careful attention to the library preparation and the alignment settings. Reads that span a repeat will have one end in the repeat and the other end in the flanking unique sequence. If the repeat is collapsed, the flanking sequence on both sides will be present in the assembly, and the read pairs will map to the collapsed region with the correct orientation and insert size. If the repeat is not collapsed, the read pairs will map to different copies of the repeat, and the insert size will appear to be larger than expected. The Galaxy Training Network provides tutorials on analyzing read pairs and insert sizes that can help you interpret this information.
Assembly Statistics and Completeness Metrics
Standard assembly statistics such as N50, L50, and total assembly size can provide indirect evidence of collapse. If the assembly is significantly smaller than the expected genome size based on flow cytometry or other estimates, collapse is a likely cause. Completeness metrics such as BUSCO scores can identify missing conserved genes, but they do not detect collapse of repetitive regions because BUSCO genes are typically single-copy and located in unique regions. The NCBI Data Resources provide reference genomes and annotation that can be used to estimate expected genome size and repeat content for comparison.
The comparison of assembly size to expected genome size is complicated by the fact that some genomes have large amounts of repetitive DNA that may be difficult to assemble even with the best strategies. A small assembly size does not necessarily indicate collapse if the genome is naturally compact. You should therefore compare the assembly size to the expected size based on multiple lines of evidence, including flow cytometry, k-mer analysis, and comparison to closely related species. The k-mer analysis is particularly useful because it provides an estimate of the genome size and the repeat content that is independent of the assembly.
Records and Documentation for Assembly Troubleshooting
Maintaining an Assembly Log
When you are troubleshooting a collapsed assembly, it is essential to keep a detailed log of the assembly parameters, the input data, and the quality metrics for each assembly attempt. The log should include the assembler version, the exact command line used, the parameter values, the sequencing platform and read length distribution, and the resulting assembly statistics. This log allows you to compare different assembly attempts and to identify which parameter changes had the desired effect. The The Carpentries Lessons provide training on reproducible research practices that include maintaining detailed computational records.
The assembly log should be maintained in a format that is easy to search and compare. A spreadsheet or a structured text file with columns for each parameter and metric is a practical choice. You should also record the date and the reason for each assembly attempt, such as testing a new parameter or incorporating new data. This information is valuable when you need to explain your assembly strategy to collaborators or reviewers.
Recording Depth Profiles
For each assembly attempt, you should record the depth profile across the assembly, including the genome-wide median depth, the variance in depth, and the locations of regions with abnormal depth. This information is essential for comparing the extent of collapse between different assembly attempts. You should also record the copy number estimates for any repeat families of interest, based on the depth ratio analysis. The Bioconductor project provides packages for generating and visualizing depth profiles.
The depth profile should be recorded in a format that allows you to compare the same genomic regions across different assembly attempts. This requires a common coordinate system, which can be established by aligning the different assemblies to a reference genome or by using a whole-genome alignment tool. The comparison of depth profiles across assembly attempts can reveal whether a parameter change improved the representation of a specific repeat family or shifted the collapse to a different region.
Documenting Validation Experiments
If you perform targeted validation experiments such as quantitative PCR, Southern blot, or fluorescence in situ hybridization to confirm copy number, you should document the results in the assembly log. This documentation provides the ground truth against which you can evaluate the assembly quality. The EMBL-EBI Training portal provides resources on experimental validation of genome assemblies.
The validation documentation should include the experimental protocol, the primers or probes used, the results for each replicate, and the interpretation of the results. You should also record any discrepancies between the experimental results and the assembly-based estimates, as these discrepancies can reveal limitations in either the assembly or the experimental method. The protocol for patient-derived organotypic tumor spheroids provides an example of detailed protocol documentation that includes troubleshooting guidance and quality control details, which is a useful model for documenting validation experiments.
Common Failure Patterns in Repeat Resolution
Failure Pattern 1: All Copies Collapsed into One
The most common failure pattern is that all copies of a repeat family are collapsed into a single copy in the assembly. This is diagnosed by a depth ratio equal to the expected copy number, with the entire repeat family represented by one locus. This pattern occurs when the reads are shorter than the repeat unit and the assembler has no information to distinguish the copies. The corrective action is to use longer reads or to use a hybrid assembly approach that incorporates long-range information.
The all-copies-collapsed pattern is particularly common for transposable elements that are present in high copy number and have conserved sequences. The study of transposable elements in parasitic wasps found that TE abundance and diversity were highly variable across Braconidae species, and this variability would be obscured if the assemblies collapsed TE copies. If you are studying a repeat family that is known to be present in high copy number, you should be particularly alert to the possibility of collapse.
Failure Pattern 2: Partial Collapse of Tandem Arrays
Tandem arrays such as ribosomal RNA genes, histone genes, and satellite DNA are often partially collapsed, with the assembly containing fewer copies than the true genome but more than one. This pattern is diagnosed by a depth ratio that is lower than the expected copy number but greater than one. Partial collapse occurs when the assembler can resolve some copies but not others, often because of sequence divergence between copies or because of the position of the array relative to unique sequence. The corrective action is to increase the overlap identity threshold or to use an assembler that is specifically designed for tandem repeat resolution.
Partial collapse of tandem arrays is often caused by the presence of variant copies within the array. If some copies have diverged from the consensus sequence, the assembler may be able to resolve those copies but collapse the identical copies. The depth ratio will then reflect the number of resolved copies plus the collapsed copies, which is difficult to interpret without additional information. You should examine the sequence variation within the array to determine whether the collapse is due to sequence divergence or to limitations of the assembly algorithm.
Failure Pattern 3: Chimeric Merging of Dispersed Repeats
Dispersed repeats such as transposable elements can cause chimeric merges, where the assembler joins sequence from two different genomic locations that share a repeat element. This produces a contig that is a mosaic of two distinct genomic regions. Chimeric merges are diagnosed by inspecting the assembly graph for nodes with multiple connections that do not correspond to a known repeat structure. The corrective action is to use an assembler with better graph cleaning algorithms or to break the assembly at the chimeric junctions and re-scaffold using long-range information.
Chimeric merges are particularly problematic because they can produce false gene fusions and incorrect gene order. The chimeric contig may contain sequence from two different chromosomes or from two distant regions of the same chromosome. This can lead to incorrect conclusions about gene family evolution, synteny, and genome structure. You should therefore inspect the assembly graph carefully for evidence of chimeric merges, particularly in regions that contain dispersed repeats.
Failure Pattern 4: Complete Absence of Repeat Families
In some cases, a repeat family that is known to be present in the genome is entirely absent from the assembly. This occurs when the reads from the repeat family are filtered out during assembly, either because of repeat masking or because of coverage cutoff parameters that remove low-depth regions. The corrective action is to re-assemble without masking and to adjust the coverage cutoff to retain low-depth regions. The NCBI Data Resources provide databases of repeat families that can help you identify which families are expected in your organism.
The complete absence of a repeat family is often the most difficult failure pattern to diagnose because the absence is not apparent from the assembly itself. You must have external evidence that the repeat family is present in the genome, such as PCR amplification, Southern blot, or comparison to a closely related species. The ENGRAM protocol demonstrates the importance of considering repetitive regions in genomic analysis, as the prime editing-mediated insertions used in this system can be affected by the repeat content of the target region.
Limitations of Depth-Based Collapse Detection
GC Bias and Sequencing Artifacts
Depth-based collapse detection assumes that sequencing depth is uniform across the genome, but this assumption is violated by GC bias, which causes regions with extreme GC content to have lower depth. If a repeat family has unusual GC content, the depth ratio may be misleading. You should therefore interpret depth ratios in the context of the GC content of the region and use additional evidence such as read pair information or assembly graph structure to confirm collapse.
GC bias is particularly problematic for repeat families that are AT-rich or GC-rich, such as satellite DNA and some transposable elements. The depth at these regions may be lower than the genome-wide median even if the repeats are not collapsed, leading to a false negative result. Conversely, if the repeat family has a GC content that is similar to the genome-wide average, the depth ratio may be accurate. You should calculate the GC content of the candidate collapsed regions and compare it to the genome-wide distribution to assess the potential impact of GC bias.
Copy Number Variation That Is Real
Not all regions with high depth represent collapsed repeats. Some regions have genuinely high copy number in the genome, such as ribosomal RNA genes and other highly amplified gene families. The depth ratio in these regions reflects the true copy number, not assembly collapse. You should therefore compare the depth ratio against external evidence for copy number before concluding that collapse occurred.
The distinction between real copy number variation and assembly collapse is important because the corrective actions differ. If the high copy number is real, re-assembling with different parameters will not change the depth ratio. If the high copy number is due to collapse, re-assembling with longer reads or different parameters may resolve the copies. You should therefore validate the copy number estimate with an independent method before investing time in re-assembly.
Heterozygosity Confounding Depth Estimates
In diploid organisms, the two alleles at a locus may have different sequences, and the assembler may represent them as two separate contigs or as a single collapsed contig. If the alleles are collapsed, the depth at the locus will be approximately twice the genome-wide median, which could be misinterpreted as two collapsed copies of a repeat. Distinguishing between allelic collapse and repeat collapse requires additional evidence such as the presence of heterozygous variants in the reads. The Bioconductor project provides packages for variant calling and haplotype analysis that can help with this distinction.
The distinction between allelic collapse and repeat collapse is particularly important in organisms with high heterozygosity, such as many plant and invertebrate species. In these organisms, the assembly may represent both alleles at a locus as a single haploid sequence, which is a form of collapse that affects all regions, beyond repeats. The depth at these regions will be approximately twice the genome-wide median, which could be misinterpreted as two collapsed copies of a repeat. You should therefore examine the variant density in the candidate collapsed regions to determine whether the high depth is due to allelic collapse or repeat collapse.
Quality Controls for Assembly Validation
BUSCO and Completeness Assessment
BUSCO (Benchmarking Universal Single-Copy Orthologs) is a standard tool for assessing the completeness of a genome assembly by searching for a set of conserved single-copy genes. A high BUSCO score indicates that most conserved genes are present in the assembly, but it does not detect collapse of repetitive regions. You should use BUSCO as a general quality check but not as a substitute for repeat-specific analysis. The Galaxy Training Network provides tutorials on BUSCO analysis.
The BUSCO analysis should be interpreted in the context of the expected completeness of the assembly. If the BUSCO score is low, the assembly may have missing regions that include both unique genes and repeats. If the BUSCO score is high but the assembly is missing known repeat families, the collapse is specific to repetitive regions. The combination of BUSCO and repeat-specific analysis provides a more complete picture of assembly quality than either metric alone.
Merqury and K-mer-Based Validation
Merqury is a tool that uses k-mers from the raw reads to assess the completeness and correctness of an assembly. It compares the k-mer spectrum of the reads to the k-mer spectrum of the assembly to identify missing k-mers and to estimate the base-level accuracy. Merqury can detect collapse because collapsed repeats will have missing k-mers that are present in the reads but absent from the assembly. The nf-core Documentation describes pipelines that include Merqury as a quality control step.
The k-mer-based validation provided by Merqury is complementary to the depth-based analysis described in this article. Merqury identifies missing k-mers, which indicate regions that are absent from the assembly, while the depth analysis identifies regions that are present but have higher than expected depth. The combination of these two approaches can distinguish between complete absence of a repeat family and collapse of a repeat family into a single copy.
Read Depth Distribution Analysis
The distribution of read depth across the assembly can reveal collapse even without a reference for expected copy number. In a correctly assembled genome, the depth distribution should be approximately normal with a single peak at the genome-wide median. If there are secondary peaks at higher depth, these indicate regions with higher copy number, which may be collapsed repeats or genuine amplifications. The Bioconductor project provides packages for analyzing depth distributions.
The depth distribution analysis should be performed on the raw depth values before any normalization or filtering. The distribution can be visualized as a histogram or a density plot, and the peaks can be identified using standard peak-calling algorithms. The presence of multiple peaks in the depth distribution is a strong indicator of collapse or amplification, and the position of the peaks provides an estimate of the copy number.
Safety and Regulatory Context for Genome Assembly
Data Management and Privacy Considerations
Genome assembly projects often involve data from organisms that are subject to regulatory oversight, including pathogens, agricultural species, and endangered species. You should ensure that your data management practices comply with applicable regulations and institutional policies. The NCBI Data Resources provide guidance on data submission and data sharing that can help you comply with regulatory requirements.
The data management requirements for genome assembly projects depend on the organism and the intended use of the data. If you are assembling the genome of a pathogen, you may be subject to biosafety regulations that require specific containment and handling procedures. If you are assembling the genome of an endangered species, you may be subject to permits and reporting requirements. You should consult with your institutional biosafety committee and your institutional review board before starting a genome assembly project.
Ethical Use of Genomic Data
If you are assembling the genome of a eukaryotic organism, you should consider the ethical implications of your work, particularly if the organism is a model organism or has cultural significance. The EMBL-EBI Training portal provides resources on the ethical use of genomic data that can guide your decisions.
The ethical considerations for genome assembly projects include the potential for misuse of the data, the impact on the organism or its habitat, and the rights of any communities that have cultural or economic ties to the organism. You should consider these factors when deciding whether to publish the assembly and how to share the data. The protocol for detecting genomic insulators in Drosophila provides an example of a research protocol that includes ethical considerations for working with model organisms.
Reproducibility and Documentation Standards
Regulatory and funding agencies increasingly require that genome assemblies be accompanied by detailed documentation of the methods used and the quality metrics achieved. The The Carpentries Lessons provide training on reproducible research practices that can help you meet these standards. You should document all assembly parameters, quality metrics, and validation experiments in a format that can be shared with reviewers and regulators.
The documentation should include the version numbers of all software tools, the exact commands used, and the parameter values. It should also include the quality metrics for the final assembly, including the N50, the BUSCO score, and the results of the repeat-specific analysis. The documentation should be stored in a version-controlled repository so that changes can be tracked and reviewed.
Professional Escalation Criteria
When to Seek Expert Assistance
If you have followed the diagnostic workflow described in this article and you are still unable to resolve the collapsed repeats, you should consider seeking assistance from a bioinformatics core facility or a collaborator with expertise in genome assembly. This is particularly important if the collapsed repeats are in a region of biological interest, such as a gene family involved in disease resistance or a transposable element family that is the focus of your research. The Galaxy Training Network provides a community forum where you can ask for advice from experienced assemblers.
The decision to seek expert assistance should be based on the importance of the collapsed regions to your research questions and the resources available for re-assembly. If the collapsed regions are peripheral to your research, you may be able to proceed with the current assembly and note the limitations. If the collapsed regions are central to your research, you should invest the time and resources to resolve them properly.
When to Consider a New Sequencing Strategy
If the depth analysis reveals that the reads do not contain sufficient information to resolve the repeats, you should consider generating new sequencing data with a different platform or strategy. This may involve long-read sequencing, linked-read sequencing, or Hi-C sequencing, depending on the repeat architecture of your organism. The decision to generate new data should be based on a cost-benefit analysis that considers the importance of the collapsed regions to your research questions. The EMBL-EBI Training portal provides resources on sequencing strategy design.
The cost-benefit analysis should consider the cost of the new sequencing data, the time required to generate and analyze the data, and the potential impact of resolving the collapsed regions on your research. If the collapsed regions are likely to affect the conclusions of your study, the investment in new data is justified. If the collapsed regions are unlikely to affect the conclusions, you may be able to proceed with the current assembly.
When to Report Assembly Limitations
If you are unable to resolve the collapsed repeats and you plan to publish or deposit the assembly, you should report the limitations of the assembly in the associated documentation. This includes describing the repeat families that are likely collapsed, the evidence for collapse, and the implications for downstream analyses. Transparent reporting of assembly limitations is essential for the scientific community to interpret your results correctly. The NCBI Data Resources provide guidance on assembly submission and annotation that includes recommendations for reporting limitations.
The reporting of assembly limitations should be specific and actionable. You should identify the repeat families that are likely collapsed, the evidence for collapse, and the methods that were used to detect the collapse. You should also describe the implications of the collapse for downstream analyses, such as gene family counting, transposable element annotation, and comparative genomics. This information allows other researchers to interpret your results correctly and to design follow-up studies that address the limitations.
Frequently Asked Questions
How can I tell if my assembly collapsed repeats?
The most direct way is to align your raw reads back to the assembly and examine the depth of coverage. Regions with depth significantly higher than the genome-wide median are candidates for collapsed repeats. The depth ratio, calculated as observed depth divided by median depth, estimates the number of collapsed copies. You should validate candidate regions with external evidence such as quantitative PCR or comparison to a closely related reference genome.
What read length do I need to resolve a specific repeat?
The read length must be longer than the repeat unit plus sufficient flanking unique sequence to anchor the read to a specific genomic location. As a rule of thumb, the read should extend at least a few hundred base pairs into the unique flanking sequence on both sides of the repeat. If your reads are shorter than the repeat unit, you will not be able to resolve the copy number with sequence data alone.
Why did my assembler merge two different repeat copies?
Assemblers merge repeat copies when they cannot distinguish reads from different copies. This happens when the copies are identical or nearly identical and the reads are shorter than the repeat unit. The assembler has no information to separate the copies and therefore collapses them into a single consensus sequence. Increasing the overlap identity threshold can help if the copies differ by a small number of variants.
Can I rescue collapsed repeats from my existing assembly?
In some cases, you can
Related Bioinformatics Guides
- Metagenomic Assembly Overview: Challenges and Applications
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- How to Interpret Gene Set Enrichment Analysis Results
- Evaluating Genome Assembly Quality: Metrics and Tools
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Protocol for the preparation and analysis of patient-derived organotypic tumor spheroids.. 2026.
- Multichannel genomic recording of biological information with ENGRAM.. 2026.
- Protocol for detecting genomic insulators in Drosophila using insulator-seq, a massively parallel reporter assay.. 2024.
- The highly diverse repertoire of transposable elements within the genomes of parasitic wasps (Hymenoptera: Braconidae).. 2025.
- TARPON-A Telomere Analysis and Research Pipeline Optimized for Nanopore.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.