Resolving Repeats in De Novo Assembly: How Graph Complexity Impacts Contiguity and Accuracy

By Dr. Zubair Khalid, DVM, MS, PhD ·

Resolving Repeats in De Novo Assembly: How Graph Complexity Impacts Contiguity and Accuracy

Key Takeaways

  • Repeated sequences are the primary impediment to contiguous and accurate de novo genome assembly, creating ambiguous paths in assembly graphs that lead to contig fragmentation or misjoins.
  • Long-read sequencing (e.g., PacBio HiFi, Oxford Nanopore) is crucial for resolving repeats by spanning them, providing direct evidence for correct flanking region connections, a capability short reads cannot achieve for repeats exceeding insert size.
  • Assembly graph complexity, characterized by branching nodes and cycles, directly impacts contiguity metrics like N50 and NGA50; graph-aware assemblers (e.g., Flye, Canu) preserve repeat structure for downstream analysis rather than forcing potentially incorrect linearizations.
  • Repeat collapse, where multiple copies are merged into one, and misjoins, where disparate genomic regions are erroneously linked via repeats, are common failure patterns that compromise downstream analyses like gene annotation and variant calling.
  • Practical strategies involve assessing repeat content via k-mer analysis before sequencing, selecting platforms (long-read vs. short-read) based on repeat profile, choosing appropriate assemblers (e.g., Flye for error-prone long reads, GTasm for HiFi graphs), and validating assemblies with independent data like Hi-C or optical maps.

De novo genome assembly reconstructs a genome sequence without a reference, and repeated sequences are the primary obstacle to contiguous and accurate results. Repeats create ambiguous paths in assembly graphs, causing contigs to break at repeat boundaries or misjoin across unrelated genomic regions. This article explains how repeat-induced graph complexity affects assembly outcomes and describes practical strategies for mitigating these effects using long-read sequencing, graph-aware assemblers, and post-assembly validation.

The Core Problem: Why Repeats Break Assemblies

Genomes contain repeated sequences at multiple scales, from short tandem repeats to large segmental duplications spanning hundreds of kilobases. When an assembler encounters a repeat, it cannot determine which copy of the repeat a given read originates from. This ambiguity manifests in the assembly graph as a branching structure where a single node connects to multiple possible successors.

The severity of the problem depends on the relationship between read length and repeat length. Short reads that are shorter than the repeat unit cannot span the repeat to connect flanking unique sequence. Long reads that span the entire repeat can resolve the ambiguity by providing evidence for which flanking regions belong together. This fundamental relationship explains why long-read sequencing has transformed assembly contiguity for complex genomes.

The assembly graph represents all possible paths through the sequence data. In a genome without repeats, the graph is a simple linear path. With repeats, the graph becomes a tangle of branching nodes and cycles. The assembler must decide which path represents the true genome sequence, and incorrect decisions produce misassemblies that are often more damaging than simple fragmentation.

At a Glance: Repeat Resolution Strategies and Expected Outcomes

StrategyData RequirementEffect on RepeatsPrimary Limitation
Short-read assembly with paired-end libraries100 to 300 bp reads with insert sizes up to 20 kbResolves repeats shorter than the insert size, collapses larger repeatsCannot span large segmental duplications or recent repeat expansions
Long-read assembly with error correction10 to 100 kb reads from PacBio or Oxford NanoporeSpans most repeats shorter than read length, resolves complex structuresHigher per-base cost and higher error rates require correction steps
Hybrid assembly combining short and long readsBoth short-read and long-read data from the same sampleUses short reads for base accuracy and long reads for repeat spanningRequires coordinated library preparation and increased computational load
Graph-based assembly with repeat-aware algorithmsLong reads with assembly graph output in GFA formatPreserves repeat structure for downstream analysis and scaffoldingRequires additional analysis tools to interpret graph complexity
Linked-read or Hi-C scaffoldingShort reads with barcode information or proximity ligation dataProvides long-range contiguity information to order and orient contigsDoes not resolve the underlying repeat sequence, only orders assembled fragments

How Assembly Graphs Represent Repeats

Nodes, Edges, and Paths in Assembly Graphs

Assembly graphs represent sequence relationships as nodes connected by edges. In a de Bruijn graph, nodes are fixed-length k-mers and edges connect k-mers that overlap by k-1 bases. In an overlap graph, nodes are reads or contigs and edges represent detected overlaps between sequences. Both representations face the same challenge when repeats are present.

A unique region of the genome produces a simple path through the graph. A repeat that appears twice in the genome produces a node with two incoming and two outgoing edges. The assembler cannot determine which incoming edge connects to which outgoing edge without additional information. This creates a branch point where four possible paths exist, but only two represent true genome sequence.

The repeat graph concept extends this idea to represent all copies of a repeat and their flanking sequences. An assembler that constructs a repeat graph explicitly can preserve the ambiguity for later resolution instead of forcing a potentially incorrect decision. This approach underlies the design of Flye, which builds a repeat graph from error-prone long reads and uses it to generate assembly paths [<a href="#ref-1">1</a>].

Collapsed Repeats and Their Consequences

When an assembler cannot distinguish between copies of a repeat, it may collapse them into a single sequence. This produces a contig that is shorter than the true genome and loses the distinction between paralogous regions. Collapsed repeats are particularly problematic for downstream analysis because they create false homozygous variants in what are actually heterozygous or duplicated regions.

Repeat collapse also affects gene annotation and comparative genomics. If two gene copies are collapsed into one, the assembly appears to have a single gene when the organism actually has two. This can lead to incorrect conclusions about gene family size, copy number variation, and evolutionary history.

Misjoins and Chimeric Contigs

The alternative to collapse is misjoin, where the assembler connects sequence from two different genomic locations through a shared repeat. This produces a chimeric contig that does not exist in the true genome. Misjoins are especially dangerous because they can appear valid in assembly statistics while containing large-scale structural errors.

Detecting misjoins requires comparing the assembly to independent evidence such as genetic maps, optical maps, or Hi-C contact data. Without such validation, a misassembled genome can propagate errors into every downstream analysis that relies on the reference sequence.

Long-Read Sequencing as the Primary Repeat Resolution Tool

Read Length and Repeat Spanning

Long-read sequencing platforms produce reads that routinely exceed 10 kb and can reach 100 kb or more. These reads can span repeats that are shorter than the read length, providing direct evidence for the correct connection between flanking regions. The ability to span repeats is the primary reason long-read assemblies achieve substantially better contiguity than short-read assemblies.

The relationship between read length and repeat resolution is straightforward. A repeat of length L requires a read that covers the repeat plus sufficient flanking unique sequence on both sides to anchor the read to a specific genomic location. Reads that are shorter than the repeat plus its flanking anchors cannot resolve the repeat.

Error Rates and the Need for Correction

Long reads have higher error rates than short reads. Early PacBio reads had error rates around 15 percent, and Oxford Nanopore reads had similar or higher error rates. These errors complicate assembly because they create false variation in the overlap graph and reduce the effective overlap length between reads.

Modern long-read platforms have improved accuracy substantially. HiFi reads from PacBio achieve accuracy above 99 percent with read lengths of 10 to 25 kb. Oxford Nanopore has introduced improved basecalling that reduces error rates, though raw accuracy remains lower than HiFi. The choice between accuracy and read length is a central tradeoff in long-read assembly design.

Assembler Adaptations for Noisy Long Reads

Assemblers designed for long reads must handle high error rates while preserving repeat information. Canu introduced adaptive k-mer weighting and repeat separation strategies to address these challenges. The assembler uses a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes, preserving structural information that would be lost with more aggressive graph simplification [<a href="#ref-2">2</a>].

Flye takes a different approach by constructing a repeat graph directly from error-prone reads. The algorithm generates disjointigs, which are arbitrary paths through the unknown repeat graph, and then builds an accurate repeat graph from these error-riddled sequences. This approach allows Flye to resolve repeats that would fragment other assemblers [<a href="#ref-1">1</a>].

Graph Complexity and Assembly Quality Metrics

N50 and NGA50 as Contiguity Measures

Assembly contiguity is commonly measured with N50, the contig length at which half the assembled bases are in contigs of that length or longer. N50 provides a simple summary of fragmentation but does not account for accuracy. A misassembled genome can have a high N50 while containing large-scale errors.

NGA50 addresses this limitation by aligning the assembly to a reference genome and breaking contigs at misjoin points before calculating the statistic. This metric penalizes misassemblies and provides a more realistic assessment of usable contiguity. The distinction between N50 and NGA50 is important when evaluating assembly quality for downstream applications.

The Impact of Repeats on Contiguity Metrics

Repeats affect contiguity metrics differently depending on their size and distribution. Small repeats that are shorter than read length have minimal impact on long-read assemblies. Large repeats that exceed read length create breaks at every repeat boundary, fragmenting the assembly into pieces that cannot be ordered without additional information.

The human genome provides a useful example. Flye nearly doubled the NGA50 of the human genome assembly compared with existing assemblers, demonstrating that repeat-aware graph construction can substantially improve usable contiguity [<a href="#ref-1">1</a>]. This improvement came from better resolution of the repeat structures that fragmented previous assemblies.

Beyond Contiguity: Accuracy and Completeness

Contiguity metrics do not capture all aspects of assembly quality. A genome can be contiguous but contain collapsed repeats, misjoins, or missing sequences. Comprehensive assembly evaluation requires multiple metrics, including base accuracy, gene completeness, and structural correctness.

Graph-based assembly outputs in GFA format allow researchers to examine the unresolved repeat structure directly. This information can guide additional sequencing or scaffolding efforts to resolve specific problematic regions. Canu provides such graph-based outputs specifically to support analysis of assembly structures that cannot be linearly represented [<a href="#ref-2">2</a>].

Practical Workflow for Repeat-Aware Assembly

Step 1: Assess Repeat Content Before Sequencing

Genome size and repeat content can be estimated from k-mer frequency distributions in shallow sequencing data. A genome with high repeat content will show a distinct k-mer spectrum with multiple peaks corresponding to different copy numbers. This information guides sequencing strategy and expected assembly difficulty.

Flow cytometry or other genome size estimates provide a baseline for expected assembly size. If the assembly is substantially smaller than the estimated genome size, repeat collapse is likely. If the assembly is larger, contamination or haplotype separation may be present.

Step 2: Select Sequencing Platforms Based on Repeat Profile

Genomes with few repeats can be assembled adequately with short reads. Genomes with substantial repeat content benefit from long-read sequencing. The choice between HiFi and traditional long reads depends on the specific repeat structures and the accuracy requirements of downstream analyses.

HiFi reads provide high accuracy that simplifies assembly and reduces the need for polishing. Traditional long reads provide greater read length at lower accuracy, which can resolve larger repeats but requires more sophisticated assembly algorithms and additional correction steps.

Step 3: Choose an Assembler Matched to Your Data Type

Different assemblers are optimized for different data types and repeat profiles. Canu handles noisy long reads with adaptive overlap strategies [<a href="#ref-2">2</a>]. Flye constructs repeat graphs from error-prone reads [<a href="#ref-1">1</a>]. GTasm uses graph transformer networks to find optimal paths in assembly graphs built from HiFi reads [<a href="#ref-3">3</a>].

The choice of assembler should consider the expected repeat structure, the error profile of the sequencing data, and the computational resources available. Benchmarking multiple assemblers on a subset of the data can identify the best performer before committing to a full assembly run.

Step 4: Examine the Assembly Graph for Unresolved Repeats

After assembly, examine the graph output to identify unresolved repeat structures. GFA files contain the graph topology, including nodes with multiple connections that indicate potential repeats. Tools that visualize assembly graphs can help identify problematic regions that require additional attention.

Graph-based repeat detection methods can classify sequences into repetitive and nonrepetitive categories using the assembly graph structure. GraSSRep uses graph neural networks in a self-supervised framework to achieve this classification, providing a systematic approach to identifying repeats that may cause assembly problems [<a href="#ref-4">4</a>].

Step 5: Validate the Assembly with Independent Evidence

Assembly validation requires comparing the assembled sequence to independent data. This can include genetic maps, optical maps, Hi-C contact maps, or comparisons to closely related reference genomes. Validation should check both contiguity and accuracy, identifying misjoins and collapsed repeats that may not be apparent from assembly statistics alone.

For clinical or diagnostic applications, validation is especially critical. Long-read sequencing can detect and clarify disease-associated variants that are missed by short-read workflows, including structural variants, tandem repeat expansions, and regions of high sequence similarity [<a href="#ref-5">5</a>]. The additional diagnostic yield depends on the quality of the assembly and the completeness of repeat resolution.

Options and Tradeoffs in Repeat Resolution

Short-Read Only Assembly

Short-read assembly remains the most cost-effective option for small genomes with limited repeat content. Bacterial genomes and small eukaryotic genomes can often be assembled to completion with short reads alone. The primary limitation is the inability to span repeats longer than the library insert size.

Paired-end and mate-pair libraries extend the effective span of short-read assemblies. Insert sizes up to 20 kb allow resolution of repeats up to that length, though the gap between read pairs is not directly sequenced. This approach cannot resolve larger repeats or complex repeat structures.

Long-Read Only Assembly

Long-read assembly provides the best repeat resolution for most genomes. The ability to span repeats directly with single reads eliminates the ambiguity that fragments short-read assemblies. Modern long-read assemblers can produce near-complete eukaryotic chromosomes with appropriate data.

The primary tradeoff is cost. Long-read sequencing remains more expensive per base than short-read sequencing, though the gap has narrowed substantially. For large genomes, the cost of achieving sufficient long-read coverage can be substantial.

Hybrid Assembly Approaches

Hybrid assembly combines short and long reads to leverage the strengths of both. Short reads provide high base accuracy and low cost. Long reads provide repeat spanning and structural information. Hybrid assemblers use long reads to build the assembly structure and short reads to correct errors and polish the final sequence.

The main challenge in hybrid assembly is coordinating the different data types. Coverage requirements differ, and the assembler must handle the different error profiles appropriately. When implemented well, hybrid approaches can achieve high contiguity and accuracy at lower cost than long-read-only approaches.

Graph-Based Assembly and Post-Assembly Scaffolding

Some assemblers produce graph-based outputs that preserve unresolved repeat structure for downstream analysis. These graphs can be combined with scaffolding information from Hi-C or optical mapping to order and orient contigs across repeats. This approach does not resolve the repeat sequence itself but can produce chromosome-scale assemblies despite unresolved repeats.

The combination of highly resolved assembly graphs with long-range scaffolding information has been proposed as a path to complete and automated assembly of complex genomes [<a href="#ref-2">2</a>]. This approach acknowledges that some repeats cannot be resolved with current sequencing technology and uses external information to place contigs in their correct genomic context.

Records and Measurements for Assembly Quality

Coverage Statistics and Their Interpretation

Sequencing coverage is the average number of reads covering each base in the genome. Coverage requirements depend on the sequencing platform and the assembly strategy. Long-read assemblies typically require lower coverage than short-read assemblies because each read provides more information per base.

Coverage variation across the genome can indicate problems. Regions with unusually high or low coverage may represent collapsed repeats, assembly errors, or biological variation such as copy number differences. Examining coverage along the assembly can identify problematic regions that require additional attention.

K-mer Analysis for Repeat Detection

K-mer frequency analysis provides a rapid assessment of repeat content and assembly completeness. Unique sequence produces k-mers at the expected coverage level. Repeats produce k-mers at multiples of the expected coverage, with the multiplicity indicating copy number.

Comparing k-mer spectra between the raw reads and the assembly can identify collapsed repeats. If the assembly contains fewer high-copy k-mers than the raw data, repeats have been collapsed. This analysis can be performed before and after assembly to assess repeat resolution.

Alignment-Based Quality Metrics

Aligning the assembly to a reference genome, when available, provides detailed quality information. The alignment can identify misjoins, collapsed repeats, and missing sequence. Metrics such as NGA50 and genome fraction quantify the usable contiguity and completeness of the assembly.

For organisms without a close reference, alignment to related species can still provide useful information. Conserved synteny can identify large-scale structural errors, and conserved gene content can assess completeness. These analyses require careful interpretation because genuine biological differences can be mistaken for assembly errors.

Common Failure Patterns in Repeat Resolution

Repeat Collapse in Assemblies

Repeat collapse occurs when the assembler merges multiple copies of a repeat into a single sequence. This produces an assembly that is smaller than the true genome and loses the distinction between paralogous regions. Collapse is most common for recent repeats with high sequence identity, where the assembler cannot distinguish between copies.

Detection of collapse requires comparing assembly size to expected genome size or examining coverage patterns. Collapsed regions show elevated coverage because reads from multiple genomic locations map to the same assembled sequence. K-mer analysis can also identify collapse by comparing k-mer frequencies between reads and assembly.

Chimeric Contigs from Misjoins

Misjoins occur when the assembler connects sequence from different genomic locations through a shared repeat. The resulting chimeric contig contains sequence from two or more genomic regions that are not adjacent in the true genome. Misjoins are more difficult to detect than collapses because the assembled sequence may be internally consistent.

Detection of misjoins requires comparison to independent evidence. Hi-C contact maps can identify incorrect joins because sequences that are far apart in the genome should show low contact frequency. Optical maps provide similar validation at lower resolution. Alignment to a reference genome, when available, directly identifies misjoins.

Overly Aggressive Graph Simplification

Some assemblers simplify the assembly graph aggressively to produce longer contigs. This simplification can collapse diverged repeats and haplotypes, losing biological information. The tradeoff between contiguity and accuracy is central to assembly algorithm design.

Canu specifically addresses this issue with a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes [<a href="#ref-2">2</a>]. This approach preserves structural information that can be used for downstream phasing and scaffolding, at the cost of reduced contiguity in the primary assembly output.

Unresolved Haplotypes in Diploid Genomes

Diploid genomes contain two copies of each chromosome that differ by sequence variants. Assemblers must decide whether to collapse these haplotypes into a single sequence or represent them separately. Collapsing haplotypes produces a haploid representation that may contain false variants. Representing haplotypes separately produces a more complex assembly that is difficult to interpret.

Repeat resolution and haplotype resolution are related problems because both involve ambiguous graph paths. Assemblers that preserve graph complexity can represent both repeats and haplotypes, but the resulting assembly requires specialized tools for interpretation.

Quality Control and Validation Strategies

Assembly Polishing and Error Correction

Polishing uses additional sequencing data to correct errors in the assembled sequence. Short reads are commonly used to polish long-read assemblies because they provide high base accuracy. The polishing process aligns reads to the assembly and corrects discrepancies, improving base accuracy without changing the assembly structure.

Multiple rounds of polishing may be required to achieve the desired accuracy. Each round should be evaluated to determine whether it improves accuracy or introduces new errors. Over-polishing can introduce errors by overcorrecting genuine sequence variation.

Completeness Assessment with Benchmarking Sets

Benchmarking sets of conserved genes provide a standardized assessment of assembly completeness. These sets contain genes that are expected to be present in all members of a taxonomic group. The proportion of these genes found in the assembly indicates completeness, and the proportion found as complete single copies indicates both completeness and the absence of collapse.

The results of completeness assessment should be interpreted in the context of the organism and the assembly strategy. Some genes may be genuinely absent, and some may be too repetitive to assemble. The assessment provides a lower bound on completeness instead of a definitive measure.

Structural Validation with Independent Data

Structural validation compares the assembly to independent evidence about genome structure. This can include genetic maps, optical maps, Hi-C contact maps, or comparisons to related genomes. Structural validation identifies misjoins and incorrect ordering that are not apparent from sequence-level analysis.

For clinical applications, structural validation is essential. Long-read sequencing can detect structural variants and repeat expansions that are missed by short-read workflows, but the accuracy of these detections depends on the quality of the assembly [<a href="#ref-5">5</a>]. Validation with independent data provides confidence in the clinical interpretation.

Limitations and Interpretation Boundaries

Repeats That Cannot Be Resolved with Current Technology

Some repeats cannot be resolved with any current sequencing technology. Very large repeats, such as centromeric satellite arrays, may exceed the length of any single read. Highly identical repeats with no unique flanking sequence cannot be anchored to specific genomic locations.

These unresolved repeats should be represented as gaps or ambiguous regions in the assembly instead of forcing incorrect connections. The assembly graph can represent the ambiguity, and the graph structure can be reported to downstream users who need to interpret the assembly.

The Difference Between Assembly and True Genome Structure

An assembly is a model of the genome, not the genome itself. The assembly may contain errors, collapsed repeats, and unresolved regions. The quality of the assembly depends on the sequencing data, the assembly algorithm, and the validation performed.

Researchers should interpret assembly-based analyses with appropriate caution. Variants identified in repetitive regions should be validated with additional evidence. Gene content should be assessed with completeness metrics. Structural conclusions should be supported by independent data.

Computational Resource Requirements

Repeat-aware assembly with long reads requires substantial computational resources. The overlap and assembly steps are computationally intensive, and the memory requirements scale with genome size and repeat content. Graph-based assembly outputs require additional storage and analysis tools.

Researchers should plan for the computational requirements of their assembly project. This includes the assembly itself but also the validation, polishing, and downstream analysis steps. Cloud computing and institutional clusters can provide the necessary resources for large genome projects.

Professional Escalation Criteria

When to Seek Additional Sequencing

Additional sequencing is warranted when the assembly contains unresolved repeats that are important for downstream analysis. This includes repeats that span genes of interest, repeats in regions associated with disease, or repeats that prevent chromosome-scale assembly.

The decision to sequence more should be based on the specific requirements of the project. If the unresolved repeats do not affect the planned analyses, additional sequencing may not be justified. If the repeats are critical, additional sequencing with longer reads or different platforms may resolve them.

When to Consult a Bioinformatics Specialist

A bioinformatics specialist should be consulted when assembly results are unexpected or when standard approaches fail. This includes assemblies with unusually low contiguity, evidence of extensive collapse or misassembly, or difficulty interpreting the assembly graph.

Specialists can provide expertise in advanced assembly strategies, graph analysis, and validation approaches. They can also help interpret assembly results in the context of the specific organism and research question.

When to Reconsider the Assembly Strategy

The assembly strategy should be reconsidered when the current approach is not producing acceptable results. This may involve changing the sequencing platform, the assembler, or the overall approach. Benchmarking multiple strategies on a subset of the data can identify the best approach before committing to a full assembly run.

The decision to change strategy should be based on evidence instead of frustration. If the current approach is producing good results with minor issues, targeted improvements may be more appropriate than a complete restart. If the current approach is fundamentally flawed, a new strategy may be necessary.

A Practical Decision Framework for Choosing Repeat Resolution Strategies

Selecting the right repeat resolution approach requires a structured evaluation of the specific genome, the repeat structures present, and the downstream analysis goals. Researchers often default to the most expensive or most recent technology without systematically assessing whether the investment will address the actual assembly problems. This section provides a decision framework that connects repeat characteristics to concrete sequencing and assembly choices, along with a record system for tracking decisions and outcomes across assembly projects.

Step 1: Characterize the Repeat Landscape Before Committing Resources

The first decision point occurs before sequencing begins. A shallow sequencing pass, typically 5 to 10 fold coverage with short reads, provides enough data to estimate genome size and repeat content through k-mer frequency analysis. The k-mer spectrum reveals distinct peaks corresponding to different copy numbers. A single dominant peak at the expected coverage level indicates a low-repeat genome. Multiple peaks at integer multiples of the haploid coverage indicate repeated sequences, with the multiplicity revealing copy number.

For example, a genome with a haploid coverage of 30 fold will show a main peak at 30 and secondary peaks at 60, 90, and 120 if substantial diploid, triploid, and tetraploid repeats are present. The proportion of k-mers in each peak estimates the fraction of the genome occupied by each repeat copy number class. This information directly informs whether long reads are necessary and what read length is required.

Genome size estimates from flow cytometry or other independent methods provide a baseline for comparison. If the k-mer-based genome size estimate is substantially larger than the assembled size later, repeat collapse has occurred. Recording these estimates before assembly creates a reference point for evaluating assembly completeness.

Step 2: Match Read Length to Repeat Length Distribution

The relationship between read length and repeat length determines whether a given sequencing platform can resolve specific repeats. A repeat of length L requires a read that spans the repeat plus flanking unique sequence on both sides. The flanking sequence must be long enough to anchor the read uniquely, typically at least several hundred bases on each side for reliable mapping.

For practical decision making, classify repeats into three categories based on the available read length. Repeats shorter than half the read length are generally resolvable because a single read can span the repeat with substantial flanking sequence on both sides. Repeats between half and the full read length may be resolvable depending on the exact read length distribution and the quality of the flanking sequence. Repeats longer than the read length cannot be spanned by a single read and require alternative strategies such as linked reads, Hi-C scaffolding, or targeted sequencing.

This classification should be recorded for each major repeat family in the genome. The record should include the repeat length estimate, the copy number, the sequence identity between copies, and the expected resolvability with the chosen platform. This record becomes the basis for evaluating whether the assembly achieved the expected repeat resolution.

Step 3: Select the Assembly Strategy Based on Repeat Complexity

The repeat characterization from Step 1 and the read length analysis from Step 2 combine to guide assembler selection. For genomes where most repeats are shorter than the available read length, standard long-read assemblers such as Canu or Flye are appropriate. These assemblers handle the error profiles of noisy long reads and construct graphs that resolve repeats within the read length range [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

For genomes with substantial repeat content that exceeds read length, consider assemblers that preserve graph complexity for downstream analysis. Canu provides graph-based assembly outputs in GFA format specifically to support analysis of assembly structures that cannot be linearly represented [<a href="#ref-2">2</a>]. This output allows researchers to identify unresolved repeats and plan targeted resolution strategies.

For HiFi data with high accuracy, graph transformer-based approaches such as GTasm can find optimal paths through the assembly graph using learned features [<a href="#ref-3">3</a>]. These methods are particularly useful when the assembly graph contains complex branching structures that confuse simpler path-finding algorithms.

The assembler choice should be recorded with the rationale, including the expected repeat profile, the data type, and the specific features of the assembler that address the identified challenges. This record supports reproducibility and provides a basis for comparing assembler performance across projects.

Step 4: Establish a Repeat Resolution Tracking System

A repeat resolution tracking system documents each repeat family, its characteristics, and the outcome of the assembly. This system serves multiple purposes. It identifies which repeats were resolved and which remain problematic. It provides a basis for deciding whether additional sequencing or alternative strategies are needed. It creates a record that supports interpretation of downstream analyses.

The tracking system should include the following fields for each repeat family. The repeat name or identifier, the estimated length, the copy number, the sequence identity between copies, the genomic context if known, the expected resolvability with the chosen platform, the actual resolution outcome in the assembly, and any notes on why resolution succeeded or failed.

The resolution outcome should be classified into one of three categories. Resolved means the repeat copies are correctly represented as distinct sequences in the assembly. Collapsed means multiple copies were merged into a single sequence. Ambiguous means the repeat structure remains unresolved in the assembly graph and requires additional information.

This tracking system should be maintained throughout the assembly project and updated as additional data or analysis becomes available. The record provides a quantitative basis for deciding whether the assembly is complete or whether additional sequencing is warranted.

Step 5: Apply a Decision Matrix for Unresolved Repeats

When the tracking system identifies unresolved repeats, a decision matrix guides the response. The matrix considers the importance of the repeat for downstream analysis, the feasibility of resolving it with additional data, and the cost of the additional effort.

For repeats that are important for downstream analysis and feasible to resolve, additional sequencing is warranted. This may involve longer reads, different sequencing platforms, or targeted approaches such as CRISPR-based enrichment or optical mapping. The specific approach depends on the repeat characteristics and the available technology.

For repeats that are important but not feasible to resolve with current technology, the assembly should represent the ambiguity explicitly in the graph structure. The GFA output preserves the unresolved structure for downstream users who need to interpret the assembly [<a href="#ref-2">2</a>]. This approach avoids forcing incorrect connections that would create misassemblies.

For repeats that are not important for the planned analyses, document the limitation and proceed. The tracking system records the unresolved repeat for future reference, but the assembly is considered adequate for the intended purpose.

Step 6: Document Decisions and Outcomes for Reproducibility

The decision framework produces a record that supports reproducibility and interpretation. Each decision should be documented with the evidence that informed it, the alternative options considered, and the rationale for the chosen approach. This documentation should be maintained alongside the assembly files and analysis scripts.

The record should include the repeat characterization data, the read length analysis, the assembler selection rationale, the tracking system entries, and the decision matrix outcomes. This documentation allows other researchers to understand why specific choices were made and to evaluate whether those choices were appropriate for the research question.

Reproducible workflows support this documentation process. Community standards such as those provided by nf-core emphasize reproducible pipeline usage and configuration [<a href="#ref-6">6</a>]. Training resources from The Carpentries provide foundational skills for managing data and code in reproducible research [<a href="#ref-7">7</a>]. These resources support the documentation practices that make assembly decisions transparent and auditable.

Common Failure Patterns in Decision Making

Several recurring failure patterns undermine repeat resolution decisions. The most common is proceeding with short-read sequencing for a genome with substantial repeat content, despite evidence from k-mer analysis that long reads are necessary. This pattern produces fragmented assemblies that require complete resequencing, wasting time and resources.

Another pattern is selecting a read length that is marginally too short for the dominant repeat families. The assembly resolves most repeats but breaks at the largest repeats, producing a fragmented result that could have been avoided with longer reads. The read length analysis in Step 2 prevents this failure by explicitly comparing read length to repeat length distributions.

A third pattern is failing to examine the assembly graph after assembly. Researchers rely on contiguity statistics such as N50 without inspecting the graph for unresolved structures. This pattern misses collapsed repeats and misjoins that are not apparent from summary statistics. The tracking system and graph examination steps prevent this failure by making repeat resolution an explicit part of the assembly workflow.

A fourth pattern is over-investing in resolving repeats that are not important for the downstream analysis. Researchers spend substantial resources attempting to resolve every repeat when only a subset affects the research question. The decision matrix in Step 5 prevents this failure by explicitly weighing the importance of each repeat against the cost of resolution.

Records and Measurements for Decision Evaluation

The decision framework generates specific records that support evaluation and improvement. The repeat characterization record provides the baseline data for all subsequent decisions. The read length analysis record documents the expected resolvability for each repeat family. The assembler selection record documents the rationale for the chosen approach. The tracking system records the actual outcomes for each repeat family. The decision matrix records the response to each unresolved repeat.

These records should be reviewed after each assembly project to identify patterns and improve future decisions. Questions for review include whether the repeat characterization accurately predicted assembly outcomes, whether the read length analysis correctly identified resolvable and unresolvable repeats, whether the assembler selection achieved the expected results, and whether the decision matrix responses were appropriate.

This review process builds institutional knowledge that improves assembly outcomes across projects. The records also support communication with collaborators and reviewers who need to understand the assembly decisions and their rationale.

Professional Escalation Criteria for Decision Framework

The decision framework includes specific criteria for escalating to professional support. If the repeat characterization reveals an unexpected repeat landscape that does not match the organism's known biology, consult a bioinformatics specialist before proceeding. If the read length analysis indicates that no available platform can resolve the dominant repeats, consult a specialist about alternative strategies such as linked reads, optical mapping, or targeted sequencing.

If the assembly graph reveals structures that cannot be interpreted with available tools, consult a specialist who can provide expertise in graph analysis. If the tracking system shows widespread collapse or misassembly despite following the decision framework, consult a specialist to identify whether the framework was applied correctly or whether the genome presents unusual challenges.

If the decision matrix indicates that important repeats cannot be resolved with any available approach, consult a specialist about whether the research question can be answered despite the limitation or whether alternative experimental approaches are needed. These escalation criteria ensure that complex assembly problems receive appropriate expertise instead of repeated failed attempts.

Frequently Asked Questions

Why do repeats cause assembly breaks even with long reads?

Repeats cause assembly breaks when they are longer than the reads available. A read must span the entire repeat plus sufficient flanking unique sequence to anchor the read to a specific genomic location. If the repeat exceeds this length, the assembler cannot determine which flanking regions belong together and must break the assembly at the repeat boundary. Even with long reads, very large repeats such as centromeric arrays or recent segmental duplications can exceed read length and cause breaks.

What is the difference between N50 and NGA50 in assembly evaluation?

N50 measures contiguity by calculating the contig length at which half the assembled bases are in contigs of that length or longer. NGA50 aligns the assembly to a reference genome and breaks contigs at misjoin points before calculating the same statistic. NGA50 therefore penalizes misassemblies and provides a more accurate assessment of usable contiguity. A high N50 with a low NGA50 indicates that the assembly contains substantial misjoins.

How does repeat collapse affect downstream variant calling?

Repeat collapse merges multiple copies of a repeat into a single sequence, creating false homozygous variants in regions that are actually heterozygous or duplicated. Variants in collapsed regions may appear homozygous when they are actually heterozygous, or may be missed entirely if the collapsed sequence does not represent either true copy accurately. This can lead to incorrect conclusions about genotype, copy number, and gene content.

What are the advantages of graph-based assembly outputs?

Graph-based assembly outputs in GFA format preserve the unresolved repeat structure for downstream analysis. Instead of forcing a potentially incorrect linear representation, the graph shows the ambiguity explicitly. This information can guide additional sequencing, scaffolding, or analysis efforts. Graph-based outputs also support integration with complementary phasing and scaffolding techniques for complex genomes [<a href="#ref-2">2</a>].

How do HiFi reads differ from traditional long reads for repeat resolution?

HiFi reads provide high accuracy above 99 percent with read lengths of 10 to 25 kb. Traditional long reads from PacBio or Oxford Nanopore provide longer reads but with higher error rates. HiFi reads simplify assembly because the high accuracy reduces the need for error correction and polishing. Traditional long reads can span larger repeats but require more sophisticated assembly algorithms to handle the error profile.

What is the role of polishing in long-read assembly?

Polishing uses additional sequencing data to correct errors in the assembled sequence. Short reads are commonly used to polish long-read assemblies because they provide high base accuracy. The polishing process aligns reads to the assembly and corrects discrepancies, improving base accuracy without changing the assembly structure. Multiple rounds of polishing may be required to achieve the desired accuracy.

How can I detect collapsed repeats in my assembly?

Collapsed repeats can be detected by comparing assembly size to expected genome size, examining coverage patterns, or analyzing k-mer frequencies. Collapsed regions show elevated coverage because reads from multiple genomic locations map to the same assembled sequence. K-mer analysis compares the frequency spectrum between raw reads and the assembly to identify high-copy k-mers that are underrepresented in the assembly.

When should I use a hybrid assembly approach instead of long-read only?

Hybrid assembly is appropriate when long-read coverage is limited by cost or sample availability. The approach uses short reads for base accuracy and long reads for repeat spanning, potentially achieving high quality at lower cost than long-read-only approaches. Hybrid assembly is also useful when the long-read data has high error rates that require short-read correction. The main challenge is coordinating the different data types and handling their different error profiles appropriately.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Assembly of long, error-prone reads using repeat graphs.](https://pubmed.ncbi.nlm.nih.gov/30936562). Nature biotechnology, 2019. [2] [Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation.](https://pubmed.ncbi.nlm.nih.gov/28298431). Genome research, 2017. [3] [GTasm: a genome assembly method using graph transformers and HiFi reads.](https://pubmed.ncbi.nlm.nih.gov/39525812). Frontiers in genetics, 2024. [4] [Graph-based self-supervised learning for repeat detection in metagenomic assembly.](https://pubmed.ncbi.nlm.nih.gov/39029947). Genome research, 2024. [5] [The additional diagnostic yield of long-read sequencing in undiagnosed rare diseases.](https://pubmed.ncbi.nlm.nih.gov/39900460). Genome research, 2025. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.