Evaluating Completeness in Complex Assemblies: How to Use BUSCO and K-mer Spectra to Detect Missing Repeats

By Dr. Zubair Khalid, DVM, MS, PhD ·

Evaluating Completeness in Complex Assemblies: How to Use BUSCO and K-mer Spectra to Detect Missing Repeats

Key Takeaways

  • BUSCO completeness, while a standard metric, can mask significant missing sequence content if that content is primarily repetitive and lacks conserved orthologs.
  • K-mer spectrum analysis, using tools like GenomeScope and Merqury, complements BUSCO by assessing the presence of all short sequence motifs in the sequencing reads within the assembly, thereby capturing repetitive regions.
  • A substantial gap between high BUSCO completeness and lower k-mer completeness directly indicates that the assembly is missing repetitive elements, such as transposable elements or satellite DNA, which are crucial for understanding genome architecture and evolution.
  • The practical workflow involves estimating genome properties (size, repeat content) from raw reads using k-mer analysis, followed by running BUSCO on the assembly and then comparing assembly k-mers to read k-mers using tools like Merqury to identify missing sequence.
  • Decisions on assembly acceptance or the need for additional sequencing (e.g., higher coverage long reads) or alternative assembly strategies should be informed by the magnitude of the BUSCO-Merqury completeness gap and the biological context of the intended downstream analyses.

Genome assembly completeness is often reported as a single BUSCO percentage, but that number can hide a critical problem: repetitive regions that are absent from the assembly yet invisible to gene-based completeness checks. This article explains how to combine BUSCO assessments with k-mer spectrum analysis using tools such as GenomeScope and Merqury to detect missing repeats, interpret the results correctly, and decide when an assembly requires additional sequencing or alternative assembly strategies.

The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who generate or evaluate genome assemblies. The practical outcome is a workflow that distinguishes between assemblies that are genuinely complete and those that merely appear complete because their missing content is repetitive instead of genic.

The Completeness Blind Spot in Gene-Based Assessment

BUSCO (Benchmarking Universal Single-Copy Orthologs) has become the default completeness metric for genome assemblies. The method searches an assembly for a curated set of single-copy orthologous genes that are expected to be present in a given taxonomic group. When a high percentage of these genes are found complete, the assembly is typically described as high quality.

The limitation of this approach is structural. BUSCO genes are conserved protein-coding sequences, and they are overwhelmingly located in gene-rich, often repeat-poor regions of the genome. Repetitive elements, including transposable elements, tandem repeats, satellite DNA, and ribosomal RNA gene clusters, are frequently excluded from BUSCO gene sets. An assembly can therefore lose substantial portions of its repetitive content without any corresponding drop in the BUSCO completeness score.

Consider the burrowing sea anemone Paracondylactis sinensis genome. The published assembly reports 95.91% BUSCO completeness, a value that confirms high assembly quality. Yet repetitive elements account for 26.43% of that genome, with transposable elements representing 20.47%. If a researcher assembled only the gene-rich fraction of this genome and discarded most of the repeats, the BUSCO score might remain near 95%, while the actual genome size and repeat content would be severely underestimated.

The consequence is an assembly that passes the standard quality gate while being biologically incomplete. Repeat content matters for genome biology. Transposable elements influence genome structure, gene regulation, and evolutionary dynamics. Satellite repeats are central to centromere function. Ribosomal DNA arrays are essential for cellular function. An assembly missing these regions cannot support robust downstream analysis of genome architecture, comparative genomics, or population genetics.

How K-mer Spectra Reveal What BUSCO Cannot See

K-mer spectrum analysis approaches completeness from a different direction. Instead of searching for known genes, this method examines the frequency distribution of all short sequences of length k that occur in the sequencing reads. The resulting spectrum provides a direct measurement of the sequence content present in the read set and allows comparison with the sequence content present in the assembly.

The core logic is straightforward. If sequencing coverage is uniform across the genome, each unique k-mer from a haploid region should appear at a frequency proportional to the sequencing depth. K-mers from repetitive regions appear at higher frequencies because multiple copies contribute to the read set. When an assembly is complete, nearly all k-mers from the reads should be present in the assembly. When an assembly is missing sequence, the missing k-mers are absent.

GenomeScope uses the k-mer spectrum from raw reads to estimate genome size, heterozygosity, and repeat content before assembly. Merqury uses k-mer spectra to compare an assembly against the read set and produces a completeness metric that reflects the fraction of read k-mers present in the assembly. This k-mer-based completeness value captures all sequence content, including repeats, and therefore provides a complement to BUSCO.

The distinction matters for interpretation. A BUSCO score of 98% with a Merqury completeness of 85% indicates that the assembly is missing a substantial fraction of the genome, and that the missing fraction is likely repetitive. A BUSCO score of 98% with a Merqury completeness of 97% indicates that the assembly captures nearly all sequence content present in the reads.

At a Glance: Completeness Metrics and What They Measure

MetricInput DataWhat It DetectsBlind SpotTypical Use
BUSCO completenessAssemblyPresence of conserved single-copy orthologsRepetitive regions without conserved genesStandard assembly quality reporting
GenomeScope genome size estimateRaw readsTotal genome size, heterozygosity, repeat contentDoes not assess assembly directlyPre-assembly planning and read set evaluation
Merqury completenessAssembly plus raw readsFraction of read k-mers present in assemblyDoes not identify which specific regions are missingPost-assembly completeness verification
Assembly size comparisonAssemblyTotal assembled basesDoes not distinguish collapsed from missing repeatsFirst-pass sanity check against expected genome size

The table above summarizes the four metrics that form the basis of a repeat-aware completeness assessment. No single metric is sufficient. BUSCO alone cannot detect missing repeats. GenomeScope alone cannot assess an assembly. Merqury alone cannot distinguish biologically absent sequence from assembly error. The combination provides the necessary evidence.

The Practical Workflow for Repeat-Aware Completeness Assessment

The workflow described here assumes that raw sequencing reads and a draft assembly are available. The steps are ordered so that each stage produces evidence that informs the next decision.

Step 1: Estimate Genome Properties from Raw Reads

Before assessing the assembly, establish the expected genome size and repeat content from the raw reads. This step uses GenomeScope or an equivalent k-mer-based estimator. The input is a set of sequencing reads, typically long reads for eukaryotic genome projects.

The output includes an estimated genome size, a heterozygosity estimate, and a repeat content estimate. Record these values. The estimated genome size becomes the reference against which the assembly size is compared. The repeat content estimate provides context for interpreting k-mer completeness values.

For the Trypoxylon clavicerum wasp genome, the final assembly reached 270.89 megabases. A k-mer-based estimate from the reads should have predicted a genome of approximately this size before assembly. If the read-based estimate had been substantially larger than the assembly, the discrepancy would have indicated missing sequence.

Step 2: Assemble and Measure Assembly Size

After assembly, record the total assembled length. Compare this value against the GenomeScope estimate. A complete assembly should be close to the estimated genome size. The acceptable difference depends on the organism. Highly repetitive genomes may assemble to less than the estimated size because repeats collapse during assembly. Highly heterozygous genomes may assemble to more than the haploid estimate because both haplotypes are represented.

The Miniopterus schreibersii bat genome illustrates the haplotype issue. The assembly contains two haplotypes with total lengths of 1,850.57 megabases and 1,699.63 megabases. A haploid genome estimate would be approximately 1.7 to 1.9 gigabases, and the assembly correctly represents both haplotypes. A researcher comparing the total assembly size against a haploid estimate would see an apparent excess, not a deficit.

Step 3: Run BUSCO and Record the Gene Completeness Score

Run BUSCO on the assembly using the appropriate lineage dataset for the organism. Record the complete, fragmented, and missing percentages. This value is the gene-based completeness metric.

For the Rhinolophus hipposideros lesser horseshoe bat, gene annotation identified 18,700 protein-coding genes. The BUSCO assessment for this assembly would be expected to show high completeness because the assembly is chromosome-scale and generated by the Darwin Tree of Life project. The BUSCO score alone, however, would not reveal whether the assembly captured all repetitive content.

Step 4: Run Merqury and Record the K-mer Completeness Value

Run Merqury with the assembly and the raw reads. The output includes a completeness value that represents the fraction of read k-mers present in the assembly. Record this value alongside the BUSCO score.

The comparison between the two values is the diagnostic step. A large gap between BUSCO completeness and Merqury completeness indicates missing sequence that is not genic. The missing sequence is most likely repetitive because BUSCO genes are present while total sequence content is not.

Step 5: Compare Repeat Content Estimates

If the assembly includes a repeat annotation, compare the annotated repeat fraction against the GenomeScope repeat estimate. The Paracondylactis sinensis genome reports repetitive elements at 26.43% of the genome and transposable elements at 20.47%. If the read-based estimate had predicted a higher repeat fraction, the assembly would be missing repeat copies.

This comparison is particularly important for species with recent repeat expansions. Young transposable element insertions are often highly similar to each other, and assemblers may collapse them into a single consensus sequence. The result is an assembly that contains the repeat sequence but at a fraction of its true copy number.

Step 6: Decide Whether Additional Work Is Required

The decision to accept an assembly or pursue additional sequencing depends on the gap between the metrics. A small gap between BUSCO and Merqury completeness, with assembly size close to the read-based estimate, supports acceptance. A large gap indicates that the assembly is missing sequence content, and the missing content is likely repetitive.

The decision also depends on the intended use of the assembly. A gene-focused analysis may tolerate missing repeats because the genes are present. A comparative genomics analysis of transposable elements, a centromere study, or a genome architecture investigation requires the missing repeats and would be compromised by their absence.

Options and Tradeoffs for Recovering Missing Repeats

When the assessment reveals missing repeats, several options exist. Each has distinct costs, benefits, and limitations.

Additional Long-Read Sequencing

The most direct solution is additional sequencing. Missing repeats are often absent because coverage was insufficient to assemble them. Long reads that span entire repeat arrays can resolve their structure. Increasing coverage from 30x to 50x or higher can provide the depth needed to assemble repetitive regions that were previously collapsed or fragmented.

The tradeoff is cost. High-coverage long-read sequencing remains expensive, particularly for large genomes. The decision to add sequencing should be based on the magnitude of the completeness gap and the biological importance of the missing repeats.

Alternative Assembly Strategies

Some assemblers handle repeats better than others. If the initial assembly used a particular assembler, testing an alternative may recover additional repeat content. The choice of assembler interacts with read type, coverage, and genome characteristics.

The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including assembly workflows. Using a standardized pipeline can help ensure that assembly parameters are appropriate and that results are reproducible across runs.

Hi-C or Optical Mapping for Scaffolding

If the missing repeats are present but fragmented, Hi-C or optical mapping data can help place them within the assembly. The Paracondylactis sinensis genome used Hi-C data to anchor 93.44% of the assembly to 19 pseudo-chromosomes. This approach does not recover missing sequence, but it does improve the structural completeness of the assembly.

Accepting the Assembly with Documented Limitations

For some projects, the missing repeats are acceptable. The assembly may be sufficient for gene discovery, phylogenetic analysis, or variant calling in genic regions. The decision to accept should be documented, and the limitations should be reported in any publication or database submission.

The NCBI Data Resources provide the primary repository for genome assemblies. When submitting an assembly with known missing repeats, the submission should include the relevant quality metrics so that downstream users can interpret the assembly appropriately.

Records and Measurements for Completeness Assessment

Maintaining systematic records is essential for reproducibility and for defending assembly quality decisions. The following records should be kept for every assembly project.

Read Set Records

Record the sequencing platform, read length, coverage, and quality metrics for each read set. Include the k-mer-based genome size estimate and the repeat content estimate. These values establish the expected properties of the genome before assembly.

Assembly Records

Record the assembler name and version, the parameter settings, and the assembly statistics including total length, contig N50, scaffold N50, and GC content. Record the BUSCO lineage dataset and version, and the complete, fragmented, and missing percentages.

K-mer Comparison Records

Record the Merqury completeness value, the k-mer size used, and the read set used for the comparison. Record the GenomeScope parameters and the resulting genome size, heterozygosity, and repeat content estimates.

Decision Records

Record the rationale for accepting the assembly or pursuing additional work. Include the specific metric values that informed the decision. This record is valuable when the assembly is published or submitted to a database, because it provides context for the reported quality metrics.

Common Failure Patterns in Completeness Assessment

Several recurring patterns indicate problems in completeness assessment. Recognizing these patterns helps avoid incorrect conclusions.

High BUSCO with Low K-mer Completeness

This pattern indicates missing repetitive sequence. The BUSCO genes are present, but a substantial fraction of the read k-mers are absent from the assembly. The gap between the two metrics quantifies the missing content.

This pattern is common in assemblies generated from moderate-coverage short reads, where repetitive regions are difficult to assemble. It can also occur in long-read assemblies when coverage is insufficient to resolve highly repetitive arrays.

Assembly Size Below the Read-Based Estimate

When the assembly is substantially smaller than the GenomeScope estimate, sequence is missing. The missing sequence may be repetitive, or it may represent haplotypes that were collapsed. The distinction matters for interpretation.

For a haploid assembly, the size should be close to the haploid genome estimate. For a diploid assembly that represents both haplotypes, the size should be approximately twice the haploid estimate. The Colias erate butterfly assembly contains two haplotypes at 324.54 and 323.68 megabases, consistent with a haploid genome of approximately 324 megabases represented in both haplotypes.

Assembly Size Above the Read-Based Estimate

An assembly larger than the read-based estimate may contain contamination or may represent a genuine underestimate by the k-mer method. K-mer-based genome size estimates can be inaccurate for genomes with very high heterozygosity or very high repeat content. The Galaxy Training Network provides accessible tutorials on genome assembly and quality assessment that cover these interpretation issues.

BUSCO Fragmented but K-mer Complete

This pattern indicates that genes are present but fragmented across multiple contigs or scaffolds. The sequence content is present, but the assembly structure is incomplete. This pattern is less severe than missing sequence, but it still affects gene annotation and downstream analysis.

Limitations of K-mer-Based Completeness Assessment

K-mer methods have their own limitations, and these must be understood to interpret results correctly.

Coverage Variation

K-mer spectrum methods assume relatively uniform sequencing coverage. If coverage varies substantially across the genome, the spectrum becomes distorted. Regions with very low coverage may produce k-mers that fall below the detection threshold, leading to an underestimate of completeness. Regions with very high coverage may produce k-mers that are incorrectly classified as repetitive.

Sequencing Error

Sequencing errors create k-mers that are not present in the true genome. These error k-mers are typically present at low frequency and can be filtered by setting a minimum frequency threshold. However, aggressive filtering can also remove legitimate low-coverage k-mers, particularly in regions with low sequencing depth.

K-mer Size Selection

The choice of k affects the results. Small k values produce more k-mers but increase the chance of spurious matches. Large k values reduce spurious matches but may miss short repeats. The optimal k depends on the read length and the genome composition. Most tools provide guidance on k selection, and the EMBL-EBI Training resources include practical instruction on k-mer-based analysis.

Heterozygosity Confusion

In highly heterozygous genomes, k-mers from the two haplotypes appear at different frequencies. This can complicate the interpretation of the k-mer spectrum and lead to incorrect genome size estimates. The Miniopterus schreibersii and Rhinolophus hipposideros assemblies both contain two haplotypes, and their k-mer spectra would reflect this heterozygosity.

Interpreting Completeness in the Context of Assembly Purpose

The acceptable level of completeness depends on the biological questions the assembly is intended to support. A single completeness threshold does not apply across all projects.

Gene-Focused Analyses

For gene discovery, phylogenomics, or variant calling in coding regions, BUSCO completeness is the primary metric. Missing repeats are less problematic because they do not affect the genes. An assembly with high BUSCO completeness and moderate k-mer completeness may be sufficient for these purposes.

Repeat-Focused Analyses

For transposable element biology, centromere evolution, or satellite DNA studies, k-mer completeness is the critical metric. Missing repeats directly compromise these analyses. An assembly with high BUSCO completeness but low k-mer completeness is inadequate for repeat-focused research.

Comparative Genomics

Comparative analyses across species require consistent completeness. If one species has a complete assembly and another has missing repeats, the comparison is biased. The missing repeats in one assembly will appear as lineage-specific losses, which is a false biological conclusion.

Reference Genome Production

Reference-quality assemblies, such as those produced by the Darwin Tree of Life project, require both high BUSCO completeness and high k-mer completeness. The Trypoxylon clavicerum, Miniopterus schreibersii, Rhinolophus hipposideros, and Colias erate assemblies are all generated within this framework, and their quality assessments include both gene-based and sequence-based metrics.

Professional Escalation Criteria

Certain findings should trigger consultation with a bioinformatics specialist or a genome assembly expert. The following situations warrant escalation.

Large Completeness Gap

If the Merqury completeness is substantially below the BUSCO completeness, and the gap exceeds what can be explained by known limitations of the methods, consult an expert. The gap may indicate a systematic assembly problem that requires a different approach.

Unexpected Genome Size

If the assembly size differs dramatically from the read-based estimate, and the difference cannot be explained by heterozygosity or known repeat content, escalate. The discrepancy may indicate contamination, sample mix-up, or a fundamental problem with the read set.

Inconsistent Results Across Tools

If different k-mer-based tools produce conflicting estimates, escalate. The conflict may indicate a problem with the read set, the k-mer parameters, or the underlying assumptions of the tools.

Assembly Failure in Known Repeat Regions

If the assembly is missing specific repeat families that are known to be present in the species or related species, escalate. The missing repeats may require specialized assembly strategies or additional sequencing.

Reproducibility and Training Considerations

Completeness assessment should be reproducible. The nf-core documentation describes community standards for pipeline usage and configuration that support reproducibility. Using version-controlled pipelines and recording all parameters ensures that the assessment can be repeated and verified.

The Bioconductor project provides R packages for genomic analysis, including tools for working with assembly and read data. These packages support reproducible analysis workflows and provide documentation for their use.

Training is available through multiple sources. The Galaxy Training Network offers accessible tutorials on genome assembly and quality assessment. The EMBL-EBI Training provides learning pathways for bioinformatics data resources. The Carpentries lessons cover foundational computing skills that support reproducible analysis.

Safety and Data Management Context

Genome assembly projects involve large data volumes and require careful data management. Raw sequencing reads and assemblies should be stored in appropriate repositories. The NCBI Data Resources provide the primary archive for sequence data and genome assemblies.

Data management decisions affect reproducibility. Record the location of all raw data, intermediate files, and final assemblies. Document the software versions and parameters used at each step. This documentation ensures that the completeness assessment can be reproduced and that the assembly can be updated if better methods become available.

A Decision Framework for Triaging Assemblies with Missing Repeats

When the combined BUSCO and k-mer assessment reveals a completeness gap, the next question is practical: what do you do with this assembly now? The answer depends on the magnitude of the gap, the biological purpose of the assembly, and the resources available for additional sequencing or assembly work. This section provides a structured decision framework that converts the metric values into concrete management actions, with explicit thresholds for triage, escalation, and documentation.

The Three-Tier Triage System

The decision framework uses three tiers based on the relationship between BUSCO completeness, Merqury completeness, and the assembly size relative to the read-based genome size estimate. Each tier leads to a distinct set of actions.

Tier 1: Accept and Proceed

An assembly falls into this tier when the Merqury completeness is within five percentage points of the BUSCO completeness, and the assembly size is within ten percent of the read-based genome size estimate. This pattern indicates that the assembly captures both the genic and repetitive content of the genome at comparable rates. The assembly is suitable for most downstream analyses, including gene annotation, comparative genomics, and repeat family characterization.

For the Paracondylactis sinensis sea anemone genome, the reported 95.91% BUSCO completeness and 26.43% repeat content would need to be paired with a Merqury completeness value in the low 90s to qualify for Tier 1. The assembly size of 210.63 Mb would need to be close to the k-mer-based genome size estimate from the raw reads.

Tier 2: Conditional Acceptance with Documented Limitations

An assembly falls into this tier when the Merqury completeness is five to fifteen percentage points below the BUSCO completeness, or when the assembly size differs from the read-based estimate by ten to twenty percent. This pattern indicates that the assembly is missing a measurable fraction of sequence content, and that the missing content is likely repetitive.

Conditional acceptance requires three actions. First, document the specific metric values and the calculated gap. Second, identify which repeat families are underrepresented by comparing the repeat annotation in the assembly against the repeat content predicted from the reads. Third, state clearly in any publication or database submission that the assembly is incomplete for repetitive regions and specify which analyses are and are not supported.

A gene-focused analysis can proceed with a Tier 2 assembly because the BUSCO genes are present. A transposable element census or a centromere study cannot proceed because the missing repeats would produce false negative results.

Tier 3: Escalate and Reassemble

An assembly falls into this tier when the Merqury completeness is more than fifteen percentage points below the BUSCO completeness, or when the assembly size differs from the read-based estimate by more than twenty percent. This pattern indicates substantial missing sequence that will compromise most downstream analyses.

Tier 3 requires escalation to a bioinformatics specialist or genome assembly expert. The assembly should not be used for publication or database submission until the missing sequence is addressed. The escalation should include the complete record of metrics, the read set information, and the assembly parameters so that the specialist can diagnose the cause and recommend a path forward.

The Decision Matrix for Assembly Acceptance

The following decision matrix translates the metric relationships into specific actions. The matrix uses the gap between BUSCO and Merqury completeness as the primary axis and the assembly size relative to the read-based estimate as the secondary axis.

Completeness GapAssembly Size vs Read EstimateDecisionRequired Actions
Less than 5%Within 10%AcceptStandard quality reporting
Less than 5%10-20% largerInvestigate heterozygosityConfirm ploidy and haplotype representation
Less than 5%10-20% smallerInvestigate repeat collapseCompare repeat annotations against read estimates
5-15%Within 10%Conditional acceptDocument limitations and identify missing repeat families
5-15%10-20% differentConditional accept with cautionAdditional validation before publication
More than 15%AnyEscalateConsult specialist and plan additional sequencing
AnyMore than 20% differentEscalateVerify read set and assembly parameters

The matrix accounts for the fact that a small completeness gap with a size discrepancy points to a different problem than a large completeness gap with a normal size. The former suggests collapsed repeats or haplotype issues, while the latter suggests missing sequence content.

Applying the Framework to Published Assemblies

The framework can be applied retrospectively to published assemblies to illustrate how the decision process works in practice.

**The Trypoxylon clavicerum Wasp Assembly**

This assembly reached 270.89 Mb with 85.18% of the sequence scaffolded into 9 chromosomal pseudomolecules. The Darwin Tree of Life project generated this assembly as part of a program that produces reference genomes for eukaryotic species. The project standards require both high BUSCO completeness and high k-mer completeness.

Under the decision framework, this assembly would be evaluated by comparing its BUSCO score against its Merqury completeness value. If the gap is less than five percentage points, the assembly qualifies for Tier 1 acceptance. The chromosome-scale scaffolding provides additional confidence in the structural completeness of the assembly.

**The Miniopterus schreibersii Bat Assembly**

This assembly contains two haplotypes at 1,850.57 Mb and 1,699.63 Mb, with 97.59% of haplotype 1 scaffolded into 24 chromosomal pseudomolecules. The two-haplotype structure requires careful interpretation of the assembly size relative to the read-based estimate.

A haploid genome estimate for this species would be approximately 1.7 to 1.9 Gb. The assembly contains both haplotypes, so the total assembly size is approximately twice the haploid estimate. The decision framework must account for this by comparing each haplotype separately against the read-based estimate, instead of comparing the total assembly size.

**The Colias erate Butterfly Assembly**

This assembly contains two haplotypes at 324.54 Mb and 323.68 Mb, with 99.67% of haplotype 1 scaffolded into 31 chromosomal pseudomolecules. The close agreement between the two haplotype sizes suggests that the assembly is complete and that both haplotypes are represented at similar coverage.

Under the decision framework, this assembly would qualify for Tier 1 acceptance if the Merqury completeness is within five percentage points of the BUSCO completeness. The high scaffold percentage and the consistent haplotype sizes provide additional evidence of completeness.

Record Keeping for the Decision Framework

The decision framework requires systematic record keeping to support the triage decision and to provide evidence for publication or database submission. The following record template captures the essential information for each assembly.

Assembly Identification Record

Record the species name, the sample identifier, the sequencing platform, the read length, and the coverage. Record the assembler name and version, the assembly date, and the assembly version. This information establishes the provenance of the assembly and supports reproducibility.

Metric Comparison Record

Record the BUSCO completeness percentage, the BUSCO lineage dataset and version, and the BUSCO run date. Record the Merqury completeness percentage, the k-mer size used, and the read set used for the comparison. Record the GenomeScope genome size estimate, the heterozygosity estimate, and the repeat content estimate.

Decision Record

Record the tier assignment, the calculated gap between BUSCO and Merqury completeness, and the assembly size relative to the read-based estimate. Record the rationale for the decision, including any biological considerations that influenced the triage. Record the date of the decision and the name of the person making the decision.

Action Record

Record the actions taken after the decision. For Tier 1, record the publication or database submission details. For Tier 2, record the documented limitations and the specific analyses that are and are not supported. For Tier 3, record the escalation details and the plan for additional sequencing or reassembly.

Common Failure Patterns in the Decision Process

Several recurring patterns indicate problems in the decision process itself. Recognizing these patterns helps avoid incorrect triage decisions.

Overweighting BUSCO Completeness

The most common failure is accepting an assembly based on a high BUSCO score without checking the Merqury completeness. This pattern produces assemblies that pass the standard quality gate while missing substantial repetitive content. The decision framework prevents this failure by requiring both metrics to be recorded and compared.

Misinterpreting Assembly Size Discrepancies

A second failure pattern is misinterpreting the assembly size relative to the read-based estimate. An assembly that is larger than the estimate may be interpreted as complete when it actually contains both haplotypes. An assembly that is smaller than the estimate may be interpreted as incomplete when it actually represents a haploid assembly of a highly repetitive genome.

The decision framework addresses this by requiring the ploidy and haplotype structure to be recorded alongside the size comparison. The Miniopterus schreibersii and Colias erate assemblies illustrate the importance of this context.

Ignoring Repeat Family Composition

A third failure pattern is treating all missing repeats as equivalent. The biological impact of missing repeats depends on which families are absent. Missing satellite repeats affect centromere studies. Missing transposable elements affect genome evolution analyses. Missing ribosomal DNA arrays affect functional genomics.

The decision framework addresses this by requiring the repeat family composition to be examined when the completeness gap exceeds five percentage points. The comparison between the assembly repeat annotation and the read-based repeat estimate identifies which families are underrepresented.

Failing to Document Limitations

A fourth failure pattern is publishing an assembly with known missing repeats without documenting the limitations. This pattern misleads downstream users who may not realize that the assembly is incomplete for repetitive regions.

The decision framework addresses this by requiring documentation for Tier 2 assemblies. The documentation must state the specific metric values, the calculated gap, and the analyses that are and are not supported.

Escalation Criteria and Specialist Consultation

The decision framework includes explicit escalation criteria that trigger consultation with a bioinformatics specialist or genome assembly expert. The following situations require escalation regardless of the tier assignment.

Conflicting Metric Values

If different k-mer-based tools produce conflicting estimates of genome size or completeness, escalate. The conflict may indicate a problem with the read set, the k-mer parameters, or the underlying assumptions of the tools. A specialist can diagnose the cause and recommend appropriate parameters.

Unexpected Repeat Content

If the read-based repeat content estimate is substantially higher than the repeat content annotated in the assembly, and the difference cannot be explained by known limitations of the annotation method, escalate. The discrepancy may indicate that the assembly is missing specific repeat families that require specialized assembly strategies.

Assembly Failure in Known Repeat Regions

If the assembly is missing specific repeat families that are known to be present in the species or related species, escalate. The missing repeats may require additional sequencing, alternative assemblers, or specialized repeat-aware assembly methods.

Ploidy Confusion

If the assembly structure is inconsistent with the expected ploidy of the species, escalate. The inconsistency may indicate sample contamination, sample mix-up, or a fundamental problem with the read set.

Training and Reproducibility for the Decision Framework

The decision framework should be applied consistently across assembly projects within a research group or consortium. Consistency requires training and reproducible workflows.

The Galaxy Training Network provides accessible tutorials on genome assembly and quality assessment that cover the metrics used in this framework. The EMBL-EBI Training resources include practical instruction on k-mer-based analysis and genome assembly evaluation.

The nf-core documentation describes community standards for reproducible bioinformatics pipelines. Using a standardized pipeline for completeness assessment ensures that the metrics are calculated consistently and that the results can be compared across projects.

The Bioconductor project provides R packages for genomic analysis that support reproducible workflows. These packages can be used to automate the metric comparison and the decision record generation.

The Carpentries lessons cover foundational computing skills that support reproducible analysis. These skills include shell scripting, version control with Git, and data management practices that are essential for maintaining the records required by the decision framework.

Integrating the Decision Framework into Project Management

The decision framework should be integrated into the project management plan for any genome assembly project. The framework provides a structured way to evaluate assembly quality at each stage of the project and to decide when additional sequencing or assembly work is required.

The framework should be applied at three points in the project lifecycle. First, after the initial assembly is generated, to determine whether the assembly meets the quality standards for the project. Second, after any additional sequencing or reassembly work, to determine whether the revised assembly addresses the identified gaps. Third, before publication or database submission, to document the final quality assessment and any remaining limitations.

The NCBI Data Resources provide the primary repository for genome assemblies. When submitting an assembly, include the completeness metrics and the decision record so that downstream users can interpret the assembly quality appropriately.

The decision framework does not replace the need for expert judgment. The tier assignments and escalation criteria provide structure, but the final decision should always consider the specific biological questions the assembly is intended to support. A Tier 2 assembly may be perfectly adequate for a gene-focused study, while a Tier 1 assembly may be inadequate for a repeat-focused study if the repeat content is not fully characterized.

Frequently Asked Questions

Why does my assembly have a high BUSCO score but a low Merqury completeness value?

The BUSCO score measures the presence of conserved single-copy genes, which are mostly located in gene-rich regions. The Merqury completeness value measures the fraction of read k-mers present in the assembly, which includes all sequence content. A large gap between the two indicates that the assembly is missing sequence that is not genic, and that missing sequence is most likely repetitive. The repeats may be absent because they were collapsed during assembly or because coverage was insufficient to assemble them.

What is the difference between GenomeScope and Merqury?

GenomeScope analyzes the k-mer spectrum of raw reads to estimate genome size, heterozygosity, and repeat content before assembly. Merqury compares an assembly against raw reads to measure the fraction of read k-mers present in the assembly. GenomeScope provides expectations for the genome, while Merqury assesses the assembly against those expectations.

How much of a gap between BUSCO and Merqury completeness is acceptable?

There is no universal threshold. The acceptable gap depends on the organism, the assembly purpose, and the repeat content of the genome. A small gap of a few percent may be acceptable for many purposes. A large gap of ten percent or more indicates substantial missing sequence and should be investigated. The decision should be based on the biological questions the assembly is intended to support.

Can missing repeats be recovered by polishing the assembly?

Polishing corrects base errors but does not recover missing sequence. If repeats are absent from the assembly, polishing will not add them. Recovering missing repeats requires additional sequencing, a different assembly strategy, or both.

How do I know if my assembly is missing repeats or if the repeats were never in the reads?

Compare the repeat content estimate from GenomeScope against the repeat content annotated in the assembly. If the read-based estimate is higher, the assembly is missing repeats that were present in the reads. If the estimates are similar, the repeats were likely not present in the read set, which indicates a sequencing coverage problem.

What should I do if my assembly is missing repeats that I need for my analysis?

The first step is to document the gap using the metrics described in this article. The second step is to decide whether additional sequencing is justified. If the missing repeats are essential for the analysis, additional long-read sequencing is the most direct solution. If the repeats are not essential, the assembly may be sufficient with documented limitations.

How do I report completeness metrics in a genome paper?

Report both BUSCO and k-mer-based completeness values. Include the assembly size, the read-based genome size estimate, and the repeat content estimates. Describe the gap between the metrics and explain whether the gap is acceptable for the intended use of the assembly. This reporting standard allows readers to interpret the assembly quality appropriately.

Are there standardized pipelines for completeness assessment?

The nf-core documentation describes community pipelines that include assembly quality assessment. The Galaxy Training Network provides tutorials that cover completeness assessment workflows. These resources support reproducible assessment and reduce the risk of methodological errors.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.