# Why Your Assembly Fails BUSCO: Troubleshooting Incomplete and Fragmented Genomes


## Key Takeaways

- BUSCO failure, indicated by low completeness or high fragmentation, is rarely due to a single error but stems from one of five root causes: contaminated input data, high heterozygosity, repetitive content, insufficient sequencing depth, or incorrect workflow parameters.
- Contamination is a primary driver of missing BUSCOs, detectable by aligning contigs to NCBI databases or using k-mer based taxonomic classification on raw reads; corrective action involves filtering reads by taxonomic classification prior to reassembly.
- High heterozygosity leads to fragmented BUSCOs due to haplotype divergence, diagnosed by k-mer spectrum analysis showing dual peaks; resolution involves using haplotype-aware assemblers or purging haplotigs.
- Repetitive content causes fragmentation by breaking contigs, identifiable through k-mer based repeat content estimation; increasing sequencing depth or switching to longer reads (e.g., PacBio HiFi, Oxford Nanopore) is crucial for resolution.
- Insufficient sequencing depth results in missing or fragmented BUSCOs, directly impacting the assembler's ability to resolve genome structure; resequencing to achieve adequate coverage is the primary corrective action.
- Incorrect workflow parameters, such as inappropriate k-mer sizes or coverage cutoffs, can lead to poor assemblies; reviewing and adjusting these parameters based on assembler recommendations and genome characteristics is essential.

---

BUSCO (Benchmarking Universal Single-Copy Orthologs) scores are the most widely used metric for assessing genome assembly completeness. When your assembly returns low completeness scores, the cause is rarely a single error. The failure typically traces to one of five root causes: contaminated input data, high heterozygosity, repetitive content, sequencing depth problems, or incorrect workflow parameters. This article provides a systematic diagnostic path to identify which cause applies to your data and the specific corrective actions available at each stage.

The diagnostic approach presented here follows the evidence available from published genome assembly projects across diverse taxa, including a chromosome-level plant genome that achieved 98.4 percent BUSCO completeness, a single-specimen insect genome assembled from ultra-low input DNA, and a classroom-based mammalian genome project using long-read sequencing. These cases demonstrate that successful assembly outcomes depend on matching your workflow to the biological characteristics of your organism and the limitations of your sequencing platform.

## At a Glance: BUSCO Failure Diagnosis and Response

| Symptom Pattern | Most Likely Cause | Primary Diagnostic Check | First Corrective Action |
|---|---|---|---|
| Low completeness with many missing BUSCOs | Contamination or incorrect taxonomic lineage | Run BUSCO against multiple lineage datasets and inspect contig taxonomy | Filter reads with taxonomic classification before reassembly |
| High fragmentation with many fragmented BUSCOs | Repetitive content or insufficient read depth | Check k-mer spectra and repeat content estimates | Increase sequencing depth or switch to longer reads |
| Low completeness in a heterozygous organism | Haplotype divergence splitting orthologs | Examine k-mer peak patterns for heterozygosity | Use haplotype-aware assemblers or purge haplotigs |
| Low completeness after polishing | Over-polishing collapsed repeats | Compare pre and post polishing BUSCO scores | Revert to pre-polish assembly or reduce polishing rounds |
| Complete but low single-copy with high duplicated | Uncollapsed haplotypes or recent polyploidy | Check duplication ratio and k-mer spectra | Run haplotype purging tools |

## Understanding What BUSCO Actually Measures

BUSCO evaluates assembly completeness by searching for a set of single-copy orthologs that are expected to be present in nearly all genomes within a particular taxonomic group. The method compares your assembly against these conserved gene sets and classifies each expected gene as complete, fragmented, duplicated, or missing. The resulting percentages provide a standardized measure of how much of the expected gene content is present in your assembly.

The interpretation of BUSCO results depends on the lineage dataset you select. NCBI maintains comprehensive taxonomic databases that define the expected gene sets for different evolutionary groups, and the choice of lineage directly affects your results. If you run BUSCO with the wrong lineage, such as using a vertebrate dataset for an insect genome, you will obtain misleading completeness scores because the expected gene set does not match your organism. The NCBI taxonomy resources provide the framework for selecting appropriate lineage datasets, and cross-referencing your organism's taxonomic position against these databases is the first step in interpreting any BUSCO result.

BUSCO scores reflect the presence of conserved single-copy genes, not the overall quality of every region in your genome. A genome can achieve high BUSCO completeness while still containing misassemblies in repetitive regions, structural errors, or incorrect gene models. Conversely, a genome with low BUSCO scores may still be useful for certain analyses if the missing genes are not relevant to your research question. The metric is best understood as a standardized completeness check, not a comprehensive quality assessment.

The relationship between BUSCO completeness and assembly contiguity is important to understand. A highly fragmented assembly with a low contig N50 can still achieve high BUSCO completeness if the conserved genes are intact within individual contigs. The watering-pot shell genome assembly, for example, achieved 96.5 percent anchoring onto pseudochromosomes with a contig N50 of 5.33 megabases, demonstrating that chromosome-scale scaffolding can be achieved when the underlying contigs are complete. Conversely, an assembly with long contigs can have low BUSCO completeness if conserved genes are missing due to assembly errors or gaps.

## The Five Root Causes of BUSCO Failure

### Contamination and Its Detection

Contamination is the most common cause of unexpectedly low BUSCO completeness, particularly in assemblies generated from environmental samples, clinical specimens, or samples processed in shared laboratory facilities. Foreign DNA from bacteria, fungi, or other organisms dilutes the reads from your target species, and the assembler may incorporate contaminant sequences into your assembly or fail to assemble your target genome to sufficient depth.

The first indication of contamination often appears in the BUSCO results themselves. If you observe a pattern where some expected genes are completely missing instead of fragmented, contamination is a likely explanation. The assembler may have spent coverage on contaminant DNA, leaving insufficient depth to assemble regions of your target genome. Alternatively, contaminant contigs may be present in your assembly, and these contigs will not contain the expected BUSCO genes for your target lineage.

Detection of contamination requires examining the taxonomic composition of your reads and contigs. Several approaches are available. You can align your assembled contigs against the NCBI nucleotide database using BLAST searches to identify sequences that match unexpected taxa. You can also use k-mer based approaches to identify sequences that do not match the expected genome composition of your target species. The NCBI sequence databases provide the reference data for these comparisons, and the search systems available through NCBI allow you to identify the taxonomic origin of individual contigs.

The TaxaScope workstation demonstrates a practical approach to this problem by integrating genome quality assessment tools within a graphical interface that does not require command-line expertise. The platform uses containerized execution environments to provide standardized analysis workflows, and it includes tools for genome quality assessment that can identify contamination. This approach is particularly valuable for laboratory scientists who need to assess genome quality without extensive bioinformatics training.

Once contamination is identified, the corrective action is to filter your reads before reassembly. Taxonomic classification tools can assign each read to a taxonomic group, and you can remove reads that do not match your target species. This filtering step is most effective when performed before assembly, as it prevents contaminant sequences from consuming assembly coverage. After filtering, you should reassemble and rerun BUSCO to confirm that completeness improves.

### Heterozygosity and Haplotype Divergence

Heterozygous organisms present a fundamental challenge to genome assembly. In a diploid genome, the two parental haplotypes differ at many positions, and these differences can cause assemblers to produce separate contigs for each haplotype. The result is an assembly that contains both haplotypes as separate sequences, which BUSCO detects as duplicated genes instead of single-copy genes.

The k-mer spectrum provides the most direct diagnostic for heterozygosity. When you plot the frequency of k-mers in your sequencing data, a homozygous genome produces a single major peak, while a heterozygous genome produces two peaks: one at lower coverage representing heterozygous k-mers and one at higher coverage representing homozygous k-mers. The position and shape of these peaks indicate the level of heterozygosity and the appropriate assembly strategy.

The Rhododendron simsii genome project provides a relevant example of how heterozygosity affects assembly outcomes. The study reported incomplete divergence among coastal-island populations with moderate genetic diversity and gene flow, and the chromosome-level assembly achieved 98.4 percent BUSCO completeness. This result demonstrates that high-quality assembly is possible in organisms with population structure and genetic diversity, but it requires appropriate assembly strategies that account for heterozygosity.

For highly heterozygous organisms, several assembly strategies are available. Haplotype-aware assemblers such as those that use the HiFi error-corrected long-read data can separate haplotypes during assembly. Alternatively, you can assemble without haplotype separation and then use haplotype purging tools to remove alternative haplotigs from the assembly. The choice between these approaches depends on your research goals. If you need a single representative haplotype, purging is appropriate. If you need both haplotypes for population genomics or haplotype-resolved analyses, you should use haplotype-aware assembly.

The single-specimen Culicoides stellifer genome assembly demonstrates that even challenging heterozygous insect genomes can be assembled from minimal input material. The study used an ultra-low input DNA protocol with PacBio sequencing and produced a 119 megabase assembly with a contig N50 of 479.3 kilobases. The assembly contained 11 percent repeat sequences and 18,895 annotated protein-coding genes, and the workflow specifically considered the challenges of insect genome assembly, including heterozygosity and repetitive content.

### Repetitive Content and Assembly Fragmentation

Repetitive sequences pose a distinct challenge to genome assembly. When a genome contains many copies of similar or identical sequences, the assembler cannot determine the correct arrangement of these repeats, leading to breaks in the assembly or incorrect joins. The result is fragmentation, where BUSCO genes that span repeat boundaries are broken into multiple contigs and classified as fragmented instead of complete.

The repeat content of a genome varies widely across taxa. The Culicoides stellifer genome contained 11 percent repeat sequences, while many plant and vertebrate genomes contain 50 percent or more repetitive content. The watering-pot shell genome, at approximately 507 megabases, required chromosome-scale assembly to resolve its repetitive regions, and the study anchored 96.5 percent of sequences onto 19 pseudochromosomes.

Diagnosing repeat-related BUSCO failure requires estimating the repeat content of your genome before assembly. K-mer based genome size estimation tools can provide an estimate of genome size and repeat content from your sequencing data. If the estimated genome size is substantially larger than the assembled size, you likely have collapsed repeats in your assembly. If the assembly is fragmented with a low N50, you likely have unresolved repeats that broke the assembly.

The corrective actions for repeat-related failures depend on your sequencing platform. Short-read assemblies are most susceptible to repeat-induced fragmentation because the reads cannot span long repetitive regions. Long-read sequencing, particularly with Oxford Nanopore or PacBio platforms, provides reads that can span many repetitive regions, reducing fragmentation. The Przewalski's horse classroom project demonstrated this approach by using Oxford Nanopore long-read sequencing to produce a high-quality genome with only $4,000 of materials. The resulting genome statistics far exceeded the previous assembly and matched the quality of the domestic horse reference genome.

For genomes with very high repeat content, additional strategies may be necessary. These include using ultra-long reads to span the largest repeats, incorporating optical mapping or Hi-C data for scaffolding, and using assembly algorithms specifically designed for repetitive genomes. The choice of strategy depends on the repeat content of your genome and the resources available for your project.

### Sequencing Depth and Coverage Issues

Insufficient sequencing depth is a common cause of BUSCO failure, particularly in projects with limited budgets or challenging samples. When coverage is too low, the assembler cannot resolve the genome structure, leading to gaps and fragmentation. The result is an assembly with missing or fragmented BUSCO genes.

The relationship between sequencing depth and assembly quality is not linear. Below a certain depth threshold, assembly quality degrades rapidly because the assembler lacks sufficient evidence to make correct joins. Above the threshold, additional depth provides diminishing returns. The optimal depth depends on the genome size, repeat content, heterozygosity, and sequencing platform.

Estimating the required sequencing depth requires knowing your genome size. Flow cytometry, k-mer analysis, or published estimates for related species can provide this information. Once you know the genome size, you can calculate the coverage provided by your sequencing data and determine whether you have sufficient depth for assembly.

The Przewalski's horse classroom project provides a practical example of managing sequencing depth on a limited budget. The project used Oxford Nanopore long-read sequencing with $4,000 of materials and produced a genome on par with the domestic horse reference. This outcome demonstrates that sufficient depth for a mammalian genome can be achieved with modest resources when the workflow is carefully designed.

If you have already assembled your genome and obtained low BUSCO scores due to insufficient depth, the corrective action is to sequence additional data and reassemble. Adding more sequencing data to an existing assembly through polishing or gap-filling is generally less effective than reassembling with the complete dataset. The reassembly should use all available reads to maximize coverage and assembly quality.

### Workflow Parameter Errors

The final common cause of BUSCO failure is incorrect workflow parameters. Assembly tools have many parameters that affect the outcome, and default parameters are not optimal for all genomes. Incorrect k-mer sizes, coverage cutoffs, or error correction settings can all lead to poor assemblies with low BUSCO scores.

The k-mer size is particularly important for short-read assemblies. Small k-mers provide more coverage but are more affected by repeats, while large k-mers provide less coverage but are more specific. The optimal k-mer size depends on the sequencing depth and genome complexity. Many assembly tools include k-mer optimization features that test multiple k-mer sizes and select the best assembly, but these features increase computational cost.

Coverage cutoffs are another common source of errors. Many assemblers use coverage-based filtering to remove low-coverage sequences, which are often contaminants or sequencing errors. However, if the coverage cutoff is set too high, legitimate genomic regions with low coverage may be removed, leading to missing BUSCO genes. Conversely, if the cutoff is set too low, contaminants may be retained.

The nf-core documentation provides guidance on reproducible workflow configuration, and the community standards emphasize the importance of documenting and versioning workflow parameters. Reproducible workflows ensure that assembly parameters are recorded and can be revisited if problems are identified. The Galaxy Training Network similarly provides accessible training on assembly workflows that emphasize parameter selection and quality assessment.

If you suspect parameter errors, the corrective action is to review your assembly parameters against the recommendations for your assembler and genome type. The Galaxy Training Network and EMBL-EBI Training provide practical guidance on assembly parameter selection, and the Bioconductor project offers packages for assembly quality assessment that can help identify parameter-related problems.

## Diagnostic Workflow for BUSCO Failure

### Step 1: Verify Your BUSCO Lineage Selection

The first diagnostic step is to confirm that you used the correct BUSCO lineage dataset. The lineage determines which genes are expected in your assembly, and using the wrong lineage produces misleading results. Check the taxonomic position of your organism and select the most specific lineage dataset available.

The NCBI taxonomy database provides the authoritative classification for your organism. Cross-reference your organism's taxonomic position against the NCBI taxonomy to identify the appropriate lineage. If you are uncertain which lineage to use, run BUSCO with multiple related lineages and compare the results. A consistent pattern across lineages suggests that the BUSCO failure is real, while widely varying results suggest a lineage selection problem.

### Step 2: Examine the BUSCO Output Categories

The distribution of BUSCO results across the four categories (complete, fragmented, duplicated, missing) provides diagnostic information. A high proportion of missing BUSCOs suggests contamination or insufficient depth. A high proportion of fragmented BUSCOs suggests repetitive content or assembly breaks. A high proportion of duplicated BUSCOs suggests haplotype uncollapsing or recent polyploidy.

Record the BUSCO output for your assembly and compare the category distribution against the expected patterns for different failure modes. This comparison guides your subsequent diagnostic steps and helps you prioritize corrective actions.

### Step 3: Assess Contamination

Contamination assessment should be performed early in the diagnostic process because it is the most common cause of low completeness and the easiest to correct. Several approaches are available:

1. Align your assembled contigs against the NCBI nucleotide database and examine the taxonomic distribution of the best matches. Contigs that match unexpected taxa indicate contamination.
2. Use k-mer based taxonomic classification on your raw reads to identify reads from unexpected organisms.
3. Examine the GC content distribution of your contigs. A bimodal GC distribution often indicates contamination from a distantly related organism.

The TaxaScope workstation provides an integrated environment for these assessments, with containerized tools that run consistently across computing platforms. This approach is particularly useful for laboratory scientists who need to assess genome quality without extensive command-line experience.

### Step 4: Estimate Heterozygosity and Repeat Content

K-mer analysis provides estimates of genome size, heterozygosity, and repeat content from your raw sequencing data. The k-mer spectrum shows peaks that correspond to homozygous and heterozygous k-mers, and the shape of the spectrum indicates the level of heterozygosity. The estimated genome size, calculated from the total number of k-mers divided by the peak coverage, indicates whether your assembly is complete or collapsed.

Compare the k-mer estimated genome size against your assembly size. If the assembly is substantially smaller than the estimated genome size, you likely have collapsed repeats or haplotypes. If the assembly is larger than the estimated genome size, you likely have contamination or uncollapsed haplotypes.

### Step 5: Review Assembly Parameters and Workflow

Review the parameters used in your assembly workflow. Check the k-mer size, coverage cutoffs, error correction settings, and any other parameters that affect assembly outcome. Compare your parameters against the recommendations for your assembler and genome type.

The nf-core documentation provides standards for reproducible workflow configuration, and the Galaxy Training Network offers practical training on assembly parameter selection. The Carpentries lessons provide foundational computing skills that are useful for managing assembly workflows and troubleshooting parameter issues.

### Step 6: Decide on Corrective Action

Based on your diagnostic findings, select the appropriate corrective action:

| Diagnostic Finding | Corrective Action |
|---|---|
| Contamination detected | Filter reads by taxonomic classification, reassemble |
| High heterozygosity | Use haplotype-aware assembler or purge haplotigs |
| High repeat content | Add long-read data, use repeat-aware assembler |
| Insufficient depth | Sequence additional data, reassemble |
| Parameter errors | Adjust parameters, reassemble |
| Multiple causes | Address each cause in order of impact |

## Practical Implementation: Correcting Your Assembly

### Read Filtering for Contamination

If contamination is identified, the first corrective action is to filter your reads. Taxonomic classification tools assign each read to a taxonomic group based on sequence similarity to reference databases. You can remove reads that do not match your target species and retain reads that match your target or are unclassified.

The filtering threshold requires careful consideration. A strict threshold that removes all reads not matching your target species may also remove legitimate reads from your target if the reference database is incomplete. A lenient threshold may retain contaminant reads. The optimal threshold depends on the contamination level and the quality of the reference database for your taxonomic group.

After filtering, reassemble your genome and rerun BUSCO. Compare the new BUSCO scores against the original scores to confirm that contamination was the cause of the failure. The improvement in completeness should be substantial if contamination was the primary issue.

### Haplotype Purging for Heterozygous Genomes

If heterozygosity is the cause of BUSCO failure, you have two options: reassemble with a haplotype-aware assembler or purge haplotigs from your existing assembly. Haplotype purging is generally faster and less expensive than reassembly, and it is appropriate when you need a single representative haplotype.

Haplotype purging tools identify and remove alternative haplotigs from an assembly. These tools use sequence similarity and coverage information to distinguish primary contigs from alternative haplotigs. After purging, the assembly contains a single haplotype, and BUSCO duplicated genes should be reduced.

The Rhododendron simsii genome project demonstrates that high BUSCO completeness is achievable in organisms with population structure and genetic diversity. The study reported 98.4 percent BUSCO completeness in a chromosome-level assembly, indicating that appropriate assembly strategies can overcome the challenges of heterozygosity.

### Adding Long-Read Data for Repetitive Genomes

If repetitive content is causing assembly fragmentation, the most effective corrective action is to add long-read sequencing data. Long reads can span repetitive regions that break short-read assemblies, and the resulting assembly will have improved contiguity and fewer fragmented BUSCO genes.

The Przewalski's horse classroom project demonstrates the feasibility of long-read assembly with modest resources. The project used Oxford Nanopore sequencing with $4,000 of materials and produced a genome on par with the domestic horse reference. This outcome shows that long-read assembly is accessible to research groups with limited budgets.

When adding long-read data, consider the tradeoff between read length and accuracy. Oxford Nanopore reads are longer but have higher error rates, while PacBio HiFi reads are shorter but more accurate. The optimal choice depends on your genome complexity and the resources available for your project.

### Reassembly with Adjusted Parameters

If parameter errors are the cause of BUSCO failure, reassembly with adjusted parameters is the corrective action. The specific parameter adjustments depend on the assembler and the identified problem. Common adjustments include changing the k-mer size, adjusting coverage cutoffs, or modifying error correction settings.

The Galaxy Training Network provides accessible training on assembly workflows that emphasize parameter selection and quality assessment. The EMBL-EBI Training program offers learning pathways for bioinformatics that include assembly and quality assessment modules. These resources can help you identify appropriate parameters for your genome type.

## Records and Measurements for Assembly Quality

### Essential Records for Every Assembly Project

Maintaining detailed records of your assembly workflow is essential for troubleshooting and reproducibility. The following records should be maintained for every assembly project:

1. Sequencing platform and chemistry version
2. Read length distribution and quality metrics
3. Sequencing depth and coverage estimates
4. Assembly tool and version
5. Assembly parameters and settings
6. BUSCO lineage and version
7. BUSCO results for each assembly iteration
8. Assembly statistics (N50, total length, GC content)
9. K-mer analysis results
10. Contamination assessment results

The nf-core documentation emphasizes the importance of reproducible workflow configuration, and the community standards require documentation of all workflow parameters. The Carpentries lessons provide foundational skills for managing computational workflows and maintaining reproducible records.

### Measuring Assembly Quality Beyond BUSCO

BUSCO completeness is one metric among many that should be used to assess assembly quality. Additional metrics include:

1. Contig N50 and L50, which measure assembly contiguity
2. Total assembly size compared to estimated genome size
3. GC content distribution
4. Coverage uniformity across the assembly
5. Completeness of specific genomic features (telomeres, centromeres, etc.)
6. Agreement with genetic or physical maps

The Bioconductor project provides packages for comprehensive assembly quality assessment that integrate multiple metrics. These packages can generate reports that combine BUSCO results with other quality measures, providing a more complete picture of assembly quality.

### Comparing Assembly Iterations

When you make corrective changes to your assembly workflow, you should compare the new assembly against the previous version using multiple metrics. The comparison should include BUSCO scores, assembly statistics, and any other quality measures relevant to your research question.

The comparison should be systematic and documented. Record the changes made between iterations and the effect on each quality metric. This documentation helps identify which changes were effective and provides a basis for future troubleshooting.

## Common Failure Patterns and Their Resolution

### Pattern 1: High Missing, Low Fragmented

This pattern, where many BUSCO genes are completely absent instead of fragmented, typically indicates contamination or insufficient sequencing depth. The missing genes are regions of your target genome that were not assembled, either because contaminant reads consumed coverage or because the sequencing depth was too low to assemble the entire genome.

Resolution: Filter reads for contamination and reassemble, or add sequencing data and reassemble. After reassembly, rerun BUSCO to confirm that the missing genes are now present.

### Pattern 2: High Fragmented, Low Missing

This pattern, where many BUSCO genes are present but broken into fragments, typically indicates repetitive content or assembly breaks. The genes are present in your sequencing data but the assembler could not join the regions spanning the genes, resulting in fragmented assemblies.

Resolution: Add long-read data to span repetitive regions, or adjust assembly parameters to improve repeat resolution. After reassembly, rerun BUSCO to confirm that the fragmented genes are now complete.

### Pattern 3: High Duplicated, Low Single-Copy

This pattern, where many BUSCO genes are present in duplicate instead of as single copies, typically indicates uncollapsed haplotypes or recent polyploidy. The assembly contains both haplotypes as separate sequences, and BUSCO detects the duplicated genes.

Resolution: Purge haplotigs from the assembly or reassemble with a haplotype-aware assembler. After purging, rerun BUSCO to confirm that the duplicated genes are now single-copy.

### Pattern 4: Low Completeness Across All Categories

This pattern, where all BUSCO categories are low, typically indicates a fundamental problem with the assembly, such as severe contamination, extremely low sequencing depth, or incorrect lineage selection. The assembly may contain very little of the target genome.

Resolution: Perform a comprehensive diagnostic assessment, including contamination screening, k-mer analysis, and parameter review. Address the identified causes in order of impact and reassemble.

### Pattern 5: Good BUSCO but Poor Assembly Statistics

This pattern, where BUSCO scores are acceptable but assembly statistics are poor, indicates that the conserved genes are intact but the assembly is fragmented or incomplete in other regions. This pattern is common in assemblies that focus on gene space completeness but have unresolved repetitive regions.

Resolution: This pattern may be acceptable for gene-focused analyses but requires improvement for whole-genome analyses. Add long-read data or scaffolding data to improve contiguity.

## Limitations of BUSCO and Assembly Quality Assessment

### What BUSCO Cannot Detect

BUSCO has several limitations that should be understood when interpreting results. The metric only assesses the presence of conserved single-copy genes, and it cannot detect:

1. Misassemblies that do not break conserved genes
2. Errors in repetitive regions that do not contain conserved genes
3. Incorrect gene models or annotation errors
4. Structural rearrangements that preserve gene order
5. Contamination that does not affect conserved gene representation

The watering-pot shell genome project illustrates this limitation. The study produced a chromosome-scale assembly with 96.5 percent anchoring onto pseudochromosomes, but the morphological interpretations required additional analysis beyond the genome sequence. The genome sequence alone could not explain the adaptive significance of the species' unusual body form.

### The Importance of Multiple Quality Metrics

Because BUSCO cannot detect all assembly errors, you should use multiple quality metrics to assess your assembly. The combination of BUSCO completeness, assembly statistics, k-mer analysis, and alignment-based assessments provides a more complete picture of assembly quality than any single metric.

The Bioconductor project provides packages that integrate multiple quality metrics into comprehensive reports. These reports can be used to document assembly quality for publications and to guide further improvement efforts.

### Context-Dependent Quality Requirements

The required assembly quality depends on your research question. A gene-focused analysis may be possible with a fragmented assembly that has high BUSCO completeness, while a structural genomics analysis requires a chromosome-scale assembly with high contiguity. The watering-pot shell genome project demonstrates that chromosome-scale assembly is achievable for challenging genomes, but the effort required may not be justified for all research questions.

The Przewalski's horse classroom project demonstrates that high-quality assembly is achievable with modest resources, but the project benefited from a streamlined workflow and simplified methods. The methods were conveyed in markdown for complete recording and use in the classroom, and all students were authors on the resulting manuscript. This approach shows that genome assembly can be taught and performed in educational settings while producing high-quality results.

## Professional Escalation Criteria

### When to Seek Expert Assistance

Some assembly problems require specialized expertise beyond what is available in a typical laboratory. You should consider seeking expert assistance when:

1. You have exhausted standard troubleshooting approaches without improvement
2. Your genome has unusual characteristics (extreme heterozygosity, very high repeat content, or polyploidy)
3. You need chromosome-scale assembly and lack experience with scaffolding approaches
4. Your contamination problem is severe or involves closely related species
5. You are working with a non-model organism with no closely related reference genome

The Galaxy Training Network and EMBL-EBI Training provide learning pathways that can help you develop the skills needed to address complex assembly problems. The Carpentries lessons provide foundational computing skills that are useful for managing assembly workflows.

### Collaborative Approaches to Assembly Problems

Collaboration with bioinformatics core facilities or experienced genome assembly groups can accelerate problem resolution. These groups have experience with diverse genome types and can provide guidance on appropriate assembly strategies. The nf-core community provides a forum for sharing workflow expertise, and the Bioconductor project offers packages and support for genomic analysis.

The TaxaScope workstation demonstrates the value of accessible analysis platforms for laboratory scientists. The platform provides a graphical interface for genome quality assessment and analysis, reducing the barrier to entry for researchers without extensive command-line experience. The containerized execution environments ensure that analyses run consistently across computing platforms.

## Frequently Asked Questions

### Why does my BUSCO score vary when I use different lineage datasets?

BUSCO lineage datasets define the expected gene sets for different taxonomic groups. Using a more specific lineage, such as a genus-level dataset instead of a family-level dataset, changes the expected gene set and can produce different completeness scores. The more specific lineage typically provides a more accurate assessment for your organism, but it requires that the lineage dataset is available and well-curated. Cross-reference your organism's taxonomic position against the NCBI taxonomy to select the most appropriate lineage.

### Can I improve BUSCO scores by polishing my assembly?

Polishing can improve BUSCO scores if the assembly contains errors that break conserved genes, but over-polishing can reduce scores by introducing errors in repetitive regions. The effect of polishing on BUSCO scores depends on the assembly quality and the polishing approach. Compare BUSCO scores before and after polishing to determine whether polishing improved or degraded your assembly.

### How much sequencing depth do I need for a good assembly?

The required sequencing depth depends on your genome size, repeat content, heterozygosity, and sequencing platform. Short-read assemblies typically require higher depth than long-read assemblies because the reads are shorter and provide less overlap information. The Przewalski's horse classroom project demonstrated that a mammalian genome can be assembled with modest resources using long-read sequencing, but the optimal depth for your genome depends on its specific characteristics.

### What is the difference between complete and fragmented BUSCO genes?

A complete BUSCO gene is present in your assembly as a single contiguous sequence that covers the expected gene length. A fragmented BUSCO gene is present but broken into multiple contigs, indicating that the assembly has a break within the gene region. Fragmented genes typically indicate repetitive content or assembly errors that prevent correct joining of the gene sequence.

### Why does my assembly have high duplicated BUSCO genes?

High duplicated BUSCO genes typically indicate that your assembly contains both haplotypes of a heterozygous genome as separate sequences. The assembler could not collapse the two haplotypes into a single sequence, and BUSCO detects the duplicated genes. Haplotype purging or haplotype-aware assembly can resolve this issue.

### Can contamination from closely related species be detected?

Contamination from closely related species is more difficult to detect than contamination from distantly related species because the sequences are similar. K-mer based approaches may not distinguish closely related species, and alignment-based approaches may produce ambiguous results. In these cases, coverage analysis can help identify contaminant contigs, as contaminant sequences typically have different coverage than the target genome.

### Should I use BUSCO as the only quality metric for my assembly?

BUSCO should not be used as the only quality metric because it cannot detect all assembly errors. Use BUSCO in combination with assembly statistics, k-mer analysis, and alignment-based assessments to obtain a complete picture of assembly quality. The Bioconductor project provides packages that integrate multiple quality metrics into comprehensive reports.

### How do I know if my BUSCO failure is caused by the assembler or the input data?

The cause of BUSCO failure can be identified through systematic diagnostic testing. Run the same assembly with different assemblers or parameters to determine whether the problem is specific to one tool or consistent across tools. Examine the input data for contamination, heterozygosity, and depth issues. The diagnostic workflow described in this article provides a systematic approach to identifying the root cause of BUSCO failure.

## Related Bioinformatics Guides

- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Binning in Metagenomics: From Contigs to Genomes](/knowledge/bioinformatics/binning-in-metagenomics-from-contigs-to-genomes)
- [Metagenomic Assembly Overview: Challenges and Applications](/knowledge/bioinformatics/metagenomic-assembly-overview-challenges-and-applications)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [TaxaScope: a container-native, visualization-centric workstation for genome-based bacterial taxonomy.](https://doi.org/10.3389/fmicb.2026.1809734). 2026.
- [Sequencing and assembling the genome of Przewalski's horse in the classroom.](https://doi.org/10.1016/j.jevs.2025.105383). 2025.
- [Single specimen genome assembly of Culicoides stellifer shows evidence of a non-retroviral endogenous viral element.](https://doi.org/10.1186/s12864-025-11449-5). 2025.
- [Genome of the enigmatic watering-pot shell and morphological adaptations for anchoring in sediment.](https://doi.org/10.1186/s12864-025-11622-w). 2025.
- [Chromosome-level genome and population genomics reveal demographic history, incomplete divergence and coastal adaptation of Rhododendron simsii var. putuoense in East China](https://doi.org/10.1016/j.pld.2025.12.002). Plant Diversity, 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.