# How to Use BUSCO for Genome Assembly Completeness: A Step-by-Step Tutorial


## Key Takeaways

- BUSCO quantifies genome assembly completeness by searching for conserved, single-copy orthologous genes expected within a specific taxonomic lineage, providing a standardized metric distinct from contiguity (e.g., N50).
- BUSCO classifies genes as Complete (single-copy or duplicated), Fragmented, or Missing, with high Complete percentages (e.g., >90%) and low Fragmented/Missing percentages indicating robust gene-space representation.
- The selection of an appropriate lineage dataset is paramount; using a lineage that is too broad (e.g., Eukaryota for fungi) or too narrow can lead to misleading completeness scores.
- BUSCO's output (C/D/F/M scores) directly informs critical assembly decisions, such as the need for polishing (indicated by high fragmentation) or haplotype purging (indicated by high duplication).
- Integration into automated pipelines (e.g., Puzzler, AquaaG) highlights BUSCO's role as a standard quality control step for both raw assemblies and annotated gene sets, ensuring reproducibility and defensible reporting.
- Running BUSCO in both genome mode (assessing the assembly) and protein mode (assessing the annotation) provides a comprehensive evaluation of gene-space integrity.

---

## Direct Answer and Scope

BUSCO (Benchmarking Universal Single-Copy Orthologs) is a computational tool that assesses genome assembly and annotation completeness by searching for a curated set of single-copy orthologous genes expected to be present in a given taxonomic lineage. This tutorial provides a practical workflow for running BUSCO, selecting appropriate lineage datasets, and interpreting the completeness scores that the tool reports. The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who need to evaluate whether a genome assembly contains the expected gene space. The primary outcome is the ability to run BUSCO correctly, choose the right parameters, and translate the C/D/F/M scores into defensible decisions about assembly quality, polishing needs, and publication-ready reporting.

Genome assembly completeness is a distinct quality metric from contiguity. A highly contiguous assembly can still lack genes that are difficult to assemble due to repetitive content, high heterozygosity, or sequencing gaps. BUSCO addresses this gap by providing a standardized, reproducible measure of gene-space completeness that can be compared across assemblies, species, and studies. The tool is widely integrated into genome assembly pipelines, including automated workflows for chromosome-scale assembly and genome annotation quality assessment.

## What BUSCO Measures and Why It Matters

### The Biological Basis of BUSCO

BUSCO relies on the observation that certain genes are conserved as single copies across broad evolutionary lineages. These orthologs perform essential cellular functions and are under purifying selection that maintains their presence and copy number. The BUSCO software compares a genome assembly or annotation against a lineage-specific set of these expected genes and classifies each one as complete, duplicated, fragmented, or missing.

The classification scheme is central to interpretation. A complete single-copy BUSCO indicates that the gene is present in full and appears once in the assembly. A complete duplicated BUSCO suggests that the gene is present in full but appears more than once, which can indicate a recent whole-genome duplication, an assembly artifact such as a haplotype duplication, or a genuine paralog. A fragmented BUSCO means that only part of the gene was found, often because the assembly has a break within the gene or the gene spans a gap. A missing BUSCO means that no portion of the gene was detected, which can result from assembly absence, extreme divergence, or contamination.

### Why Completeness Scores Matter for Assembly Decisions

Completeness scores directly inform whether an assembly is suitable for downstream analysis. A genome assembly with high contiguity but low BUSCO completeness may still be inadequate for comparative genomics, gene family evolution studies, or annotation projects. Conversely, an assembly with moderate contiguity but high completeness may be sufficient for certain analyses that depend on gene content instead of chromosome structure.

The practical implication is that BUSCO scores should be evaluated alongside other assembly metrics such as N50, L50, total length, and GC content. No single metric determines assembly quality. The decision to polish, scaffold, or re-sequence should be based on the pattern of BUSCO failures. For example, a high proportion of fragmented BUSCOs may indicate that polishing with additional sequencing data could resolve assembly errors, while a high proportion of missing BUSCOs may indicate deeper problems that polishing cannot fix.

### BUSCO in Automated Pipelines

Modern genome assembly pipelines integrate BUSCO as a standard quality control step. The Puzzler pipeline for chromosome-scale assembly from HiFi and Hi-C data includes BUSCO as part of its integrated quality control, alongside Hi-C contact maps, k-mer completeness checks, and contamination screening. This integration reflects the expectation that a platinum-quality genome assembly must demonstrate high gene-space completeness, beyond high contiguity.

Similarly, the AquaaG pipeline for genome annotation includes BUSCO as a gene-space completeness evaluation step after annotation with Prokka for prokaryotes or BRAKER3 for eukaryotes. This placement is important because BUSCO can assess both the raw assembly and the annotated gene set. Running BUSCO on an annotation tells you whether the gene prediction step captured the expected genes, which is a different question from whether the assembly contains them.

## Understanding BUSCO Lineage Datasets

### What Lineage Datasets Contain

BUSCO uses lineage-specific datasets that contain the expected single-copy orthologs for a particular taxonomic group. These datasets are derived from OrthoDB, a database of orthologous groups across the tree of life. Each dataset includes the gene sequences and the associated metadata needed to run the search.

The choice of lineage dataset is the single most important parameter in a BUSCO run. Using a lineage that is too broad, such as eukaryota for a fungal genome, will include genes that are not expected in fungi and may produce misleading completeness scores. Using a lineage that is too narrow, such as a species-specific dataset for a genus with few sequenced representatives, may exclude genes that are genuinely present in the target genome.

### How to Select the Correct Lineage

The BUSCO software provides a lineage selection tool that can automatically identify the most appropriate dataset for a given genome. This tool uses the assembly itself to determine the taxonomic placement and recommends a lineage. For well-studied organisms, the choice is usually straightforward. For less-studied organisms or those with uncertain taxonomy, the automated selection should be reviewed manually.

The lineage datasets are organized hierarchically. The broadest datasets cover major divisions such as bacteria, archaea, and eukaryota. More specific datasets cover phyla, classes, orders, families, and genera. The number of BUSCO genes in a dataset varies by lineage, with more specific datasets typically containing fewer genes because they represent a narrower evolutionary window.

### Lineage Selection for Non-Model Organisms

Non-model organisms present a particular challenge for lineage selection. A genome from an understudied group may not have a dedicated lineage dataset. In this case, the appropriate choice is the most specific dataset that includes the organism's taxonomic group. For example, a genome from a poorly characterized insect family would use the insecta dataset if no more specific dataset exists.

The tradeoff is between sensitivity and specificity. A broader dataset will contain more genes that may not be single-copy in the target lineage, potentially inflating the missing or duplicated counts. A narrower dataset will contain fewer genes but may be more accurate for the target lineage. The BUSCO documentation and the automated lineage selection tool provide guidance for these decisions.

## Installing BUSCO and Dependencies

### Installation Methods

BUSCO can be installed through several methods, including conda, Docker, and source compilation. The conda installation is the most common and is recommended for most users because it handles dependencies automatically. The [Bioconductor project](https://bioconductor.org/) provides documentation for reproducible genomic-analysis workflows, and the same principles of environment management apply to BUSCO installations.

The conda command for installing BUSCO is straightforward, but the exact command depends on the channel configuration and the version of BUSCO being installed. Users should consult the official BUSCO documentation for the current installation instructions. The key point is that BUSCO has several dependencies, including Python, the HMMER suite for hidden Markov model searches, and the Augustus gene prediction tool for some analysis modes.

### Containerized Installation

Containerized installation is an alternative that provides reproducibility across systems. The [nf-core documentation](https://nf-co.re/docs) describes community standards for reproducible workflows, and the same containerization principles apply to individual tools like BUSCO. A container image includes the software and all dependencies in a single package that can be run on any system with the appropriate container runtime.

The advantage of containerized installation is that the exact software versions are fixed, which is important for reproducibility. A BUSCO run performed with version 5.4.3 in a container will produce the same results on any system, whereas a conda installation may resolve to different dependency versions over time.

### Verifying the Installation

After installation, the BUSCO version should be verified to ensure that the expected version is being used. The version number is important for reporting because BUSCO results can vary between versions due to changes in the underlying databases and search algorithms. The version should be recorded in the methods section of any publication or report.

The installation should also be tested with a small example dataset to confirm that the software runs correctly. This test run should use a small genome or a subset of a genome to verify that the lineage datasets are downloaded correctly and that the search pipeline executes without errors.

## Running BUSCO on a Genome Assembly

### Basic Command Structure

The basic BUSCO command requires three pieces of information: the input genome file, the lineage dataset, and the output directory. The input genome file should be in FASTA format. The lineage dataset is specified by name, and the output directory is where the results will be written.

The command structure is consistent across BUSCO versions, though the specific flags may vary. The core flags include the input file, the lineage, the output name, and the mode. The mode specifies whether BUSCO is assessing a genome assembly or a protein set from an annotation.

### Genome Mode versus Protein Mode

BUSCO has two primary modes. Genome mode searches the nucleotide sequence of an assembly and uses gene prediction to identify candidate regions before comparing them to the BUSCO dataset. Protein mode searches a set of protein sequences, typically from a genome annotation, and compares them directly to the BUSCO dataset.

The choice of mode depends on the question being asked. Genome mode assesses the raw assembly and does not depend on the quality of an annotation. Protein mode assesses the annotation and can reveal whether gene prediction missed expected genes. Running both modes provides a complete picture: genome mode tells you what the assembly contains, and protein mode tells you what the annotation captured.

### Transcriptome Mode

BUSCO also has a transcriptome mode for assessing RNA-seq assemblies. This mode is useful for evaluating the completeness of a transcriptome assembly, which is a different question from genome assembly completeness. Transcriptome mode is less commonly used in genome assembly projects but is relevant for studies that rely on transcriptome data.

The transcriptome mode requires a transcriptome assembly in FASTA format and uses the same lineage datasets as genome mode. The interpretation of the scores is similar, but the biological meaning differs because a transcriptome reflects expressed genes instead of the complete gene space.

## Selecting Parameters for the BUSCO Run

### The Lineage Parameter

The lineage parameter is the most consequential choice in a BUSCO run. The automated lineage selection tool can identify the appropriate lineage, but manual review is recommended. The lineage name must match the available datasets, and the BUSCO software will download the dataset if it is not already present locally.

For a genome from a well-characterized species, the species-level or genus-level lineage is appropriate. For a genome from a less-characterized species, the family-level or order-level lineage may be the best choice. The lineage selection should be documented in the methods, including the version of the lineage dataset.

### The Mode Parameter

The mode parameter determines whether BUSCO runs in genome, protein, or transcriptome mode. The mode must match the input file type. Genome mode expects a nucleotide FASTA file. Protein mode expects a protein FASTA file. Transcriptome mode expects a nucleotide FASTA file of transcripts.

The mode parameter also affects the computational requirements. Genome mode is the most computationally intensive because it includes gene prediction steps. Protein mode is faster because it skips gene prediction and searches the protein sequences directly.

### The Output Parameters

The output parameters control where the results are written and how the run is labeled. The output directory should be named descriptively so that results from different assemblies or lineages can be distinguished. The output name is used as a prefix for the result files.

The BUSCO run produces several output files, including a short summary text file, a full table of results, and a log file. The summary file contains the completeness percentages that are typically reported. The full table contains the classification of each BUSCO gene and can be used for detailed analysis.

## Interpreting BUSCO Completeness Scores

### The C/D/F/M Classification

The BUSCO summary reports four categories: complete (C), complete and duplicated (D), fragmented (F), and missing (M). The complete category includes both single-copy and duplicated BUSCOs. The duplicated category is a subset of the complete category. The sum of complete, fragmented, and missing should equal the total number of BUSCO genes in the lineage dataset.

The percentages are calculated relative to the total number of BUSCO genes in the lineage dataset. A typical report might show C:95.2%[S:90.1%,D:5.1%],F:2.3%,M:2.5%,n:1234. The n value is the total number of BUSCO genes searched. The S value is the percentage of complete single-copy BUSCOs, and the D value is the percentage of complete duplicated BUSCOs.

### What the Scores Mean for Assembly Quality

The complete percentage is the primary indicator of assembly quality. A high complete percentage, typically above 90% for a good assembly, indicates that the assembly contains most of the expected gene space. The fragmented and missing percentages indicate the proportion of genes that are incomplete or absent.

The duplicated percentage requires careful interpretation. A low duplicated percentage, typically below 5%, is expected for most diploid genomes. A high duplicated percentage can indicate that the assembly contains both haplotypes, which is a common artifact in assemblies from heterozygous samples. This artifact can be addressed by haplotype purging, which is a standard step in chromosome-scale assembly pipelines.

### Comparing Scores Across Assemblies

BUSCO scores are comparable across assemblies only when the same lineage dataset and the same BUSCO version are used. Different lineage datasets contain different numbers of genes, so the percentages are not directly comparable. Different BUSCO versions may use different search algorithms or updated lineage datasets, which can change the results.

For this reason, publications should report the BUSCO version, the lineage dataset name and version, and the mode used. This information allows readers to interpret the scores in context. The n value, or total number of BUSCO genes searched, should also be reported because it indicates the size of the lineage dataset.

## At a Glance: BUSCO Decision Table

| Assembly Scenario | Expected BUSCO Pattern | Recommended Action |
| --- | --- | --- |
| High-quality diploid assembly | C > 95%, D < 5%, F < 3%, M < 3% | Proceed to annotation and downstream analysis |
| Assembly with haplotype duplication | C > 90%, D > 10%, F low, M low | Run haplotype purging and re-run BUSCO |
| Assembly with many fragmented genes | C moderate, F > 10%, M moderate | Consider polishing with additional sequencing data |
| Assembly with many missing genes | C < 80%, M > 10%, F variable | Investigate sequencing depth, contamination, and lineage choice |

## Practical Workflow for BUSCO Assessment

### Step 1: Prepare the Input Assembly

The input assembly should be in FASTA format and should represent the final or near-final version of the assembly. Running BUSCO on an intermediate assembly can be useful for tracking progress, but the final assessment should be performed on the assembly that will be used for downstream analysis.

The assembly file should be checked for common issues before running BUSCO. The file should not contain duplicate sequence names, and the sequences should be in valid FASTA format. The assembly should also be checked for contamination, as contaminant sequences can affect BUSCO results by introducing foreign genes or diluting the representation of the target genome.

### Step 2: Select the Lineage Dataset

The lineage dataset should be selected based on the taxonomic placement of the target organism. The automated lineage selection tool can provide a recommendation, but the recommendation should be reviewed manually. The lineage dataset version should be recorded for reporting purposes.

For a genome from a well-studied organism, the species-level or genus-level lineage is appropriate. For a genome from a less-studied organism, the most specific available lineage that includes the organism should be used. The lineage selection should be documented in the methods section of any report or publication.

### Step 3: Run BUSCO in Genome Mode

The initial BUSCO run should be performed in genome mode on the raw assembly. This run provides the baseline completeness assessment. The command should include the input assembly, the lineage dataset, the output directory, and the genome mode flag.

The run may take several hours depending on the genome size and the computational resources available. The BUSCO software provides progress updates, and the log file can be monitored for errors. The run should be allowed to complete without interruption to ensure that the results are valid.

### Step 4: Review the Summary Results

The summary results should be reviewed after the run completes. The complete, duplicated, fragmented, and missing percentages should be recorded. The n value should be checked to confirm that the expected number of BUSCO genes was searched.

The results should be compared to the expectations for the target organism. A genome from a species with a well-assembled reference should have completeness scores in the same range as the reference. A genome from a species without a reference should be compared to related species if available.

### Step 5: Run BUSCO in Protein Mode on the Annotation

If an annotation is available, BUSCO should be run in protein mode on the annotated protein set. This run assesses whether the annotation captured the expected gene space. The protein mode results can reveal annotation errors even when the assembly completeness is high.

The protein mode results should be interpreted in the context of the genome mode results. If the genome mode shows high completeness but the protein mode shows lower completeness, the annotation pipeline may have missed genes. If both modes show low completeness, the assembly itself may be the problem.

### Step 6: Document the Results

The BUSCO results should be documented with the version, lineage, mode, and date of the run. The summary statistics should be recorded in a consistent format that can be compared across assemblies. The full results table should be archived for detailed analysis if needed.

The documentation should include the command used for the run, the input file names, and the output directory. This information supports reproducibility and allows the run to be repeated if needed.

## Common Failure Patterns and Their Causes

### High Duplicated Percentage

A high duplicated percentage, typically above 10%, often indicates that the assembly contains both haplotypes of a heterozygous genome. This pattern is common in assemblies from samples with high heterozygosity, such as many plant and fungal species. The duplicated BUSCOs represent genes that are present in both haplotypes.

The solution is haplotype purging, which removes the redundant haplotype sequences from the assembly. The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) includes duplicate purging as a standard step in chromosome-scale assembly, and this step should be applied before the final BUSCO assessment. After purging, the duplicated percentage should decrease and the single-copy percentage should increase.

### High Fragmented Percentage

A high fragmented percentage, typically above 10%, indicates that many genes are present but incomplete in the assembly. This pattern can result from assembly errors, sequencing gaps, or repetitive regions that are difficult to assemble. The fragmented genes may be split across multiple contigs or scaffolds.

The solution depends on the cause. If the fragmentation results from sequencing gaps, additional sequencing data and polishing may resolve the issue. If the fragmentation results from repetitive regions, the assembly may require specialized approaches such as long-read sequencing or optical mapping. The decision to invest in additional sequencing should be based on the proportion of fragmented BUSCOs and the importance of the affected genes.

### High Missing Percentage

A high missing percentage, typically above 10%, indicates that many expected genes are absent from the assembly. This pattern can result from contamination, where the assembly contains sequences from another organism, or from genuine absence due to biological factors. The missing genes may be absent because they are difficult to assemble or because they are not present in the target genome.

The first step in addressing a high missing percentage is to verify the lineage selection. An incorrect lineage can produce a high missing percentage because the expected genes are not relevant to the target organism. The second step is to check for contamination using tools such as BlobTools, which is included in the [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168). The third step is to review the assembly for systematic issues such as low sequencing depth in specific regions.

## Records and Measurements for BUSCO Assessment

### Essential Records

The essential records for a BUSCO assessment include the BUSCO version, the lineage dataset name and version, the mode, the input file, and the date. These records should be maintained for every BUSCO run to support reproducibility and comparison.

The summary statistics should be recorded in a standardized format. The complete percentage, single-copy percentage, duplicated percentage, fragmented percentage, missing percentage, and total number of BUSCO genes should be recorded. The full results table should be archived for detailed analysis.

### Comparative Measurements

Comparative measurements are useful for tracking assembly improvement over time. If an assembly is polished or purged, the BUSCO scores should be measured before and after the process. The comparison shows whether the process improved gene-space completeness.

The comparison should use the same lineage dataset and BUSCO version for both runs. If the lineage dataset or BUSCO version changes, the comparison is not valid. The before and after scores should be recorded in a table or spreadsheet for easy reference.

### Quality Control Thresholds

Quality control thresholds should be established before the BUSCO assessment, not after. The thresholds depend on the target organism and the intended use of the assembly. A genome intended for comparative genomics may require higher completeness than a genome intended for a specific gene family analysis.

The thresholds should be documented and applied consistently. If an assembly fails the threshold, the decision to polish, purge, or re-sequence should be based on the BUSCO pattern and the available resources. The thresholds should be reviewed periodically as new lineage datasets and BUSCO versions become available.

## Limitations of BUSCO Assessment

### Lineage Dataset Limitations

The BUSCO lineage datasets are limited by the taxonomic coverage of the underlying OrthoDB database. Organisms from understudied groups may have limited representation in the lineage datasets, which can affect the accuracy of the completeness assessment. The lineage datasets are updated periodically, and the version should be recorded.

The lineage datasets are also limited by the assumption of single-copy orthology. Genes that are genuinely duplicated in a lineage will be classified as duplicated BUSCOs, which can inflate the duplicated percentage. This limitation is particularly relevant for lineages that have experienced whole-genome duplications.

### Assembly-Specific Limitations

BUSCO assesses gene-space completeness, not assembly correctness. An assembly can have high BUSCO completeness but contain structural errors such as misjoins or misassemblies. The BUSCO scores should be interpreted alongside other quality metrics such as contiguity, read mapping, and Hi-C contact maps.

BUSCO also cannot detect errors that do not affect the BUSCO gene set. A misassembly in a gene-poor region will not be detected by BUSCO. The completeness scores should not be used as the sole indicator of assembly quality.

### Interpretation Limitations

The BUSCO scores are relative to the lineage dataset, not to the true gene content of the target organism. A high completeness score indicates that the assembly contains the expected BUSCO genes, but it does not guarantee that all genes are present or correctly assembled. The scores should be interpreted in the context of the target organism and the intended use of the assembly.

The scores are also dependent on the BUSCO version and the lineage dataset version. Comparisons across versions should be made with caution. The methods section of any report should include the version information to support interpretation.

## BUSCO in the Context of Genome Assembly Pipelines

### Integration with Assembly Workflows

BUSCO is a standard component of modern genome assembly workflows. The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) for chromosome-scale assembly includes BUSCO as an integrated quality control step, alongside Hi-C contact maps, k-mer completeness checks, and contamination screening. This integration ensures that the final assembly meets the expected gene-space completeness standards.

The placement of BUSCO in the workflow matters. Running BUSCO after haplotype purging and scaffolding provides a final assessment of the assembly. Running BUSCO before these steps can identify issues that need to be addressed before proceeding. The workflow should include BUSCO at multiple stages if the assembly is being iteratively improved.

### Integration with Annotation Workflows

BUSCO is also a standard component of genome annotation workflows. The [AquaaG pipeline](https://doi.org/10.1016/j.mex.2026.103955) includes BUSCO as a gene-space completeness evaluation step after annotation. This placement allows the pipeline to assess whether the annotation captured the expected gene space.

The protein mode BUSCO results are particularly useful for annotation quality assessment. If the protein mode shows lower completeness than the genome mode, the annotation pipeline may have missed genes. The annotation can then be refined to improve the gene-space completeness.

### Reproducibility Considerations

Reproducibility is a key consideration for BUSCO assessments. The BUSCO version, lineage dataset version, and parameters should be recorded for every run. The [nf-core documentation](https://nf-co.re/docs) describes community standards for reproducible workflows, and these standards apply to BUSCO assessments as well.

Containerized execution provides the highest level of reproducibility. A container image with a fixed BUSCO version and lineage dataset will produce the same results on any system. The container image should be archived with the assembly and the BUSCO results.

## Professional Escalation Criteria

### When to Seek Expert Assistance

BUSCO assessments can usually be performed by a competent bioinformatician without expert assistance. However, certain situations warrant escalation to a specialist. These situations include persistent high missing percentages that do not respond to polishing or purging, unusual BUSCO patterns that suggest contamination, and assemblies from organisms with complex genomic features such as polyploidy or extreme heterozygosity.

The decision to escalate should be based on the BUSCO results and the intended use of the assembly. If the assembly is intended for publication or for a large-scale comparative analysis, the BUSCO results should be reviewed by an expert before proceeding. The expert can provide guidance on lineage selection, parameter choices, and interpretation.

### When to Reconsider the Assembly Strategy

A BUSCO assessment that shows persistent low completeness may indicate that the assembly strategy needs to be reconsidered. The decision to re-sequence, use a different assembler, or apply additional polishing should be based on the BUSCO pattern and the available resources.

The escalation criteria should be documented before the BUSCO assessment. The criteria should specify the completeness thresholds that trigger a reassessment of the assembly strategy. The criteria should also specify the steps to be taken when the thresholds are not met.

## Frequently Asked Questions

### What is the difference between complete and duplicated BUSCOs?

Complete BUSCOs are genes that are present in full in the assembly. Duplicated BUSCOs are a subset of complete BUSCOs that appear more than once. A duplicated BUSCO can indicate a genuine gene duplication, a whole-genome duplication, or an assembly artifact where both haplotypes are present. The duplicated percentage is reported separately from the single-copy percentage to help distinguish these scenarios.

### How do I choose the right lineage dataset for my genome?

The lineage dataset should match the taxonomic placement of your target organism. The BUSCO software includes an automated lineage selection tool that can recommend a dataset based on the assembly sequence. For well-studied organisms, the species-level or genus-level dataset is appropriate. For less-studied organisms, use the most specific dataset that includes your organism's taxonomic group.

### What is a good BUSCO completeness score?

A good BUSCO completeness score depends on the organism and the intended use of the assembly. For most diploid genomes, a complete percentage above 90% is considered good, and above 95% is considered excellent. The duplicated percentage should typically be below 5% for a diploid assembly. The fragmented and missing percentages should be low, typically below 5% each.

### Why does my assembly have a high duplicated BUSCO percentage?

A high duplicated percentage often indicates that the assembly contains both haplotypes of a heterozygous genome. This is a common artifact in assemblies from samples with high heterozygosity. The solution is haplotype purging, which removes the redundant haplotype sequences. After purging, the duplicated percentage should decrease and the single-copy percentage should increase.

### Can I compare BUSCO scores across different assemblies?

BUSCO scores are comparable across assemblies only when the same lineage dataset and the same BUSCO version are used. Different lineage datasets contain different numbers of genes, so the percentages are not directly comparable. Different BUSCO versions may use different search algorithms or updated lineage datasets. Always report the BUSCO version and lineage dataset version when comparing scores.

### Should I run BUSCO on the assembly or the annotation?

You should run BUSCO on both. Genome mode assesses the raw assembly and tells you whether the expected genes are present in the assembly. Protein mode assesses the annotated protein set and tells you whether the annotation captured the expected genes. Running both modes provides a complete picture of gene-space completeness.

### How long does a BUSCO run take?

The runtime depends on the genome size, the lineage dataset size, and the computational resources available. A small bacterial genome can be assessed in minutes. A large eukaryotic genome can take several hours or longer. Genome mode is more computationally intensive than protein mode because it includes gene prediction steps.

### What should I do if my BUSCO completeness is low?

First, verify that you selected the correct lineage dataset. An incorrect lineage can produce misleading results. Second, check for contamination using tools such as BlobTools. Third, review the assembly for systematic issues such as low sequencing depth. If the completeness remains low after these checks, consider polishing with additional sequencing data or re-assembling with a different approach.

## BUSCO in Comparative Genomics Applications

### Gene Family Evolution Studies

BUSCO completeness scores are a prerequisite for comparative genomics studies that examine gene family evolution. The [black soldier fly comparative analysis](https://doi.org/10.1038/s41437-025-00805-6) demonstrates how genome assemblies are used to study gene family expansions related to digestive, immunity, and olfactory functions. In such studies, the completeness of each genome assembly directly affects the reliability of gene family copy number estimates.

A genome assembly with low BUSCO completeness will produce unreliable gene family counts because missing genes will be interpreted as gene losses. A genome assembly with high duplicated BUSCOs may inflate gene family sizes due to haplotype duplication artifacts. The BUSCO assessment should be performed before any comparative gene family analysis to ensure that the observed differences reflect biology instead of assembly artifacts.

### Polar Adaptation Genomics

The [Antarctic sea-ice diatom study](https://doi.org/10.3389/fmicb.2026.1755917) illustrates how BUSCO completeness supports ecological and evolutionary interpretations. The study examined genomic features related to oxidative stress response, DNA repair, protein quality control, osmotic regulation, and nutrient acquisition. These functional interpretations depend on the completeness of the gene space in the assembly.

When a genome assembly is used to identify shared genes across species or to detect gene family contractions, the BUSCO scores provide confidence that the observed patterns are not artifacts of incomplete assembly. A high-quality draft genome assembly with strong BUSCO completeness supports conclusions about gene content and functional potential.

### Chromosome-Level Assembly Standards

The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) establishes a standard for platinum-quality chromosome-scale genome assemblies that includes BUSCO as an integrated quality metric. The pipeline was validated on genomes ranging from 24 Mbp to 6.5 Gbp, demonstrating that BUSCO assessment scales across genome sizes. The integration of BUSCO with Hi-C contact maps, k-mer completeness, and contamination screening provides a multi-faceted quality assessment.

For researchers aiming to produce chromosome-level assemblies, the BUSCO score is one of several metrics that define platinum quality. The pipeline automates contig assembly, duplicate purging, Hi-C-based scaffolding, and chromosome assignment via synteny, with BUSCO providing the gene-space completeness check at the end.

## Training and Skill Development for BUSCO Users

### Foundational Bioinformatics Skills

BUSCO usage requires foundational bioinformatics skills including command-line navigation, file management, and basic scripting. The [Carpentries lessons](https://carpentries.org/lessons) provide structured training in shell, Git, and programming that form the basis for running tools like BUSCO effectively. These skills are essential for managing input files, interpreting output directories, and troubleshooting errors.

The [EMBL-EBI training](https://www.ebi.ac.uk/training) resources offer learning pathways for bioinformatics data analysis that complement the practical skills needed for BUSCO assessment. Understanding how to navigate sequence databases and interpret biological data formats supports the correct preparation of input files and the interpretation of results.

### Reproducible Workflow Training

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that demonstrates how tools like BUSCO can be integrated into reproducible analysis pipelines. The graphical interface of Galaxy lowers the barrier for researchers who are not comfortable with command-line tools, while still producing reproducible results.

The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline development that emphasize reproducibility and portability. These standards are relevant for researchers who want to integrate BUSCO into larger automated workflows or who need to run BUSCO across many assemblies in a batch processing context.

### Practical Considerations for Training

Training should include hands-on practice with example datasets before applying BUSCO to real assemblies. The practice runs should cover lineage selection, parameter choices, and interpretation of results. The training should also cover common errors and how to troubleshoot them.

The training should emphasize the importance of documentation. Every BUSCO run should be recorded with the version, lineage, mode, and parameters. This documentation practice supports reproducibility and allows results to be compared across time and across assemblies.

## BUSCO Output Files and Their Uses

### The Short Summary File

The short summary file contains the key completeness statistics in a compact format. This file is the primary output for reporting purposes. The summary includes the complete, single-copy, duplicated, fragmented, and missing percentages, along with the total number of BUSCO genes searched.

The summary file should be archived for every BUSCO run. The summary provides the information needed for methods sections in publications and for quality control records. The summary should be reviewed immediately after the run completes to identify any obvious problems.

### The Full Table File

The full table file contains the classification of each individual BUSCO gene. This file is useful for detailed analysis of which genes are missing or fragmented. The full table can be used to identify patterns in the missing genes, such as whether they cluster in specific genomic regions.

The full table is also useful for comparing BUSCO results across assemblies at the gene level. If two assemblies have similar overall completeness scores but different missing genes, the full table can reveal which genes differ. This information can guide targeted improvements to the assembly.

### The Log File

The log file contains the detailed progress of the BUSCO run, including any warnings or errors. The log file should be reviewed after the run completes to confirm that the run executed without issues. The log file can also be used to estimate the runtime and to identify any steps that were unusually slow.

The log file is particularly important for troubleshooting. If a BUSCO run fails or produces unexpected results, the log file provides the first source of diagnostic information. The log file should be archived with the other output files.

## Batch Processing and Scalability

### Running BUSCO on Multiple Assemblies

Many research projects require BUSCO assessment of multiple assemblies. The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) demonstrates a sample sheet input structure that supports scalable batch processing. This approach allows many assemblies to be processed with minimal user input.

For batch processing, the output directory naming convention becomes important. Each assembly should have a distinct output name that includes the assembly identifier and the lineage dataset. The results should be collected into a summary table that allows easy comparison across assemblies.

### Computational Resource Planning

The computational resources required for BUSCO depend on the genome size and the mode. Genome mode is more resource-intensive than protein mode. Large eukaryotic genomes may require significant memory and multiple CPU cores. The resource requirements should be estimated before starting a batch run.

The runtime for a batch of assemblies can be substantial. The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) includes a checkpointing system that ensures previously completed tasks are not re-executed. This feature is valuable for long-running batch processes where interruptions are possible.

### Scaling to Large Genomes

The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) was validated on genomes ranging from 24 Mbp to 6.5 Gbp, demonstrating that BUSCO assessment scales to very large genomes. However, the runtime and resource requirements increase with genome size. Researchers working with large genomes should plan for longer runtimes and allocate sufficient computational resources.

For very large genomes, the lineage dataset size also matters. A broader lineage dataset with more genes will require more computational time than a narrower dataset. The lineage selection should balance the need for accuracy with the computational cost.

## Quality Control Integration with Other Metrics

### Combining BUSCO with Assembly Statistics

BUSCO completeness should be evaluated alongside standard assembly statistics such as N50, L50, total length, and GC content. The [AquaaG pipeline](https://doi.org/10.1016/j.mex.2026.103955) integrates QUAST for assembly quality assessment alongside BUSCO for gene-space completeness. This combination provides a more complete picture of assembly quality than either metric alone.

A high N50 with low BUSCO completeness may indicate that the assembly is contiguous but missing genes. A low N50 with high BUSCO completeness may indicate that the assembly is fragmented but contains the expected gene space. The interpretation depends on the intended use of the assembly.

### Combining BUSCO with Read Mapping

Read mapping statistics provide another layer of quality assessment. The proportion of reads that map back to the assembly can indicate whether the assembly represents the sequenced sample. Low mapping rates may indicate contamination or assembly errors.

The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) includes k-mer completeness checks alongside BUSCO. The k-mer completeness provides an independent measure of whether the sequencing data is fully represented in the assembly. Combining these metrics with BUSCO provides a robust quality assessment.

### Combining BUSCO with Contamination Screening

Contamination can affect BUSCO results by introducing foreign sequences or diluting the representation of the target genome. The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) includes BlobTools contamination screening alongside BUSCO. This integration allows researchers to identify and address contamination before finalizing the assembly.

If contamination is detected, the contaminant sequences should be removed and the BUSCO assessment repeated. The contamination screening should be performed before the final BUSCO assessment to ensure that the completeness scores reflect the target genome.

## Reporting BUSCO Results in Publications

### Methods Section Reporting

The methods section of a publication should include the BUSCO version, the lineage dataset name and version, the mode, and the parameters used. This information allows readers to interpret the results and to compare them with other studies. The [BUSCO methods paper](https://doi.org/10.1007/978-1-4939-9173-0_14) provides the methodological foundation for this reporting standard.

The methods should also describe the input assembly version and any preprocessing steps. If the assembly was polished or purged before the BUSCO assessment, this should be stated. The date of the BUSCO run should be included because lineage datasets are updated periodically.

### Results Section Reporting

The results section should report the complete, single-copy, duplicated, fragmented, and missing percentages. The total number of BUSCO genes searched should be included. The results should be presented in a format that allows comparison with other assemblies.

The results should be interpreted in the context of the assembly strategy and the intended use of the assembly. A discussion of the BUSCO results should address any unusual patterns, such as high duplication or high missing percentages, and explain how they were resolved.

### Supplementary Materials

The full BUSCO results table should be included in the supplementary materials. This table allows readers to examine the classification of individual BUSCO genes. The full table also supports meta-analyses that combine BUSCO results across multiple studies.

The supplementary materials should also include the BUSCO command used for the run. This information supports reproducibility and allows other researchers to repeat the assessment if needed.

## Troubleshooting Common BUSCO Errors

### Lineage Download Failures

The BUSCO software downloads lineage datasets on demand. If the download fails, the run will terminate with an error. The download can fail due to network issues or because the lineage name is incorrect. The lineage name should be checked against the available datasets.

If the download fails repeatedly, the lineage dataset can be downloaded manually and placed in the appropriate directory. The BUSCO documentation provides instructions for manual lineage dataset installation. The lineage dataset version should be recorded for reporting purposes.

### Memory Errors

Large genomes can require substantial memory for BUSCO runs. If the run terminates with a memory error, the memory allocation should be increased. The memory requirements depend on the genome size and the lineage dataset size.

The memory allocation can be adjusted through the BUSCO parameters or through the job submission system if running on a cluster. The memory requirements should be estimated before starting the run to avoid failures.

### Gene Prediction Failures

Genome mode uses gene prediction to identify candidate regions before comparing them to the BUSCO dataset. If the gene prediction step fails, the run will terminate with an error. The gene prediction step can fail due to issues with the input assembly or due to missing dependencies.

The log file should be reviewed to identify the cause of the gene prediction failure. The input assembly should be checked for format issues. The dependencies should be verified to ensure that all required tools are installed correctly.

## Future Directions and Updates

### Lineage Dataset Updates

The BUSCO lineage datasets are updated periodically as new genomes are sequenced and added to OrthoDB. The updated datasets may contain additional genes or revised gene sets. The lineage dataset version should be recorded for every BUSCO run to support comparisons across versions.

Researchers should check for updated lineage datasets before starting a new BUSCO assessment. The updated datasets may provide more accurate completeness assessments for well-studied organisms. The decision to use an updated dataset should be documented.

### Software Version Updates

The BUSCO software is updated periodically with new features and bug fixes. The updated versions may produce different results due to changes in the search algorithms or the underlying databases. The software version should be recorded for every BUSCO run.

Researchers should be aware that BUSCO results from different versions may not be directly comparable. When comparing results across studies, the BUSCO version should be checked. The methods section should include the version information to support interpretation.

### Integration with Emerging Technologies

The [Puzzler pipeline](https://pubmed.ncbi.nlm.nih.gov/41573168) demonstrates how BUSCO is integrated into modern assembly workflows that use HiFi and Hi-C data. The integration of BUSCO with emerging sequencing technologies will continue to evolve as new assembly methods are developed.

The [AquaaG pipeline](https://doi.org/10.1016/j.mex.2026.103955) demonstrates how BUSCO is integrated into automated annotation workflows. The integration of BUSCO with automated pipelines reduces the manual effort required for quality assessment and supports high-throughput genome projects.

## Frequently Asked Questions

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Gene Set Enrichment Analysis in R: A Practical Tutorial for Interpreting Omics Data](/knowledge/bioinformatics/gene-set-enrichment-analysis-in-r-a-practical-tutorial-for-interpreting-omics-data)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Puzzler: scalable one-command platinum-quality genome assembly from HiFi and Hi-C.](https://pubmed.ncbi.nlm.nih.gov/41573168). Bioinformatics advances, 2026.
- [AquaaG: A comprehensive pipeline for quality assessment and annotation of genomes.](https://doi.org/10.1016/j.mex.2026.103955). 2026.
- [Comparative genomics of an Antarctic sea-ice diatom (&lt,i&gt,Nitzschia&lt,/i&gt, sp.) provides insights into potential polar adaptation.](https://doi.org/10.3389/fmicb.2026.1755917). 2026.
- [Comparative analysis of gene family evolution demonstrates expansion of digestive, immunity and olfactory functions in the black soldier fly (Hermetia illucens) lineage.](https://doi.org/10.1038/s41437-025-00805-6). 2026.
- [BUSCO: Assessing Genome Assembly and Annotation Completeness.](https://doi.org/10.1007/978-1-4939-9173-0_14). Methods in molecular biology, 2019.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.