# How to Choose the Right BUSCO Lineage for Your Genome Assembly


## Key Takeaways

- The selection of the appropriate BUSCO lineage is paramount for accurate genome assembly completeness assessment, as scores are only meaningful when the reference gene set aligns taxonomically with the organism under study. Mismatched lineages can lead to erroneous conclusions regarding assembly quality and evolutionary inferences.

- BUSCO evaluates genome completeness by identifying conserved single-copy orthologs within a defined taxonomic group, with scores reflecting the proportion of these expected genes found in complete, fragmented, or missing states within the assembly. The interpretation of these scores is inherently relative to the chosen lineage's taxonomic scope.

- BUSCO lineages are structured according to the NCBI taxonomy, ranging from broad domains (e.g., Eukaryota) to highly specific clades (e.g., Primates). Choosing a more specific lineage offers a more stringent completeness test but increases the risk of misinterpreting genuine gene loss or divergence as assembly deficiencies.

- The auto-lineage workflow automates lineage selection, particularly useful for unknown taxonomic origins, but can be misled by assembly contamination or may select a lineage that is too specific for comparative analyses. It should be used as a starting point, with manual verification of the selected lineage's biological appropriateness.

- Verification of lineage choice involves examining the lineage dataset's species composition, performing quick test assessments, and scrutinizing BUSCO output for patterns of missing or fragmented genes that might indicate lineage mismatch rather than assembly gaps. Documenting the exact lineage, BUSCO version, and assembly date is crucial for reproducibility.

- Interpreting BUSCO results requires biological context; a low completeness score (<90%) for a well-characterized species warrants investigation, while similar scores may be acceptable for poorly characterized organisms or metagenome-assembled genomes. Genuine gene loss, as observed in mycoheterotrophic plants, can explain missing BUSCOs.

---

Selecting the correct BUSCO lineage is the single most important decision you will make when assessing genome assembly completeness, because BUSCO scores are only meaningful when the reference gene set matches the taxonomic position of your organism. A mismatched lineage produces misleading completeness percentages that can lead you to over-polish a finished assembly, under-report a fragmented one, or draw false evolutionary conclusions. This article provides a practical decision framework for choosing lineages, interpreting results within their biological context, and building custom lineage sets when standard options do not fit your study organism.

## Understanding What BUSCO Lineages Actually Measure

BUSCO (Benchmarking Universal Single-Copy Orthologs) evaluates genome assembly quality by searching for a set of near-universal single-copy orthologs derived from the OrthoDB database. These orthologs are genes expected to be present in a single copy in virtually all species within a given taxonomic group. When you run BUSCO, you are essentially asking a simple question: of the genes that should be present in every member of this taxonomic group, how many are complete, fragmented, duplicated, or missing in your assembly?

The biological rationale for this approach is straightforward. If your assembly represents a complete genome, it should contain nearly all of the conserved single-copy genes expected for that lineage. A completeness score of 95 percent means that 95 percent of the expected single-copy orthologs were found in a complete state. The remaining 5 percent may be missing due to assembly gaps, sequencing errors, or genuine biological gene loss. The [BUSCO protocol documentation](https://pubmed.ncbi.nlm.nih.gov/34936221) emphasizes that these biologically meaningful metrics complement technical measures such as contiguity statistics including N50 values and contig counts.

The critical implication for lineage choice is that BUSCO completeness is always relative to a specific taxonomic reference set. A score of 90 percent against the Metazoa dataset means something very different from a score of 90 percent against the Arthropoda dataset. The Metazoa set includes genes conserved across all animals, which is a smaller and more conserved set. The Arthropoda set includes genes specific to arthropods, which is a larger set that includes more lineage-specific genes. Your assembly might score 95 percent against Metazoa but only 85 percent against Arthropoda if it is missing some arthropod-specific genes.

This relativity is not a flaw in BUSCO but rather a feature that allows researchers to assess completeness at different biological scales. However, it creates a practical problem: you must choose the lineage that provides the most informative assessment for your specific organism and research question. The [BUSCO protocol](https://pubmed.ncbi.nlm.nih.gov/34936221) describes running modes that include batch analysis on multiple inputs and an auto-lineage workflow that runs assessments without requiring you to specify a dataset manually.

## The Taxonomic Hierarchy of BUSCO Lineages

BUSCO lineages are organized according to the NCBI taxonomy hierarchy, which provides a standardized framework for naming and organizing organisms. The [NCBI taxonomy resources](https://www.ncbi.nlm.nih.gov/) maintain the authoritative classification that BUSCO uses to structure its lineage datasets. Understanding this hierarchy is essential for choosing the appropriate lineage for your organism.

At the broadest level, BUSCO offers lineages for major domains and kingdoms. These include Bacteria, Archaea, Eukaryota, and within Eukaryota, separate lineages for major groups such as Metazoa, Fungi, Viridiplantae, and Protists. These broad lineages contain genes that are conserved across the entire group. For example, the Metazoa lineage contains genes conserved across all animals, from sponges to vertebrates.

Moving down the hierarchy, BUSCO provides increasingly specific lineages. For vertebrates, there are lineages for Mammalia, Aves, Reptilia, Actinopterygii, and others. Within Mammalia, there are lineages for Primates, Rodentia, Cetartiodactyla, and additional subgroups. Each level down the hierarchy includes the conserved genes from the broader lineage plus additional genes that are conserved within the more specific group.

The practical consequence of this hierarchy is that you must decide how specific your lineage should be. A more specific lineage provides a more stringent test of completeness because it includes more genes that are expected to be present. However, a more specific lineage also carries a higher risk of including genes that your organism has genuinely lost or that have diverged beyond recognition. A broader lineage is more forgiving but provides less information about lineage-specific gene content.

The [BUSCO protocol](https://pubmed.ncbi.nlm.nih.gov/34936221) notes that BUSCO is capable of assessing genome assemblies, transcriptomes, annotated gene sets, and metagenome-assembled genomes where the taxonomic origin of the species may be unknown. This versatility means that the lineage choice depends also on your organism but also on the type of data you are assessing.

## Why Lineage Choice Directly Affects Your Completeness Scores

The relationship between lineage choice and completeness scores is also a technical detail. It has direct consequences for how you interpret your assembly quality and what decisions you make based on that interpretation. Consider the example of the rough periwinkle Littorina saxatilis, a marine snail used as a model system for studying speciation and local adaptation. The first reference genome for this species was assembled from Illumina data and was highly fragmented with an N50 of 44 kilobases. This assembly had a BUSCO completeness of 80.1 percent against the Metazoa dataset. A newer reference genome achieved an N50 of 67 megabases and a BUSCO completeness of 94.1 percent against the same Metazoa dataset, as reported in the [chromosome-scale genome assembly study](https://pubmed.ncbi.nlm.nih.gov/38584387).

This example illustrates two important points. First, the same lineage dataset was used for both assemblies, allowing a direct comparison of improvement. Second, the choice of the Metazoa dataset instead of a more specific mollusk or gastropod lineage meant that the completeness assessment was based on genes conserved across all animals. If the researchers had used a more specific lineage, the absolute scores might have been different, but the relative improvement between the two assemblies would likely have been similar.

The deeper issue is that completeness scores are only interpretable when you know what gene set they are based on. A score of 94.1 percent against Metazoa tells you that the assembly contains nearly all animal-conserved single-copy genes. It does not tell you whether the assembly contains all genes expected for a gastropod mollusk. If you are interested in lineage-specific gene content, you need a more specific lineage.

The [codon usage study across the Saccharomycotina subphylum](https://pubmed.ncbi.nlm.nih.gov/39213398) provides another perspective on how BUSCO metrics are used in comparative genomics. The researchers characterized codon usage across 1,154 strains from 1,051 species and found that codon usage bias was influenced by a combination of features including the number of coding sequences, BUSCO count, and genome length. This study used BUSCO metrics as one of several genomic features for comparative analysis, demonstrating that BUSCO scores can serve as inputs to downstream analyses beyond simple quality assessment.

## The Auto-Lineage Workflow and Its Limitations

BUSCO version 5 introduced an auto-lineage workflow that automatically selects the most appropriate lineage for your input data. This workflow is designed to handle cases where you do not know the taxonomic origin of your sequence data, such as metagenome-assembled genomes or transcriptomes from poorly characterized organisms. The [BUSCO protocol](https://pubmed.ncbi.nlm.nih.gov/34936221) describes this auto-lineage mode as a way to run assessments without specifying a dataset manually.

The auto-lineage workflow works by first running a quick assessment against a set of broad lineages to determine the approximate taxonomic position of your input sequences. Based on this initial screen, it then selects the most specific lineage that appears to match your data and runs the full assessment against that lineage. This two-step process is computationally efficient because the initial screen uses a reduced set of markers.

The auto-lineage workflow has several practical advantages. It removes the need for you to know the exact taxonomic position of your organism, which is particularly useful for environmental samples or poorly studied groups. It also reduces the risk of accidentally choosing a lineage that is too specific or too broad for your data. For researchers working with metagenome-assembled genomes where the taxonomic origin is unknown, the auto-lineage workflow is often the most practical starting point.

However, the auto-lineage workflow has limitations that you should understand. The initial taxonomic screen can be misled by contamination in your assembly. If your assembly contains sequences from multiple organisms, the auto-lineage workflow may select a lineage based on the dominant contaminant instead of your target organism. This is a particular concern for metagenome-assembled genomes, which may contain fragments from multiple species.

The auto-lineage workflow also tends to select the most specific lineage that matches your data, which may not always be the most appropriate choice for your research question. If you are comparing your assembly to published genomes that were assessed against a broader lineage, you may want to use the same broader lineage for consistency. The auto-lineage workflow does not know about your comparison set and will select the most specific match regardless.

For these reasons, the auto-lineage workflow is best used as a starting point instead of a final decision. You should examine the lineage that the auto-lineage workflow selects and verify that it makes biological sense for your organism. If you are working with a well-characterized species, you should consider whether a more specific or more appropriate lineage is available.

## Building a Decision Tree for Lineage Selection

A practical decision tree can help you select the appropriate BUSCO lineage for your genome assembly. This decision tree is based on the taxonomic position of your organism, the type of data you are assessing, and the purpose of your assessment.

The first decision point is whether you know the taxonomic position of your organism. If you do not know the taxonomic position, use the auto-lineage workflow as your starting point. If you do know the taxonomic position, proceed to the second decision point.

The second decision point is whether your organism belongs to a major group for which BUSCO provides a dedicated lineage. Major groups include Bacteria, Archaea, Eukaryota, Metazoa, Fungi, Viridiplantae, and Protists. If your organism belongs to one of these groups, you should use the most specific lineage available within that group that matches your organism.

The third decision point is whether a more specific lineage is available for your organism. For example, if your organism is a mammal, you should use the Mammalia lineage instead of the Metazoa lineage. If your organism is a primate, you should use the Primates lineage instead of the Mammalia lineage. The most specific lineage that matches your organism provides the most stringent and informative completeness assessment.

The fourth decision point is whether you need to compare your assembly to published genomes that were assessed against a specific lineage. If you are making such comparisons, you should use the same lineage as the published genomes to ensure that your completeness scores are directly comparable. This is particularly important for studies that compare completeness across multiple species or assemblies.

The fifth decision point is whether your organism is a hybrid, polyploid, or has an unusual genome structure that might affect the interpretation of single-copy orthologs. For polyploid organisms, you may need to interpret duplication scores differently, and you may want to use a lineage that is appropriate for the ploidy level of your organism.

The final decision point is whether you need to create a custom lineage for your organism. Custom lineage creation is appropriate when your organism belongs to a group that is not well represented by existing BUSCO lineages, or when you need to assess completeness against a gene set that is specific to your study system.

## Practical Steps for Selecting and Verifying Your Lineage

Once you have used the decision tree to select a candidate lineage, you should verify that your choice is appropriate before running the full BUSCO assessment. This verification process involves several practical steps that can save you from wasting computational resources on an inappropriate lineage.

The first step is to check the lineage dataset itself. BUSCO lineage datasets are available for download from the BUSCO website, and each dataset includes a list of the orthologs it contains and the species used to construct the dataset. You should verify that your organism or a close relative is included in the species list for the lineage. If your organism is not represented and no close relative is included, the lineage may not be appropriate for your assessment.

The second step is to run a quick test assessment on a small sample of your data. This test can be a subset of your assembly or a single chromosome or scaffold. The goal is to confirm that the lineage produces sensible results before committing to a full assessment. If the test assessment produces an unexpectedly low completeness score, you should investigate whether the lineage is appropriate or whether there are issues with your assembly.

The third step is to examine the BUSCO output files in detail. The full BUSCO output includes also the summary statistics but also the locations of complete, fragmented, duplicated, and missing BUSCOs in your assembly. You should examine the missing and fragmented BUSCOs to determine whether they are concentrated in specific genomic regions or distributed throughout the assembly. Concentrated missing BUSCOs may indicate assembly gaps or misassemblies, while distributed missing BUSCOs may indicate a lineage mismatch.

The fourth step is to compare your results against expectations for your organism. If your organism is a well-studied species with a published reference genome, you should compare your completeness scores against the published values. Large discrepancies may indicate problems with your assembly or with your lineage choice.

The fifth step is to document your lineage choice and the rationale behind it. This documentation is important for reproducibility and for interpreting your results in the context of published studies. Your methods section should state the exact lineage dataset and version used, as well as the BUSCO version and parameters.

## Records and Measurements for Lineage Decisions

Keeping systematic records of your lineage selection process and BUSCO results is essential for reproducible genome assembly assessment. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that emphasizes the importance of documenting analysis steps and parameters for reproducibility. Similarly, the [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that include consistent parameter recording and reporting.

Your records should include the following information for each BUSCO assessment you run:

The BUSCO version number and the exact command used to run the assessment. This includes all parameters such as the lineage dataset, the mode (genome, transcriptome, protein), and any options for batch analysis or auto-lineage selection.

The lineage dataset name and version. BUSCO lineage datasets are versioned, and different versions may contain different sets of orthologs. You should record the exact version you used so that your results can be compared to future assessments.

The date of the assessment and the version of your assembly that was assessed. Genome assemblies are often updated, and you need to know which version of the assembly produced which BUSCO results.

The full BUSCO summary statistics including the number and percentage of complete, complete and single-copy, complete and duplicated, fragmented, and missing BUSCOs. These statistics should be recorded for each lineage you test.

The taxonomic information for your organism including the NCBI taxonomy ID if available. The [NCBI taxonomy resources](https://www.ncbi.nlm.nih.gov/) provide the authoritative classification that you should use to document your organism's taxonomic position.

Any observations about the distribution of missing or fragmented BUSCOs in your assembly. These observations can help you identify assembly problems or lineage mismatches.

The [Bioconductor project](https://bioconductor.org/) provides packages and workflows for reproducible genomic analysis that can help you manage and document your BUSCO results. While BUSCO itself is not a Bioconductor package, you can use Bioconductor tools to analyze BUSCO output files and integrate completeness metrics with other genomic analyses.

## Common Failure Patterns in Lineage Selection

Several common failure patterns can lead to misleading BUSCO results. Recognizing these patterns can help you avoid them and interpret results correctly when they occur.

The first failure pattern is using a lineage that is too broad for your organism. This occurs when you use a lineage such as Eukaryota or Metazoa for an organism that belongs to a more specific group with an available lineage. The result is a completeness score that is artificially high because the broad lineage contains only the most conserved genes. Your assembly may appear complete against the broad lineage but be missing many lineage-specific genes.

The second failure pattern is using a lineage that is too specific for your organism. This occurs when you use a lineage for a group that your organism does not actually belong to, or when you use a lineage that includes genes your organism has genuinely lost. The result is a completeness score that is artificially low because the lineage includes genes that are not expected to be present in your organism.

The third failure pattern is using a lineage that is mismatched due to contamination. This occurs when your assembly contains sequences from multiple organisms, and the dominant contaminant drives the lineage selection. The result is a completeness score that reflects the contaminant instead of your target organism. This is a particular concern for metagenome-assembled genomes and for assemblies from organisms that live in close association with symbionts or parasites.

The fourth failure pattern is using different lineages for different assemblies that you are comparing. This occurs when you assess one assembly against a broad lineage and another assembly against a specific lineage, then compare the completeness scores directly. The comparison is invalid because the scores are based on different gene sets.

The fifth failure pattern is ignoring the duplication score when interpreting completeness. A high duplication score can indicate a polyploid genome, a recent whole-genome duplication, or assembly artifacts such as haplotypic duplication. Each of these has different biological implications, and you need to investigate the cause of duplication before interpreting your results.

The sixth failure pattern is using BUSCO results without considering the quality of the underlying assembly. BUSCO completeness is one metric among many, and it should be interpreted alongside contiguity metrics such as N50, assembly size, and GC content. A highly contiguous assembly with a low BUSCO score may have a different problem than a fragmented assembly with a high BUSCO score.

## Interpreting BUSCO Results in Biological Context

BUSCO completeness scores are only meaningful when interpreted in the context of your organism's biology and your research question. The same numerical score can have different implications depending on the organism and the purpose of the assessment.

For a well-characterized species with a high-quality reference genome, a BUSCO completeness score below 90 percent against the appropriate lineage should trigger investigation. The missing or fragmented BUSCOs may indicate assembly gaps, sequencing errors, or genuine biological gene loss. You should examine the specific BUSCOs that are missing or fragmented to determine whether they cluster in particular genomic regions or represent particular functional categories.

For a poorly characterized species or a metagenome-assembled genome, a lower BUSCO completeness score may be acceptable. The [BUSCO protocol](https://pubmed.ncbi.nlm.nih.gov/34936221) notes that BUSCO can assess metagenome-assembled genomes where the taxonomic origin of the species is unknown. In these cases, the auto-lineage workflow can help you select an appropriate lineage, but you should interpret the results with caution because the taxonomic assignment may be uncertain.

The [mycoheterotrophic monocot study](https://pubmed.ncbi.nlm.nih.gov/36582124) provides an instructive example of how BUSCO results can reflect genuine biological gene loss. The researchers studied mycoheterotrophic plants that obtain sugars from soil fungi and have lost photosynthesis. They found that 174 of 1,375 plant benchmark universally conserved orthologous genes were undetected in any mycoheterotroph transcriptome or the genome of the mycoheterotrophic orchid Gastrodia, but were expressed in green relatives. This finding provides evidence for massively convergent gene loss in nonphotosynthetic lineages.

This example illustrates that missing BUSCOs are not always assembly artifacts. They can represent genuine biological gene loss that is relevant to your research question. If you are studying a lineage that has undergone reductive evolution, such as parasites, symbionts, or mycoheterotrophs, you should expect some BUSCOs to be missing and interpret these absences as biological findings instead of assembly failures.

The [Saccharomycotina codon usage study](https://pubmed.ncbi.nlm.nih.gov/38826271) provides another example of how BUSCO metrics can be used in comparative genomics. The researchers found that codon usage bias was influenced by a combination of genome features and assembly metrics that included the number of coding sequences, BUSCO count, and genome length. This study used BUSCO count as a proxy for genome completeness in a comparative analysis across more than 1,000 species, demonstrating that BUSCO metrics can serve as inputs to downstream analyses.

## Custom Lineage Creation for Specialized Applications

When existing BUSCO lineages do not adequately represent your organism or research question, you may need to create a custom lineage. Custom lineage creation is appropriate in several situations.

The first situation is when your organism belongs to a taxonomic group that is not well represented by existing BUSCO lineages. This can occur for poorly studied groups or for groups that have undergone rapid diversification. If the existing lineage for your group contains few species or few orthologs, a custom lineage based on a broader sampling of your group may provide a more informative assessment.

The second situation is when you need to assess completeness against a gene set that is specific to your study system. For example, if you are studying a family of genes that is specific to your organism or group, you may want to create a custom lineage that includes these genes. This allows you to assess whether your assembly contains the complete set of genes relevant to your research question.

The third situation is when you are working with a group that has undergone extensive gene loss or gene family expansion. Existing BUSCO lineages may include genes that your organism has lost or exclude genes that your organism has duplicated. A custom lineage based on the actual gene content of your group can provide a more accurate assessment.

Creating a custom lineage requires you to identify a set of single-copy orthologs that are expected to be present in your organism. This typically involves selecting a set of species that represent your group, identifying orthologous genes across these species, and filtering for genes that are present in single copy in most or all of the selected species. The [OrthoDB database](https://pubmed.ncbi.nlm.nih.gov/34936221) provides the underlying ortholog data that BUSCO uses, and you can use this resource to identify appropriate ortholog sets for custom lineages.

The [BUSCO protocol](https://pubmed.ncbi.nlm.nih.gov/34936221) describes the use of BUSCO plugin workflows for performing common operations in genomics using BUSCO results, such as building phylogenomic trees and visualizing syntenies. These plugins can be useful for validating custom lineages and for integrating BUSCO results with other genomic analyses.

Custom lineage creation is a significant undertaking that requires careful validation. You should validate your custom lineage by running BUSCO on a set of reference genomes from your group and confirming that the completeness scores are reasonable. You should also compare your custom lineage results against results from existing BUSCO lineages to ensure that your custom lineage is providing additional information instead of simply reproducing existing results.

## Quality Controls and Reproducibility Considerations

Reproducibility is a core requirement for genome assembly assessment. Your BUSCO results should be reproducible by other researchers, which means you must document your lineage choice, parameters, and data versions. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in computing and data skills that emphasize reproducible practices. The [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that includes documentation and reproducibility practices.

Several quality controls should be in place before you run BUSCO. First, you should verify that your assembly file is in the correct format and that the sequence headers are properly formatted. BUSCO requires FASTA format for genome assemblies and transcriptomes, and protein files for protein mode. Incorrect formatting can cause BUSCO to fail or produce incorrect results.

Second, you should verify that your assembly does not contain excessive contamination. Contamination can be detected by examining the GC content distribution, by running taxonomic classification tools, or by examining the BUSCO results for unexpected duplication or missing patterns. If you suspect contamination, you should clean your assembly before running BUSCO.

Third, you should verify that your assembly has an appropriate level of completeness before running BUSCO. If your assembly is highly fragmented, BUSCO may report low completeness simply because the genes are split across multiple contigs. In this case, you may want to run BUSCO in genome mode with the appropriate parameters for fragmented assemblies.

Fourth, you should run BUSCO with the appropriate mode for your data type. Genome mode is used for genome assemblies, transcriptome mode for assembled transcriptomes, and protein mode for annotated protein sets. Using the wrong mode can produce misleading results.

Fifth, you should consider running BUSCO with multiple lineages to understand how your completeness scores vary with lineage choice. This is particularly useful when you are uncertain about the appropriate lineage or when you are comparing your assembly to published genomes that used different lineages.

The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that include consistent parameter recording and reporting. If you are using nf-core pipelines for your genome assembly workflow, you should ensure that BUSCO parameters are recorded consistently across runs.

## Professional Escalation Criteria for BUSCO Results

Knowing when to escalate a BUSCO result to a more detailed investigation is important for efficient genome assembly workflows. The following criteria indicate that you should investigate your assembly or your lineage choice further.

If your BUSCO completeness score is below 90 percent against the appropriate lineage for a well-characterized organism, you should investigate the cause. This may indicate assembly problems, sequencing errors, or lineage mismatch. You should examine the missing and fragmented BUSCOs to determine whether they cluster in specific genomic regions.

If your BUSCO duplication score is above 5 percent for a haploid or diploid organism, you should investigate the cause. High duplication can indicate haplotypic duplication in the assembly, a recent whole-genome duplication, or contamination from a closely related species. Each of these causes requires different corrective actions.

If your BUSCO completeness score varies dramatically between different lineages, you should investigate the cause. This may indicate that your organism has undergone gene loss or gene family expansion that affects the interpretation of specific lineages. You should examine which BUSCOs are missing or duplicated in each lineage to understand the biological basis for the variation.

If your BUSCO results are inconsistent with other quality metrics for your assembly, you should investigate the cause. For example, if your assembly has a high N50 but a low BUSCO completeness score, there may be systematic errors in the assembly that are not reflected in contiguity metrics. Conversely, if your assembly has a low N50 but a high BUSCO completeness score, the assembly may be fragmented but contain most of the expected genes.

If you are working with a metagenome-assembled genome and the auto-lineage workflow selects a lineage that seems inconsistent with your expected organism, you should investigate the cause. This may indicate contamination or a misassembly that has mixed sequences from multiple organisms.

If you are comparing your assembly to published genomes and your BUSCO completeness scores are substantially lower, you should investigate whether the difference is due to assembly quality or lineage choice. You should verify that you are using the same lineage and BUSCO version as the published studies.

## Safety and Ethical Considerations in Genome Assembly Assessment

While genome assembly assessment does not involve the same safety considerations as laboratory work with hazardous materials, there are ethical and data management considerations that you should address. The [EMBL-EBI training resources](https://www.ebi.ac.uk/training) provide guidance on responsible data management and analysis practices in bioinformatics.

Data management considerations include ensuring that your assembly data are stored securely and backed up appropriately. Genome assemblies can be large files, and you should have a data management plan that addresses storage, backup, and sharing. The [NCBI data resources](https://www.ncbi.nlm.nih.gov/) provide repositories for genome assemblies and associated metadata, and you should consider depositing your assembly and BUSCO results when you publish your work.

Ethical considerations include ensuring that you have the appropriate permissions to use the sequence data you are analyzing. If you are working with data from endangered species, protected species, or human subjects, you should ensure that your analysis complies with relevant regulations and ethical guidelines.

Reproducibility considerations include documenting your analysis steps and parameters so that other researchers can reproduce your results. The [Carpentries lessons](https://carpentries.org/lessons) provide training in reproducible research practices, and the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that emphasizes reproducibility.

## Integrating BUSCO Results with Other Assembly Quality Metrics

BUSCO completeness is one of several metrics that you should use to assess genome assembly quality. A comprehensive assessment should integrate BUSCO results with contiguity metrics, assembly size, GC content, and other quality indicators.

Contiguity metrics such as N50, L50, and the number of contigs or scaffolds provide information about the fragmentation of your assembly. A highly contiguous assembly with a high N50 is generally easier to analyze than a fragmented assembly with a low N50, even if both have similar BUSCO completeness scores. The [Littorina saxatilis genome study](https://pubmed.ncbi.nlm.nih.gov/38584387) illustrates this point: the first reference genome had an N50 of 44 kilobases and a BUSCO completeness of 80.1 percent, while the newer reference genome had an N50 of 67 megabases and a BUSCO completeness of 94.1 percent. Both metrics improved substantially, but the improvement in contiguity was more dramatic than the improvement in completeness.

Assembly size relative to the expected genome size provides information about whether your assembly is complete in terms of total sequence content. If your assembly is substantially smaller than the expected genome size, you may be missing genomic regions even if BUSCO completeness is high. Conversely, if your assembly is substantially larger than the expected genome size, you may have contamination or duplication.

GC content distribution can reveal contamination or systematic sequencing errors. Unexpected peaks in the GC content distribution may indicate contamination from other organisms or biases in sequencing or assembly.

The number of coding sequences predicted from your assembly can be compared to expectations for your organism. The [Saccharomycotina codon usage study](https://pubmed.ncbi.nlm.nih.gov/39213398) used the number of coding sequences as one of several genomic features in their comparative analysis, demonstrating that gene content metrics can complement BUSCO completeness.

## At a Glance

| Decision Point | Recommended Action | Rationale |
| --- | --- | --- |
| Taxonomic position known | Select the most specific BUSCO lineage that matches your organism | More specific lineages provide more stringent and informative completeness assessments |
| Taxonomic position unknown | Use the auto-lineage workflow as a starting point | The auto-lineage workflow screens against broad lineages and selects the most specific match |
| Comparing to published genomes | Use the same lineage and BUSCO version as the published studies | Direct comparability requires identical assessment parameters |
| Organism has undergone gene loss | Expect some missing BUSCOs and interpret them as biological findings | Reductive evolution can cause genuine gene loss that is biologically meaningful |
| Contamination suspected | Clean the assembly before running BUSCO | Contamination can drive lineage selection and produce misleading results |
| Existing lineages inadequate | Create a custom lineage based on your study system | Custom lineages can assess completeness against gene sets relevant to your research question |

## Common Failure Patterns and Corrective Actions

| Failure Pattern | Symptom | Corrective Action |
| --- | --- | --- |
| Lineage too broad | Artificially high completeness score | Select a more specific lineage that matches your organism |
| Lineage too specific | Artificially low completeness score | Select a broader lineage or verify that your organism belongs to the lineage |
| Contamination driving lineage selection | Unexpected lineage selected by auto-lineage | Clean the assembly and rerun the assessment |
| Inconsistent lineages across comparisons | Non-comparable completeness scores | Standardize lineage choice across all assemblies being compared |
| Ignoring duplication score | Misinterpretation of polyploidy or assembly artifacts | Investigate the cause of duplication before interpreting results |
| Ignoring assembly quality context | Overinterpretation of completeness scores | Integrate BUSCO results with contiguity and other quality metrics |

## Frequently Asked Questions

### What is the difference between the Metazoa and Vertebrata BUSCO lineages?

The Metazoa lineage contains single-copy orthologs that are conserved across all animals, from sponges to vertebrates. The Vertebrata lineage contains the Metazoa-conserved genes plus additional genes that are conserved specifically within vertebrates. Using the Vertebrata lineage provides a more stringent test of completeness for vertebrate genomes because it includes more genes that are expected to be present. A genome that scores 95 percent against Metazoa might score lower against Vertebrata if it is missing some vertebrate-specific genes.

### How do I know which BUSCO lineage to use for a newly sequenced species?

Start by determining the taxonomic position of your species using the NCBI taxonomy. Identify the most specific BUSCO lineage that includes your species or a close relative. If you are uncertain, run the auto-lineage workflow to get a suggested lineage, then verify that the suggestion makes biological sense. You can also run BUSCO against multiple lineages to understand how completeness scores vary with lineage choice.

### Can I use the same BUSCO lineage for a genome assembly and a transcriptome assembly?

Yes, you can use the same lineage for both genome and transcriptome assessments, but you should use the appropriate mode for each data type. Genome mode is used for genome assemblies, and transcriptome mode is used for assembled transcriptomes. The lineage dataset is the same, but the assessment parameters differ between modes.

### What does a high BUSCO duplication score indicate?

A high duplication score can indicate several things. It may indicate that your organism is polyploid or has undergone a recent whole-genome duplication. It may also indicate haplotypic duplication in your assembly, where both haplotypes of a diploid organism are assembled as separate sequences. Contamination from a closely related species can also produce duplication. You should investigate the cause of duplication before interpreting your results.

### How do I create a custom BUSCO lineage for my organism?

Creating a custom lineage requires you to identify a set of single-copy orthologs that are expected to be present in your organism. Select a set of species that represent your group, identify orthologous genes across these species, and filter for genes that are present in single copy in most or all of the selected species. The OrthoDB database provides the underlying ortholog data that BUSCO uses. Validate your custom lineage by running BUSCO on reference genomes from your group and comparing results against existing lineages.

### Why does my BUSCO completeness score differ between the auto-lineage and a manually selected lineage?

The auto-lineage workflow selects the most specific lineage that matches your data based on an initial taxonomic screen. A manually selected lineage may be broader or narrower than the auto-lineage suggestion. The completeness score will differ because different lineages contain different sets of expected genes. You should choose the lineage that is most appropriate for your research question and document your choice.

### Should I use the same BUSCO lineage for all species in a comparative study?

For direct comparability, you should use the same lineage for all species in a comparative study. However, you may also want to run each species against its most specific lineage to assess lineage-specific completeness. The choice depends on your research question. If you are comparing overall genome quality across species, use a common lineage. If you are assessing whether each species contains its expected gene content, use the most specific lineage for each species.

### What should I do if my BUSCO completeness score is unexpectedly low?

First, verify that you are using the appropriate lineage for your organism. Second, examine the missing and fragmented BUSCOs to determine whether they cluster in specific genomic regions. Third, check your assembly for contamination and other quality issues. Fourth, compare your results against published genomes for your organism or close relatives. If the low score persists, consider whether your organism has undergone genuine gene loss that explains the missing BUSCOs.

## Related Bioinformatics Guides

- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Long-Read Genome Assembly and Polishing Strategies](/knowledge/bioinformatics/long-read-genome-assembly-and-polishing-strategies)
- [Transcriptome Assembly Without a Reference Genome](/knowledge/bioinformatics/transcriptome-assembly-without-a-reference-genome)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [BUSCO: Assessing Genomic Data Quality and Beyond.](https://pubmed.ncbi.nlm.nih.gov/34936221). Current protocols, 2021.
- [Chromosome-scale Genome Assembly of the Rough Periwinkle Littorina saxatilis.](https://pubmed.ncbi.nlm.nih.gov/38584387). Genome biology and evolution, 2024.
- [Genomic factors shaping codon usage across the Saccharomycotina subphylum.](https://pubmed.ncbi.nlm.nih.gov/39213398). G3 (Bethesda, Md.), 2024.
- [Genomic factors shaping codon usage across the Saccharomycotina subphylum.](https://pubmed.ncbi.nlm.nih.gov/38826271). bioRxiv : the preprint server for biology, 2024.
- [Phylotranscriptomic Analyses of Mycoheterotrophic Monocots Show a Continuum of Convergent Evolutionary Changes in Expressed Nuclear Genes From Three Independent Nonphotosynthetic Lineages.](https://pubmed.ncbi.nlm.nih.gov/36582124). Genome biology and evolution, 2023.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.