How to Assess the Quality of Metagenome-Assembled Genomes (MAGs): Completeness, Contamination, and Strain Heterogeneity
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Completeness and Contamination are Marker-Based Estimates: MAG completeness is inferred from the presence of expected single-copy marker genes, while contamination is estimated by the frequency of multi-copy markers, indicating foreign DNA. Tools like CheckM and BUSCO utilize distinct marker sets, potentially leading to differing estimates for the same MAG.
- Strain Heterogeneity is Crucial for Variant Analysis: High strain heterogeneity, indicated by nearly identical multi-copy markers (e.g., CheckM score near 100%), suggests a MAG contains multiple closely related strains, which can confound variant calling and phylogenetic analyses by mixing signals.
- MIMAG Standards Define Quality Tiers: The Minimum Information about a Metagenome-Assembled Genome (MIMAG) standard categorizes MAGs into high (≥90% complete, <5% contamination), medium (≥50% complete, <10% contamination), and low quality tiers, providing a framework for consistent reporting and data sharing.
- CheckM and BUSCO Offer Complementary Assessments: CheckM provides lineage-specific marker analysis and an explicit strain heterogeneity score, while BUSCO uses universal orthologs and is valuable for cross-validation of completeness and as a contamination indicator through duplicated gene counts.
- Quality Assessment is Essential for Downstream Validity: Inaccurate MAG quality (low completeness, high contamination, or strain heterogeneity) directly leads to misleading biological conclusions in gene presence-absence analyses, functional annotations, and comparative genomics, underscoring the necessity of rigorous quality control.
- Remediation Strategies Address Specific Issues: Low-quality MAGs can often be improved through rebinning to separate chimeric assemblies, assembly polishing to address fragmentation, or long-read scaffolding to enhance contiguity, depending on the nature of the quality deficit.
Metagenome-assembled genomes are microbial genomes reconstructed from DNA sequence data derived from environmental or host-associated samples. Unlike genomes obtained from isolated pure cultures, MAGs are computational reconstructions that carry inherent uncertainty about which sequences truly belong to the target organism and which originated from co-assembled contaminants. Before using a MAG for downstream analysis, researchers must evaluate three core quality dimensions: completeness, which estimates how much of the expected genome is present, contamination, which estimates how much foreign sequence is mixed in, and strain heterogeneity, which indicates whether the assembly contains multiple closely related variants. This article explains how to compute and interpret these metrics using CheckM and BUSCO, how to apply the Minimum Information about a Metagenome-Assembled Genome (MIMAG) standards, and how to make defensible inclusion or exclusion decisions for your dataset.
Why MAG Quality Assessment Matters
A MAG that enters downstream analysis without rigorous quality assessment can produce misleading biological conclusions. If a genome is only 60 percent complete, gene presence-absence analyses will falsely report that metabolic pathways are missing. If a genome is 15 percent contaminated, functional annotations will attribute genes from unrelated taxa to the organism of interest. If strain heterogeneity is high, variant calling and phylogenetic placement will mix signals from multiple populations.
The practical consequence is that every downstream result, from metabolic reconstruction to taxonomic assignment to comparative genomics, inherits the errors of the input MAG. Quality assessment is the step that determines whether your dataset supports the claims you intend to make. Researchers who skip this step risk publishing results that cannot be reproduced or that misrepresent the biology of the system under study.
Quality assessment also matters for data sharing and reproducibility. Public sequence repositories and journals increasingly expect MAG submissions to include standardized quality metadata. The NCBI Data Resources provide official documentation on sequence database submission requirements and search systems that can help you understand what quality information is expected when depositing genomes. Similarly, the EMBL-EBI Training portal offers structured learning pathways for bioinformatics data analysis that cover genome assembly and quality assessment in the context of reproducible research practices.
Core Quality Metrics Defined
Completeness
Completeness estimates the fraction of the expected genome that is present in the assembly. The estimate is indirect. It relies on marker genes that are expected to occur in single copy across a broad phylogenetic group. If a MAG contains most of these markers, the genome is likely mostly assembled. If many markers are missing, the assembly is incomplete.
Two widely used approaches exist. CheckM uses lineage-specific marker sets that are selected based on the phylogenetic placement of the MAG. BUSCO uses universal single-copy orthologs organized into lineage datasets such as Bacteria, Archaea, or Eukaryota. Both tools report completeness as a percentage, but they may produce different values for the same MAG because they use different marker collections.
Contamination
Contamination estimates the fraction of sequence in the assembly that originates from organisms other than the target. The estimate is also marker-based. When a single-copy marker appears multiple times in the assembly, the extra copies are presumed to come from a contaminating genome. CheckM reports contamination as a percentage based on the frequency of multi-copy markers. BUSCO reports duplicated markers, which serve a similar diagnostic role.
Contamination is not always evenly distributed. A small contaminant genome may contribute only a few extra markers, producing a low contamination estimate even though the foreign sequence is biologically significant. Conversely, a closely related contaminant may share markers with the target, causing the contamination estimate to understate the true mixing.
Strain Heterogeneity
Strain heterogeneity describes the situation where a MAG contains sequence from two or more closely related strains of the same species. These strains share most of their genome, so standard contamination markers may not detect them as foreign. CheckM provides a strain heterogeneity score that indicates the proportion of multi-copy markers that are identical or nearly identical. A high strain heterogeneity score suggests that the duplicated markers come from closely related genomes instead of from distantly related contaminants.
Strain heterogeneity matters because it affects variant calling, gene content comparisons, and phylogenetic analyses. If two strains are mixed in one MAG, single-nucleotide variants will appear heterozygous, and gene presence-absence calls will be ambiguous.
MIMAG Standards for Reporting
The Minimum Information about a Metagenome-Assembled Genome (MIMAG) standard provides a common language for describing MAG quality. The standard defines three quality tiers based on completeness and contamination estimates.
| Quality Tier | Completeness | Contamination | Strain Heterogeneity |
|---|---|---|---|
| High quality | Greater than 90 percent | Less than 5 percent | Should be reported when available |
| Medium quality | Greater than 50 percent | Less than 10 percent | Should be reported when available |
| Low quality | Less than 50 percent | Any value | Should be reported when available |
The MIMAG standard also requires reporting the methods used to generate the MAG, including the assembler, the binning approach, and the quality assessment tools. This reporting requirement supports reproducibility and allows other researchers to interpret quality metrics in context. Public databases increasingly expect MIMAG-compliant metadata for MAG submissions, and journals may require this information in methods sections.
CheckM Workflow
Installation and Input Requirements
CheckM is available through standalone distribution channels and can be installed via package managers commonly used in bioinformatics environments. The tool requires a set of marker genes that are bundled with the installation. Input is a directory of bins in FASTA format. Each bin should represent one putative genome.
The Bioconductor project provides documentation for installing and using genomic analysis packages in the R environment. While CheckM itself is a standalone tool, the Bioconductor ecosystem offers complementary packages for downstream analysis of MAGs, including phylogenetic placement and functional annotation. Researchers who prefer working in R can integrate CheckM output with these packages for further analysis.
Running CheckM
The standard workflow has two stages. The first stage, lineage_wf, places each bin in a phylogenetic context and selects the appropriate marker set. The second stage, tree_qa, assesses the quality of the phylogenetic tree used for marker selection.
The command structure is:
checkm lineage_wf -x fa --tab_table -t 8 bins/ checkm_output/
The -x flag specifies the file extension of the bin files. The --tab_table flag produces a tab-delimited output that is easy to import into spreadsheet software. The -t flag sets the number of threads.
The output file contains columns for Bin Id, Marker lineage, number of genomes, number of markers, number of marker sets, marker copy number categories (0, 1, 2, 5, 10, 15, 20), Completeness, Contamination, and Strain heterogeneity. The columns labeled 0 through 20 represent the number of markers with that many copies. A marker with two copies contributes to the contamination estimate.
Interpreting CheckM Output
The Completeness column is the primary inclusion criterion. The Contamination column is the primary exclusion criterion. The Strain heterogeneity column provides additional context for interpreting contamination.
A MAG with 95 percent completeness and 2 percent contamination is generally acceptable for most downstream analyses. A MAG with 70 percent completeness and 0 percent contamination may be acceptable for some analyses but will produce biased gene content results. A MAG with 90 percent completeness and 12 percent contamination should be excluded or subjected to additional bin refinement.
When examining the marker copy number columns, pay attention to the distribution. A MAG with many markers at two copies but few at higher copy numbers suggests a single contaminant genome. A MAG with markers distributed across multiple copy number categories may contain several contaminants or may represent a mixed population.
BUSCO Workflow
Installation and Dataset Selection
BUSCO is distributed as a standalone tool and is also integrated into some workflow managers. The tool requires a lineage dataset that matches the taxonomic group of your MAGs. Common datasets include bacteria_odb10, archaea_odb10, and eukaryota_odb10.
The Galaxy Training Network offers accessible tutorials for running BUSCO and other genome quality assessment tools in a web-based environment. These tutorials are useful for researchers who prefer not to work exclusively on the command line and for those who want to learn the underlying concepts before implementing their own pipelines.
Running BUSCO
The command structure is:
busco -i bin.fasta -l bacteria_odb10 -o busco_output -m genome
The -i flag specifies the input FASTA file. The -l flag specifies the lineage dataset. The -o flag specifies the output directory name. The -m flag specifies the mode, which should be genome for MAGs.
The output includes a short summary file and a full table of results. The summary reports the number of complete, complete and single-copy, complete and duplicated, fragmented, and missing BUSCOs.
Interpreting BUSCO Output
The percentage of complete BUSCOs is the completeness estimate. The percentage of duplicated BUSCOs is a contamination indicator. Fragmented BUSCOs suggest assembly breaks. Missing BUSCOs may indicate true gene absence or assembly incompleteness.
BUSCO and CheckM often produce similar completeness estimates for high-quality MAGs. Discrepancies are more common for medium-quality MAGs. When the two tools disagree, examine the specific markers that differ and consider whether the discrepancy reflects a real biological difference or a methodological artifact.
Comparing CheckM and BUSCO
CheckM and BUSCO use different marker collections and different computational approaches. CheckM selects lineage-specific markers based on the phylogenetic placement of each bin. BUSCO uses fixed lineage datasets that are updated periodically.
| Feature | CheckM | BUSCO |
|---|---|---|
| Marker selection | Lineage-specific, based on phylogenetic placement | Fixed lineage datasets |
| Completeness estimate | Based on single-copy marker presence | Based on complete BUSCO percentage |
| Contamination estimate | Based on multi-copy marker frequency | Based on duplicated BUSCO percentage |
| Strain heterogeneity | Explicit score provided | Not directly reported |
| Output format | Tab-delimited table | Summary file and full table |
For most MAG quality assessment tasks, running both tools provides a more complete picture than running either alone. When the tools agree, confidence in the quality estimate increases. When they disagree, the discrepancy itself is informative and should be investigated.
Practical Workflow for MAG Quality Assessment
Step 1: Organize Your Bins
Create a directory containing one FASTA file per bin. Use a consistent naming convention that includes the sample identifier and a bin number. Record the binning method and the assembler used for each bin in a separate metadata file. This organization step prevents confusion when you have hundreds of bins from multiple samples.
Step 2: Run CheckM
Run the CheckM lineage workflow on the entire bin directory. Review the output table and identify bins that fall into the high-quality, medium-quality, and low-quality tiers. The lineage workflow can take several hours for large datasets, so plan your compute resources accordingly.
Step 3: Run BUSCO
Run BUSCO on each bin that passes an initial completeness threshold. The EMBL-EBI Training portal provides structured learning pathways for genome assembly and quality assessment that can help you interpret BUSCO results in a broader bioinformatics context. Running BUSCO on every bin may be computationally expensive, so a two-stage approach is often practical.
Step 4: Compare Results
Create a combined table that includes CheckM completeness, CheckM contamination, CheckM strain heterogeneity, BUSCO completeness, and BUSCO duplication. Flag bins where the two tools disagree by more than 10 percentage points on completeness.
Step 5: Apply Inclusion Criteria
Decide on your inclusion thresholds before you examine the results. This pre-registration step prevents bias. A common threshold for downstream comparative analysis is 90 percent completeness and 5 percent contamination. For exploratory analyses, 70 percent completeness and 10 percent contamination may be acceptable.
Step 6: Document Your Decisions
Record the quality metrics for every bin, including those that fail your thresholds. This documentation supports transparency and allows other researchers to reanalyze your data with different thresholds.
Records and Measurements
Quality assessment produces a set of records that should be preserved alongside the MAG sequences. The minimum record set includes the CheckM output table, the BUSCO summary files, the assembly statistics for each bin, and the metadata describing how each bin was generated.
The nf-core documentation describes community standards for reproducible bioinformatics workflows. Many nf-core pipelines include MAG quality assessment modules that produce standardized output files. Using a workflow manager ensures that your quality assessment is reproducible and that the records are complete.
For each MAG, record the following measurements:
- Number of contigs
- Total assembly length
- N50 contig length
- GC content
- CheckM completeness and contamination
- BUSCO completeness and duplication
- Number of tRNAs and rRNAs if predicted
- Taxonomic assignment
These measurements provide context for interpreting the quality metrics. A MAG with high completeness but very low N50 may be fragmented into many small contigs, which can affect downstream analyses that depend on contig order or local genomic context.
Common Failure Patterns
Failure Pattern 1: Overly Stringent Filtering
Some researchers exclude all MAGs below 90 percent completeness, even when their study question does not require near-complete genomes. This practice reduces statistical power and can bias comparative analyses toward easily assembled taxa. If your study focuses on gene content, medium-quality MAGs may be sufficient for detecting the presence of conserved pathways.
Failure Pattern 2: Ignoring Strain Heterogeneity
A MAG with 5 percent contamination and 100 percent strain heterogeneity may contain two nearly identical strains. The contamination estimate is low because the two strains share most markers. However, the mixed assembly will produce ambiguous variant calls and may inflate estimates of gene content diversity. Always check the strain heterogeneity score before proceeding with variant-based analyses.
Failure Pattern 3: Using Only One Quality Tool
CheckM and BUSCO measure overlapping but not identical properties. Relying on a single tool can miss contamination that the other tool would detect. The Carpentries lessons include foundational training in shell scripting and data management that can help you automate running both tools across large datasets.
Failure Pattern 4: Applying Eukaryotic Standards to Bacterial MAGs
Eukaryotic MAGs have different marker expectations and different contamination dynamics. The MIMAG standards were developed primarily for bacterial and archaeal MAGs. If you are working with eukaryotic MAGs, consult the lineage-specific BUSCO datasets and consider additional quality metrics such as completeness of the core eukaryotic gene set.
Failure Pattern 5: Failing to Document Quality Metrics
A published MAG that lacks quality metrics cannot be evaluated by other researchers. Journals increasingly require MIMAG-compliant reporting. Document your quality metrics in the methods section and provide the full quality table as supplementary material.
Limitations of Quality Metrics
Marker-Based Estimates Are Indirect
Completeness and contamination are estimated from marker genes, not from direct measurement of the assembled genome. A MAG could be missing genes that are not in the marker set, or it could contain contaminating sequence that does not include any marker genes. The estimates are useful but not exact.
Lineage-Specific Bias
CheckM selects markers based on the phylogenetic placement of the bin. If the placement is incorrect, the marker set will be inappropriate, and the quality estimates will be unreliable. BUSCO lineage datasets are fixed and may not be optimal for deeply branching or poorly sampled lineages.
Genome Size Variation
Completeness estimates assume that the target genome has a typical gene content for its lineage. Organisms with reduced genomes, such as obligate symbionts, will have fewer markers and may be scored as incomplete even when the assembly is correct. Organisms with expanded genomes may have duplicated markers that inflate contamination estimates.
Chimeric Bins
A bin that contains sequence from two distantly related organisms will have high contamination and low strain heterogeneity. A bin that contains sequence from two closely related organisms will have low contamination and high strain heterogeneity. Both situations produce unreliable MAGs, but they require different remediation strategies.
Remediation Strategies for Low-Quality MAGs
Rebinning
If a MAG has high contamination, the binning step may have merged two genomes. Review the coverage profile and GC content of the contigs in the bin. Contigs with distinct coverage or GC values may belong to different organisms. Separate them and create new bins.
Assembly Polishing
If a MAG has low completeness but low contamination, the assembly may be fragmented. Consider whether the sequencing depth was sufficient and whether the assembler parameters were appropriate. Some assemblers benefit from parameter tuning or from using a different k-mer size.
Long-Read Scaffolding
If short-read assemblies produce fragmented MAGs, consider generating long-read sequencing data to scaffold the assembly. Long reads can bridge repetitive regions and produce more complete genomes. The NCBI Data Resources provide access to sequence data and analysis tools that can support hybrid assembly approaches.
Removing Contaminating Contigs
For MAGs with moderate contamination, it may be possible to identify and remove contaminating contigs. Tools that calculate coverage and GC content per contig can help identify outliers. However, manual curation is time-consuming and may not be feasible for large datasets.
Reproducibility and Workflow Management
Version Control and Environment Management
Quality assessment results depend on the specific versions of CheckM, BUSCO, and their underlying databases. A completeness estimate of 92 percent from one version of a marker set may differ from the estimate produced by a newer version. Record the exact software versions and database versions in your methods. The Carpentries lessons provide foundational training in version control with Git and reproducible computing practices that apply directly to managing bioinformatics workflows.
Containerization
Container technologies package software with all dependencies and database versions. A containerized CheckM or BUSCO run produces results that can be reproduced exactly on any system. The nf-core documentation describes how community pipelines use containers to ensure reproducibility across computing environments. If you are developing your own workflow, consider containerizing the quality assessment steps.
Workflow Managers
Workflow managers track inputs, outputs, and parameters for each step in an analysis. They also handle reruns when inputs change. The Galaxy Training Network offers tutorials on using workflow managers for genome assembly and quality assessment. A workflow manager becomes valuable when you have hundreds of bins and need to track which version of each tool processed each bin.
Quality Control in Large Datasets
Scaling Considerations
Running CheckM and BUSCO on hundreds of bins requires substantial compute resources. The lineage workflow in CheckM builds a phylogenetic tree for each bin, which is computationally intensive. BUSCO runs are parallelizable across bins. Plan your compute allocation before starting a large-scale quality assessment.
Batch Effects
If you process bins from multiple sequencing runs or multiple studies, batch effects can influence quality metrics. Sequencing depth, library preparation, and assembly parameters can differ between batches. Compare quality metric distributions across batches to identify systematic differences that are unrelated to true biological variation.
Automated Flagging
For large datasets, automated flagging based on quality thresholds is practical. The Bioconductor project provides R packages for reading and manipulating tabular data that can help you build automated quality flagging workflows. Define your thresholds in a configuration file so that the flagging logic is transparent and reproducible.
Reporting Quality Metrics in Publications
When you publish MAGs, report the following information for each genome:
- Completeness estimate and the tool used to compute it
- Contamination estimate and the tool used to compute it
- Strain heterogeneity score if CheckM was used
- Number of contigs and N50
- Total genome size
- Sequencing platform and assembly method
- Binning method
- Quality assessment tool versions
This reporting standard allows other researchers to reproduce your quality assessment and to apply different thresholds if they wish.
Professional Escalation Criteria
You should escalate a MAG quality problem to a supervisor, collaborator, or bioinformatics specialist when:
- CheckM and BUSCO produce conflicting completeness estimates that differ by more than 15 percentage points
- A MAG passes quality thresholds but produces biologically implausible results in downstream analysis
- You are uncertain whether strain heterogeneity is affecting your variant calls
- You need to decide whether to generate additional sequencing data to improve MAG quality
- You are preparing MAGs for submission to a public database and need to ensure MIMAG compliance
A Decision Framework for MAG Quality Triage Across Large Datasets
When you have hundreds or thousands of MAGs from a large environmental or host-associated study, running individual quality assessments and manually reviewing each output table becomes impractical. The core problem shifts from understanding what completeness and contamination mean to deciding how to allocate finite compute time and analytical effort across a large bin set. A structured triage framework helps you process bins in stages, apply effort only where it changes your downstream conclusions, and document decisions in a way that survives peer review.
Stage 1: Coarse Filtering with Assembly Statistics
Before running CheckM or BUSCO on every bin, use assembly statistics that are cheap to compute and available directly from your binning output. The goal is to remove bins that cannot possibly meet quality thresholds and to prioritize compute resources for bins that are plausible genomes.
Compute for every bin: number of contigs, total assembly length, N50, GC content, and coverage. These values come from the assembler or from simple sequence statistics tools. The Bioconductor project provides R packages for reading FASTA files and computing these statistics programmatically, which is useful when you have thousands of bins and need a reproducible record of the filtering step.
Apply coarse exclusion rules based on biological plausibility. A bacterial MAG with a total assembly length below 400 kilobases is unlikely to represent a complete genome for most free-living bacteria, although obligate symbionts and some archaea have genuinely small genomes. A bin with more than 2000 contigs for a 3 megabase genome suggests extreme fragmentation that will likely produce low completeness scores. A bin with GC content that differs by more than 15 percentage points from the expected value for its taxonomic assignment may contain contamination or may be misassigned.
Record the number of bins removed at this stage and the reason for each removal. This record is important because reviewers will ask why certain bins were excluded from downstream analysis. The coarse filter is not a quality assessment. It is a resource allocation step that prevents you from spending hours running CheckM on bins that are obviously not usable genomes.
Stage 2: CheckM Lineage Workflow on the Filtered Set
Run the CheckM lineage workflow on the bins that survive the coarse filter. The lineage workflow is the most computationally expensive step because it places each bin in a phylogenetic context and selects appropriate marker sets. The Galaxy Training Network provides tutorials on running CheckM in a reproducible manner, which is valuable when you need to document the exact command and parameters used.
The output table gives you completeness, contamination, and strain heterogeneity for every bin. Sort the table by completeness and examine the distribution. A typical large dataset shows a bimodal distribution: a cluster of high-quality bins above 90 percent completeness and a long tail of fragmented bins below 70 percent. The shape of this distribution tells you about your assembly and binning quality overall, beyond about individual bins.
Apply your pre-registered thresholds at this stage. For most comparative genomics studies, bins below 70 percent completeness or above 10 percent contamination are excluded from downstream analysis. Bins between 70 and 90 percent completeness are flagged for BUSCO confirmation. Bins above 90 percent completeness and below 5 percent contamination proceed to the next stage.
Stage 3: BUSCO Confirmation for Borderline and High-Quality Bins
Run BUSCO on bins that pass the CheckM thresholds and on bins that fall in the borderline zone where CheckM and BUSCO are most likely to disagree. The EMBL-EBI Training portal provides structured learning pathways for genome quality assessment that explain how to select the appropriate BUSCO lineage dataset for your organism group.
BUSCO serves two purposes in the triage framework. First, it confirms the CheckM completeness estimate. When the two tools agree within 5 percentage points, you can report either value with confidence. When they disagree by more than 10 percentage points, the bin needs individual investigation. Second, BUSCO provides a standardized completeness value that is comparable across studies. Many journals and databases expect BUSCO completeness in addition to CheckM metrics.
The BUSCO duplication percentage serves as a contamination cross-check. A bin with low CheckM contamination but high BUSCO duplication may contain a closely related contaminant that CheckM markers did not detect. This situation is exactly where the strain heterogeneity score from CheckM becomes important.
Stage 4: Strain Heterogeneity Review for Contamination Interpretation
The CheckM strain heterogeneity score tells you whether multi-copy markers are identical or divergent. A score near 100 percent means the duplicated markers are nearly identical, which indicates the bin contains closely related strains. A score near 0 percent means the duplicated markers are divergent, which indicates the bin contains distantly related contaminants.
This distinction changes your remediation decision. A bin with 8 percent contamination and 100 percent strain heterogeneity likely contains two strains of the same species. The contamination estimate may overstate the problem because the two strains share most of their genome. However, the bin is still problematic for variant calling and for any analysis that assumes a single haploid genome. A bin with 8 percent contamination and 0 percent strain heterogeneity likely contains a distinct contaminant genome that can potentially be removed by rebinning.
Record the strain heterogeneity score for every bin that passes the completeness threshold. Even if you do not use the score as an exclusion criterion, it provides essential context for interpreting contamination and for deciding whether a bin is suitable for population genetics or variant-based analyses.
Stage 5: Taxonomic Consistency Checks
Quality metrics alone do not tell you whether a bin is biologically coherent. A bin can pass completeness and contamination thresholds yet contain contigs from two different species that happen to share marker genes. Taxonomic consistency checks add a layer of validation that marker-based metrics cannot provide.
For each bin that passes the quality thresholds, assign a taxonomy using a rapid classifier. Compare the taxonomic assignments of the contigs within the bin. If the bin contains contigs assigned to different phyla, the bin is chimeric regardless of its completeness and contamination scores. If the bin contains contigs assigned to different genera within the same family, investigate further before proceeding.
The NCBI Data Resources provide access to taxonomic databases and classification tools that support this consistency check. The NCBI taxonomy is the reference framework used by most classification tools, and understanding its structure helps you interpret the output of taxonomic assignment programs.
Stage 6: Final Inclusion Decision and Documentation
The final decision to include or exclude a MAG from downstream analysis should be based on the complete evidence set: CheckM completeness and contamination, BUSCO completeness and duplication, strain heterogeneity, assembly statistics, and taxonomic consistency. No single metric determines the decision.
Create a decision table with one row per bin and columns for each metric. Add a column for the final decision: include, exclude, or include with caveats. The caveats column records specific limitations, such as high strain heterogeneity that affects variant calling or moderate contamination that is acceptable for presence-absence analysis but not for quantitative comparisons.
The nf-core documentation describes community standards for reproducible bioinformatics workflows that include MAG quality assessment modules. Using a workflow manager to produce this decision table ensures that the process is reproducible and that the documentation is complete. The decision table itself becomes a supplementary data file for your publication.
A Record System for MAG Quality Decisions
The record system for MAG quality assessment should capture also the final metrics but also the reasoning behind each decision. A spreadsheet with metric columns is necessary but not sufficient. You need a structured record that links each bin to its source sample, its assembly and binning parameters, its quality metrics, and the rationale for its inclusion or exclusion.
Design a record with the following fields for each bin:
- Bin identifier and source sample identifier
- Assembler and version, binning tool and version
- Number of contigs, total length, N50, GC content, coverage
- CheckM completeness, contamination, strain heterogeneity, marker lineage
- BUSCO completeness, duplication, fragmentation, missing, lineage dataset
- Taxonomic assignment and confidence
- Final decision and rationale
- Date of assessment and software versions
Store this record in a version-controlled file. The Carpentries lessons provide foundational training in version control with Git and reproducible data management that applies directly to maintaining this record across the duration of a project. When you update a bin or rerun a quality assessment with a new software version, the version control history preserves the previous record.
The record system serves three purposes. First, it supports reproducibility. A reviewer can trace any MAG in your publication back to its raw quality metrics and the decisions that led to its inclusion. Second, it supports reanalysis. If a new quality assessment tool becomes available or if a reviewer requests a different threshold, you can reapply the framework without redoing the entire analysis. Third, it supports error detection. A bin that passes all quality thresholds but produces anomalous downstream results can be traced back to its quality record to identify potential causes.
Troubleshooting Common Triage Problems
Problem 1: CheckM and BUSCO Completeness Disagree by More Than 10 Percentage Points
This disagreement usually has one of three causes. The bin may be from a lineage that is poorly represented in one marker set. The bin may contain a large amount of sequence from mobile genetic elements or plasmids that inflate one estimate. The bin may be chimeric, with different regions supporting different marker sets.
Investigate by examining which markers are present and absent in each tool. CheckM reports the marker lineage it selected. BUSCO reports the lineage dataset used. If the CheckM marker lineage is different from the expected taxonomy of the bin, the phylogenetic placement may be incorrect. Consider running CheckM with a specific taxonomy instead of the lineage workflow.
Problem 2: Low Contamination but High Strain Heterogeneity
This combination indicates the bin contains multiple closely related strains. The contamination estimate is low because the strains share most marker genes. The strain heterogeneity score is high because the duplicated markers are nearly identical.
For downstream analysis, this bin is problematic for any variant-based analysis. The mixed strains will produce heterozygous calls that do not reflect a single genome. For gene content analysis, the bin may be acceptable if the strains have similar gene content. Document the strain heterogeneity in the record and note the limitation in the final decision.
Problem 3: High Contamination but Low Strain Heterogeneity
This combination indicates the bin contains a distinct contaminant genome. The contaminant is distantly related to the target, so its markers are divergent from the target markers. The contamination estimate is high because the contaminant contributes many multi-copy markers.
Consider rebinning the contaminated bin. Examine the coverage and GC content of the contigs. Contigs with distinct coverage or GC values may belong to the contaminant. Separate them and create new bins. The Bioconductor project provides R packages for visualizing coverage and GC distributions that support this manual curation step.
Problem 4: A Bin Passes All Quality Thresholds but Produces Implausible Downstream Results
Quality metrics are necessary but not sufficient for biological validity. A bin can pass completeness and contamination thresholds yet contain a genome rearrangement, a misassembled region, or a chimeric sequence that does not trigger marker-based flags.
Investigate the specific downstream result that seems implausible. If a metabolic pathway appears present but the genes are scattered across contigs in an unusual order, examine the assembly graph. If a phylogenetic placement is unexpected, check the taxonomic consistency of the contigs. The NCBI Data Resources provide tools for comparing your MAG against reference genomes that can help identify misassembled regions.
Problem 5: The Triage Framework Removes Too Many Bins
If the coarse filter removes a large fraction of your bins, the problem is likely in the assembly or binning stage, not in the quality assessment. Review the assembly parameters and consider whether the sequencing depth was sufficient. The Galaxy Training Network offers tutorials on assembly and binning that can help you diagnose why bins are fragmented or contaminated.
If the quality thresholds themselves are too stringent for your study question, revisit the thresholds. A study of gene presence in a poorly characterized environment may tolerate lower completeness than a study of core genome phylogeny. The thresholds should be justified by the downstream analysis requirements, not by convention.
Applying the Framework to Different Study Types
Comparative Genomics Studies
For comparative genomics, the inclusion threshold should be high because gene content comparisons are sensitive to both completeness and contamination. A missing gene due to incomplete assembly is indistinguishable from a true gene absence. A foreign gene due to contamination is indistinguishable from a true gene presence. Use the 90 percent completeness and 5 percent contamination thresholds as the default, and apply BUSCO confirmation to every included bin.
Metabolomic or Functional Potential Studies
For studies of metabolic potential, medium-quality MAGs may be sufficient for detecting the presence of conserved pathways. A 70 percent complete genome will contain most conserved metabolic genes, and the missing 30 percent is more likely to include strain-specific or accessory genes. However, absence calls are unreliable at this completeness level. If your study claims that a pathway is absent, you need high-quality MAGs to support that claim.
Phylogenetic and Taxonomic Studies
Phylogenetic placement can tolerate lower completeness because phylogenetic markers are a small subset of the genome. A 50 percent complete MAG may contain enough phylogenetic markers for a robust placement. However, contamination is more problematic because foreign markers will distort the phylogenetic signal. Apply a stricter contamination threshold for phylogenetic studies than for gene content studies.
Population Genetics and Variant Calling Studies
Population genetics requires MAGs with low strain heterogeneity because mixed strains produce false heterozygous calls. A MAG with 95 percent completeness, 2 percent contamination, and 100 percent strain heterogeneity is unsuitable for variant calling even though it passes standard quality thresholds. Document this limitation explicitly in the record and exclude such bins from variant-based analyses.
Escalation Criteria for the Triage Framework
Escalate to a bioinformatics specialist or supervisor when the triage framework produces results that cannot be resolved with standard troubleshooting. Specific escalation triggers include:
- More than 50 percent of bins fail the coarse filter, indicating a systematic assembly or binning problem
- CheckM and BUSCO disagree by more than 15 percentage points on multiple bins, suggesting a marker set or lineage assignment issue
- A bin passes all quality thresholds but consistently produces implausible results across multiple downstream analyses
- You need to decide whether to generate additional sequencing data to rescue a large fraction of low-quality bins
- You are preparing MAGs for submission to a public database and need to ensure the quality assessment meets repository requirements
The EMBL-EBI Training portal provides advanced learning pathways that can help you build the skills to resolve these issues independently. The nf-core documentation describes community pipelines that may already implement the triage framework you need, saving you from building it from scratch.
Frequently Asked Questions
What is the difference between completeness and genome size?
Completeness is an estimate of the fraction of the expected gene content that is present in the assembly. Genome size is the total number of base pairs in the assembly. A MAG can have a large genome size but low completeness if it contains substantial contaminating sequence. A MAG can have a small genome size but high completeness if the organism has a small genome.
How do I choose between CheckM and BUSCO for MAG quality assessment?
Run both tools and compare the results. CheckM provides a strain heterogeneity score that BUSCO does not offer. BUSCO provides a standardized lineage dataset that is updated regularly. When the tools agree, you can report either or both. When they disagree, investigate the cause before proceeding.
What completeness threshold should I use for my study?
The threshold depends on your research question. For comparative genomics and gene content analysis, 90 percent completeness is a common minimum. For presence-absence detection of conserved pathways, 70 percent may be sufficient. For phylogenetic placement, even medium-quality MAGs can be informative. Decide your threshold before examining your results.
How does strain heterogeneity affect downstream analysis?
Strain heterogeneity indicates that the MAG contains sequence from multiple closely related strains. This mixing produces ambiguous variant calls, inflates estimates of gene content diversity, and can distort phylogenetic placement. If strain heterogeneity is high, consider whether your analysis can tolerate this ambiguity or whether you need to separate the strains.
Can I improve a low-quality MAG?
Yes, in some cases. Rebinning can separate contaminated bins. Additional sequencing, especially long-read sequencing, can improve completeness. Parameter tuning of the assembler can reduce fragmentation. However, some MAGs cannot be improved with available data, and you may need to exclude them from downstream analysis.
What is the MIMAG standard and why does it matter?
MIMAG is the Minimum Information about a Metagenome-Assembled Genome standard. It defines quality tiers based on completeness and contamination and specifies what metadata should be reported. Journals and public databases increasingly require MIMAG-compliant reporting. Following the standard makes your MAGs usable by the broader research community.
How do I report MAG quality in a publication?
Report the completeness and contamination estimates with the tool versions used to compute them. Include the strain heterogeneity score if available. Provide the number of contigs, N50, and total genome size. State the assembly and binning methods. Include the full quality table as supplementary material.
What should I do if my MAGs fail quality thresholds?
First, determine whether the failure is due to contamination, incompleteness, or strain heterogeneity. Then consider remediation options such as rebinning, additional sequencing, or contig removal. If remediation is not feasible, exclude the MAGs from analyses that require high quality and document the exclusion in your methods.
Related Bioinformatics Guides
- Metagenome Assembled Genome Analysis: From Bins to Biological Insights
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Evaluating Genome Assembly Quality: Metrics and Tools
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Binning in Metagenomics: From Contigs to Genomes
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Sleep and the athlete: narrative review and 2021 expert consensus recommendations.. British journal of sports medicine, 2020.
- Inotersen Treatment for Patients with Hereditary Transthyretin Amyloidosis.. The New England journal of medicine, 2018.
- Gerontechnology.. IEEE engineering in medicine and biology magazine : the quarterly magazine of the Engineering in Medicine & Biology Society, 2008.
- Effects of repetitive transcranial magnetic stimulation of the left dorsolateral prefrontal cortex on symptom domains in neuropsychiatric disorders: a systematic review and cross-diagnostic meta-analysis.. The lancet. Psychiatry, 2023.
- X-Linked Hypophosphatemia Management in Children: An International Working Group Clinical Practice Guideline.. The Journal of clinical endocrinology and metabolism, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.