Assembly Quality Metrics for Metagenomes: Adapting N50 and BUSCO for Mixed Communities
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Metagenome assembly quality assessment necessitates metrics beyond single-genome standards like N50, which can be misleading due to varying organism abundance and genomic complexity within mixed communities. A high N50 might reflect a dominant genome, not overall community representation.
- Completeness and contamination are the primary quality axes for Metagenome-Assembled Genomes (MAGs), typically estimated using single-copy marker genes. CheckM and BUSCO employ lineage-specific or universal marker sets, respectively, to infer these metrics.
- Reference-free misassembly detection, such as the deep learning approach ResMiCo, is crucial as standard completeness/contamination metrics do not identify structural errors within contigs, which can arise from sequencing errors or repetitive regions.
- Contiguity (e.g., N50) alone does not predict biological utility; long-read assemblers may produce contiguous assemblies but fail to recover certain MAGs, highlighting the need for evaluating biological completeness and contamination.
- Establishing pre-assembly quality targets based on downstream analysis requirements, and implementing a two-stage quality gate (whole-assembly and MAG-level), are critical for ensuring the assembly output is fit for purpose.
- Troubleshooting poor assembly quality requires a systematic approach, verifying read quality, coverage depth, appropriate assembly strategy for sequencing technology, and the effectiveness of binning and refinement steps before reconsidering quality targets.
Metagenome assembly quality assessment requires metrics that account for mixed species, uneven coverage, and the absence of a single reference genome. Standard single-genome metrics like N50 and BUSCO completeness scores lose their straightforward interpretation when applied to a community of organisms with varying abundance and genomic complexity. This article explains how to adapt assembly quality metrics for metagenomic data, focusing on the practical use of CheckM, BUSCO with metagenomic mode, and contig-level metrics to evaluate and improve metagenome assemblies.
The Problem with Single-Genome Metrics in Metagenomes
Why N50 Misleads in Mixed Communities
The N50 statistic, defined as the contig length at which half of the assembled bases are contained in contigs of that length or longer, was designed for single-genome assemblies. In a metagenome, the assembly graph contains multiple genomes simultaneously, each with different coverage depths and repeat structures. A high N50 in a metagenome assembly may reflect the successful assembly of one dominant high-coverage genome while the rest of the community remains fragmented. Conversely, a low N50 may indicate either genuine fragmentation across the community or the presence of many low-abundance organisms that are inherently difficult to assemble.
The 2023 benchmark of 19 assembly tools applied to metagenomic datasets from simulations, mock communities, and human gut microbiomes found that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality metagenome-assembled genomes (MAGs) [<a href="#ref-1">1</a>]. This observation demonstrates that contiguity alone does not predict biological utility. A metagenome assembly with excellent N50 values can still miss entire genomes from less abundant community members.
The Reference-Free Challenge
Single-genome assembly quality is often assessed by comparing the assembly to a known reference genome. Metagenomes lack this luxury. The community composition is often unknown, and many organisms present have no cultured representatives or reference genomes. This means that quality assessment must rely on reference-free approaches that infer completeness and contamination from the assembly itself.
The challenge of detecting misassemblies in taxonomically novel genomic data is substantial. Sequencing errors, variable coverage, repetitive genomic regions, and other factors can produce misassemblies that are difficult to detect when no reference genome exists [<a href="#ref-2">2</a>]. The state of the art in reference-free misassembly prediction has historically shown limited accuracy, which motivated the development of deep learning approaches like ResMiCo [<a href="#ref-2">2</a>].
Core Principles of Metagenome Assembly Quality
Completeness and Contamination as Primary Axes
For metagenome-assembled genomes, the two most important quality axes are completeness and contamination. Completeness measures what fraction of a genome is present in the assembly. Contamination measures what fraction of the assembled sequence comes from other organisms. Both metrics are typically estimated using single-copy marker genes that are expected to be present exactly once in a bacterial or archaeal genome.
CheckM and BUSCO both use this marker gene approach, but they differ in their marker sets and methodology. CheckM uses lineage-specific marker sets that are selected based on the phylogenetic placement of the genome being assessed. BUSCO uses a fixed set of universal single-copy orthologs for the domain of interest, such as Bacteria or Archaea.
The Minimum Information about a Metagenome-Assembled Genome Standard
The genomics community has established thresholds for classifying MAG quality. A high-quality MAG is generally defined as having greater than 90 percent completeness and less than 5 percent contamination. A medium-quality MAG has greater than 50 percent completeness and less than 10 percent contamination. These thresholds provide a common language for comparing assemblies across studies, but they are minimum standards instead of guarantees of biological accuracy.
The 2020 discussion of genome curation from metagenomes emphasized that gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of MAGs [<a href="#ref-3">3</a>]. Even high-quality MAGs by the standard metrics may contain errors that affect downstream metabolic and evolutionary analyses. The authors verified the value of cumulative GC skew in combination with other metrics to establish bacterial genome sequence accuracy, identifying potential misassemblies in some reference genomes of isolated bacteria [<a href="#ref-3">3</a>].
At a Glance: Metagenome Assembly Quality Metrics
| Metric | What It Measures | Metagenome Adaptation | Primary Limitation |
|---|---|---|---|
| N50 / L50 | Contig length distribution | Report separately for total assembly and for binned MAGs | Does not indicate biological completeness or contamination |
| CheckM completeness and contamination | Presence of lineage-specific single-copy marker genes | Applied to individual bins or MAGs, not the whole assembly | Requires taxonomic placement, may misestimate for novel lineages |
| BUSCO completeness | Presence of universal single-copy orthologs | Metagenomic mode uses reduced marker sets and accounts for missing lineages | Fixed marker set may not represent all community members |
| Number of MAGs recovered | Count of bins meeting quality thresholds | Assesses biological utility of the assembly | Depends on binning success, beyond assembly quality |
| Misassembly rate | Fraction of contigs with structural errors | Reference-free tools like ResMiCo can estimate this | No perfect reference-free method exists |
Practical Workflow for Metagenome Assembly Quality Assessment
Step 1: Assess Raw Assembly Statistics
Before evaluating biological completeness, examine the basic assembly statistics. These include total assembled length, number of contigs, N50, L50, and the length distribution of contigs. For metagenomes, also record the number of contigs longer than 1 kilobase and longer than 5 kilobases, as these are the contigs most likely to be useful for downstream analysis.
The 2026 methods chapter on metagenomic assembly and gene prediction outlines core assembly strategies including per-sample versus co-assembly and short-read versus hybrid approaches, and highlights key parameters and metrics for evaluating assembly quality [<a href="#ref-4">4</a>]. The choice of assembly strategy directly affects the contig statistics you will observe. Co-assembling multiple samples can improve contiguity for shared community members but may reduce the ability to resolve strain-level variation.
Step 2: Bin the Assembly into Population Genomes
Raw assembly statistics describe the entire community as a single unit, which is not biologically meaningful. The next step is to bin the contigs into groups that represent individual population genomes. This binning step is typically performed using coverage depth and nucleotide composition features. The resulting bins are candidate MAGs.
The 2025 MetaflowX workflow integrates short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality MAG identification, and bin refinement and reassembly [<a href="#ref-5">5</a>]. This integrated approach demonstrates that assembly and binning are not independent steps. The quality of the assembly directly constrains the quality of the bins that can be recovered from it.
Step 3: Evaluate Completeness and Contamination for Each MAG
For each bin, run CheckM or BUSCO to estimate completeness and contamination. These estimates are based on the presence and multiplicity of single-copy marker genes. A complete genome should have each marker gene present exactly once. Missing markers indicate incompleteness. Markers present more than once suggest contamination from closely related organisms.
The 2023 benchmark study found that linked-read assemblers obtained the highest number of overall near-complete MAGs from human gut microbiomes, while hybrid assemblers using both short- and long-read sequencing were promising methods to improve both total assembly length and the number of near-complete MAGs [<a href="#ref-1">1</a>]. These findings illustrate that the choice of sequencing technology and assembly strategy affects also contiguity but also the number of biologically useful MAGs that can be recovered.
Step 4: Check for Misassemblies
Completeness and contamination metrics do not detect all assembly errors. A contig can be chimeric, joining sequence from two different organisms, while still containing the expected set of marker genes. Misassemblies can also occur within a single genome, joining regions that are not adjacent in the true genome.
Reference-free misassembly detection is an active area of research. The ResMiCo deep learning approach was developed specifically for reference-free identification of misassembled contigs and was shown to be substantially more accurate than the state of the art, with robustness to novel taxonomic diversity and varying assembly methods [<a href="#ref-2">2</a>]. ResMiCo estimated 7 percent misassembled contigs per metagenome across multiple real-world datasets [<a href="#ref-2">2</a>]. This finding suggests that misassemblies are common enough to warrant routine screening.
Step 5: Refine and Reassemble Where Needed
MAGs that fall below quality thresholds can sometimes be improved through refinement and reassembly. The MetaflowX workflow includes a dedicated reassembly module that improved MAG quality, increasing completeness by 5.6 percent and reducing contamination by 53 percent on average [<a href="#ref-5">5</a>]. This refinement step can rescue MAGs that would otherwise be discarded.
Refinement typically involves removing contigs that appear to be contaminants, splitting bins that contain multiple genomes, and reassembling the remaining reads to close gaps. The decision to refine a MAG should be based on its intended use. A MAG intended for metabolic reconstruction requires higher quality than one intended for taxonomic assignment.
Options and Tradeoffs in Assembly Strategies
Short-Read versus Long-Read Assembly
Short-read sequencing, such as Illumina, has been the workhorse of metagenomics due to its low cost and high accuracy. However, short reads cannot span long repetitive regions, leading to fragmented assemblies. Long-read sequencing, such as PacBio and Oxford Nanopore, provides long-range DNA connectedness that can resolve repeats and produce more contiguous assemblies.
The 2023 benchmark found that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality MAGs [<a href="#ref-1">1</a>]. This counterintuitive result may reflect the higher error rates of long-read sequencing or the difficulty of binning long contigs that span multiple genomes. The benchmark also found that hybrid assemblers using both short- and long-read sequencing were promising methods to improve both total assembly length and the number of near-complete MAGs [<a href="#ref-1">1</a>].
Per-Sample versus Co-Assembly
Per-sample assembly treats each sequencing library independently. Co-assembly combines reads from multiple samples before assembly. Co-assembly can improve contiguity for organisms shared across samples by increasing coverage depth. However, co-assembly can also create chimeric contigs that join sequences from closely related but distinct strains present in different samples.
The 2026 methods chapter on metagenomic assembly and gene prediction outlines per-sample versus co-assembly as a core assembly strategy decision [<a href="#ref-4">4</a>]. The choice depends on the research question. If the goal is to compare community composition across samples, per-sample assembly may be more appropriate. If the goal is to recover high-quality MAGs from a shared community, co-assembly may be more effective.
Assembly Tool Selection
The 2023 benchmark evaluated 19 assembly tools and found substantial variation in their performance across different sequencing technologies and datasets [<a href="#ref-1">1</a>]. No single tool performed best across all criteria. The benchmark also discussed running time and peak memory consumption, providing practical guidance on tool selection [<a href="#ref-1">1</a>].
When selecting an assembly tool, consider the sequencing technology used, the expected community complexity, and the computational resources available. Tools optimized for single-genome assembly may not perform well on metagenomes. Tools designed for metagenomes may have different strengths and weaknesses depending on whether they were optimized for short reads, long reads, or hybrid data.
Observations and Measurements for Quality Control
Recording Assembly Statistics
Maintain a structured record of assembly statistics for each sample or dataset. This record should include the assembly tool and version, the parameters used, the input read counts and quality metrics, and the output assembly statistics. This information is essential for reproducibility and for comparing assemblies across studies.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analysis [<a href="#ref-6">6</a>]. Following reproducible workflow practices ensures that assembly quality metrics can be interpreted in the context of the specific methods used to generate them.
Tracking MAG Quality Across the Workflow
For each MAG, record the completeness and contamination estimates from CheckM or BUSCO, the number of contigs, the total length, and the N50. Also record the taxonomic assignment and the coverage depth. This information allows you to track how MAG quality changes through refinement and reassembly steps.
The nf-core documentation describes community pipeline standards for reproducible workflow usage and configuration [<a href="#ref-7">7</a>]. Applying these standards to metagenome assembly ensures that quality metrics are computed consistently across samples and studies.
Benchmarking Against Known Communities
Mock communities with known composition provide a valuable control for metagenome assembly quality assessment. By assembling a mock community with known genome sequences, you can directly measure the sensitivity and precision of your assembly and binning pipeline. The 2023 benchmark used simulated and mock community datasets for this purpose [<a href="#ref-1">1</a>].
If mock community data are not available, consider using a subset of your data for parameter optimization. The ResMiCo study demonstrated how misassembly prediction can be used to optimize metagenome assembly hyperparameters to improve accuracy instead of optimizing solely for contiguity [<a href="#ref-2">2</a>]. This approach can be applied to your own data to find assembly parameters that minimize misassemblies.
Common Failure Patterns in Metagenome Assembly Quality
Overestimating Completeness
Completeness estimates based on marker genes can be inflated by contamination. If a bin contains two closely related genomes, each contributing a different set of marker genes, the combined marker set may appear complete even though neither genome is fully represented. This failure mode is particularly problematic for closely related strains that are difficult to separate during binning.
Underestimating Contamination
Contamination estimates can be underestimated when the contaminating DNA comes from a distantly related organism that does not share marker genes with the target genome. In this case, the contaminating contigs may not be flagged by marker gene analysis. Reference-free misassembly detection methods like ResMiCo can help identify these cases [<a href="#ref-2">2</a>].
Confusing Contiguity with Quality
A highly contiguous assembly is not necessarily a high-quality assembly. Long contigs can contain misassemblies that are not apparent from length statistics alone. The 2023 benchmark found that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality MAGs [<a href="#ref-1">1</a>]. This finding underscores the importance of evaluating biological completeness and contamination in addition to contiguity.
Ignoring Strain-Level Variation
Metagenome assemblies typically collapse closely related strains into a single consensus genome. This collapse can create chimeric contigs that combine sequences from multiple strains. The resulting MAG may have inflated contamination estimates or may fail to represent any single strain accurately. Strain-level variation is a fundamental limitation of metagenome assembly that cannot be fully resolved with current methods.
Limitations of Current Quality Metrics
Marker Gene Bias
Both CheckM and BUSCO rely on marker genes that are conserved across broad phylogenetic groups. Organisms that lack these markers, such as highly reduced genomes or novel lineages, will have underestimated completeness. The fixed marker set used by BUSCO may not represent all community members, particularly in environments with high novel diversity.
Taxonomic Placement Dependence
CheckM requires taxonomic placement of the genome being assessed to select appropriate lineage-specific marker sets. For taxonomically novel genomes, this placement may be uncertain, leading to inaccurate completeness and contamination estimates. The 2020 discussion of genome curation noted that the vast majority of microbial life has not been cultured, and genome-resolved metagenomics can circumvent this limitation by obtaining MAGs [<a href="#ref-3">3</a>].
Reference-Free Misassembly Detection Limits
Reference-free misassembly detection has historically shown limited accuracy. The ResMiCo study noted that accuracy for the state of the art in reference-free misassembly prediction did not exceed an AUPRC of 0.57, and it was not clear how well these models generalize to real-world data [<a href="#ref-2">2</a>]. While ResMiCo substantially improved on this performance, no reference-free method can detect all misassemblies.
The Completeness-Contamination Tradeoff
There is an inherent tradeoff between completeness and contamination in metagenome binning. Aggressive binning that includes more contigs can increase completeness but also increases the risk of including contaminating sequence. Conservative binning reduces contamination but may leave out legitimate genome sequence. The quality thresholds for MAG classification represent a compromise between these competing goals.
Safety and Reproducibility Context
Computational Reproducibility
Metagenome assembly quality assessment should be reproducible. This requires documenting the exact software versions, parameters, and input data used for each analysis step. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-8">8</a>]. Following these practices ensures that quality metrics can be verified and compared across studies.
Data Management
Metagenome assembly produces large intermediate files, including assembly graphs, binning results, and quality assessment outputs. These files should be managed systematically to ensure that quality metrics can be traced back to the specific assembly and binning steps that produced them. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support data management and sharing [<a href="#ref-9">9</a>].
Training and Skill Development
Metagenome assembly quality assessment requires specialized skills in bioinformatics. The EMBL-EBI Training program provides bioinformatics learning pathways, data-resource training, and practical analysis education [<a href="#ref-10">10</a>]. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training context [<a href="#ref-11">11</a>]. Investing in these skills improves the reliability of quality assessment and reduces the risk of misinterpretation.
Professional Escalation Criteria
When to Seek Expert Consultation
If your metagenome assembly shows consistently poor quality metrics across multiple assembly strategies, or if the quality metrics are contradictory, consider seeking consultation from a bioinformatics specialist. Specific situations that warrant escalation include:
- Completeness and contamination estimates that are inconsistent with the expected community composition
- A high proportion of contigs flagged as misassembled by reference-free methods
- MAGs that fail to improve after multiple rounds of refinement and reassembly
- Uncertainty about the appropriate marker gene set for novel or unusual lineages
When to Reconsider the Experimental Design
Poor assembly quality can sometimes be traced to the experimental design instead of the assembly methods. If coverage depth is too low for the community complexity, no assembly strategy will produce high-quality MAGs. If the community contains many closely related strains, assembly may be fundamentally limited regardless of the tools used.
The 2023 benchmark provided practical guidance on selecting appropriate metagenome assembly tools based on sequencing technology and dataset characteristics [<a href="#ref-1">1</a>]. If your assembly quality is consistently poor, review whether the sequencing technology and depth are appropriate for your research question and community type.
A Decision Framework for Setting Assembly Quality Targets Before You Assemble
Defining Quality Requirements from the Downstream Analysis Plan
Assembly quality metrics are only meaningful when interpreted against the biological questions the assembly must answer. A metagenome assembly destined for taxonomic profiling of dominant community members has different quality requirements than one intended for metabolic reconstruction of rare organisms or for comparative genomics of closely related strains. Setting explicit quality targets before assembly prevents the common failure of optimizing for generic metrics like N50 without regard to whether the assembly can support the intended analyses.
The 2026 methods chapter on metagenomic assembly and gene prediction emphasizes that assembly connects quality-controlled reads to downstream microbiome analyses, and that the resulting contigs and gene sets provide essential input for MAG reconstruction as well as taxonomic and functional annotation [<a href="#ref-4">4</a>]. This connection means the quality bar must be set by the downstream analysis, not by the assembler output. For example, a study aiming to recover near-complete genomes for pangenome analysis requires a higher proportion of MAGs meeting the high-quality threshold than a study aiming only to profile community functional potential through gene catalogs.
Define three quality tiers before starting assembly. The first tier is the minimum acceptable quality for any contig or MAG to be included in downstream analysis. The second tier is the target quality for the primary biological conclusions. The third tier is the quality level that would allow the most demanding analyses, such as structural variant detection or strain-level comparisons. Write these tiers down with specific numeric thresholds for completeness, contamination, and contiguity. The Minimum Information about a Metagenome-Assembled Genome standard provides a common language with its high-quality threshold of greater than 90 percent completeness and less than 5 percent contamination, and its medium-quality threshold of greater than 50 percent completeness and less than 10 percent contamination. These community thresholds are the starting point for defining your own tiers, but they are minimum standards instead of guarantees of biological accuracy [<a href="#ref-3">3</a>].
Matching Sequencing Strategy to Quality Targets
The sequencing technology and depth chosen before assembly determine the ceiling for achievable assembly quality. The 2023 benchmark of 19 assembly tools applied to metagenomic datasets from simulations, mock communities, and human gut microbiomes found that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality MAGs [<a href="#ref-1">1</a>]. This finding has direct implications for target setting. If your quality targets emphasize contiguity for resolving repetitive genomic regions, long-read sequencing may be appropriate. If your targets emphasize recovering the maximum number of high-quality MAGs, the benchmark found that linked-read assemblers obtained the highest number of overall near-complete MAGs from human gut microbiomes, while hybrid assemblers using both short- and long-read sequencing were promising methods to improve both total assembly length and the number of near-complete MAGs [<a href="#ref-1">1</a>].
The same benchmark also discussed running time and peak memory consumption of assembly tools, providing practical guidance on selecting them [<a href="#ref-1">1</a>]. Computational resources constrain the feasible assembly strategy and therefore the achievable quality targets. A laboratory with limited compute capacity may need to accept lower contiguity targets or focus on a subset of samples for deep assembly. Document these resource constraints alongside the quality targets so that assembly choices can be traced back to explicit tradeoffs.
Coverage depth planning should be tied to the abundance distribution of the community. Low-abundance organisms require more sequencing depth to achieve the coverage necessary for assembly and binning. If the research question targets rare community members, the quality targets for those organisms must account for the fact that they will have lower coverage and therefore lower expected completeness. The 2023 benchmark used datasets from simulation, mock communities, and human gut microbiomes generated with mainstream sequencing platforms including Illumina and BGISEQ short-read sequencing, 10x Genomics linked-read sequencing, and PacBio and Oxford Nanopore long-read sequencing [<a href="#ref-1">1</a>]. These diverse datasets illustrate that sequencing platform choice affects which organisms can be assembled to a given quality level.
Building a Pre-Assembly Quality Target Document
Create a written quality target document before running any assembler. This document serves as the reference against which assembly output is judged. Include the following elements for each sample or sample group.
First, state the biological questions the assembly must answer. Write these as specific analysis tasks, such as recovering genomes of organisms above 1 percent relative abundance, identifying virulence genes in the community, or comparing metabolic pathways across samples. Each task implies different quality requirements.
Second, specify the quality tiers with numeric thresholds. For each tier, define the minimum completeness, maximum contamination, minimum contig length for inclusion, and the minimum number of MAGs that must meet the tier. Use the community standards as anchors but adjust them based on the biological questions. For example, a study of a poorly characterized environment with high novel diversity may need to accept medium-quality MAGs as the primary output because the marker gene sets may underestimate completeness for novel lineages.
Third, document the sequencing strategy and expected coverage. Record the sequencing platform, read length, insert size, and planned depth. Note the expected tradeoffs based on published benchmarks. The 2023 benchmark provides a basis for these expectations, including the finding that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality MAGs [<a href="#ref-1">1</a>].
Fourth, specify the assembly and binning tools to be used, including versions and parameters. The nf-core documentation describes community pipeline standards for reproducible workflow usage and configuration [<a href="#ref-7">7</a>]. Following these standards ensures that the quality target document can be executed consistently.
Fifth, define the decision rules for when to accept, refine, or reject an assembly. These rules should specify what happens when the assembly falls short of each quality tier. For example, if the assembly fails to recover any MAGs meeting the high-quality threshold, the rule might trigger a switch to hybrid assembly or additional sequencing.
Implementing a Two-Stage Quality Gate
A two-stage quality gate provides a structured way to evaluate assembly output against the pre-defined targets. The first stage evaluates the whole-assembly statistics. The second stage evaluates the MAG-level quality after binning. This separation prevents premature rejection of an assembly based on whole-assembly metrics that may be misleading for mixed communities.
The first gate uses whole-assembly statistics including total assembled length, number of contigs, N50, L50, and the number of contigs longer than 1 kilobase and longer than 5 kilobases. These statistics describe the overall assembly but do not indicate biological completeness or contamination. The 2026 methods chapter highlights key parameters and metrics for evaluating assembly quality, including the distinction between per-sample and co-assembly strategies [<a href="#ref-4">4</a>]. At this first gate, compare the observed statistics to the expected ranges from the quality target document. If the assembly falls far outside these ranges, investigate whether the sequencing depth was sufficient or whether the assembly parameters need adjustment.
The second gate evaluates individual MAGs after binning. For each bin, run CheckM or BUSCO to estimate completeness and contamination. Compare each MAG against the quality tiers defined in the target document. The 2025 MetaflowX workflow integrates short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality MAG identification, and bin refinement and reassembly [<a href="#ref-5">5</a>]. This integrated approach demonstrates that the second gate is not a single evaluation but a process that can include refinement and reassembly to bring MAGs up to quality targets.
The two-stage gate also provides a natural point for misassembly screening. The ResMiCo deep learning approach was developed for reference-free identification of misassembled contigs and was shown to be substantially more accurate than the state of the art, with robustness to novel taxonomic diversity and varying assembly methods [<a href="#ref-2">2</a>]. ResMiCo estimated 7 percent misassembled contigs per metagenome across multiple real-world datasets [<a href="#ref-2">2</a>]. Apply misassembly screening at both gates. At the first gate, screen a sample of contigs to estimate the overall misassembly rate. At the second gate, screen all contigs in MAGs that will be used for detailed analysis.
Recording Quality Decisions and Outcomes
Maintain a structured record of every quality decision made during the assembly and evaluation process. This record should include the sample identifier, the assembly tool and version, the parameters used, the quality tier being targeted, the observed metrics, and the decision to accept, refine, or reject. This record serves multiple purposes.
First, it provides the data needed to compare assembly strategies across samples. If one assembly strategy consistently fails to meet quality targets while another succeeds, the record shows which strategy works for which sample types. The 2023 benchmark provided practical guidance on selecting appropriate metagenome assembly tools based on sequencing technology and dataset characteristics [<a href="#ref-1">1</a>]. Your own records extend this guidance to your specific sample types and research questions.
Second, the record supports reproducibility. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-8">8</a>]. Following these practices ensures that quality decisions can be verified and compared across studies. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics analysis [<a href="#ref-6">6</a>]. These resources support the development of reproducible quality assessment workflows.
Third, the record documents the limitations of the final dataset. When publishing results, the quality record allows you to state precisely which MAGs meet which quality tiers and which analyses are supported by the assembly quality. The 2020 discussion of genome curation from metagenomes emphasized that gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of MAGs [<a href="#ref-3">3</a>]. A transparent quality record communicates these limitations to downstream users of the data.
Troubleshooting When Quality Targets Are Not Met
When an assembly fails to meet the pre-defined quality targets, work through a systematic troubleshooting sequence instead of immediately changing assembly parameters at random.
First, verify that the input reads passed quality control. Poor quality reads produce poor assemblies regardless of the assembler or parameters. Check the read quality metrics and repeat the quality control step if necessary. The 2025 MetaflowX workflow includes short-read quality control as the first module in its integrated framework [<a href="#ref-5">5</a>]. This placement reflects the fundamental importance of read quality for assembly outcomes.
Second, check whether the coverage depth is sufficient for the community complexity. Low coverage of low-abundance organisms will limit their assembly completeness. If the quality targets include rare organisms, additional sequencing may be necessary. The 2023 benchmark datasets included simulations, mock communities, and human gut microbiomes, which have different complexity and abundance distributions [<a href="#ref-1">1</a>]. The appropriate coverage depth depends on the community type and the quality targets.
Third, evaluate whether the assembly strategy matches the sequencing technology. The 2023 benchmark found substantial variation in performance across 19 assembly tools applied to different sequencing technologies [<a href="#ref-1">1</a>]. A tool optimized for short reads may perform poorly on long-read data and vice versa. Review the benchmark findings for guidance on tool selection for your sequencing platform.
Fourth, examine whether the binning step is limiting MAG quality. Assembly quality and binning quality are intertwined. The 2025 MetaflowX workflow demonstrated that a dedicated reassembly module improved MAG quality, increasing completeness by 5.6 percent and reducing contamination by 53 percent on average [<a href="#ref-5">5</a>]. If MAGs are falling short of quality targets, the problem may be in the binning or refinement steps instead of the assembly itself.
Fifth, consider whether the quality targets themselves are realistic for the data. If the community contains many closely related strains, assembly may be fundamentally limited regardless of the tools used. If the community has high novel diversity, marker gene based completeness estimates may be inaccurate. The 2020 discussion of genome curation noted that the vast majority of microbial life has not been cultured, and genome-resolved metagenomics can circumvent this limitation by obtaining MAGs [<a href="#ref-3">3</a>]. However, novel lineages may lack the marker genes used for quality assessment, leading to underestimated completeness.
Comparing Assembly Strategies Against Quality Targets
The quality target document provides the basis for systematic comparison of assembly strategies. Instead of comparing assemblies solely on N50 or other contiguity metrics, compare them on their ability to meet the pre-defined quality tiers.
For each assembly strategy under consideration, run the full assembly and binning workflow and evaluate the output against the quality targets. Record the number of MAGs meeting each quality tier, the completeness and contamination distributions across all MAGs, and the misassembly rate. The 2023 benchmark used this approach to evaluate 19 assembly tools against many criteria, revealing that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality MAGs [<a href="#ref-1">1</a>]. This finding would not have emerged from a comparison based solely on contiguity metrics.
The ResMiCo study demonstrated how misassembly prediction can be used to optimize metagenome assembly hyperparameters to improve accuracy instead of optimizing solely for contiguity [<a href="#ref-2">2</a>]. This approach can be applied to your own data to find assembly parameters that minimize misassemblies while meeting completeness and contamination targets. The quality target document provides the objective function for this optimization.
When comparing strategies, also record the computational cost. The 2023 benchmark discussed running time and peak memory consumption of assembly tools [<a href="#ref-1">1</a>]. These practical constraints may determine which strategy is feasible for your laboratory. A strategy that produces slightly better quality metrics but requires ten times the compute resources may not be the best choice for routine use.
Setting Realistic Expectations for Novel Communities
Communities with high novel diversity present special challenges for quality target setting. The marker gene sets used by CheckM and BUSCO are based on known genomes. Organisms from poorly characterized environments may lack these markers, leading to underestimated completeness. The 2020 discussion of genome curation noted that the bottleneck imposed by the requirement for isolates precluded genomic insights for the vast majority of microbial life, and that shotgun sequencing of microbial communities can circumvent this limitation by obtaining MAGs [<a href="#ref-3">3</a>]. However, the quality assessment tools themselves depend on the accumulated knowledge of microbial genomes.
For novel communities, set quality targets that account for the expected limitations of marker gene based assessment. Consider using multiple assessment methods and comparing their results. If CheckM and BUSCO give substantially different completeness estimates for the same MAG, investigate the cause. The discrepancy may indicate that the MAG contains novel lineages that are not well represented in one marker set.
The ResMiCo study noted that accuracy for the state of the art in reference-free misassembly prediction did not exceed an AUPRC of 0.57, and it was not clear how well these models generalize to real-world data [<a href="#ref-2">2</a>]. While ResMiCo substantially improved on this performance, no reference-free method can detect all misassemblies. For novel communities, the uncertainty in quality metrics is higher, and the quality targets should reflect this uncertainty.
Integrating Quality Targets with Publication and Deposition Requirements
The quality target document also serves as the foundation for meeting publication and data deposition requirements. Many journals and databases require quality metrics for deposited MAGs. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that support data management and sharing [<a href="#ref-9">9</a>]. Depositing MAGs with their quality metrics allows other researchers to assess the suitability of the data for their own analyses.
The Minimum Information about a Metagenome-Assembled Genome standard provides a common framework for reporting MAG quality. Following this standard ensures that your quality metrics are interpretable by other researchers. The quality target document provides the context for interpreting these metrics, showing which targets were set before assembly and how the final assembly met or fell short of those targets.
The EMBL-EBI Training program provides bioinformatics learning pathways, data-resource training, and practical analysis education [<a href="#ref-10">10</a>]. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training context [<a href="#ref-11">11</a>]. These resources support the development of the skills needed to implement a rigorous quality target framework and to communicate quality decisions effectively.
Professional Escalation Criteria for Quality Target Failures
Define specific situations that warrant escalation to a bioinformatics specialist or a reconsideration of the experimental design. These criteria should be written into the quality target document before assembly begins.
Escalate when the assembly consistently fails to meet the minimum quality tier across multiple assembly strategies and parameter sets. This pattern suggests a fundamental problem with the sequencing data or the community type instead of a fixable assembly issue. The 2023 benchmark provided practical guidance on selecting appropriate metagenome assembly tools based on sequencing technology and dataset characteristics [<a href="#ref-1">1</a>]. If none of the recommended tools meet your quality targets, the problem may be in the experimental design.
Escalate when completeness and contamination estimates are contradictory across assessment methods. If CheckM reports high completeness while BUSCO reports low completeness for the same MAG, the discrepancy may indicate a novel lineage or a binning artifact. The 2020 discussion of genome curation emphasized that gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of MAGs [<a href="#ref-3">3</a>]. Contradictory quality estimates warrant investigation before proceeding with downstream analysis.
Escalate when the misassembly rate is high and cannot be reduced through parameter optimization. The ResMiCo study estimated 7 percent misassembled contigs per metagenome across multiple real-world datasets [<a href="#ref-2">2</a>]. If your assembly shows a substantially higher misassembly rate, the sequencing strategy or assembly approach may need to be reconsidered.
Escalate when the quality targets themselves appear to be unachievable for the community type. This situation requires a conversation about whether the research questions can be answered with the achievable assembly quality or whether the experimental design needs to change. The quality target document provides the framework for this conversation, showing which targets were set and why they may not be achievable.
Frequently Asked Questions
Why is N50 not sufficient for evaluating metagenome assembly quality?
N50 measures contig length distribution but does not indicate whether the assembled contigs are biologically correct or complete. In a metagenome, a high N50 can result from the successful assembly of one dominant genome while the rest of the community remains fragmented. The 2023 benchmark found that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality MAGs [<a href="#ref-1">1</a>]. N50 must be interpreted alongside completeness, contamination, and misassembly metrics.
How does BUSCO metagenomic mode differ from standard BUSCO mode?
BUSCO metagenomic mode uses reduced marker sets and accounts for the possibility that some lineages present in the community are not represented in the reference marker set. Standard BUSCO mode assumes a single genome from a known lineage. In metagenomic mode, the assessment is applied to individual bins or MAGs instead of the entire assembly, and the results are interpreted with the understanding that some community members may lack the universal single-copy orthologs used for assessment.
What is the difference between CheckM and BUSCO for MAG quality assessment?
CheckM uses lineage-specific marker sets that are selected based on the phylogenetic placement of the genome being assessed. BUSCO uses a fixed set of universal single-copy orthologs for the domain of interest. CheckM can be more sensitive for well-characterized lineages but may be less accurate for novel lineages where taxonomic placement is uncertain. BUSCO provides a consistent assessment across all genomes but may underestimate completeness for organisms that lack some universal markers.
How can I detect misassemblies in metagenome assemblies without a reference genome?
Reference-free misassembly detection methods use features of the assembly itself to identify contigs that are likely to be chimeric or otherwise incorrectly assembled. The ResMiCo deep learning approach was developed for this purpose and was shown to be substantially more accurate than previous methods, with robustness to novel taxonomic diversity and varying assembly methods [<a href="#ref-2">2</a>]. ResMiCo estimated 7 percent misassembled contigs per metagenome across multiple real-world datasets [<a href="#ref-2">2</a>].
What completeness and contamination thresholds should I use for MAG classification?
The community standard defines high-quality MAGs as having greater than 90 percent completeness and less than 5 percent contamination, and medium-quality MAGs as having greater than 50 percent completeness and less than 10 percent contamination. These thresholds are minimum standards for reporting. The appropriate thresholds for your analysis depend on the intended use of the MAGs. Metabolic reconstruction requires higher quality than taxonomic assignment.
How does co-assembly affect assembly quality metrics?
Co-assembly combines reads from multiple samples before assembly, which can improve contiguity for organisms shared across samples by increasing coverage depth. However, co-assembly can also create chimeric contigs that join sequences from closely related but distinct strains present in different samples. The 2026 methods chapter on metagenomic assembly and gene prediction outlines per-sample versus co-assembly as a core assembly strategy decision [<a href="#ref-4">4</a>]. Quality metrics should be interpreted in the context of the assembly strategy used.
Can assembly quality be improved after the initial assembly is complete?
Yes, MAG quality can often be improved through refinement and reassembly. The MetaflowX workflow includes a dedicated reassembly module that improved MAG quality, increasing completeness by 5.6 percent and reducing contamination by 53 percent on average [<a href="#ref-5">5</a>]. Refinement typically involves removing contaminating contigs, splitting bins that contain multiple genomes, and reassembling the remaining reads to close gaps.
What should I do if my assembly quality metrics are poor across all strategies?
If assembly quality is consistently poor across multiple assembly strategies, review the experimental design. Consider whether the sequencing depth is sufficient for the community complexity, whether the sequencing technology is appropriate for the research question, and whether the community contains closely related strains that are difficult to resolve. The 2023 benchmark provided practical guidance on selecting appropriate metagenome assembly tools based on sequencing technology and dataset characteristics [<a href="#ref-1">1</a>].
Related Bioinformatics Guides
- Evaluating Genome Assembly Quality: Metrics and Tools
- Metagenome Co-Assembly: Strategies for Multi-Sample Data
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
- Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Benchmarking genome assembly methods on metagenomic sequencing data.](https://pubmed.ncbi.nlm.nih.gov/36917471). Briefings in bioinformatics, 2023. [2] [ResMiCo: Increasing the quality of metagenome-assembled genomes with deep learning.](https://pubmed.ncbi.nlm.nih.gov/37126495). PLoS computational biology, 2023. [3] [Accurate and complete genomes from metagenomes.](https://pubmed.ncbi.nlm.nih.gov/32188701). Genome research, 2020. [4] [Metagenomic Assembly and Gene Prediction.](https://pubmed.ncbi.nlm.nih.gov/42108291). Methods in molecular biology (Clifton, N.J.), 2026. [5] [MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.](https://pubmed.ncbi.nlm.nih.gov/41036626). Nucleic acids research, 2025. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [9] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [11] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.