# How to Refine Metagenome-Assembled Genomes (MAGs): A Guide to Improving Purity and Completeness


## Key Takeaways

- MAG refinement is critical for improving genome purity and completeness, addressing inherent errors from assembly and binning that can lead to chimeric genomes or incomplete representation of the target organism.
- Quality assessment relies on metrics like completeness and contamination, typically estimated using single-copy marker genes via tools such as CheckM and BUSCO, with MIMAG standards defining quality tiers (e.g., >90% complete, <5% contamination for high-quality drafts).
- Contamination removal tools like RefineM and MAGpurify leverage genomic properties (GC content, coverage, tetranucleotide frequency) and taxonomic assignments to identify and excise non-target contigs, requiring careful consideration to avoid removing true genome sequence.
- Completeness improvement strategies include re-binning with adjusted parameters, iterative clustering refinement based on contig abundance and marker genes, and resource-intensive reassembly of reads mapped back to refined bins.
- Independent validation, including Average Nucleotide Identity (ANI) comparisons to reference genomes and analysis of inter- and intra-contig relationships, is essential to detect chimerism that marker gene analysis may miss, particularly for high-resolution applications like outbreak investigations.
- A robust refinement decision framework should integrate contig-level evidence (sequence composition, coverage, taxonomic assignment) using a "three-evidence rule" before contig removal, supplemented by a decision matrix that considers contig length and marker gene presence to mitigate common failure patterns like over-removal or under-removal of sequences.

---

Metagenome-assembled genomes (MAGs) are microbial genomes reconstructed directly from shotgun metagenomic sequencing data. Initial binning often produces genomes with contamination from co-assembled sequences of related taxa and incomplete coverage of the true genome. Refinement is the process of improving MAG purity and completeness after initial binning, using tools such as RefineM and MAGpurify, re-binning with adjusted parameters, and validating results with marker gene analysis. This guide covers the practical workflow for refining MAGs, the decisions required at each step, and the quality standards needed for downstream analysis.

## Context and Scope of MAG Refinement

Shotgun metagenomics enables recovery of microbial genomes directly from environmental or host-associated communities without cultivation. The scale of publicly available metagenomic data is substantial, with over 600,000 metagenomes deposited in repositories such as NCBI's Sequence Read Archive (SRA) [<a href="#ref-1">1</a>]. This data availability has driven development of numerous bioinformatics pipelines for MAG reconstruction, each with different computational requirements, interfaces, workflow managers, and modular structures [<a href="#ref-2">2</a>].

The central challenge in MAG construction is that assembly and binning are imperfect processes. Assembly algorithms reconstruct contigs from short reads, and binning algorithms group contigs into genome bins based on sequence composition and coverage patterns. Both steps introduce errors. Contigs from different taxa can be joined during assembly, particularly when organisms share similar nucleotide composition. Binning can assign contigs from related species to the same bin, creating chimeric genomes. Conversely, contigs from the target genome can be assigned to different bins, reducing completeness.

Refinement addresses these errors through several strategies. One approach is to subject low-quality MAGs to iterative clustering based on contig abundance, increasing bin purity through validated universal marker genes [<a href="#ref-3">3</a>]. Another approach is to use multiple complementary binning tools with a unified refinement strategy, which has been shown to yield more high-quality MAGs than single-tool approaches [<a href="#ref-1">1</a>]. A third approach involves reassembly, where reads are mapped back to refined bins and reassembled to improve completeness and reduce contamination [<a href="#ref-4">4</a>].

The practical outcome of refinement is a set of MAGs that meet quality thresholds for downstream analysis. These thresholds depend on the intended use. Taxonomic classification, functional annotation, comparative genomics, and public health investigations each have different requirements. For high-resolution objectives such as outbreak pathogenomics, quality checking for congeneric chimerism is especially important [<a href="#ref-5">5</a>].

## Core Principles of MAG Quality Assessment

### Completeness and Contamination Metrics

Completeness measures the fraction of the true genome present in the MAG. Contamination measures the fraction of sequence in the MAG that originates from other organisms. Both are estimated using marker genes that are expected to be single-copy in most microbial genomes. The presence of multiple copies of a marker gene suggests contamination, while the absence of expected markers suggests incompleteness.

These estimates are statistical inferences, not direct measurements. They depend on the marker gene set used, the taxonomic breadth of the reference database, and the quality of the assembly. Different tools may report different completeness and contamination values for the same MAG because they use different marker sets and inference methods.

### CheckM and BUSCO

CheckM is the most widely used tool for estimating MAG completeness and contamination. It uses lineage-specific marker gene sets identified from the placement of the MAG in a reference tree. BUSCO (Benchmarking Universal Single-Copy Orthologs) uses a different approach, assessing completeness against a set of universal single-copy orthologs for a given taxonomic group.

Both tools provide useful but imperfect estimates. CheckM can underestimate contamination when the contaminating sequences come from closely related taxa that share marker genes. BUSCO can overestimate completeness when the MAG contains sequences from multiple organisms that collectively cover the BUSCO set.

### Minimum Information about a Metagenome-Assembled Genome (MIMAG)

The MIMAG standards define quality tiers for MAGs. A high-quality draft MAG is estimated to be more than 90% complete with less than 5% contamination. A medium-quality draft MAG is estimated to be more than 50% complete with less than 10% contamination. These thresholds provide a common language for reporting MAG quality across studies.

Meeting these thresholds is necessary but not sufficient for all downstream applications. A MAG that meets MIMAG high-quality criteria can still contain chimeric sequences from closely related taxa that do not affect marker gene estimates. Detection of such chimerism requires additional analysis, such as comparing average nucleotide identity (ANI) between contigs within the MAG or examining coverage patterns across the genome [<a href="#ref-5">5</a>].

## At a Glance: Refinement Decision Table

| Situation | Recommended Action | Expected Outcome | Key Consideration |
|-----------|-------------------|------------------|-------------------|
| MAG has >10% contamination | Run contamination removal with RefineM or MAGpurify, then re-estimate quality | Reduced contamination, possible reduction in completeness | Removing contigs can remove true genome sequence if contamination is misidentified |
| MAG has <50% completeness | Re-bin with different parameters or use iterative clustering refinement | Improved completeness through recovery of missed contigs | Re-binning may introduce new contamination from related taxa |
| MAG has 5-10% contamination and 50-90% completeness | Apply iterative refinement combining abundance clustering and marker gene validation | Improved purity with moderate completeness gains | Multiple refinement rounds may be needed for complex communities |
| MAG has high completeness but suspected chimerism | Perform inter- and intra-contig comparisons to identify contaminating contigs | Removal of chimeric sequences, improved ANI to reference genomes | Conservative removal is valuable for outbreak investigations [<a href="#ref-5">5</a>] |
| Multiple MAGs from same sample have similar quality | Use multi-tool binning with unified refinement strategy | Higher number of high-quality MAGs overall | Different tools recover different genomes, integration increases yield [<a href="#ref-1">1</a>] |

## Practical Workflow for MAG Refinement

### Step 1: Assess Initial Bin Quality

Before refinement, establish a baseline. Run CheckM or BUSCO on each bin from the initial binning step. Record completeness, contamination, and strain heterogeneity for each MAG. This baseline determines which refinement strategy to apply.

For bins with high contamination, focus on contamination removal. For bins with low completeness, focus on re-binning or reassembly. For bins with moderate quality, apply iterative refinement.

### Step 2: Remove Contamination

Contamination removal tools identify contigs that do not belong to the target genome. RefineM uses genomic properties such as GC content, coverage, and tetranucleotide frequency to identify outlier contigs. MAGpurify uses a similar approach with additional features such as taxonomic assignment of contigs.

The decision to remove a contig should consider the evidence. A contig with markedly different GC content and coverage from the rest of the bin is likely contamination. A contig with slightly different properties may be a true part of the genome, particularly in organisms with heterogeneous genomic regions such as prophages or genomic islands.

Conservative removal is especially valuable when the MAG will be used for high-resolution analysis. In the investigation of a Burkholderia pseudomallei outbreak, inter- and intra-contig comparisons revealed few potential contaminants of related taxa. Conservative removal of those contigs ultimately obtained an ANI of 99.9% between the refined MAG and the genome of an isolate from the contaminated product [<a href="#ref-5">5</a>].

### Step 3: Improve Completeness

Completeness improvement requires recovering contigs that were missed during initial binning. Several strategies exist.

Re-binning with different parameters can recover contigs that were excluded due to strict binning thresholds. Adjusting the minimum contig length, coverage cutoff, or composition distance can include additional contigs. However, relaxed parameters also increase the risk of incorporating contamination.

Iterative clustering refinement subjects low-quality MAGs to repeated k-means clustering based on contig abundance, increasing bin purity through validated universal marker genes [<a href="#ref-3">3</a>]. This approach can improve both completeness and purity simultaneously.

Reassembly is a more resource-intensive approach. Reads are mapped back to the refined bin, and the mapped reads are reassembled. This can improve completeness by reconstructing regions that were fragmented in the original assembly. A dedicated reassembly module in the MetaflowX workflow improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average [<a href="#ref-4">4</a>].

### Step 4: Re-assess Quality

After each refinement step, re-run quality assessment. Compare completeness and contamination estimates to the baseline. Record which refinement steps were applied and their effect on quality metrics.

Quality assessment after refinement is essential because refinement can introduce new errors. Contamination removal can remove true genome sequence. Re-binning can incorporate contaminating contigs. Reassembly can create new chimeric sequences.

### Step 5: Validate with Independent Methods

Marker gene estimates provide one view of MAG quality. Independent validation methods provide additional confidence.

Taxonomic assignment of the refined MAG should be consistent with the expected organism. ANI comparison to reference genomes can confirm the identity of the MAG and detect chimerism. Coverage analysis can identify regions with anomalous coverage that may indicate contamination or assembly errors.

For public health applications, validation against isolate genomes is the gold standard. In the Burkholderia pseudomallei investigation, the refined MAG achieved an ANI of 99.9% to the isolate genome, and the identical conclusion about the outbreak source was possible with the MAG as with the isolate genome [<a href="#ref-5">5</a>].

## Tools and Workflow Options

### RefineM

RefineM is a standalone tool for MAG refinement. It provides functions for contamination removal, completeness improvement, and genome QC. Contamination removal uses genomic properties including GC content, coverage, and tetranucleotide frequency. Completeness improvement uses a database of marker genes to identify missing genomic regions.

RefineM is appropriate when you have a small number of bins to refine and want fine-grained control over the refinement process. It requires a reference genome database for some functions, which must be downloaded separately.

### MAGpurify

MAGpurify is another standalone refinement tool. It uses a different set of features for contamination detection, including taxonomic assignment of contigs and comparison to reference genomes. MAGpurify can identify contaminating contigs that have similar GC content and coverage to the target genome but different taxonomic origins.

MAGpurify is appropriate when contamination is suspected to come from distantly related organisms. It is less effective for detecting contamination from closely related species that share taxonomic markers.

### Iterative Clustering Refinement

The Additional Clustering Refiner (ACR) uses iterative k-means clustering predicated on contig abundance to refine low-quality MAGs. It increases bin purity through validated universal marker genes. ACR demonstrated improved MAG purity and a significant increase in high- and medium-quality MAG recovery rates in synthetic and real-world metagenomic datasets, including short- and long-read sequences [<a href="#ref-3">3</a>].

ACR integrates with various binning algorithms without modifying their core features. Its multiple sequencing technology compatibilities expand its applicability [<a href="#ref-3">3</a>].

### Multi-Tool Binning with Unified Refinement

Using multiple complementary binning tools with a unified refinement strategy yields more high-quality MAGs than any single tool. The TOFU-MAaPO workflow integrates multiple complementary binning tools with a unified refinement strategy, yielding 12% to 77% more high-quality MAGs than three established metagenome software pipelines [<a href="#ref-1">1</a>].

This approach is appropriate when you have sufficient computational resources and want to maximize MAG recovery from complex communities. The tradeoff is increased computational cost and complexity.

### Integrated Workflows

Several integrated workflows include refinement as a module. MetaflowX is a modular framework encompassing short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality MAG identification, and bin refinement and reassembly. Benchmarking showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows, while recovering the highest number of high-quality and taxonomically diverse MAGs [<a href="#ref-4">4</a>].

nf-core/mag is a community pipeline that includes MAG construction and refinement. It was used successfully to construct a Burkholderia pseudomallei MAG from contaminated aromatherapy spray, demonstrating the utility of public pipelines for public health response [<a href="#ref-5">5</a>]. The nf-core documentation provides standards for pipeline usage, configuration, and reproducible workflow context [<a href="#ref-6">6</a>].

### Choosing Between Options

The choice of refinement approach depends on your data, computational resources, and downstream analysis goals. The 2Pipe decision-support web application helps users identify a suitable workflow based on input data characteristics, desired outcomes, and computational constraints. It presents a question-driven interface, a pipeline gallery, and a pipeline comparison based on key factors [<a href="#ref-2">2</a>].

For small numbers of bins, standalone tools like RefineM or MAGpurify provide control and transparency. For large datasets, integrated workflows like MetaflowX or TOFU-MAaPO provide efficiency and scalability. For public health applications, validated pipelines with documented performance are preferable.

## Observations and Measurements

### Recording Refinement Decisions

Maintain a record of refinement decisions for each MAG. This record should include the initial quality metrics, the refinement steps applied, the parameters used, and the final quality metrics. This information is essential for reproducibility and for interpreting downstream analysis results.

A refinement log should include the following fields for each MAG:

| Field | Description | Example |
|-------|-------------|---------|
| MAG ID | Unique identifier for the bin | bin_042 |
| Initial completeness | CheckM completeness estimate before refinement | 78.5% |
| Initial contamination | CheckM contamination estimate before refinement | 12.3% |
| Refinement tool | Tool used for refinement | RefineM |
| Parameters | Key parameters used | gc_window=1000, coverage_window=5000 |
| Contigs removed | Number of contigs removed during refinement | 14 |
| Final completeness | CheckM completeness estimate after refinement | 82.1% |
| Final contamination | CheckM contamination estimate after refinement | 3.2% |
| Validation method | Independent validation performed | ANI to reference genome |

### Measuring Refinement Effectiveness

Refinement effectiveness can be measured by the change in quality metrics and by the success of downstream analysis. A refinement that reduces contamination from 12% to 3% while maintaining completeness above 90% is effective. A refinement that reduces contamination but drops completeness below 50% may not be effective for downstream analysis.

The reassembly module in MetaflowX provides a benchmark for refinement effectiveness, increasing completeness by 5.6% and reducing contamination by 53% on average [<a href="#ref-4">4</a>]. These figures provide a reference point for evaluating your own refinement results.

### Detecting Chimerism

Chimerism is the presence of sequences from multiple organisms in a single MAG. It can occur when closely related species are co-assembled or co-binned. Marker gene estimates may not detect chimerism when the contaminating organism shares marker genes with the target organism.

Detection of chimerism requires comparing sequences within the MAG. Inter-contig comparisons examine the relationship between contigs. Intra-contig comparisons examine the relationship within contigs. In the Burkholderia pseudomallei investigation, inter- and intra-contig comparisons revealed few potential contaminants of related taxa, including Burkholderia cepacia, Burkholderia cenocepacia, Burkholderia multivorans, Burkholderia pseudomultivorans, Cupriavidus pauculus, and Pseudomonas aeruginosa. These related species had ANI to B. pseudomallei ranging from 72.3% to 84.8% [<a href="#ref-5">5</a>].

## Common Failure Patterns in MAG Refinement

### Over-Removal of True Genome Sequence

Aggressive contamination removal can remove contigs that are true parts of the target genome. This is particularly problematic for organisms with heterogeneous genomic regions. Prophages, genomic islands, and regions with atypical GC content can be misidentified as contamination.

Prevention requires examining the evidence for each removed contig. A contig with coverage consistent with the rest of the genome and containing genes expected in the target organism should be retained even if its GC content is atypical.

### Under-Removal of Contamination

Conservative contamination removal can leave contaminating contigs in the MAG. This is particularly problematic when the contaminating organism is closely related to the target organism. Marker gene estimates may not detect this contamination because the contaminating organism contributes expected marker genes.

Detection requires comparing the MAG to reference genomes. ANI analysis can identify contigs that are more similar to a different organism than to the target organism.

### Chimeric Reassembly

Reassembly can create chimeric sequences when reads from different organisms are assembled together. This is more likely when the organisms are closely related and share sequence similarity.

Prevention requires careful read mapping. Reads should be mapped to the refined bin with stringent parameters, and the mapped reads should be examined for evidence of chimerism before reassembly.

### Parameter Overfitting

Refinement parameters optimized for one dataset may not generalize to other datasets. A parameter set that works well for a simple community with few closely related species may fail for a complex community with many related species.

Prevention requires validating parameters on multiple datasets and documenting the rationale for parameter choices.

### Ignoring Strain Heterogeneity

Strain heterogeneity occurs when a MAG contains sequences from multiple strains of the same species. This is common in environmental and host-associated communities. Strain heterogeneity can inflate contamination estimates and complicate downstream analysis.

Detection requires examining the coverage and sequence variation within the MAG. Regions with variable coverage or high sequence diversity may indicate strain heterogeneity.

## Limitations of MAG Refinement

### Inability to Recover Everything

Refinement cannot recover genomes that are not present in the assembly. If a genome is too rare in the community, has too low coverage, or is too similar to other genomes, it may not be assembled or binned. Refinement can only improve the quality of genomes that are already partially recovered.

### Dependence on Assembly Quality

Refinement operates on the assembly. A poor assembly limits the potential for refinement. Fragmented assemblies produce fragmented MAGs that are difficult to refine. Misassembled contigs produce chimeric MAGs that are difficult to correct.

### Marker Gene Limitations

Marker gene estimates of completeness and contamination are imperfect. They can underestimate contamination from closely related taxa and overestimate completeness for genomes with unusual gene content. These limitations should be considered when interpreting quality metrics.

### Computational Requirements

Refinement can be computationally intensive, particularly for large datasets. Reassembly is the most resource-intensive refinement step. Multi-tool binning with unified refinement increases computational requirements further.

The TOFU-MAaPO workflow demonstrates that large-scale analysis is feasible with efficient implementation. It automatically downloaded 16,462 human gut metagenome samples from the SRA and taxonomically annotated them against the Genome Taxonomy Database on a high-performance cluster in less than 55 hours, including download time [<a href="#ref-1">1</a>].

### Reference Database Dependence

Some refinement tools require reference genome databases. These databases may be incomplete or biased toward well-studied organisms. Refinement of MAGs from poorly characterized lineages may be limited by the absence of appropriate references.

## Quality Controls and Validation

### Internal Quality Controls

Internal quality controls are built into the refinement process. These include marker gene analysis, coverage analysis, and composition analysis. Each provides a different view of MAG quality.

Marker gene analysis estimates completeness and contamination. Coverage analysis identifies regions with anomalous coverage. Composition analysis identifies contigs with atypical GC content or tetranucleotide frequency.

### External Validation

External validation compares the refined MAG to independent data. This can include comparison to isolate genomes, reference genomes, or other MAGs from the same sample.

ANI comparison to reference genomes is a powerful validation method. A refined MAG with high ANI to a reference genome is likely accurate. A refined MAG with low ANI to all reference genomes may represent a novel species or may contain errors.

### Reproducibility Controls

Reproducibility controls ensure that refinement results can be reproduced. This requires documenting the software versions, parameters, and reference databases used. Workflow managers such as Nextflow, used by nf-core pipelines, provide built-in reproducibility features [<a href="#ref-6">6</a>].

The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible bioinformatics analysis [<a href="#ref-7">7</a>]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-8">8</a>].

## Safety and Regulatory Context

### Public Health Applications

MAG refinement has direct public health applications. In the Burkholderia pseudomallei outbreak investigation, a MAG constructed from contaminated aromatherapy spray was used to determine the origin of the outbreak. The refined MAG achieved an ANI of 99.9% to the isolate genome, and the identical conclusion about the outbreak source was possible with the MAG as with the isolate genome [<a href="#ref-5">5</a>].

For public health applications, MAG quality requirements are higher than for basic research. Chimerism that would be acceptable for a comparative genomics study could lead to incorrect conclusions in an outbreak investigation. Conservative refinement and thorough validation are essential.

### Data Management

MAG refinement generates large amounts of data, including intermediate files, quality reports, and refined genomes. Proper data management ensures that this data is preserved and accessible. NCBI provides data resources for sequence data, including the SRA and genome databases [<a href="#ref-9">9</a>].

### Professional Escalation Criteria

Seek expert assistance when refinement results are unexpected or when downstream analysis depends critically on MAG quality. Specific situations that warrant escalation include:

- A refined MAG that does not match any known lineage despite high completeness
- A refined MAG with conflicting quality estimates from different tools
- A refined MAG that produces unexpected results in downstream analysis
- A refinement process that consistently fails to improve quality metrics

## Records and Documentation

### Refinement Log

Maintain a refinement log for each MAG. This log should document the initial quality, refinement steps, parameters, and final quality. The log should also document any validation performed and the results.

### Pipeline Documentation

Document the refinement pipeline, including software versions, parameters, and reference databases. This documentation enables reproduction of the refinement process and interpretation of the results.

### Data Archiving

Archive the refined MAGs and the data needed to reproduce the refinement. This includes the original reads, the assembly, the initial bins, and the refinement logs. NCBI provides data resources for archiving sequence data [<a href="#ref-9">9</a>].

## Building a Refinement Decision Framework Based on Contig-Level Evidence

Refinement decisions are often made on aggregate quality metrics alone, but these summary statistics can mask the underlying structure of contamination and incompleteness. A more reliable approach is to evaluate each contig within a MAG against multiple lines of evidence before deciding whether to retain, remove, or reassign it. This section presents a practical decision framework that integrates contig-level observations with bin-level quality targets, giving you a defensible basis for every refinement action.

### The Three-Evidence Rule for Contig Retention

Before removing any contig from a MAG, require at least three independent lines of evidence that the contig does not belong to the target genome. This rule prevents over-removal of true genome sequence, which is one of the most common failure patterns in refinement. The three evidence categories are sequence composition, coverage profile, and taxonomic assignment.

Sequence composition evidence includes GC content, tetranucleotide frequency, and k-mer distribution. Coverage profile evidence includes read depth relative to the rest of the bin and coverage uniformity across the contig. Taxonomic assignment evidence includes the taxonomic classification of genes on the contig and the average nucleotide identity to reference genomes.

A contig that differs from the bin median in GC content by more than 10 percentage points, has coverage less than half or more than double the bin median, and classifies to a different genus than the majority of the bin meets the three-evidence threshold for removal. A contig that differs in only one or two of these properties should be retained and flagged for further investigation.

The three-evidence rule is especially important for organisms with heterogeneous genomic regions. Prophages, genomic islands, and regions with atypical GC content can appear as outliers in composition analysis while being true parts of the target genome. Requiring multiple independent lines of evidence protects these regions from inappropriate removal.

### Contig-Level Evidence Categories and Their Interpretation

The table below summarizes the evidence categories, the measurements used, and the interpretation of outlier status for each category.

| Evidence Category | Measurement | Outlier Indicator | Interpretation |
|-------------------|-------------|-------------------|----------------|
| Sequence composition | GC content | More than 10 percentage points from bin median | Possible contamination or genuine genomic island |
| Sequence composition | Tetranucleotide frequency | Z-score greater than 3 from bin mean | Possible contamination from distantly related taxon |
| Coverage profile | Read depth | Less than half or more than double bin median | Possible contamination or differential replication |
| Coverage profile | Coverage uniformity | High variance across contig length | Possible misassembly or chimeric junction |
| Taxonomic assignment | Gene classification | Majority of genes classify to different genus | Probable contamination from related taxon |
| Taxonomic assignment | ANI to references | Higher ANI to different species than to bin consensus | Probable chimeric sequence |
| Marker gene content | Single-copy marker presence | Multiple copies of same marker on one contig | Probable contamination or strain heterogeneity |
| Linkage evidence | Paired-end read connections | Reads connect contig to sequences outside bin | Possible misbinning or assembly error |

Each evidence category has limitations. GC content is unreliable for organisms with wide genomic GC variation. Coverage is affected by replication timing and copy number. Taxonomic assignment depends on reference database completeness. The three-evidence rule mitigates these individual limitations by requiring convergence across independent measurements.

### The Refinement Decision Matrix

Once you have collected contig-level evidence, use the decision matrix below to determine the appropriate action for each contig. The matrix combines the number of evidence categories supporting removal with the contig length and the presence of marker genes.

| Evidence Categories Supporting Removal | Contig Length | Marker Genes Present | Recommended Action |
|----------------------------------------|---------------|---------------------|--------------------|
| 3 or more | Any | No | Remove contig from bin |
| 3 or more | Any | Yes | Remove contig, record marker gene loss |
| 2 | Greater than 10 kb | No | Flag for manual review |
| 2 | Greater than 10 kb | Yes | Retain contig, document uncertainty |
| 2 | Less than 10 kb | No | Remove contig if coverage is anomalous |
| 1 or fewer | Any | Any | Retain contig in bin |
| 3 or more | Any | Yes, multiple copies | Remove contig, investigate strain heterogeneity |

The presence of marker genes on a flagged contig warrants additional scrutiny. A contig with multiple copies of a single-copy marker gene may represent contamination from a closely related strain instead of a separate species. In this case, the contig may need to be placed in a separate strain bin instead of removed entirely.

For contigs with two evidence categories supporting removal, manual review is appropriate. This review should examine the specific genes on the contig, the read coverage pattern, and the genomic context. The review outcome should be documented in the refinement log.

### Implementing the Framework with Available Tools

The three-evidence rule and decision matrix can be implemented using features available in RefineM, MAGpurify, and related tools. RefineM provides functions for computing GC content, coverage, and tetranucleotide frequency for each contig. MAGpurify provides taxonomic assignment of contigs based on marker gene content.

To implement the framework, start by running RefineM to generate per-contig statistics for GC content, coverage, and tetranucleotide frequency. Export these statistics to a table. Run MAGpurify to obtain taxonomic classifications for each contig. Merge the two tables by contig identifier.

For each contig, calculate the deviation from the bin median for GC content and coverage. Calculate the tetranucleotide frequency z-score. Record the taxonomic classification. Count the number of evidence categories that support removal. Apply the decision matrix to determine the action.

This implementation requires some scripting but is straightforward with standard bioinformatics tools. The Carpentries lessons provide foundational training in shell and programming that supports this type of analysis [<a href="#ref-7">7</a>]. Bioconductor provides documentation for reproducible genomic-analysis workflows that can be adapted for this purpose [<a href="#ref-8">8</a>].

### Recording Contig-Level Decisions

The refinement log should include contig-level decisions in addition to bin-level quality metrics. For each contig that is removed, record the contig identifier, length, GC content, coverage, taxonomic classification, the evidence categories supporting removal, and the rationale for the decision.

For each contig that is retained despite flagging, record the same information plus the reason for retention. This documentation is essential for reproducibility and for interpreting downstream analysis results. A reviewer should be able to understand why each contig was retained or removed.

The table below shows the recommended fields for contig-level decision records.

| Field | Description | Example |
|-------|-------------|---------|
| Contig ID | Unique identifier from assembly | contig_001234 |
| Length | Contig length in base pairs | 45,678 |
| GC content | Percent GC of the contig | 52.3 |
| Coverage | Mean read depth across the contig | 18.5 |
| Bin median GC | Median GC of all contigs in the bin | 48.1 |
| Bin median coverage | Median coverage of all contigs in the bin | 22.0 |
| Taxonomic classification | Genus-level assignment from MAGpurify | Burkholderia |
| Evidence count | Number of categories supporting removal | 3 |
| Decision | Remove, retain, or flag for review | Remove |
| Rationale | Specific reason for the decision | GC and coverage outliers, classified to different genus |

### Applying the Framework to Strain Heterogeneity

Strain heterogeneity presents a special case for the decision framework. When a MAG contains sequences from multiple strains of the same species, the contigs from different strains may have similar GC content, coverage, and taxonomic classification. The three-evidence rule may not identify these contigs as contaminants because they match the bin in most categories.

In this situation, additional evidence is needed. Examine the coverage distribution across the bin. Multiple strains often have different coverage levels due to different abundances in the community. Examine single-nucleotide variant density across contigs. Contigs from different strains may show higher variant density than contigs from the same strain.

If strain heterogeneity is suspected, consider splitting the bin into multiple strain bins instead of removing contigs. This approach preserves the genomic information while improving the purity of each resulting bin. The decision matrix should be extended to include a split action for contigs that show evidence of strain-level variation.

### Common Failure Patterns in Contig-Level Decision Making

Several failure patterns recur when researchers apply contig-level decisions without a structured framework.

The first pattern is composition-only removal. Removing contigs based solely on GC content or tetranucleotide frequency without considering coverage or taxonomy leads to over-removal of genuine genomic islands and prophages. These regions often have atypical composition but are true parts of the genome.

The second pattern is coverage-only removal. Removing contigs based solely on coverage anomalies fails when the target organism has variable copy number regions or when the community contains closely related strains at different abundances. Coverage outliers should trigger investigation, not automatic removal.

The third pattern is reference-only removal. Removing contigs based solely on taxonomic classification fails when the reference database lacks representation of the target lineage. A contig that classifies to a different genus may simply reflect the absence of the true lineage from the database.

The fourth pattern is threshold rigidity. Applying fixed thresholds for GC content deviation or coverage ratios without considering the specific organism and community context leads to inappropriate decisions. The thresholds suggested in this framework are starting points, not universal rules.

The fifth pattern is failure to document. Making contig-level decisions without recording the evidence and rationale makes the refinement process unreproducible and complicates troubleshooting when downstream analysis produces unexpected results.

### Escalation Criteria for Contig-Level Decisions

Some contig-level decisions require expert review. Escalate to a colleague or supervisor when any of the following conditions apply.

A contig longer than 100 kb is flagged for removal. Long contigs represent substantial genomic content, and their removal can significantly reduce completeness. The evidence for removal should be exceptionally strong.

A contig contains multiple marker genes that are absent from the rest of the bin. This pattern may indicate that the contig is essential for the target genome and that its removal would create a misleading completeness estimate.

A contig shows conflicting evidence across categories. For example, a contig with anomalous GC content but coverage and taxonomy consistent with the bin requires careful interpretation. The conflict may indicate a genuine genomic island or a misassembly.

A contig is flagged in multiple MAGs from the same sample. This pattern may indicate a shared contaminant or a sequence that is genuinely shared between organisms. The decision should consider the broader community context.

A contig is flagged during refinement of a MAG that will be used for public health or clinical applications. The higher quality requirements for these applications warrant additional scrutiny and validation [<a href="#ref-5">5</a>].

### Validating Contig-Level Decisions with Read Mapping

After applying the decision framework, validate the decisions by mapping reads back to the refined bin. Read mapping provides independent confirmation that the retained contigs are supported by sequencing data and that the removed contigs were not supported by reads from the target organism.

Map the original reads to the refined bin using a stringent read mapper. Examine the coverage profile across each retained contig. Contigs with uniform coverage and no evidence of chimeric junctions are likely accurate. Contigs with variable coverage or abrupt coverage transitions may contain misassemblies.

For removed contigs, map the reads to the removed sequences. If a substantial number of reads map to the removed contigs, investigate whether those reads also map to retained contigs. Reads that map uniquely to removed contigs may indicate that the contigs were incorrectly removed.

This validation step is especially important when the refined MAG will be used for high-resolution analysis. In the Burkholderia pseudomallei outbreak investigation, inter- and intra-contig comparisons revealed few potential contaminants of related taxa, and conservative removal of those contigs was especially valuable [<a href="#ref-5">5</a>]. The read mapping validation provides similar confidence in your own refinement decisions.

### Integrating the Framework with Reassembly

The contig-level decision framework applies to the refinement of existing bins. When reassembly is used as a refinement strategy, the framework should be applied before and after reassembly.

Before reassembly, use the framework to identify and remove contaminating contigs. This prevents contaminating reads from being included in the reassembly. After reassembly, apply the framework to the new contigs to ensure that the reassembly did not introduce new contamination.

The MetaflowX workflow includes a dedicated reassembly module that improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average [<a href="#ref-4">4</a>]. Applying the contig-level decision framework before and after reassembly can further improve these outcomes by ensuring that only high-confidence sequences are used in the reassembly process.

### Comparing Framework Outcomes Across Tools

The contig-level decision framework can be used to compare the performance of different refinement tools on the same dataset. Apply the framework to the output of RefineM, MAGpurify, and an integrated workflow. Compare the contigs retained and removed by each tool.

This comparison provides insight into the strengths and limitations of each tool for your specific dataset. A tool that removes contigs that the framework identifies as true genome sequence is too aggressive for your data. A tool that retains contigs that the framework identifies as contamination is too conservative.

The 2Pipe decision-support application can help you select the appropriate workflow for your data characteristics and computational constraints [<a href="#ref-2">2</a>]. The contig-level framework adds an evidence-based layer to this selection process by allowing you to evaluate tool performance on your specific contigs.

### Documenting Framework Application in Publications

When publishing results that include refined MAGs, document the contig-level decision framework in the methods section. Describe the evidence categories used, the thresholds applied, and the decision rules. This documentation allows readers to evaluate the rigor of your refinement process.

Include the number of contigs removed, the total sequence removed, and the effect on completeness and contamination estimates. Report the number of contigs flagged for manual review and the outcome of that review. This level of detail distinguishes a defensible refinement process from an opaque one.

The Minimum Information about a Metagenome-Assembled Genome standards provide guidance on reporting MAG quality. The contig-level decision framework supplements these standards by documenting the refinement decisions that produced the reported quality metrics.

## Frequently Asked Questions

### What is the difference between MAG refinement and MAG reassembly?

Refinement improves the quality of existing bins by removing contaminating contigs and recovering missed contigs. Reassembly is a specific refinement strategy where reads are mapped back to the refined bin and reassembled to improve completeness. Reassembly is more resource-intensive than other refinement approaches but can recover genomic regions that were fragmented in the original assembly. The MetaflowX workflow includes a dedicated reassembly module that improved completeness by 5.6% and reduced contamination by 53% on average [<a href="#ref-4">4</a>].

### How do I know if my MAG needs refinement?

Run quality assessment with CheckM or BUSCO on each bin. Bins with contamination above 5% or completeness below 90% are candidates for refinement. Bins with contamination above 10% or completeness below 50% should be refined before downstream analysis. Bins that meet quality thresholds but will be used for high-resolution analysis, such as outbreak investigations, should also be examined for chimerism [<a href="#ref-5">5</a>].

### Which refinement tool should I use?

The choice depends on your data and goals. RefineM provides fine-grained control over contamination removal and completeness improvement. MAGpurify uses taxonomic assignment to identify contaminating contigs. ACR uses iterative k-means clustering based on contig abundance [<a href="#ref-3">3</a>]. Integrated workflows like MetaflowX and TOFU-MAaPO include refinement as a module [<a href="#ref-4">4</a>][<a href="#ref-1">1</a>]. The 2Pipe decision-support application can help you choose a workflow based on your input data characteristics, desired outcomes, and computational constraints [<a href="#ref-2">2</a>].

### Can refinement introduce new errors?

Yes. Contamination removal can remove true genome sequence. Re-binning can incorporate contaminating contigs. Reassembly can create chimeric sequences. Always re-assess quality after refinement and compare to the baseline. Validate refined MAGs with independent methods, such as ANI comparison to reference genomes.

### How many refinement rounds should I perform?

There is no fixed number. Perform refinement until quality metrics stabilize or until additional refinement does not improve quality. Multiple rounds may be needed for complex communities with many closely related species. Document each round and its effect on quality metrics.

### What completeness and contamination thresholds should I use for downstream analysis?

The thresholds depend on the downstream analysis. For taxonomic classification, medium-quality MAGs may be sufficient. For functional annotation and comparative genomics, high-quality MAGs are preferred. For public health applications, the highest quality MAGs with independent validation are required [<a href="#ref-5">5</a>]. The MIMAG standards define high-quality as more than 90% complete with less than 5% contamination and medium-quality as more than 50% complete with less than 10% contamination.

### How do I detect chimerism in my MAG?

Chimerism detection requires comparing sequences within the MAG. Inter-contig comparisons examine the relationship between contigs. Intra-contig comparisons examine the relationship within contigs. ANI comparison to reference genomes can identify contigs that are more similar to a different organism than to the target organism. In the Burkholderia pseudomallei investigation, inter- and intra-contig comparisons revealed few potential contaminants of related taxa [<a href="#ref-5">5</a>].

### What should I do if refinement does not improve my MAG quality?

If refinement does not improve quality, consider whether the underlying assembly is adequate. A poor assembly limits the potential for refinement. Consider re-assembling with different parameters or using a different assembler. Consider whether the community is too complex for current binning and refinement methods. Consider whether the target organism is too similar to other organisms in the community for reliable separation.

## Related Bioinformatics Guides

- [Metagenome Assembled Genome Analysis: From Bins to Biological Insights](/knowledge/bioinformatics/metagenome-assembled-genome-analysis-from-bins-to-biological-insights)
- [Binning in Metagenomics: From Contigs to Genomes](/knowledge/bioinformatics/binning-in-metagenomics-from-contigs-to-genomes)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive.](https://doi.org/10.1038/s41467-026-74033-9). 2026.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [2Pipe starts with a question: matching you with the correct pipeline for MAG reconstruction.](https://doi.org/10.1128/msystems.00844-25). 2026.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [ACR: metagenome-assembled prokaryotic and eukaryotic genome refinement tool.](https://pubmed.ncbi.nlm.nih.gov/37889119). Briefings in bioinformatics, 2023.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.](https://pubmed.ncbi.nlm.nih.gov/41036626). Nucleic acids research, 2025.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Shotgun metagenome sequencing and informatics can accurately form a metagenome-assembled genome (MAG) of the bacterial tier 1 select agent &lt,i&gt,Burkholderia pseudomallei&lt,/i&gt, for rapid public health response events.](https://doi.org/10.1128/spectrum.02926-25). 2026.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.