# Assembly-Based Strain Resolution: How to Recover Individual Strain Genomes from Metagenomic Assemblies


## Key Takeaways

- Strain resolution from metagenomic assemblies is critical for understanding transmission, functional variation, and population dynamics, as standard assembly collapses highly similar genomes into consensus sequences.
- Assembly-based strain resolution leverages coverage patterns and single-nucleotide variants (SNVs) across multiple samples to disentangle closely related genomes, a process exemplified by tools like STRONG and StrainPhlAn.
- Marker gene profiling (e.g., with StrainPhlAn) offers a computationally efficient method for strain placement relative to reference genomes, particularly useful for large cohorts, but is limited by database representation and requires sufficient marker gene coverage.
- De novo assembly graph phasing (e.g., with STRONG) reconstructs strain haplotypes on single-copy core genes directly from the assembly graph, enabling strain resolution even for species lacking reference genomes, but necessitates multi-sample data.
- Hybrid assembly approaches (e.g., OPERA-MS) integrating short and long reads can yield near-complete, strain-resolved genomes, offering superior contiguity and accuracy, but at a higher computational and resource cost.
- Insufficient sequencing depth, single-sample study designs, and limitations in reference databases are common failure points; multi-sample designs are essential for capturing the coverage variation needed for robust strain phasing.

---

Metagenomic assembly produces a mixture of contigs from all organisms present in a microbial community, but researchers often need genomes of individual strains to study transmission, functional variation, or population dynamics. Assembly-based strain resolution addresses this problem by using assembled contigs, coverage patterns, and single-nucleotide variants to separate closely related genomes that standard binning pipelines merge into one metagenome-assembled genome (MAG). This article explains how to recover individual strain genomes from metagenomic assemblies using tools such as MetaPhlAn and StrainPhlAn, and how to combine assembly with coverage and SNP information to separate strains. The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who have shotgun metagenomic data and need strain-level resolution for their biological questions.

## The Strain Resolution Problem in Metagenomics

Shotgun metagenomic sequencing captures DNA from every organism in a sample, but the resulting reads represent a mixture of genomes at varying abundances. When two or more strains of the same species are present, their genomes share high sequence identity, often exceeding 99% across large regions. Standard metagenomic assembly tools collapse these nearly identical sequences into a single consensus contig, which obscures the presence of multiple strains. This collapse creates a fundamental problem for researchers who need to know which strains are present, how abundant each strain is, and how strains differ functionally.

The challenge extends to species without well-characterized reference genomes. MetaPhlAn 4 was developed to address taxonomic profiling when many organisms in a community lack cultured isolates. The method integrates information from metagenome assemblies and microbial isolate genomes, and it defines unique marker genes for 26,970 species-level genome bins, of which 4,992 are taxonomically unidentified at the species level. This approach explains roughly 20% more reads in most international human gut microbiomes and more than 40% in less-characterized environments such as the rumen microbiome when compared with earlier versions. The practical implication is that strain resolution must work for both known and unknown species, and assembly-based approaches provide the raw material for discovering organisms that have no reference genome.

Strain-level information matters for biological interpretation. In a study of pediatric atopic dermatitis, shotgun metagenomic sequencing revealed that patients with more severe disease carried clonal Staphylococcus aureus strains, while Staphylococcus epidermidis strain communities were heterogeneous across all patients. The study then showed that S. aureus isolates from patients with more severe flares induced epidermal thickening and expansion of T helper 2 and T helper 17 cells in a mouse colonization model. This example demonstrates that strain-level differences can have functional consequences that species-level analysis would miss entirely.

## Core Principles of Assembly-Based Strain Resolution

### How Assembly Preserves Strain Information

Metagenomic assembly reconstructs longer contiguous sequences from short reads. The assembly graph contains information about sequence variation that is lost after variant simplification. STRONG (STrain Resolution ON assembly Graphs) exploits this property by performing coassembly, binning into MAGs, and storing the coassembly graph prior to variant simplification. The method then extracts subgraphs and their unitig per-sample coverages for individual single-copy core genes in each MAG. A Bayesian algorithm called BayesPaths determines the number of strains present, their haplotypes or sequences on the core genes, and their abundances. This approach validates on synthetic communities and produces haplotypes that match those observed from long Nanopore reads in a real anaerobic digestor time series.

The key principle is that coverage information across multiple samples provides the signal needed to separate strains. If two strains are present in different relative abundances across samples, their coverage ratios will vary, and this variation allows the algorithm to phase variants into distinct haplotypes. Single-sample analyses have less power because the coverage signal is a single observation, whereas multi-sample designs provide the replication needed to resolve strain haplotypes.

### Marker Genes and Strain Profiling

StrainPhlAn uses the marker genes identified by MetaPhlAn to profile strains within a species. The approach extracts the nucleotide sequences of these marker genes from each sample, then builds a multiple sequence alignment and phylogenetic tree to place each sample's strain in relation to reference genomes and other samples. This method works when the species has sufficient marker gene coverage in the metagenome, which depends on sequencing depth and the abundance of the species in the community.

The MetaPhlAn 4 marker gene collection is built from a curated set of 1.01 million prokaryotic reference and metagenome-assembled genomes. This large reference collection enables strain profiling for species that have no cultured isolates, because the marker genes are defined from assembled genomes recovered directly from environmental or host-associated samples. The application of MetaPhlAn 4 to more than 24,500 metagenomes highlighted previously undetected species as strong biomarkers for host conditions and lifestyles, and it showed that even previously uncharacterized species can be genetically profiled at the resolution of single microbial strains.

### Coverage-Based Separation of Strains

Coverage is the number of sequencing reads that map to a given genomic region, normalized by the length of that region. When two strains share a genomic region, the observed coverage at that region is the sum of the coverages of both strains. If the strains differ at variant sites, the allele frequencies at those sites reflect the relative abundances of the strains. For example, if strain A is present at 10x coverage and strain B is present at 5x coverage, a variant site where strain A carries one allele and strain B carries another will show the first allele at approximately 67% frequency and the second at approximately 33%.

This relationship between coverage and allele frequency is the foundation of strain separation. Algorithms such as BayesPaths model the observed allele frequencies across multiple samples to infer the number of strains, their haplotypes, and their abundances. The accuracy of this inference depends on sequencing depth, the number of samples, the genetic distance between strains, and the number of informative variant sites.

## At a Glance: Strain Resolution Approaches

| Approach | Input Data | Strain Resolution Capability | Best Use Case | Key Limitation |
| --- | --- | --- | --- | --- |
| Marker gene profiling with StrainPhlAn | Short reads mapped to MetaPhlAn marker genes | Strain-level placement relative to reference genomes | Large cohorts where species are known or discoverable | Requires sufficient marker gene coverage, limited to species with markers in the database |
| Assembly graph phasing with STRONG | Coassembled reads from multiple samples, stored assembly graph | De novo strain haplotypes on single-copy core genes | Multi-sample time series or cohort studies | Requires multiple samples, core gene regions only |
| Hybrid assembly with OPERA-MS | Short reads plus long reads from the same samples | Strain-resolved assembly with near-complete genomes | Complex communities with closely related strains | Requires long-read sequencing, higher cost and computational demand |

## Practical Workflow for Assembly-Based Strain Resolution

### Step 1: Quality Control and Preprocessing

Start with raw shotgun metagenomic reads. Remove adapter sequences, trim low-quality bases, and filter reads that map to the host genome if the samples come from host-associated communities. Document the number of reads before and after each filtering step, because coverage calculations depend on accurate read counts. The NCBI provides resources for understanding sequence data formats and quality metrics, and the [EMBL-EBI training materials](https://www.ebi.ac.uk/training) cover the fundamentals of sequence data handling for researchers who need to refresh their skills.

Record the following for each sample: total read count after quality control, mean read length, estimated sequencing depth, and the proportion of reads that pass host filtering. These records allow you to diagnose coverage failures later in the workflow.

### Step 2: Taxonomic Profiling with MetaPhlAn

Run MetaPhlAn 4 on the quality-controlled reads to determine which species are present and their relative abundances. The output includes a profile of species-level genome bins and their abundances. Use this profile to select species for strain-level analysis. Species with low abundance may not have sufficient coverage for strain resolution, and attempting strain profiling on these species will produce unreliable results.

The MetaPhlAn 4 database includes markers for 26,970 species-level genome bins, of which 4,992 are taxonomically unidentified at the species level. This means the tool can detect and quantify organisms that have no cultured representatives, which is important for environmental samples where reference genomes are sparse. The tool explains more reads in human gut microbiomes and substantially more reads in less-characterized environments such as the rumen microbiome, so it is appropriate for a wide range of sample types.

### Step 3: Metagenomic Assembly

Assemble the quality-controlled reads into contigs. The choice of assembler depends on your data type. Short-read assemblers such as MEGAHIT, IDBA-UD, and metaSPAdes are appropriate for Illumina data. If you have long reads from Oxford Nanopore or Pacific Biosciences, consider a hybrid assembler such as OPERA-MS, which integrates assembly-based metagenome clustering with repeat-aware exact scaffolding. OPERA-MS assembles metagenomes with greater base pair accuracy than long-read-only assemblers, higher contiguity than short-read-only assemblers, and fewer assembly errors than non-metagenomic hybrid assemblers.

The assembly step produces contigs that may represent multiple strains collapsed into consensus sequences. This collapse is expected and does not mean the assembly failed. The strain resolution step uses the assembly graph and coverage information to recover the individual strain sequences that were collapsed.

### Step 4: Binning into Metagenome-Assembled Genomes

Bin the assembled contigs into MAGs based on sequence composition and coverage across samples. This step groups contigs that likely originate from the same organism. The resulting MAGs represent species-level or genus-level bins, and they may contain sequences from multiple strains of the same species. STRONG performs binning as part of its workflow, and it stores the coassembly graph prior to variant simplification so that the subgraphs for individual single-copy core genes can be extracted later.

For hybrid assembly with OPERA-MS, the assembler provides strain-resolved assembly in the presence of multiple genomes of the same species. It also produces high-quality reference genomes for rare species at approximately 9x long-read coverage and near-complete genomes with higher coverage. This capability is valuable when the research question requires complete or near-complete genomes instead of marker gene profiles.

### Step 5: Strain Resolution with Coverage and SNP Information

Apply strain resolution algorithms to the assembled data. For marker-based strain profiling, use StrainPhlAn with the MetaPhlAn profile as input. For de novo strain resolution on assembly graphs, use STRONG, which performs coassembly, binning, and graph-based strain phasing in a unified workflow.

STRONG extracts subgraphs for individual single-copy core genes in each MAG and uses the unitig per-sample coverages as input to the BayesPaths algorithm. BayesPaths determines the number of strains present, their haplotypes on the core genes, and their abundances. The method requires multiple samples to provide the coverage variation needed for strain separation. A single sample provides only one coverage observation per unitig, which is insufficient for resolving multiple strain haplotypes.

### Step 6: Validation and Quality Assessment

Validate the strain resolution results using independent evidence. If long reads are available for a subset of samples, compare the strain haplotypes recovered from short-read assembly with those observed directly from long reads. STRONG was validated this way in an anaerobic digestor time series, where the haplotypes generated from short-read data matched those observed from long Nanopore reads.

Check the coverage of the marker genes or core genes used for strain resolution. Low coverage produces unreliable variant calls and ambiguous haplotypes. The MetaPhlAn 4 study demonstrated accurate quantification of organisms with no cultured isolates, but the accuracy depends on sufficient sequencing depth for the target species.

## Options and Tradeoffs in Strain Resolution Methods

### Marker-Based Strain Profiling versus De Novo Assembly Graph Phasing

Marker-based strain profiling with StrainPhlAn is computationally efficient and works well for large cohorts. The method maps reads to a curated set of marker genes, which reduces the computational burden compared with full assembly. The tradeoff is that the resolution is limited to the marker gene regions, and the method requires that the species of interest has markers in the MetaPhlAn database. The database is extensive, with markers for 26,970 species-level genome bins, but novel species that are not represented will not be profiled.

De novo assembly graph phasing with STRONG does not require reference genomes for the target species. The method identifies strains directly from the assembly graph, using coverage variation across samples to phase variants into haplotypes. The tradeoff is that the method requires multiple samples and produces haplotypes only for single-copy core genes, not for the full genome. The core gene regions are sufficient for many questions about strain identity and abundance, but they do not capture strain-specific accessory genes or mobile elements.

### Short-Read Assembly versus Hybrid Assembly

Short-read assembly is the standard approach for metagenomics because Illumina sequencing is widely available and cost-effective. The tradeoff is that short reads cannot resolve repetitive regions, and closely related strains may collapse into a single consensus sequence. Hybrid assembly with OPERA-MS combines short and long reads to overcome these limitations. The method produces more contiguous assemblies, with a 200-fold improvement in contiguity over short-read assemblies in a study of 28 gut metagenomes from antibiotic-treated patients. The hybrid assemblies included more than 80 closed plasmid or phage sequences and a new 263 kbp jumbo phage.

The tradeoff for hybrid assembly is cost and complexity. Long-read sequencing requires additional library preparation and sequencing runs, and the computational demands of hybrid assembly are higher than short-read-only assembly. The choice depends on the research question. If the goal is to identify strains and their relative abundances, short-read assembly with marker-based profiling may be sufficient. If the goal is to recover complete genomes, including plasmids and phages, hybrid assembly provides substantially more information.

### Single-Sample versus Multi-Sample Designs

Strain resolution is fundamentally a comparative problem. The coverage signal that allows strain separation comes from variation in strain abundances across samples. A single sample provides one coverage observation per genomic region, which is insufficient to distinguish multiple strains that are present at fixed relative abundances. Multi-sample designs, such as time series or cohort studies, provide the replication needed for strain phasing.

STRONG is explicitly designed for multiple metagenome samples. The method performs coassembly across all samples, which improves assembly quality for shared organisms, and then uses per-sample coverages to phase strains. The study that introduced STRONG validated the method using synthetic communities and a real anaerobic digestor time series, demonstrating that the multi-sample design is essential for the method's performance.

## Observations and Measurements for Strain Resolution

### Coverage Metrics

Measure the mean coverage for each MAG and for each single-copy core gene used in strain resolution. The coverage of a core gene is the sum of the coverages of all strains that carry that gene. If the total coverage is below a threshold where variant calling becomes unreliable, the strain resolution results will be noisy. The exact threshold depends on the sequencing platform, the error rate, and the variant calling method, but lower coverage consistently produces less reliable results.

Record the coverage distribution across samples for each target species. Species that are abundant in some samples and rare in others provide the coverage variation needed for strain phasing. Species that are uniformly rare across all samples will not yield reliable strain resolution.

### Variant Allele Frequencies

For each variant site in the core genes, record the allele frequency across samples. The allele frequency spectrum provides a direct view of the strain structure. A site with two alleles at approximately 50% frequency in all samples suggests two strains at roughly equal abundance. A site with a minor allele at 10% frequency suggests a minor strain or a mixed population. The pattern of allele frequencies across multiple linked sites allows the phasing algorithm to reconstruct haplotypes.

### Assembly Statistics

Record the assembly statistics for each sample or coassembly: total assembled bases, N50, number of contigs, and the number of MAGs recovered. These statistics provide context for interpreting strain resolution results. A fragmented assembly with a low N50 may indicate that the community is too complex for the available sequencing depth, and strain resolution will be correspondingly limited.

The OPERA-MS study provides a benchmark for assembly quality. The method achieved greater base pair accuracy than long-read-only assembly with Canu, higher contiguity than short-read assemblers including MEGAHIT, IDBA-UD, and metaSPAdes, and fewer assembly errors than non-metagenomic hybrid assemblers. These metrics are useful reference points for evaluating your own assemblies.

## Records and Documentation for Reproducibility

### Version Control and Environment Documentation

Record the exact versions of all software used in the strain resolution workflow, including the assembler, the taxonomic profiler, the strain resolution tool, and any dependency packages. The [nf-core documentation](https://nf-co.re/docs) provides standards for reproducible workflow configuration, and the [Bioconductor project](https://bioconductor.org/) offers guidance on reproducible genomic analysis with versioned packages. The [Carpentries lessons](https://carpentries.org/lessons) cover foundational computing practices, including version control with Git, which is essential for tracking changes to analysis scripts.

Create a computational environment specification that lists all software versions and their dependencies. This specification allows other researchers to reproduce your analysis or apply the same methods to their own data. Containerization tools, which are described in the nf-core documentation, provide a way to package the entire analysis environment for portability.

### Parameter Records

Record all non-default parameters used in each step of the workflow. For assembly, record the k-mer sizes, the minimum contig length threshold, and any coverage filtering parameters. For strain resolution, record the minimum coverage threshold for including a species, the minimum allele frequency for calling a variant, and the number of samples required for phasing.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that emphasizes the importance of parameter documentation for reproducibility. The training materials cover the practical aspects of running metagenomic analyses and recording the parameters that affect the results.

### Sample Metadata

Maintain a sample metadata table that includes the sample identifier, the source environment, the collection date, the DNA extraction method, the sequencing platform, and the sequencing depth. This metadata is essential for interpreting strain resolution results in the context of the biological question. For example, a time series of samples from the same individual or environment allows tracking of strain dynamics, while a cross-sectional cohort allows comparison of strain distributions across conditions.

## Common Failure Patterns in Strain Resolution

### Insufficient Sequencing Depth

The most common failure pattern is attempting strain resolution on species with insufficient coverage. When the total coverage of a species is low, the number of reads that map to each marker gene or core gene is small, and variant calls are dominated by sequencing error instead of biological variation. The result is either a failure to detect strains that are present or the detection of spurious strains that are artifacts of sequencing noise.

The solution is to check the coverage of the target species before running strain resolution. If the coverage is below the threshold required for reliable variant calling, either increase sequencing depth or exclude the species from strain-level analysis. The MetaPhlAn 4 study demonstrated accurate quantification of organisms with no cultured isolates, but the accuracy depends on sufficient sequencing depth for the target species.

### Single-Sample Designs

Attempting strain resolution on a single sample is a common failure pattern. The coverage signal from one sample is insufficient to phase variants into distinct strain haplotypes. Even if the sample contains multiple strains, the algorithm cannot determine which variants belong to which strain without the coverage variation provided by additional samples.

The solution is to design the study with multiple samples per condition or time point. STRONG requires multiple samples by design, and the method's performance improves with the number of samples and the variation in strain abundances across samples.

### Reference Database Limitations

Marker-based strain profiling fails when the target species is not represented in the reference database. The MetaPhlAn 4 database includes markers for 26,970 species-level genome bins, which is extensive, but environmental samples may contain organisms that are not represented. The database includes 4,992 bins that are taxonomically unidentified at the species level, which extends the coverage to previously uncharacterized organisms, but there is no guarantee that every organism in a given sample will have markers.

The solution is to use assembly-based approaches that do not require reference genomes. STRONG identifies strains de novo from the assembly graph, and OPERA-MS provides strain-resolved assembly without requiring reference genomes for the target species.

### Assembly Collapse of Closely Related Strains

Standard assembly tools collapse closely related strains into a single consensus sequence. This collapse is expected and is the reason why strain resolution requires specialized methods. If the assembly graph is simplified before the strain resolution step, the information needed to separate strains is lost. STRONG addresses this problem by storing the coassembly graph prior to variant simplification, which preserves the subgraphs needed for strain phasing.

The solution is to use strain resolution methods that operate on the assembly graph instead of on the simplified assembly output. If you are using a standard assembly pipeline, check whether the assembler preserves the graph information needed for strain resolution.

## Limitations of Assembly-Based Strain Resolution

### Resolution Limited to Core Genes

Assembly graph phasing with STRONG produces haplotypes for single-copy core genes, not for the full genome. Core genes are shared across all strains of a species and are under purifying selection, so they may not capture the strain-specific accessory genes that drive functional differences. The pediatric atopic dermatitis study demonstrated that S. aureus strains from patients with more severe flares induced different immune responses in mice, and these functional differences may be encoded in accessory genes that are not captured by core gene haplotypes.

### Abundance Estimation Uncertainty

The Bayesian algorithm in STRONG estimates strain abundances, but these estimates have uncertainty that depends on the coverage and the genetic distance between strains. Strains that are very similar, with few variant sites, are harder to separate, and their abundance estimates have wider confidence intervals. Strains that are present at very low abundance may be missed entirely if their coverage falls below the detection threshold.

### Computational Demands

Assembly-based strain resolution is computationally intensive. Coassembly of multiple samples requires substantial memory and processing time, and the assembly graph storage required by STRONG adds to the computational burden. Hybrid assembly with OPERA-MS requires even more resources because it integrates short and long reads. Researchers working with large cohorts or complex communities should plan for the computational requirements before starting the analysis.

The nf-core documentation provides guidance on configuring reproducible workflows for high-performance computing environments, and the Galaxy Training Network offers accessible training on running metagenomic analyses in distributed computing environments. These resources help researchers manage the computational demands of assembly-based strain resolution.

## Quality Controls and Validation Strategies

### Positive Controls

Include positive controls in the study design to validate the strain resolution workflow. Synthetic communities with known strain compositions provide a ground truth for evaluating the accuracy of strain detection, haplotype reconstruction, and abundance estimation. STRONG was validated using synthetic communities, which established the method's accuracy before application to real samples.

If synthetic communities are not available, use a well-characterized sample with known strain composition as a positive control. The validation of STRONG on an anaerobic digestor time series, where the haplotypes matched those observed from long Nanopore reads, provides a model for how to validate strain resolution results with independent data.

### Cross-Validation with Long Reads

Long reads provide an independent source of evidence for strain haplotypes. If long reads are available for a subset of samples, compare the strain haplotypes recovered from short-read assembly with those observed directly from long reads. The STRONG study demonstrated that this comparison is feasible and provides strong validation of the strain resolution results.

Hybrid assembly with OPERA-MS integrates long reads into the assembly process, which provides strain-resolved assembly directly. The method produces high-quality reference genomes for rare species with approximately 9x long-read coverage and near-complete genomes with higher coverage. This approach is more expensive than short-read-only analysis, but it provides the most reliable strain resolution.

### Reproducibility Checks

Run the strain resolution workflow on a subset of samples twice, using the same parameters, to confirm that the results are reproducible. Variability between runs indicates instability in the analysis, which may be caused by stochastic elements in the assembly or phasing algorithms. The Bioconductor project provides guidance on reproducible genomic analysis, and the nf-core documentation describes standards for workflow reproducibility.

## Safety and Regulatory Context

### Data Privacy for Human Samples

If the metagenomic samples come from human subjects, the sequence data may contain human DNA that is subject to privacy regulations. Host filtering should be performed early in the workflow to remove human reads, and the filtered data should be handled according to the applicable data protection regulations. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides resources for understanding data submission and access policies for human sequence data.

### Antimicrobial Resistance Monitoring

Strain resolution can reveal the distribution of antimicrobial resistance genes across strains within a species. The global metagenomic map of urban microbiomes identified antimicrobial resistance markers across 4,728 samples from mass-transit systems in 60 cities, and the profiles of resistance genes varied widely in type and density across cities. Researchers working on antimicrobial resistance should be aware that strain-level resolution can distinguish between resistance genes carried by different strains, which has implications for understanding the spread of resistance.

The urban microbiome study also identified 838,532 CRISPR arrays not found in reference databases, along with 10,928 viruses and 1,302 bacteria. This finding highlights the discovery potential of assembly-based approaches, but it also raises questions about the interpretation of novel sequences that have no reference context.

### Biosafety Considerations

Metagenomic samples may contain pathogens or organisms with unknown pathogenic potential. Researchers should follow institutional biosafety guidelines for handling and processing environmental or clinical samples. The strain resolution workflow itself does not involve culturing organisms, which reduces the biosafety risk compared with culture-based approaches, but the initial sample processing should follow appropriate safety protocols.

## Professional Escalation Criteria

### When to Seek Expert Assistance

Consult a bioinformatics specialist or a metagenomics core facility when the strain resolution results are inconsistent with biological expectations, when the computational demands exceed local capacity, or when the analysis requires specialized expertise that is not available in the research group. The EMBL-EBI training materials and the Galaxy Training Network provide pathways for building bioinformatics skills, but complex strain resolution analyses may benefit from expert review.

### When to Reconsider the Study Design

Reconsider the study design if the coverage of target species is consistently too low for strain resolution, if the number of samples is insufficient for strain phasing, or if the community complexity exceeds the capacity of the available assembly and strain resolution tools. The OPERA-MS study demonstrated that hybrid assembly can resolve complex communities, but the method requires long-read sequencing, which may not be feasible for all studies.

### When to Report Results with Caveats

Report strain resolution results with caveats when the coverage is marginal, when the strain haplotypes are based on a small number of variant sites, or when the validation evidence is limited. The scientific literature includes examples of strain-level findings that were later revised as methods improved, and transparent reporting of limitations supports the cumulative progress of the field.

## Decision Framework for Selecting a Strain Resolution Strategy

Choosing the correct strain resolution approach requires a structured evaluation of study goals, sample characteristics, and available resources. A practical decision framework helps researchers avoid the common failure of applying a method that does not match their biological question or data structure. The framework below organizes the decision process into four sequential checkpoints that can be applied before committing computational resources to a particular strategy.

### Checkpoint 1: Define the Resolution Target

The first decision separates studies that need strain identity and phylogenetic placement from studies that need full strain genomes. Marker-based profiling with StrainPhlAn places sample strains relative to reference genomes and other samples using conserved marker gene sequences. This approach answers questions about which strains are present, how they relate to known isolates, and whether the same strain appears across samples or time points. The pediatric atopic dermatitis study used this type of analysis to demonstrate clonal S. aureus strains in more severe patients and heterogeneous S. epidermidis strain communities across all patients, establishing strain-level differences that correlated with disease severity.

Assembly graph phasing with STRONG recovers strain haplotypes on single-copy core genes without requiring reference genomes. This approach suits studies of undercharacterized environments where reference genomes are sparse or absent. The method identifies strains de novo from multiple metagenome samples and determines the number of strains, their haplotypes, and their abundances using the BayesPaths algorithm. The validation of STRONG on an anaerobic digestor time series, where haplotypes matched those observed from long Nanopore reads, demonstrates that de novo strain resolution works in real environmental communities.

Hybrid assembly with OPERA-MS targets studies that require complete or near-complete strain genomes, including plasmids and phages. The method integrates short and long reads to produce strain-resolved assembly in the presence of multiple genomes of the same species. In a study of 28 gut metagenomes from antibiotic-treated patients, OPERA-MS produced more contiguous assemblies with a 200-fold improvement over short-read assemblies and recovered more than 80 closed plasmid or phage sequences. This level of genomic completeness supports functional analysis of strain-specific accessory genes, mobile elements, and resistance determinants.

### Checkpoint 2: Assess Sample Number and Design

The number of samples determines which strain resolution methods are viable. STRONG requires multiple metagenome samples because the coverage variation across samples provides the signal for phasing variants into distinct haplotypes. A single sample provides one coverage observation per genomic region, which is insufficient to distinguish multiple strains present at fixed relative abundances. Studies with time series or cohort designs that include multiple samples per condition provide the replication needed for assembly graph phasing.

Marker-based profiling with StrainPhlAn can work with fewer samples because it places each sample independently relative to reference genomes. However, the power to distinguish closely related strains improves with more samples, because the phylogenetic signal accumulates across the cohort. The MetaPhlAn 4 application to more than 24,500 metagenomes demonstrated that large cohorts enable detection of previously undetected species as biomarkers for host conditions and lifestyles, and that even uncharacterized species can be profiled at single strain resolution.

Hybrid assembly with OPERA-MS can be applied to individual samples or small sets of samples. The method produces high-quality reference genomes for rare species with approximately 9x long-read coverage and near-complete genomes with higher coverage. This capability makes hybrid assembly suitable for studies that prioritize genome completeness over cohort size, such as characterizing specific strains from clinical or environmental samples.

### Checkpoint 3: Evaluate Sequencing Depth and Community Complexity

Sequencing depth determines whether the target species has sufficient coverage for reliable strain resolution. Species with low abundance in the community may not have enough reads mapping to marker genes or core genes for confident variant calling. The MetaPhlAn 4 study demonstrated accurate quantification of organisms with no cultured isolates, but this accuracy depends on adequate sequencing depth for the target species.

Community complexity affects assembly quality and the ability to separate closely related strains. Complex communities with many species at similar abundances produce fragmented assemblies with lower N50 values, which reduces the contiguity of the assembled genomes available for strain resolution. The OPERA-MS evaluation using defined in vitro and virtual gut microbiomes showed that the method assembles metagenomes with greater base pair accuracy than long-read-only assembly and higher contiguity than short-read-only assembly, but the performance depends on the sequencing depth allocated to each species in the community.

For communities with multiple closely related strains of the same species, the assembly graph approach in STRONG provides a distinct advantage because it preserves the graph structure prior to variant simplification. This preservation allows the extraction of subgraphs for individual single-copy core genes, which standard assembly pipelines discard. The unitig per-sample coverages from these subgraphs provide the input for strain phasing.

### Checkpoint 4: Match Resources to the Selected Approach

Computational requirements differ substantially across the three approaches. Marker-based profiling with StrainPhlAn is the least computationally demanding because it maps reads to a curated set of marker genes instead of performing full assembly. This efficiency makes it suitable for large cohorts where the species of interest are represented in the MetaPhlAn database.

Assembly graph phasing with STRONG requires coassembly of all samples, which demands substantial memory and processing time. The method also stores the coassembly graph prior to variant simplification, adding to the storage requirements. The nf-core documentation provides standards for configuring reproducible workflows in high-performance computing environments, and the Galaxy Training Network offers accessible training on running metagenomic analyses in distributed computing systems.

Hybrid assembly with OPERA-MS requires the highest computational investment because it integrates short and long reads. The method also requires long-read sequencing, which adds cost and library preparation complexity. The decision to use hybrid assembly should be justified by the need for complete genomes, closed plasmids, or phages that short-read assembly cannot provide.

### Implementation Checklist for the Decision Framework

Apply the following checklist when selecting a strain resolution strategy for a new study. Record the decisions and their justifications in the study documentation to support reproducibility and interpretation of results.

1. State the biological question that requires strain resolution and identify whether strain identity, strain genomes, or both are needed.
2. Count the number of samples available and determine whether the study design includes multiple samples per condition or time point.
3. Estimate the expected abundance of target species from preliminary taxonomic profiling or prior knowledge of the community.
4. Assess the availability of long-read sequencing and the computational infrastructure for hybrid assembly.
5. Select the strain resolution approach that matches the resolution target, sample design, sequencing depth, and available resources.
6. Document the selection rationale and the parameters used for the chosen method.

### Escalation Criteria for the Decision Framework

Consult a bioinformatics specialist when the decision framework produces conflicting guidance, such as when the biological question requires complete strain genomes but the study design lacks the sequencing depth or sample number to support hybrid assembly. Seek expert assistance when the target species are not represented in the MetaPhlAn database and the study requires marker-based profiling, because alternative approaches such as STRONG may be necessary. Reconsider the study design when the expected abundance of target species is too low for strain resolution with the available sequencing depth, because additional sequencing or sample collection may be required before strain-level analysis is feasible.

## Frequently Asked Questions

### What is the difference between species-level and strain-level resolution in metagenomics?

Species-level resolution identifies which species are present in a community and their relative abundances. Strain-level resolution distinguishes between multiple strains of the same species, which may differ in functional genes, virulence factors, or antimicrobial resistance profiles. The pediatric atopic dermatitis study demonstrated this distinction by showing that S. aureus strain diversity, beyond species abundance, was associated with disease severity.

### How does MetaPhlAn 4 improve strain resolution compared with earlier versions?

MetaPhlAn 4 integrates information from metagenome assemblies and microbial isolate genomes to define unique marker genes for 26,970 species-level genome bins, including 4,992 that are taxonomically unidentified at the species level. This expanded marker collection allows the tool to explain more reads in human gut microbiomes and substantially more reads in less-characterized environments such as the rumen microbiome, which improves the foundation for strain-level profiling.

### What is the role of the assembly graph in strain resolution?

The assembly graph contains information about sequence variation that is lost after variant simplification. STRONG stores the coassembly graph prior to variant simplification, which enables the extraction of subgraphs for individual single-copy core genes in each MAG. The unitig per-sample coverages from these subgraphs provide the input for the Bayesian algorithm that determines the number of strains, their haplotypes, and their abundances.

### How many samples are needed for assembly-based strain resolution?

The number of samples depends on the variation in strain abundances across samples. STRONG is designed for multiple metagenome samples, and the method uses per-sample coverages to phase strains. A single sample provides only one coverage observation per genomic region, which is insufficient to distinguish multiple strains. More samples with greater variation in strain abundances provide more power for strain resolution.

### What is the advantage of hybrid assembly for strain resolution?

Hybrid assembly with OPERA-MS combines short and long reads to produce more contiguous assemblies with fewer errors. The method provides strain-resolved assembly in the presence of multiple genomes of the same species and produces high-quality reference genomes for rare species. In a study of 28 gut metagenomes from antibiotic-treated patients, hybrid assembly produced a 200-fold improvement in contiguity over short-read assemblies and recovered more than 80 closed plasmid or phage sequences.

### Can strain resolution work for species with no cultured isolates?

Yes. MetaPhlAn 4 defines marker genes from metagenome-assembled genomes, which allows taxonomic profiling and strain-level genetic profiling of organisms with no cultured isolates. The application of the method to more than 24,500 metagenomes demonstrated that even previously uncharacterized species can be genetically profiled at the resolution of single microbial strains.

### What are the main limitations of assembly-based strain resolution?

The main limitations are the resolution limited to core genes in assembly graph phasing, the uncertainty in abundance estimates for closely related strains, the computational demands of coassembly and graph storage, and the requirement for multiple samples. Hybrid assembly with long reads can overcome some of these limitations but adds cost and complexity.

### How do I validate strain resolution results?

Validate strain resolution results using positive controls with known strain compositions, cross-validation with long reads when available, and reproducibility checks by running the workflow multiple times on the same data. The STRONG study validated the method using synthetic communities and confirmed the haplotypes with long Nanopore reads in a real anaerobic digestor time series.

## Related Bioinformatics Guides

- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Metagenomic Assembly Overview: Challenges and Applications](/knowledge/bioinformatics/metagenomic-assembly-overview-challenges-and-applications)
- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)
- [Metagenomic Binning with Assembly Graph Embeddings: A New Frontier](/knowledge/bioinformatics/metagenomic-binning-with-assembly-graph-embeddings-a-new-frontier)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4.](https://pubmed.ncbi.nlm.nih.gov/36823356). Nature biotechnology, 2023.
- [STRONG: metagenomics strain resolution on assembly graphs.](https://pubmed.ncbi.nlm.nih.gov/34311761). Genome biology, 2021.
- [Staphylococcus aureus and Staphylococcus epidermidis strain diversity underlying pediatric atopic dermatitis.](https://pubmed.ncbi.nlm.nih.gov/28679656). Science translational medicine, 2017.
- [A global metagenomic map of urban microbiomes and antimicrobial resistance.](https://pubmed.ncbi.nlm.nih.gov/34043940). Cell, 2021.
- [Hybrid metagenomic assembly enables high-resolution analysis of resistance determinants and mobile elements in human microbiomes.](https://pubmed.ncbi.nlm.nih.gov/31359005). Nature biotechnology, 2019.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.