Combining Metagenomics with Metabolomics: A Guide to Integrative Analysis for Uncovering Microbiome-Metabolite Links

By Dr. Zubair Khalid, DVM, MS, PhD ·

Combining Metagenomics with Metabolomics: A Guide to Integrative Analysis for Uncovering Microbiome-Metabolite Links

Key Takeaways

  • Integrative analysis of metagenomics and metabolomics requires careful study design, including immediate sample freezing at -80°C to stabilize labile metabolites and preserve DNA integrity, and consideration of sample type (e.g., stool vs. plasma) based on the research question.
  • Robust integration necessitates large sample sizes (ideally 100-200 for exploratory work) to overcome the high-dimensional multiple-testing burden inherent in correlating thousands of microbial features with thousands of metabolite features.
  • Compositional data analysis methods, such as SparCC or log-ratio transformations, are critical for accurate correlation analysis, as standard methods like Pearson correlation are inappropriate for the relative abundance nature of both metagenomic and metabolomic data.
  • Functional metagenomic profiling, which annotates microbial genes and pathways (e.g., using MetaCyc or KEGG), is essential for moving beyond simple microbe-metabolite associations to generate testable mechanistic hypotheses about microbial metabolite production or consumption.
  • Multi-omics factor analysis (MOFA/MOFA+) offers a dimensionality reduction approach to identify latent factors capturing coordinated variation across metagenomic and metabolomic datasets, aiding in the discovery of multi-omic signatures, though factor interpretation can be challenging.
  • Reproducibility is paramount, requiring detailed analysis logs, version control of feature tables and parameter settings, and careful attention to sample identifier matching and uncorrected batch effects to avoid spurious correlations.

Integrative analysis of metagenomic and metabolomic datasets is a multi-stage process that pairs taxonomic and functional profiles of microbial communities with the chemical inventory of the same biological sample. The goal is to identify statistically robust associations between microbial features and metabolite abundances, then interpret those associations in a biological context that can generate testable mechanistic hypotheses. This guide provides a practical framework for researchers who have both data types in hand or are planning a study that will generate them. The workflow covers study design, data preprocessing, integration methods, quality control, and reporting standards, with emphasis on decisions that materially affect whether an association is real, reproducible, or an artifact of the analytical pipeline.

Scope and Reader Context

This guide is written for biology students, researchers, laboratory professionals, and life-science practitioners who need a concrete analytical framework for combining shotgun metagenomic sequencing data with untargeted or targeted metabolomic measurements. The focus is on human and animal gut microbiome studies, though the principles transfer to other environments such as soil, water, and food fermentation systems. The reader is assumed to have basic familiarity with metagenomic analysis concepts such as read quality filtering, taxonomic profiling, and functional annotation, as well as metabolomic concepts such as feature detection, peak alignment, and compound identification. What this guide adds is the integration layer: how to bring these two complex data types together without introducing spurious correlations, how to choose among available methods, and how to interpret results within the limits of what each technology can actually measure.

The problem this guide solves is specific. Many researchers can generate both metagenomic and metabolomic datasets, but the path from raw data to a credible claim about microbial drivers of metabolite changes is not obvious. The choices made at each step, from sample collection to correlation method to pathway mapping, can change the conclusions. This guide provides a decision-oriented workflow that treats integration as a designed experiment instead of a post hoc exercise.

At a Glance: Integration Method Selection

The table below summarizes the main integration approaches, their data requirements, and the types of questions they can answer. Use this table to match your research question to an appropriate method before diving into the detailed workflow.

MethodData RequirementsOutputBest Used ForKey Limitation
Correlation-based (SparCC, mmvec)Paired feature tables from same samplesMicrobe-metabolite association pairs with effect sizesExploratory screening for candidate associationsDoes not explain mechanism, sensitive to confounding
Pathway mapping (MetaboAnalyst, functional annotation)Functional metagenomic profiles plus metabolite identitiesPathway-level hypotheses linking microbial genes to metabolitesGenerating mechanistic hypotheses from known biochemistryIncomplete pathway databases, many metabolites unidentified
Multi-omics factor analysis (MOFA, MOFA+)Multiple omics matrices plus metadataLatent factors capturing coordinated variation across data typesDimensionality reduction and identifying multi-omic signaturesFactor interpretation can be difficult, requires larger sample sizes

Study Design Considerations for Multi-Omics Integration

The success of an integrative analysis is largely determined before any sequencing or mass spectrometry run begins. A study that collects samples without attention to matching metadata, technical replicates, and batch structure will produce integration results that are difficult to interpret regardless of the sophistication of the downstream analysis.

Sample Collection and Storage

Metagenomic and metabolomic measurements are typically performed on the same biological specimen, most commonly stool, but the sample handling requirements differ. DNA-based analyses are tolerant of freezing and storage conditions that can degrade metabolites. Metabolites, particularly short-chain fatty acids, bile acids, and other small molecules, are chemically labile and can be transformed by residual microbial activity after collection. For stool samples intended for both analyses, the collection protocol must stabilize the metabolome at the moment of sampling. Immediate freezing at minus 80 degrees Celsius is the standard approach, though some protocols use preservative solutions that stabilize both DNA and metabolites. The key decision is whether the preservative is compatible with both downstream assays. A preservative that works for DNA but interferes with mass spectrometry will compromise the metabolomic half of the study.

The choice of sample type also matters. Stool metabolomics reflects the chemical environment of the distal gut, including microbial products, host secretions, and undigested dietary components. Plasma or serum metabolomics captures host systemic metabolism, which is influenced by microbial products that have crossed the intestinal barrier. A study designed to link gut microbes to host physiology may need both stool and blood metabolomics. The Framingham Heart Study analysis that identified cholesterol-metabolizing bacteria used stool metagenomics and metabolomics from 1429 participants and linked those measurements to blood lipids and cardiovascular health indicators, demonstrating that stool measurements can be associated with systemic outcomes when the study is large enough and the metadata are rich [<a href="#ref-1">1</a>].

Sample Size and Statistical Power

Multi-omics studies are expensive, and the sample size needed to detect robust cross-omics associations is often larger than researchers anticipate. The number of features in each data type is enormous. A metagenomic profile may contain hundreds of species and thousands of functional pathways. An untargeted metabolomic run may detect thousands of metabolite features, many of which are unidentified. Testing all pairwise associations between these feature sets creates a massive multiple-testing burden. A study with 50 samples will have limited power to detect true associations after correction for the number of tests performed.

The inflammatory bowel disease multi-omics study that integrated nine metagenomic and four metabolomics cohorts used cross-cohort integrative analysis to identify consistent microbial and metabolite signatures, demonstrating that replication across independent cohorts is one strategy for dealing with the statistical challenges of high-dimensional integration [<a href="#ref-2">2</a>]. For a single-cohort study, the practical implication is that sample size should be as large as the budget allows, with a minimum in the range of 100 to 200 samples for exploratory integration work. Smaller studies can still generate hypotheses, but the results should be framed as hypothesis-generating instead of confirmatory.

Batch Structure and Confounding Variables

Metagenomic sequencing runs and metabolomic mass spectrometry runs are both subject to batch effects. Samples processed in different batches can show systematic differences that have nothing to do with biology. The integration step is particularly vulnerable because a batch effect in one data type can create spurious correlations with a batch effect in the other. The solution is to randomize sample processing order across experimental groups and to include technical replicates or pooled reference samples in every batch. These reference samples allow batch correction during data preprocessing.

Confounding variables are another threat to valid integration. Diet, medication use, age, sex, and body mass index can influence both the microbiome and the metabolome. If a confounder is unevenly distributed across the comparison groups, the integration analysis may attribute to microbes an association that is actually driven by the confounder. The inflammatory bowel disease literature illustrates this problem. The gut microbiota and metabolites are closely associated with IBD progression, but inconsistent findings across studies have impeded a comprehensive understanding of their roles, in part because of differences in cohort composition and confounding factors [<a href="#ref-2">2</a>]. Collecting detailed metadata and including confounders as covariates in the statistical model is essential.

Data Preprocessing for Metagenomic and Metabolomic Integration

The integration step assumes that both data types have been processed to a state where the features are reliable and the abundances are comparable across samples. Preprocessing choices made independently for each data type will shape the integration results.

Metagenomic Data Processing

Shotgun metagenomic data require a processing pipeline that converts raw sequencing reads into taxonomic and functional profiles. The standard steps are quality trimming, host DNA removal, and either read-based profiling or assembly-based analysis. Read-based profiling maps reads to reference databases to estimate taxonomic abundances and functional pathway coverage. Assembly-based analysis reconstructs metagenome-assembled genomes, which can provide strain-level resolution and access to genes that are absent from reference databases.

The choice between operational taxonomic unit-based and exact sequence variant-based methods has been a major topic in microbiome analysis. Best practices now favor exact sequence variants over operational taxonomic unit clustering because exact sequence variants provide finer taxonomic resolution and are more reproducible across studies [<a href="#ref-3">3</a>]. For shotgun data, the equivalent decision is whether to profile at the species level or to attempt strain-level analysis. Species-level profiling is more robust and is appropriate for most integration studies. Strain-level analysis adds resolution but requires deeper sequencing and more sophisticated computational methods.

Functional profiling is often more informative than taxonomic profiling for integration with metabolomics. A taxonomic association between a microbe and a metabolite is a starting point, but the mechanistic link is the microbial gene or pathway that produces or consumes the metabolite. The insulin resistance study that combined fecal metabolomics with metagenomics identified gut bacteria associated with insulin resistance and insulin sensitivity that showed distinct patterns of carbohydrate metabolism, and the functional analysis revealed the involvement of microbial carbohydrate metabolism pathways [<a href="#ref-4">4</a>]. This type of finding requires functional annotation of the metagenomic data, beyond taxonomic counts.

Metabolomic Data Processing

Metabolomic data processing depends on the platform. Nuclear magnetic resonance spectroscopy is highly reproducible and provides absolute quantification of a limited set of metabolites. Mass spectrometry coupled to liquid or gas chromatography detects thousands of features but requires careful alignment and normalization. Untargeted metabolomics generates feature tables where each feature is defined by a retention time and mass-to-charge ratio. The same metabolite may appear as multiple features due to adducts, isotopes, and in-source fragmentation. Feature annotation, the assignment of a chemical identity to a feature, is a major bottleneck. In the IBD study that performed untargeted metabolomic profiling of stool samples, more than half of the differentially abundant metabolite features were uncharacterized, and many could only be assigned putative roles through covariation with known metabolites [<a href="#ref-5">5</a>]. This pattern is typical and should be expected.

Data normalization is a critical preprocessing step for metabolomics. The total ion count, the sum of all detected features, can vary between samples due to technical factors such as injection volume and instrument sensitivity. Normalization methods include total ion count scaling, median scaling, and probabilistic quotient normalization. The choice of method affects the downstream integration. For stool metabolomics, the water content of the sample is a major source of variation. Two stool samples with the same metabolite concentration per gram of dry matter will show different concentrations per gram of wet weight if their water contents differ. Normalizing to dry weight or using a normalization method that accounts for sample dilution is important for valid comparisons.

Data Structure and Format for Integration

The integration step requires both data types to be in a compatible format. The standard structure is a sample-by-feature matrix for each data type, with samples as rows and features as columns. The sample identifiers must match exactly between the two matrices. A common error is a mismatch in sample naming between the metagenomic and metabolomic datasets, which can cause silent merging failures or, worse, incorrect merging of samples. A quality control step that verifies sample identifier consistency before integration is essential.

The feature tables should also be checked for sparsity. Metabolomic feature tables often contain many features that are detected in only a few samples. These sparse features are prone to spurious correlations and should be filtered before integration. A common filter is to retain features present in at least 20 to 30 percent of samples. Metagenomic profiles may also contain rare species that are present in only a few samples. Filtering rare features reduces the multiple-testing burden and improves the stability of the integration results.

Core Principles of Integrative Analysis

The integration of metagenomic and metabolomic data rests on several core principles that guide method selection and interpretation.

Correlation Does Not Imply Causation

The most fundamental principle is that a statistical association between a microbe and a metabolite does not establish that the microbe produces or consumes the metabolite. The association could arise because both are influenced by a third factor, such as diet or host genetics. The association could also be bidirectional, with the metabolite influencing the microbe instead of the reverse. The multi-omics biological correlation maps constructed in the IBD cross-cohort study highlighted gut microbial biotransformation deficiencies and significant alterations in aminoacyl-tRNA synthetases, but the authors framed these as resources for developing mechanistic hypotheses instead of as proven mechanisms [<a href="#ref-2">2</a>]. The integration analysis generates hypotheses. Experimental validation, such as in vitro culture experiments or animal models, is required to test them.

Compositional Data Analysis

Both metagenomic and metabolomic data are compositional. The relative abundances of taxa in a microbiome sample sum to a constant, and the relative abundances of metabolite features in a metabolomic sample are similarly constrained. This compositionality means that an increase in one feature necessarily corresponds to a decrease in others, even if the absolute abundance of the decreased features is unchanged. Standard correlation methods, such as Pearson correlation, are not appropriate for compositional data because they can produce spurious correlations. Best practices for microbiome analysis now emphasize the use of compositional data analysis methods [<a href="#ref-3">3</a>]. For integration, this means using correlation methods that account for compositionality, such as SparCC or Procrustes-based approaches, or transforming the data to log-ratios before correlation analysis.

The Importance of Matched Data

The power of integrative analysis comes from having both data types measured on the same biological samples. Unmatched datasets, where the metagenomic and metabolomic measurements come from different individuals or different time points, cannot support direct integration. The Framingham Heart Study analysis generated stool metagenomics and metabolomics from the same 1429 participants, and this matched design allowed the identification of microbial pathways implicated in cardiovascular disease, including flavonoid, gamma-butyrobetaine, and cholesterol metabolism [<a href="#ref-1">1</a>]. The matched design is what makes the integration meaningful.

Data Integration Methods and Tools

Several classes of methods are available for integrating metagenomic and metabolomic data. The choice of method depends on the research question, the data structure, and the computational resources available.

Correlation-Based Integration

The simplest integration approach is to calculate pairwise correlations between microbial features and metabolite features across samples. The result is a correlation matrix that can be filtered for statistically significant associations. The limitations of this approach are the multiple-testing burden and the compositional data problem. Methods such as SparCC and Procrustes analysis address the compositionality issue, but they do not solve the multiple-testing problem. The output of correlation-based integration is a list of microbe-metabolite pairs with correlation coefficients and p-values. This list is the starting point for pathway mapping and hypothesis generation.

The mmvec tool, which stands for microbe-metabolite vectors, uses a neural network approach to learn the conditional probability of a metabolite given a microbe. This method is designed to handle the sparsity and compositionality of both data types and can identify microbe-metabolite associations that are missed by simple correlation. The output is a set of vectors that can be visualized in a biplot, showing which microbes and metabolites cluster together. This tool is particularly useful for exploratory analysis of large datasets.

Pathway Mapping and Functional Integration

Correlation-based integration identifies associations but does not explain them. Pathway mapping adds a functional layer by asking whether the microbes associated with a metabolite also carry the genes for the biochemical pathways that produce or consume that metabolite. This approach requires functional annotation of the metagenomic data, typically through the use of pathway databases such as MetaCyc or KEGG. The metagenomic functional analysis in the IBD cross-cohort study revealed that a gene in the two-component system pathway was linked to fecal calprotectin, a measure of gut inflammation, demonstrating the value of functional analysis for connecting microbial genes to host phenotypes [<a href="#ref-2">2</a>].

The integration of pathway information with metabolite data can be done manually or with software tools. MetaboAnalyst provides a suite of tools for metabolomic analysis, including pathway enrichment analysis that can be combined with metagenomic functional profiles. The workflow is to identify differentially abundant metabolites, map them to metabolic pathways, and then ask whether the microbial genes for those pathways are differentially abundant in the same samples. This approach generates mechanistic hypotheses that are more specific than simple correlations.

Multi-Omics Factor Analysis

Multi-omics factor analysis methods, such as MOFA and MOFA+, identify latent factors that explain variation across multiple data types simultaneously. These methods decompose the data into a small number of factors, each of which captures a pattern of coordinated variation across the metagenomic and metabolomic features. The factors can be associated with metadata, such as disease status or treatment group, to identify the multi-omic signatures of the condition of interest. This approach is useful for reducing the dimensionality of the data and for identifying coordinated changes that would be missed by pairwise analysis.

The choice between correlation-based, pathway-based, and factor analysis methods is not either-or. A comprehensive integration workflow might use factor analysis for dimensionality reduction, correlation analysis for identifying specific microbe-metabolite pairs, and pathway mapping for mechanistic interpretation. The IBD study that constructed multi-omics biological correlation maps used a combination of approaches to identify robust associations between differentially abundant species and well-characterized differentially abundant metabolites, finding 122 such associations that indicated possible mechanistic relationships perturbed in IBD [<a href="#ref-5">5</a>].

Practical Workflow for Integrative Analysis

The following workflow provides a step-by-step procedure for integrating metagenomic and metabolomic datasets. The workflow assumes that both data types have been preprocessed to the feature table stage.

Step 1: Verify Data Quality and Sample Matching

Before any integration analysis, verify that the metagenomic and metabolomic feature tables have the same samples in the same order. Check for missing samples in either table. A sample present in the metagenomic table but absent from the metabolomic table cannot be used in the integration. Verify that the sample identifiers match exactly, including any prefixes or suffixes. Create a merged data object that contains both feature tables and the metadata.

Step 2: Filter Sparse Features

Filter features that are present in too few samples. For metabolomic features, a common threshold is to retain features present in at least 20 to 30 percent of samples. For metagenomic features, filter out species or pathways that are present in fewer than 10 to 20 percent of samples. The exact threshold depends on the sample size and the expected prevalence of the features. Document the filtering decisions and the number of features retained at each step.

Step 3: Apply Compositional Data Transformations

Transform both data types to address compositionality. For metagenomic data, options include centered log-ratio transformation or the use of compositional data analysis methods. For metabolomic data, the same considerations apply. The choice of transformation affects the downstream correlation analysis. Document the transformation method and the rationale.

Step 4: Perform Correlation Analysis

Calculate pairwise correlations between microbial features and metabolite features. Use a method that accounts for compositionality, such as SparCC or mmvec. Apply multiple-testing correction, such as the Benjamini-Hochberg false discovery rate method. Filter the results to retain associations that pass the significance threshold. The output is a list of microbe-metabolite pairs with effect sizes and adjusted p-values.

Step 5: Map Associations to Pathways

For the significant microbe-metabolite pairs, map the metabolites to biochemical pathways and map the microbial genes to the same pathways. Ask whether the microbes associated with a metabolite carry the genes for the pathways that produce or consume that metabolite. This step can be done with MetaboAnalyst for the metabolite side and with functional annotation tools for the metagenomic side. The output is a set of pathway-level hypotheses.

Step 6: Validate and Interpret

Assess the robustness of the associations. If the study has a validation cohort, test whether the associations replicate. If not, use cross-validation or bootstrap resampling to assess stability. Interpret the results within the limits of the study design. An association that is statistically robust and biologically plausible is a candidate for experimental validation. An association that is statistically robust but biologically implausible may indicate a technical artifact or a confounding variable.

Tools and Resources for Implementation

Several resources provide training and tools for the steps in this workflow. The Galaxy Training Network offers accessible workflow training and analysis tutorials that cover metagenomic and metabolomic analysis in a reproducible environment [<a href="#ref-6">6</a>]. The nf-core documentation describes community standards for bioinformatics pipelines, including pipelines for metagenomic analysis that follow reproducibility best practices [<a href="#ref-7">7</a>]. The Bioconductor project provides packages for genomic and metabolomic analysis with a focus on reproducible research [<a href="#ref-8">8</a>]. The EMBL-EBI Training portal offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-9">9</a>]. The Carpentries lessons provide foundational training in computing, data handling, shell, Git, and programming that are prerequisites for effective bioinformatics work [<a href="#ref-10">10</a>]. The NCBI provides the primary sequence databases and search systems used in metagenomic analysis [<a href="#ref-11">11</a>].

The choice of tools depends on the computational environment. Researchers with access to a high-performance computing cluster may prefer command-line tools and workflow managers such as nf-core pipelines. Researchers working on a laptop may prefer the Galaxy platform, which provides a web-based interface to many bioinformatics tools. The key is to use tools that support reproducible analysis, meaning that the analysis can be rerun with the same inputs to produce the same outputs.

Records and Measurements for Reproducibility

Reproducibility is a central concern in multi-omics integration. The complexity of the analysis means that small differences in preprocessing or parameter choices can change the results. The following records should be maintained for every integration study.

Analysis Log

Maintain a detailed log of every analysis step, including software versions, parameter settings, and input files. The log should be sufficient for another researcher to reproduce the analysis from the raw data. Workflow managers such as nf-core pipelines automatically generate such logs, but manual analyses require disciplined record keeping [<a href="#ref-7">7</a>].

Feature Tables

Retain the feature tables at every stage of preprocessing. The raw feature tables from the metagenomic and metabolomic pipelines are the starting point for integration. The filtered and transformed tables are the inputs to the correlation analysis. Each version should be saved with a version number and a description of the changes.

Parameter Settings

Record the parameter settings for every tool used in the analysis. This includes quality trimming thresholds, filtering thresholds, correlation method parameters, and multiple-testing correction methods. The choice of parameters can materially affect the results, and the parameters should be justified in the methods section of any publication.

Session Information

Record the computing environment, including the operating system, the versions of all software packages, and the random number generator seed if the analysis involves stochastic elements. This information is essential for reproducing the analysis at a later date.

Common Failure Patterns in Integrative Analysis

Several failure patterns recur in metagenomic and metabolomic integration studies. Recognizing these patterns can help researchers avoid them or diagnose problems when they arise.

Sample Identifier Mismatch

The most common and most damaging failure is a mismatch between the sample identifiers in the metagenomic and metabolomic feature tables. This can cause samples to be merged incorrectly, producing associations that are entirely spurious. The fix is to verify sample identifier consistency before any analysis. This check should be automated and should be part of the quality control workflow.

Uncorrected Batch Effects

Batch effects in either data type can create spurious cross-omics correlations. If the metagenomic samples were processed in two batches and the metabolomic samples were processed in the same two batches, the batch effect will appear as a correlation between the two data types. The fix is to include batch as a covariate in the statistical model or to apply batch correction methods before integration.

Compositional Artifacts

The use of standard correlation methods on compositional data can produce spurious correlations. This is a well-known problem in microbiome analysis, and best practices now emphasize compositional data analysis methods [<a href="#ref-3">3</a>]. The fix is to use correlation methods that account for compositionality or to transform the data to log-ratios before analysis.

Overinterpretation of Correlations

The interpretation of a microbe-metabolite correlation as evidence of a mechanistic relationship is a common failure. The correlation may be driven by a confounder, or it may be bidirectional. The fix is to frame correlation results as hypothesis-generating and to require experimental validation before claiming a mechanistic link.

Ignoring Unidentified Metabolites

Untargeted metabolomics typically produces a large fraction of unidentified features. In the IBD study, more than half of the differentially abundant metabolite features were uncharacterized [<a href="#ref-5">5</a>]. Ignoring these features loses information. The fix is to use approaches such as metabolomic guilt by association, where unidentified features are assigned putative roles based on their covariation with known metabolites [<a href="#ref-5">5</a>].

Limitations of Integrative Analysis

Integrative analysis has inherent limitations that should be acknowledged in any study report.

Technical Limitations

Metagenomic and metabolomic technologies each have blind spots. Metagenomics cannot distinguish between live and dead microbes, and it cannot measure the activity of microbial enzymes. Metabolomics measures the chemical inventory of the sample but cannot always identify the source of a metabolite. A metabolite detected in stool could be produced by the host, by the microbiota, or by the diet. The integration analysis cannot resolve this ambiguity without additional information.

Statistical Limitations

The high dimensionality of both data types creates a severe multiple-testing problem. Even with false discovery rate correction, the number of true associations that can be detected is limited by the sample size. Studies with small sample sizes will produce results that are unstable and unlikely to replicate. The cross-cohort integrative analysis approach used in the IBD study is one strategy for addressing this limitation, but it requires access to multiple independent cohorts [<a href="#ref-2">2</a>].

Biological Limitations

The biological interpretation of integration results is constrained by the current state of knowledge. Many microbial genes have unknown functions, and many metabolites have unknown biochemical roles. The pathway databases used for functional integration are incomplete. A microbe-metabolite association that maps to a known pathway is easier to interpret than one that does not, but the absence of a pathway annotation does not mean the association is not real.

Safety and Regulatory Context

Multi-omics research involving human subjects is subject to ethical and regulatory requirements that vary by jurisdiction. The collection of stool, blood, and other biological samples requires informed consent and institutional review board approval. The storage and sharing of genomic and metabolomic data are subject to data protection regulations, including the General Data Protection Regulation in the European Union and the Health Insurance Portability and Accountability Act in the United States. Researchers should consult with their institutional review board and data protection officer before beginning a multi-omics study.

The computational analysis of human microbiome data raises additional privacy concerns. Metagenomic data can contain human DNA sequences that are removed during preprocessing, but the removal is not always complete. The risk of re-identification from microbiome data is low but not zero. Researchers should follow best practices for data de-identification and should store data on secure systems.

Professional Escalation Criteria

Researchers should seek expert assistance when they encounter situations that exceed their analytical expertise. The following criteria indicate when escalation is appropriate.

Unfamiliar Data Types

A researcher who is experienced in metagenomics but new to metabolomics should seek assistance with the metabolomic preprocessing and quality control. The failure modes of mass spectrometry data are different from those of sequencing data, and the normalization and annotation steps require specialized knowledge.

Computational Resource Limits

Integration methods such as mmvec and multi-omics factor analysis can be computationally intensive. A researcher who lacks access to sufficient computational resources should seek assistance from a bioinformatics core facility or a collaborator with high-performance computing access.

Statistical Uncertainty

A researcher who is uncertain about the appropriate statistical methods for compositional data analysis or multiple-testing correction should consult a biostatistician. The choice of statistical methods can materially affect the conclusions, and getting it wrong can invalidate the results.

Reproducibility Requirements

A researcher who needs to produce reproducible analysis for regulatory or publication purposes should seek assistance with workflow management and documentation. The nf-core documentation and the Galaxy Training Network provide resources for reproducible analysis, but a researcher with complex requirements may need expert assistance [<a href="#ref-6">6</a>][<a href="#ref-7">7</a>].

Frequently Asked Questions

What is the difference between read-based and assembly-based metagenomic analysis for integration with metabolomics?

Read-based analysis maps sequencing reads to reference databases to estimate taxonomic and functional abundances. Assembly-based analysis reconstructs metagenome-assembled genomes from the reads, which can provide access to genes that are absent from reference databases. For integration with metabolomics, read-based functional profiling is often sufficient because the goal is to identify pathways that are present and differentially abundant. Assembly-based analysis adds resolution but requires deeper sequencing and more computational resources. The choice depends on the research question and the available resources.

How do I choose between correlation-based integration and pathway mapping?

Correlation-based integration identifies microbe-metabolite pairs that are statistically associated. Pathway mapping adds a functional layer by asking whether the associated microbes carry the genes for the pathways that produce or consume the associated metabolites. Correlation-based integration is appropriate for exploratory analysis and hypothesis generation. Pathway mapping is appropriate when the goal is to identify mechanistic links. A comprehensive workflow uses both approaches.

What is the minimum sample size for an integrative metagenomic and metabolomic study?

There is no single minimum sample size that applies to all studies. The required sample size depends on the effect sizes, the number of features, and the desired statistical power. In practice, studies with fewer than 100 samples are likely to have limited power to detect robust cross-omics associations after multiple-testing correction. Studies with 200 or more samples are better positioned for integration analysis. The cross-cohort integrative analysis approach used in the IBD study is one strategy for increasing effective sample size by combining multiple cohorts [<a href="#ref-2">2</a>].

How do I handle unidentified metabolites in the integration analysis?

Unidentified metabolites are a common feature of untargeted metabolomics. In the IBD study, more than half of the differentially abundant metabolite features were uncharacterized [<a href="#ref-5">5</a>]. One approach is to use metabolomic guilt by association, where unidentified features are assigned putative roles based on their covariation with known metabolites [<a href="#ref-5">5</a>]. Another approach is to focus the integration analysis on identified metabolites and treat unidentified features as a secondary analysis. The choice depends on the research question and the fraction of unidentified features in the dataset.

What is compositionality and why does it matter for integration?

Compositionality refers to the constraint that the relative abundances of features in a sample sum to a constant. Both metagenomic and metabolomic data are compositional. This constraint means that an increase in one feature necessarily corresponds to a decrease in others, even if the absolute abundance of the decreased features is unchanged. Standard correlation methods are not appropriate for compositional data because they can produce spurious correlations. Best practices for microbiome analysis now emphasize the use of compositional data analysis methods [<a href="#ref-3">3</a>].

Can I integrate metagenomic and metabolomic data from different samples?

Direct integration requires both data types to be measured on the same biological samples. Unmatched datasets cannot support direct integration because the associations between microbes and metabolites are computed across samples. If the samples are different, the associations are meaningless. The matched design is what makes the integration valid. The Framingham Heart Study analysis generated stool metagenomics and metabolomics from the same participants, and this matched design allowed the identification of microbial pathways implicated in cardiovascular disease [<a href="#ref-1">1</a>].

What is the role of experimental validation in integrative analysis?

Integrative analysis generates hypotheses about microbe-metabolite relationships, but it does not prove causation. The associations identified by integration could be driven by confounders or could be bidirectional. Experimental validation is required to test the hypotheses. The insulin resistance study identified gut bacteria associated with insulin resistance and insulin sensitivity that showed distinct patterns of carbohydrate metabolism, and the researchers demonstrated that insulin-sensitivity-associated bacteria ameliorated host phenotypes of insulin resistance in a mouse model [<a href="#ref-4">4</a>]. This type of experimental validation is the gold standard.

How do I report the results of an integrative analysis in a publication?

The methods section should describe the study design, sample collection, data generation, preprocessing, and integration methods in sufficient detail for another researcher to reproduce the analysis. The results section should report the number of features at each stage, the integration method, the multiple-testing correction, and the significant associations. The discussion should frame the results as hypotheses that require experimental validation. The data and code should be deposited in public repositories to support reproducibility.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Gut microbiome and metabolome profiling in Framingham heart study reveals cholesterol-metabolizing bacteria.](https://pubmed.ncbi.nlm.nih.gov/38569543). Cell, 2024. [2] [Microbiome and metabolome features in inflammatory bowel disease via multi-omics integration analyses across cohorts.](https://pubmed.ncbi.nlm.nih.gov/37932270). Nature communications, 2023. [3] [Best practices for analysing microbiomes.](https://pubmed.ncbi.nlm.nih.gov/29795328). Nature reviews. Microbiology, 2018. [4] [Gut microbial carbohydrate metabolism contributes to insulin resistance.](https://pubmed.ncbi.nlm.nih.gov/37648852). Nature, 2023. [5] [Gut microbiome structure and metabolic activity in inflammatory bowel disease.](https://pubmed.ncbi.nlm.nih.gov/30531976). Nature microbiology, 2019. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [9] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [10] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [11] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.