From Reads to Relative Abundance: A Beginner's Guide to Taxonomic Profiling in Shotgun Metagenomics
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Shotgun metagenomics provides relative, not absolute, microbial abundance, influenced by laboratory methods and bioinformatics choices; interpretation requires considering host DNA depletion efficiency and read depth targets (e.g., 20 million reads for cost-efficiency).
- Taxonomic profiling accuracy is critically dependent on the reference database used; absence of detection does not equate to absence of an organism due to database biases towards well-studied species.
- Quality control, including raw read assessment (e.g., FastQC) and host read depletion via mapping to host genomes, is paramount to improve microbial detection sensitivity and reduce downstream misclassification.
- Different classification approaches (marker gene, k-mer, alignment, assembly) offer distinct trade-offs in speed, resolution, and computational requirements, with pipeline selection strongly impacting precision, as demonstrated in bloodstream infection diagnostics.
- Relative abundance tables are compositional data, requiring specialized statistical methods (e.g., log-ratio transformations) for accurate correlation and differential abundance testing due to inherent dependencies between taxa.
- Reproducibility in metagenomic analysis is achieved through meticulous documentation of all workflow decisions, software/database versions, and parameters, often facilitated by containerization and provenance tracking tools.
Shotgun metagenomics generates sequencing data from all DNA present in a sample, and taxonomic profiling is the process of determining which organisms are present and in what proportions. This article walks through the complete workflow from raw sequencing reads to a relative abundance table, with attention to the decisions that affect result quality. The intended reader is a biology student, researcher, or laboratory professional who has sequenced metagenomic samples and needs a clear map of the analysis path ahead.
What Taxonomic Profiling Answers and What It Cannot
Taxonomic profiling assigns sequencing reads to taxa at ranks from domain down to species or strain. The output is typically a table where rows are taxa and columns are samples, with values representing relative abundance. Relative abundance means the proportion of the microbial community attributed to each taxon, not the absolute number of organisms in the original sample.
A multicenter assessment of shotgun metagenomics for pathogen detection found that mapped read counts could indicate relative and intra-site microbial abundance but not absolute or inter-site abundance. The same study reported that assay performance varied significantly across sites and microbial classes, with false positive reporting and site or library effects as common challenges. This means the abundance table you produce is an estimate shaped by your laboratory methods and bioinformatics choices, not a direct measurement of the original community.
Shotgun metagenomics detects DNA from bacteria, viruses, fungi, and parasites without prior knowledge of the infectious agent. This unbiased character makes it valuable for clinical diagnosis, environmental monitoring, and microbiome research. However, detection alone does not establish causation. A portable metagenomics review for livestock and poultry emphasizes that pathogen and resistance gene signals must be interpreted with clinical signs, lesions, epidemiology, controls, and confirmatory testing.
At a Glance: Workflow Stages and Key Decisions
| Workflow Stage | Primary Input | Key Decision Points | Common Output |
|---|---|---|---|
| Raw data preparation | FASTQ files from sequencing platform | Read depth target, quality thresholds, adapter removal | Cleaned reads ready for analysis |
| Host read depletion | Cleaned reads, host reference genome | Mapping tool selection, mismatch tolerance | Host-depleted microbial reads |
| Taxonomic classification | Host-depleted reads, reference database | Classifier choice, database version, parameters | Per-read taxonomic assignments |
| Abundance table generation | Classifier output | Normalization method, taxonomic rank for summarization | Relative abundance table |
| Quality assessment and interpretation | Abundance table, control samples | Comparison with controls, plausibility checks | Validated taxonomic profile |
The Input Data: What You Start With
Sequencing Platform Output
The workflow begins with raw sequencing data, typically FASTQ files from Illumina short-read platforms or Oxford Nanopore Technologies long-read platforms. Short-read shotgun metagenomics on Illumina platforms remains the most common approach for taxonomic profiling. Long-read shotgun metagenomics offers advantages for genome assembly but has distinct biases that shape the recovered community.
A critical review of soil microbiome benchmarking evidence notes that short-read 16S rRNA amplicon sequencing, full-length 16S sequencing on Oxford Nanopore Technologies platforms, and long-read shotgun metagenomics each have distinct biases. The choice of sequencing method has a strong effect on which species are detected and how the community is described. For taxonomic profiling specifically, short-read shotgun data is the standard input for most established pipelines.
Read Depth Considerations
Read depth, or sequencing depth, refers to the number of sequencing reads generated per sample. The multicenter shotgun metagenomics assessment identified a read depth of 20 million reads as a generally cost-efficient assay setting. This threshold balances detection sensitivity against sequencing cost. Lower depths may miss low-abundance organisms, while higher depths yield diminishing returns for taxonomic profiling.
For targeted applications, far fewer reads may suffice. A study of targeted next-generation sequencing for lower respiratory tract infections used 0.1 million reads for multiplex PCR-based tNGS and 1 million reads for hybrid capture-based tNGS. These targeted approaches restrict analysis to a predefined pathogen panel, which reduces the sequencing required but sacrifices the unbiased detection that defines shotgun metagenomics.
Sample Type and Host Contamination
The proportion of host DNA in your sample directly affects how many sequencing reads are useful for microbial profiling. Clinical samples such as blood or tissue may contain mostly human DNA, reducing the effective microbial read depth. The multicenter assessment recommended that laboratory-developed shotgun metagenomics tests aim to detect microbes at 500 CFU/mL in a clinically relevant host context of 10^5 human cells/mL. This context matters because host DNA consumes sequencing capacity.
For food and environmental samples, the matrix itself introduces challenges. A shotgun metagenomics workflow for vitamin-containing food products required an optimized DNA extraction protocol tailored to diverse vitamin formulations. Honey samples present a complex biological matrix containing plant-derived, microbial, and viral components. The choice of DNA extraction method is a workflow decision that propagates through the entire analysis.
Quality Control Before Profiling
Raw Read Assessment
Before any taxonomic assignment, assess the quality of your raw reads. FastQC or equivalent tools report per-base quality scores, GC content, adapter contamination, and duplication levels. Low-quality bases at read ends should be trimmed. Adapter sequences must be removed because they do not originate from the sample and will interfere with taxonomic assignment.
The Galaxy Training Network provides accessible workflow training for quality control and metagenomic analysis. The Carpentries offers foundational computing and data skills that support reproducible analysis practices. These training resources are appropriate for beginners building the computational skills needed for metagenomics.
Host Read Depletion
Removing host reads before taxonomic profiling improves sensitivity for microbial taxa. This step is especially important for clinical samples where host DNA dominates. Host depletion can be performed by mapping reads to the host reference genome and retaining unmapped reads for downstream analysis. The NCBI provides reference genome databases for this purpose.
For livestock and poultry samples, host depletion reduces the sequencing capacity consumed by animal DNA. The portable metagenomics review for veterinary applications identifies host depletion or target enrichment as a key step in sample-to-answer workflows for enteric and respiratory disease in food-producing animals.
Trimming and Filtering Decisions
Trimming parameters affect which reads survive to taxonomic assignment. Aggressive trimming removes low-quality bases but may discard short reads that carry taxonomic information. Conservative trimming preserves more data but leaves errors that can cause misclassification. The choice depends on your sequencing platform and the tolerance of your downstream classifier for mismatches.
Taxonomic Classification Approaches
Marker Gene Based Profiling
Marker gene based methods identify a set of clade-specific marker genes and map reads against these references. MetaPhlAn is the most widely used example. The bioBakery 3 platform integrates MetaPhlAn 3 for taxonomic profiling and reports increased accuracy compared to current alternatives. Marker gene approaches are computationally efficient because they compare reads against a reduced reference set instead of all available genomes.
The bioBakery platform was assessed against other publicly available shotgun metagenomics pipelines using mock community samples. In that assessment, bioBakery4 performed best on most accuracy metrics, while other pipelines had higher sensitivities. The authors noted that bioBakery is commonly used and only requires basic command line knowledge.
K-mer Based Classification
K-mer based classifiers such as Kraken assign reads by comparing k-mers, short subsequences of fixed length, against a database of labeled genomes. These methods are fast and can classify reads at high taxonomic resolution. The choice of database and k-mer length affects sensitivity and specificity.
A study comparing four bioinformatics pipelines for bloodstream infection diagnosis found that pipeline selection strongly affected the precision of findings. An optimized BLAST pipeline was superior to Kraken, MetaPhlAn, and RTG Core in that specific clinical context. This result does not mean BLAST is universally best. It demonstrates that the optimal pipeline depends on the sample type, the organisms of interest, and the reference databases available.
Alignment Based Methods
Alignment based methods map reads to reference genomes using tools such as BLAST or Bowtie2. These approaches are more sensitive than k-mer methods for divergent organisms but require more computational time. The optimized BLAST pipeline in the bloodstream infection study achieved accurate pathogen identification where other methods failed, suggesting that alignment based approaches retain value when reference databases are incomplete.
Assembly Based Profiling
Metagenome assembly reconstructs longer contiguous sequences from short reads before taxonomic assignment. This approach can recover genomes of novel organisms that lack close references. The nf-core documentation describes community pipeline standards for reproducible workflow execution, and the Shotgun-NF pipeline integrates metagenome assembly, genome binning, and metagenome-assembled genome quality assessment within a unified framework.
Assembly based profiling is computationally intensive and requires careful quality assessment. The metaGOflow workflow for marine genomic observatories data supports taxonomic inference based on ribosomal RNA genes and functional annotation using raw reads. Its modular implementation allows running the workflow partially, which is useful when only taxonomic profiles are needed.
Reference Databases and Their Limitations
Database Choice
The reference database determines what your classifier can detect. If an organism is absent from the database, its reads will remain unclassified or be assigned to the nearest relative. The NCBI maintains comprehensive sequence databases including genomes, genes, and taxonomic information. The NCBI taxonomy identifier system provides a standardized naming framework for comparing results across pipelines.
A mock community assessment included a workflow for labeling bacterial scientific names with NCBI taxonomy identifiers for better resolution in assessing results. This labeling step matters because different pipelines may report the same organism under different names or at different taxonomic ranks.
Database Version and Updates
Reference databases are updated regularly as new genomes are deposited. The database version used for analysis must be recorded because it affects results. Reanalysis with an updated database can change taxonomic assignments, especially for recently described species. The Shotgun-NF pipeline includes database manifests in its provenance tracking, which supports reproducibility by documenting exactly which database version was used.
Database Bias
Reference databases are biased toward well-studied organisms. Human pathogens, model organisms, and commercially important species are overrepresented. Environmental organisms, especially from soil and marine systems, are underrepresented. The soil microbiome benchmarking review notes that shotgun metagenomics reveals systematic biases in assembly that depend on population diversity within the sample. These biases mean that absence of detection does not prove absence of the organism.
The Relative Abundance Table
From Classifications to Counts
After classifying reads, the next step is to convert classifications into an abundance table. Each read assigned to a taxon contributes to that taxon's count. Relative abundance is calculated by dividing the count for each taxon by the total number of classified reads. This normalization makes samples comparable within a study but does not account for differences in total microbial load between samples.
The multicenter shotgun metagenomics assessment found that mapped read counts could indicate relative and intra-site microbial abundance but not absolute or inter-site abundance. This distinction is critical for interpretation. A taxon at 10 percent relative abundance in one sample and 20 percent in another may not have doubled in absolute terms if the total microbial load differs.
Normalization Methods
Several normalization approaches exist beyond simple relative abundance. Rarefaction subsamples reads to equal depth across samples, which reduces sensitivity for low-abundance taxa. Cumulative sum scaling and trimmed mean of M values are alternatives that adjust for compositional effects. The choice of normalization method affects downstream statistical analysis and should be documented.
Compositional Data Considerations
Relative abundance tables are compositional data. The sum of all taxa in each sample is fixed at 1, which creates dependencies between taxa. If one taxon increases, others must decrease even if their absolute abundances are unchanged. This property complicates correlation analysis and differential abundance testing. Statistical methods designed for compositional data, such as those based on log-ratio transformations, are recommended for downstream analysis.
Workflow Options and Tradeoffs
Individual Tools Versus Integrated Pipelines
Beginners face a choice between running individual tools and using an integrated pipeline. Individual tools offer flexibility and transparency but require the user to manage software dependencies, file formats, and parameter choices. Integrated pipelines package multiple tools into a single workflow with standardized inputs and outputs.
The nf-core documentation describes community pipeline standards that support reproducible workflow execution. The Shotgun-NF pipeline provides an end-to-end workflow from paired-end Illumina data through taxonomic and functional profiling, assembly, binning, and strain-level analysis. Its containerized execution using Conda and Apptainer facilitates portability across local workstations, high-performance computing environments, and cloud infrastructures.
Containerization and Reproducibility
Containerization packages software with its dependencies so that the same environment can be reproduced on any system. The metaGOflow workflow highlights how containerization technologies along with modern workflow languages and metadata packaging support researchers dealing with increasing volumes of biological data. Reproducibility is supported through automated provenance tracking, including execution reports, runtime traces, workflow timelines, database manifests, and directed acyclic graph visualizations.
Cloud Versus Local Computing
Taxonomic profiling of shotgun metagenomics data requires substantial computational resources. Marker gene based methods such as MetaPhlAn can run on a laptop for small datasets. K-mer based classification and assembly based approaches may require high-performance computing or cloud resources. The choice depends on dataset size, available infrastructure, and time constraints.
Graphical Interfaces Versus Command Line
Galaxy provides a web-based graphical interface for bioinformatics analysis. The Galaxy Training Network offers accessible workflow training and analysis tutorials that are appropriate for beginners. Command line approaches offer more flexibility and are required for some pipelines. The bioBakery assessment noted that the platform only requires basic command line knowledge, making it accessible to researchers without advanced computational training.
Practical Implementation Steps
Step 1: Organize Your Data and Metadata
Before starting analysis, organize your raw sequencing files and associated metadata. Metadata includes sample identifiers, collection dates, sample types, and any clinical or environmental information. The metaGOflow workflow uses Research Object Crate packaging so that relevant metadata about the sample and the details of the bioinformatics analysis are inherited to the data product. This practice ensures that results remain interpretable in the future.
Step 2: Run Quality Control
Assess raw read quality and document the results. Record the number of reads before and after trimming, the quality score thresholds used, and the adapter sequences removed. These records allow you to compare results across samples and to troubleshoot unexpected findings.
Step 3: Deplete Host Reads
Map reads to the host reference genome and retain unmapped reads for taxonomic profiling. Record the proportion of reads removed as host. This proportion is an important quality metric because it indicates how much sequencing capacity was consumed by host DNA.
Step 4: Select and Run a Classifier
Choose a taxonomic classifier based on your research question, sample type, and computational resources. Document the classifier version, the reference database version, and all parameters used. Run the classifier on your host-depleted reads.
Step 5: Generate the Abundance Table
Convert classifier output into a relative abundance table. Include taxonomic lineage information so that results can be summarized at different ranks. Record the total number of classified reads and the proportion of unclassified reads for each sample.
Step 6: Assess Quality and Interpret
Examine the abundance table for anomalies. A very high proportion of unclassified reads may indicate poor database coverage or sequencing quality. Unexpected taxa may represent contamination or misclassification. Compare your results with positive and negative controls if available.
Step 7: Document and Archive
Record all workflow decisions, software versions, database versions, and parameters. Archive the abundance table and the code used to generate it. This documentation supports reproducibility and allows you or others to revisit the analysis.
Records and Measurements to Keep
Essential Records
Maintain a laboratory notebook or electronic record that includes the following for each sample: sequencing platform and instrument, read length and depth, quality control metrics before and after trimming, host read depletion proportion, classifier and version, reference database and version, classifier parameters, total classified reads, proportion of unclassified reads, and the final abundance table.
Quality Metrics
The proportion of reads classified to the expected taxa is a key quality metric. For clinical samples with a suspected pathogen, the detection of that pathogen at reasonable abundance supports the diagnosis. For environmental samples, the overall taxonomic profile should be consistent with the expected community. The multicenter shotgun metagenomics assessment reported that false positive reporting was a common challenge, so unexpected detections should be treated with caution.
Control Samples
Positive controls with known microbial composition, such as mock communities, allow you to assess the accuracy of your workflow. The mock community assessment used 19 publicly available mock community samples to evaluate pipeline performance. Negative controls, such as extraction blanks and sequencing blanks, identify contamination introduced during sample processing. The wastewater surveillance study noted that no included dataset employed internal standards or multi-workflow comparisons on shared samples, which limited inter-study comparability.
Common Failure Patterns and How to Address Them
Low Classification Rates
If a large proportion of reads remain unclassified, the reference database may lack coverage for the organisms in your sample. This pattern is common for environmental samples with novel or underrepresented taxa. Options include using a different classifier with a broader database, assembling reads before classification, or accepting that the community contains organisms absent from current references.
Contamination Signals
Unexpected taxa in your abundance table may originate from reagents, laboratory equipment, or cross-contamination between samples. The vitamin-containing food product workflow detected unexpected biological impurities in commercial products, demonstrating that metagenomics can reveal contamination that targeted methods miss. Compare your results with negative controls and consider whether detected taxa are plausible given your sample type.
Batch Effects
Samples processed in different batches may show systematic differences in taxonomic profiles that reflect technical variation instead of biological differences. The wastewater surveillance study found that workflow decisions such as concentration method, DNA extraction, and sequencing strategy significantly impacted both microbiome and resistome composition. Workflow variation was confounded with geographic location, making it difficult to separate technical from biological effects.
Database Version Drift
Results obtained with different database versions are not directly comparable. If you reanalyze data with an updated database, document the version change and expect some taxonomic assignments to change. The Shotgun-NF pipeline includes database manifests in its provenance tracking to support this documentation.
Overinterpretation of Relative Abundance
Relative abundance changes do not necessarily reflect absolute abundance changes. A taxon may appear to decrease in relative abundance simply because another taxon increased. The multicenter assessment explicitly cautioned that mapped read counts indicate relative and intra-site abundance but not absolute or inter-site abundance. Interpret your results with this limitation in mind.
Interpretation Limits and Statistical Considerations
Detection Does Not Imply Causation
The portable metagenomics review for livestock and poultry emphasizes that detection alone does not establish causation. Pathogen and resistance gene signals must be interpreted with clinical signs, lesions, epidemiology, controls, and confirmatory testing. This principle applies across sample types. A taxon detected in a metagenomic sample may be a true member of the community, a transient contaminant, or a remnant of dead cells.
Sensitivity Limits
Shotgun metagenomics has limited sensitivity for low-abundance organisms. The vitamin-containing food product workflow reliably detected high-level impurities but had limited detection of trace-level impurities. The multicenter assessment recommended a detection threshold of 500 CFU/mL in a clinically relevant host context. Organisms below this threshold may be missed.
Strain Level Resolution
Taxonomic profiling at the species level is standard, but strain level resolution requires specialized approaches. The bioBakery 3 platform includes StrainPhlAn 3 and PanPhlAn 3 for strain level profiling. Strain level analysis can reveal phylogenetic and functional structure within a species that species level profiling misses. The bioBakery 3 study used strain level profiling to unravel the structure of Ruminococcus bromii, a common gut microbe previously described by only 15 isolate genomes.
Functional Profiling Complements Taxonomy
Taxonomic profiling answers which organisms are present. Functional profiling answers what they can do. The bioBakery 3 platform integrates HUMAnN 3 for functional potential and activity. A function based index was more useful than a taxonomy based index for distinguishing obesity groups in a study of gut microbiome analysis. Consider whether your research question requires functional information in addition to taxonomic composition.
Clinical and Diagnostic Context
Bloodstream Infections
Rapid identification of causative pathogens is essential for bloodstream infection management. Shotgun metagenomics can detect a wide range of fungal, viral, and bacterial organisms along with their antimicrobial resistance genes. The bloodstream infection pipeline comparison found that the selection of bioinformatics pipelines strongly affected the precision of findings. An optimized BLAST pipeline was the only method that accurately identified the causative pathogens in that study.
Respiratory Infections
Shotgun metagenomics is widely used to detect pathogens in bronchoalveolar lavage fluid. The targeted next-generation sequencing study found that tNGS approaches were faster and less expensive than mNGS while achieving comparable sensitivity and specificity. However, mNGS detected filamentous fungi that tNGS missed, and tNGS failed to detect anaerobic bacteria in some samples. These tradeoffs matter when choosing between unbiased shotgun metagenomics and targeted approaches.
Infectious Disease Diagnosis
Metagenomic next-generation sequencing enables simultaneous detection of bacteria, viruses, fungi, and parasites without prior knowledge of the infectious agent. The review of mNGS for infectious disease diagnosis highlights its capability to identify rare, novel, or unculturable pathogens. Challenges include data analysis complexity, high cost, and the need for optimized sample preparation protocols.
Food Safety Monitoring
Shotgun metagenomics can detect biological impurities in food products, including pathogens, allergens, and genetically modified microorganisms carrying antimicrobial resistance genes. The vitamin-containing food product workflow confirmed the presence of known biological impurities and uncovered unexpected ones. This open approach offers advantages over targeted PCR based methods that cannot detect untargeted or unknown impurities.
Veterinary Applications
Portable metagenomics for livestock and poultry offers a route to broad pathogen detection, antimicrobial resistance gene profiling, and outbreak investigation. Applications include calf diarrhea, bovine respiratory disease, poultry outbreaks, mastitis, and resistome monitoring. The review emphasizes that portable metagenomics is not a replacement for conventional diagnostics but can reduce uncertainty during time-sensitive outbreaks and support more judicious antimicrobial use.
Environmental and Ecological Applications
Marine Genomic Observatories
The metaGOflow workflow was developed for marine genomic observatories that conduct regular biological community samplings. It supports fast inference of taxonomic profiles from ribosomal RNA genes and functional annotation using raw reads. The workflow scales to the needs of projects producing large metagenomic data sets.
Soil Microbiomes
Soil contains thousands of microbial species at vastly different abundances. The choice of sequencing method has a strong effect on which species are detected and how the community is described. The soil microbiome benchmarking review argues that method choice should be framed as an important part of study design, with the biases of the chosen method acknowledged and controlled where possible.
Honey and Food Matrices
Untargeted shotgun metagenomics can simultaneously characterize botanical origin, microbial communities, and viral content in honey. A pilot study detected plant-derived sequences assigned to sunflower and eucalyptus, bacterial taxa including Paenibacillus larvae and Apilactobacillus kunkeei, and Apis mellifera filamentous virus. These findings were exploratory given the limited sample size but demonstrate the integrative potential of the approach.
Wastewater Surveillance
Wastewater based surveillance enables community level disease monitoring. Untargeted shotgun metagenomics enables assessment of all antimicrobial resistance genes in a sample, allowing novel gene detection and surveillance of thousands of genes simultaneously. The wastewater surveillance study found that workflow decisions significantly impacted both microbiome and resistome composition, highlighting the urgent need for workflow standards and positive controls.
Disease Association Studies
Gut Microbiome and Disease
Meta-analysis of gut microbiome profiles across 11 diseases identified shared and distinct microbial signatures. The study reinforced the generalization of binary classifiers for Crohn's disease and colorectal cancer to hold-out cohorts and identified key microbes driving these classifications. High microbial similarity was found in disease pairs such as Crohn's disease versus ulcerative colitis and Parkinson's disease versus type 2 diabetes.
Aging and Microbiota Transfer
Fecal microbiota transfer between young and aged mice demonstrated that microbiota composition profiles and key species are successfully transferred. The transfer of aged donor microbiota into young mice accelerated age-associated inflammation, while transfer of young donor microbiota reversed detrimental effects. Whole metagenomic shotgun sequencing was used to analyze changes in gut microbiota composition and metabolic potential.
Obesity Indices
A gut microbiome obesity index was developed using taxa and biological functions correlated with body mass index. A function based index differentiated between all obesity groups, while a taxonomy based index did not differentiate between healthy controls and the lower obesity group. This finding supports the value of integrating functional information into metagenomic analysis.
Professional Escalation Criteria
When to Seek Expert Assistance
Seek assistance from a bioinformatics specialist or core facility when you encounter any of the following situations: classification rates below 20 percent for reasons you cannot identify, unexpected detection of high consequence pathogens, results that conflict with clinical or epidemiological expectations, or the need to compare results across studies that used different workflows.
When to Repeat the Analysis
Repeat the analysis with different parameters or tools when results are unexpected or when you suspect a technical artifact. The bloodstream infection study demonstrated that different pipelines produced different results on the same samples. If a clinically important finding depends on the choice of pipeline, confirm it with an independent method.
When to Question the Data
Question the data when you observe patterns that are biologically implausible. A taxon detected at high abundance in a sample type where it is never found may indicate contamination. A complete absence of expected taxa may indicate a problem with DNA extraction, sequencing, or analysis. The multicenter assessment reported that false positive reporting was a common challenge, so unexpected detections warrant scrutiny.
When to Consult Reference Materials
Consult the NCBI for database information and taxonomy resources. Use the EMBL-EBI Training for bioinformatics learning pathways and data resource training. The Bioconductor project provides official package and workflow documentation for reproducible genomic analysis. The Galaxy Training Network offers accessible workflow training. The Carpentries provides foundational computing and data skills. The nf-core documentation describes community pipeline standards.
Safety and Regulatory Context
Clinical Diagnostic Use
Laboratory developed shotgun metagenomics tests for pathogen detection should aim to detect microbes at 500 CFU/mL in a clinically relevant host context within a 24 hour turnaround time and with an efficient read depth of 20 million reads. These performance targets come from the multicenter assessment and provide a benchmark for laboratory validation.
Antimicrobial Resistance Monitoring
Shotgun metagenomics can detect antimicrobial resistance genes in clinical, environmental, and agricultural samples. The livestock and poultry review emphasizes that resistance gene signals must be interpreted with clinical signs and confirmatory testing. Detection of a resistance gene does not mean the organism carrying it is causing disease.
Food Safety
Shotgun metagenomics for food safety monitoring can detect pathogens, allergens, and genetically modified microorganisms. The vitamin-containing food product workflow provides a proof of concept for metagenomics based impurity surveillance. Results that suggest contamination of commercial products should be confirmed with targeted methods before regulatory action.
Data Sharing and Privacy
Metagenomic data from human samples contains human DNA sequences that may raise privacy concerns. Host read depletion reduces but does not eliminate human sequence content. Follow institutional and regulatory requirements for data storage, sharing, and deidentification.
A Decision Framework for Choosing Between Read-Based and Assembly-Based Taxonomic Profiling
After generating a relative abundance table, many beginners face a question that the standard workflow description does not answer directly: should the taxonomic profile come from classifying individual reads against reference databases, or should the reads first be assembled into longer contigs before taxonomic assignment? The choice between read-based and assembly-based profiling is not a matter of one being universally better. It is a decision that depends on your sample type, your research question, your computational resources, and the organisms you expect to find. This section provides a practical decision framework that you can apply before committing to a pipeline.
The Core Distinction
Read-based taxonomic profiling classifies each sequencing read independently by comparing it against a reference database. Marker gene methods such as MetaPhlAn and k-mer methods such as Kraken operate this way. Assembly-based profiling first reconstructs longer contiguous sequences from overlapping reads, then assigns taxonomy to the assembled contigs or to genomes binned from those contigs. The Shotgun-NF pipeline integrates both approaches, offering read-based taxonomic profiling alongside metagenome assembly, genome binning, and metagenome-assembled genome quality assessment within a single workflow.
The distinction matters because each approach answers a different question. Read-based profiling asks which known organisms are present and in what proportions. Assembly-based profiling asks what genomes can be reconstructed from this sample, including organisms that may lack close references. A soil microbiome benchmarking review notes that shotgun metagenomics reveals systematic biases in assembly that depend on population diversity within the sample. This means assembly results are shaped by how many similar species coexist, beyond by the sequencing technology.
Decision Criteria for Read-Based Profiling
Choose read-based taxonomic profiling as your primary approach when you meet any of the following conditions.
Your research question concerns known organisms with good reference genome coverage. Clinical samples with suspected pathogens, gut microbiome studies, and surveillance of well-characterized foodborne organisms fall into this category. The bloodstream infection pipeline comparison found that an optimized BLAST pipeline accurately identified causative pathogens in blood samples from patients with hematological malignancies, demonstrating that read-based approaches can perform well in clinical contexts when references are adequate.
Your sample contains mostly host DNA. Read-based methods can classify the microbial reads that survive host depletion without requiring the depth needed for assembly. The multicenter shotgun metagenomics assessment recommended a read depth of 20 million reads as generally cost efficient, and this depth supports read-based classification. Assembly of host-dominated samples would waste computational resources on host contigs.
You need results quickly. Read-based classification is computationally faster than assembly. The targeted next-generation sequencing study reported that multiplex PCR-based tNGS took 10.3 hours and hybrid capture-based tNGS took 16 hours, but these targeted approaches are not shotgun metagenomics. For unbiased shotgun data, read-based profiling completes in hours on a standard workstation, while assembly can take days.
You are comparing many samples. Large cohort studies such as the meta-analysis of gut microbiome profiles across 11 diseases benefit from the consistency and speed of read-based profiling. The bioBakery 3 platform, which performed best on most accuracy metrics in a mock community assessment, uses marker gene based read classification and scales to thousands of metagenomes.
Decision Criteria for Assembly-Based Profiling
Choose assembly-based taxonomic profiling when you meet any of the following conditions.
You expect novel or underrepresented organisms. Environmental samples from soil, marine systems, and other poorly characterized habitats often contain organisms absent from reference databases. The metaGOflow workflow for marine genomic observatories supports taxonomic inference based on ribosomal RNA genes, but assembly can recover genomes of organisms that read-based methods cannot classify. The soil microbiome benchmarking review emphasizes that method choice should be framed as an important part of study design, with the biases of the chosen method acknowledged.
You need strain-level resolution or genome-resolved analysis. The bioBakery 3 platform includes StrainPhlAn 3 and PanPhlAn 3 for strain-level profiling of metagenomes, and these methods operate on read-based data. However, full genome reconstruction requires assembly and binning. The Shotgun-NF pipeline integrates metagenome-assembled genome quality assessment, which allows you to evaluate whether reconstructed genomes meet quality thresholds for downstream analysis.
You are investigating functional potential that requires complete genes or operons. While read-based functional profiling such as HUMAnN 3 can identify functional potential, assembly provides contiguous sequences that span entire genes. This is valuable for detecting antimicrobial resistance genes, biosynthetic gene clusters, and other features that require full-length context.
You have high sequencing depth and a sample with low complexity. Assembly performs best when reads overlap sufficiently to reconstruct long contigs. A sample dominated by a few abundant organisms assembles more successfully than a highly diverse community where many species exist at low abundance. The soil benchmarking review notes that assembly biases depend on population diversity within the sample.
A Practical Scoring Approach
To make the decision systematic, score your situation against the following criteria. Assign one point for each condition that applies.
Read-based profiling scores one point for each of these conditions: your organisms of interest have good reference coverage, your sample is host-dominated, you need results within days, you are analyzing more than 20 samples, your computational resources are limited to a laptop or standard workstation, and your research question is about community composition instead of genome reconstruction.
Assembly-based profiling scores one point for each of these conditions: you expect novel organisms, you need strain-level or genome-resolved analysis, you are investigating functional features requiring full gene context, your sample has low to moderate diversity, you have access to high-performance computing or cloud resources, and your sequencing depth exceeds 20 million reads per sample.
If read-based scoring is higher, start with a read-based pipeline such as bioBakery or Kraken. If assembly-based scoring is higher, consider a workflow such as Shotgun-NF that includes assembly and binning. If the scores are equal, run read-based profiling first to get an initial community overview, then decide whether assembly is justified based on the proportion of unclassified reads.
Recording Your Decision
Document the rationale for your choice in your laboratory notebook or electronic record. Include the following for each project: the scoring results for both approaches, the sample types and expected organisms, the sequencing depth and platform, the computational resources available, and the research question that motivated the choice. This record allows you to revisit the decision if results are unexpected.
The wastewater surveillance study found that workflow decisions such as concentration method, DNA extraction, and sequencing strategy significantly impacted both microbiome and resistome composition. The choice between read-based and assembly-based profiling is another workflow decision that shapes your results. Recording it supports reproducibility and helps you interpret differences between your results and published studies that may have used a different approach.
Common Failure Patterns in the Decision Process
One common failure is choosing assembly-based profiling for a highly diverse environmental sample with modest sequencing depth. The resulting assembly is fragmented, and taxonomic assignment of short contigs offers no advantage over read-based classification. If your assembly produces few contigs longer than 1000 base pairs, consider switching to read-based profiling for your taxonomic question.
Another failure is choosing read-based profiling when your sample contains a novel pathogen or organism of interest that is absent from reference databases. The reads remain unclassified, and you miss the organism entirely. The vitamin-containing food product workflow detected unexpected biological impurities that targeted methods missed, but this required a workflow designed for untargeted detection. If your proportion of unclassified reads exceeds 50 percent, assembly may recover organisms that read-based methods cannot identify.
A third failure is assuming that assembly-based results are always more accurate. The bloodstream infection study found that an optimized BLAST pipeline was superior to alternatives for pathogen identification, and BLAST is a read-based alignment method. Assembly does not automatically improve accuracy. It changes which organisms can be detected and how their abundances are estimated.
When to Use Both Approaches
For many projects, the best decision is to run read-based profiling for the full dataset and assembly-based analysis on a subset of samples or on samples with high proportions of unclassified reads. The metaGOflow workflow supports modular implementation, allowing you to run the workflow partially. This flexibility means you can start with read-based taxonomic profiles and add assembly only where it adds value.
The bioBakery 3 platform demonstrates that read-based and assembly-based approaches can complement each other. MetaPhlAn 3 provides accurate taxonomic profiling, while StrainPhlAn 3 and PanPhlAn 3 provide strain-level resolution. The platform detected novel disease-microbiome links in applications to colorectal cancer and inflammatory bowel disease using these integrated methods. Your workflow can similarly combine approaches to answer different aspects of your research question.
Escalation Criteria for the Decision
Seek expert assistance if you cannot determine whether your sample contains organisms absent from reference databases, if your assembly produces poor quality metrics that you cannot interpret, or if your read-based classification leaves more than half of reads unclassified and you do not understand why. A bioinformatics core facility can run both approaches on a test sample and compare the results. The nf-core documentation describes community pipeline standards that can help you evaluate whether a workflow is appropriate for your data. The Galaxy Training Network offers tutorials that compare different analysis approaches on the same datasets.
Frequently Asked Questions
What is the difference between taxonomic profiling and metagenome assembly?
Taxonomic profiling assigns sequencing reads to known taxa by comparing them against reference databases. Metagenome assembly reconstructs longer contiguous sequences from short reads without requiring a reference. Assembly can recover genomes of novel organisms but is computationally intensive and requires careful quality assessment. Many workflows perform taxonomic profiling on raw reads and assembly as a separate step for genome resolved analysis.
How many reads do I need for taxonomic profiling?
A read depth of 20 million reads is generally cost efficient for shotgun metagenomics, according to a multicenter assessment. Lower depths may miss low abundance organisms. Targeted approaches can work with far fewer reads, such as 0.1 to 1 million reads, but they restrict analysis to a predefined pathogen panel. The optimal depth depends on your sample type, the organisms of interest, and your research question.
Why do different pipelines give different results on the same data?
Different pipelines use different classification methods, reference databases, and parameters. A study of bloodstream infection diagnosis found that pipeline selection strongly affected the precision of findings. Marker gene based methods such as MetaPhlAn compare reads against clade specific markers. K-mer based methods such as Kraken compare short subsequences against labeled genomes. Alignment based methods such as BLAST map reads to reference genomes. Each approach has distinct biases.
What does relative abundance mean and why is it limited?
Relative abundance is the proportion of the microbial community attributed to each taxon. It is calculated by dividing the count for each taxon by the total number of classified reads. The multicenter shotgun metagenomics assessment found that mapped read counts could indicate relative and intra-site abundance but not absolute or inter-site abundance. A taxon at 10 percent relative abundance in one sample and 20 percent in another may not have doubled in absolute terms.
How do I know if my taxonomic profile is accurate?
Use mock community samples with known composition to assess the accuracy of your workflow. The mock community assessment used 19 publicly available mock community samples to evaluate pipeline performance. Compare your results with positive and negative controls. A very high proportion of unclassified reads may indicate poor database coverage. Unexpected taxa may represent contamination or misclassification.
What should I do if most of my reads are unclassified?
If a large proportion of reads remain unclassified, the reference database may lack coverage for the organisms in your sample. This pattern is common for environmental samples with novel or underrepresented taxa. Options include using a different classifier with a broader database, assembling reads before classification, or accepting that the community contains organisms absent from current references.
Can shotgun metagenomics detect everything in my sample?
No. Shotgun metagenomics has limited sensitivity for low abundance organisms. The vitamin-containing food product workflow reliably detected high level impurities but had limited detection of trace level impurities. The multicenter assessment recommended a detection threshold of 500 CFU/mL in a clinically relevant host context. Organisms below this threshold may be missed. Reference database bias also means that organisms absent from databases will not be detected.
How do I make my taxonomic profiling analysis reproducible?
Document all workflow decisions, software versions, database versions, and parameters. Use containerization to package software with its dependencies. The Shotgun-NF pipeline includes automated provenance tracking with execution reports, runtime traces, workflow timelines, database manifests, and directed acyclic graph visualizations. The metaGOflow workflow uses Research Object Crate packaging so that metadata and analysis details are inherited to the data product.
Related Bioinformatics Guides
- Metagenomics Pipeline: From Raw Reads to Taxonomic and Functional Profiles
- Metagenomics Data Analysis: From Raw Reads to Biological Insights
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- Metagenomics Functional Profiling: Tools and Databases for Pathway Analysis
- Metagenomics vs Metatranscriptomics: Choosing the Right Approach for Functional Profiling
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Application of metagenomic next-generation sequencing in the diagnosis of infectious diseases.. Frontiers in cellular and infection microbiology, 2024.
- metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data.. GigaScience, 2022.
- Enhancing lower respiratory tract infection diagnosis: implementation and clinical assessment of multiplex PCR-based and hybrid capture-based targeted next-generation sequencing.. EBioMedicine, 2024.
- Fecal microbiota transfer between young and aged mice reverses hallmarks of the aging gut, eye, and brain.. Microbiome, 2022.
- Meta-analysis of the human gut microbiome uncovers shared and distinct microbial signatures between diseases.. mSystems, 2024.
- Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3.. eLife, 2021.
- Diagnostic Accuracy of Shotgun Metagenomics for Bloodstream Infections Is Influenced by Bioinformatics Workflow Selection.. MicrobiologyOpen, 2025.
- Multicenter assessment of shotgun metagenomics for pathogen detection.. EBioMedicine, 2021.
- Development of a shotgun metagenomics workflow for the comprehensive surveillance of biological impurities in vitamin-containing food products. 2025.
- Shotgun-NF: A reproducible Nextflow pipeline for end-to-end shotgun metagenomics analysis. 2026.
- Portable metagenomics for preventive surveillance and outbreak control in livestock and poultry: Pathogen detection, resistome profiling, and antimicrobial stewardship.. 2026.
- Quantifying Workflow Bias in Antimicrobial Resistance Gene Wastewater Surveillance via Metagenomics Workflows. 2026.
- Choosing Between Short-Read 16S, Full-Length ONT 16S, and Long-Read Shotgun Metagenomics for Soil Microbiome Studies: A Critical Review of the Benchmarking Evidence.. 2026.
- A pilot proof-of-concept study of microbial and botanical diversity in honey samples from Necochea, Argentina.. 2026.
- Development of a shotgun metagenomics workflow for the comprehensive surveillance of biological impurities in vitamin-containing food products. LWT, 2025.
- The Gut Microbiome Obesity Index: A New Analytical Tool in the Metagenomics Workflow for the Evaluation of Gut Dysbiosis in Obese Humans. Nutrients, 2025.
- Mock community taxonomic classification performance of publicly available shotgun metagenomics pipelines. Scientific Data, 2024.
- CViewer: a Java-based statistical framework for integration of shotgun metagenomics with other omics datasets. bioRxiv, 2023.
- Detection of parasites in food and water matrices by shotgun metagenomics: A narrative review. Food and Waterborne Parasitology, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.