Essential Quality Control Metrics for Shotgun Metagenomic Sequencing Data

By Dr. Zubair Khalid, DVM, MS, PhD ·

Essential Quality Control Metrics for Shotgun Metagenomic Sequencing Data

Key Takeaways

  • Per-base quality scores, particularly the median Q30, are critical for assessing sequencing accuracy; however, in metagenomic data, it's crucial to inspect quality distributions across all reads and positions, not just averages, due to varying GC content and potential quality drops at read ends.
  • Read duplication rates must be interpreted in the context of community composition; high duplication can inflate apparent abundance of dominant taxa and obscure rare ones, necessitating careful distinction between biological and PCR-induced duplication.
  • Adapter contamination, common in short-insert metagenomic libraries, must be rigorously assessed and trimmed, as unremoved adapter sequences can lead to alignment failures, false positive taxonomic assignments, and chimeric contigs.
  • GC content distribution provides insights into library composition and potential bias; a broad distribution is expected for diverse communities, while a narrow distribution may indicate single-taxon dominance, extraction bias, or contamination.
  • Host DNA content significantly reduces effective microbial sequencing depth in host-associated samples, requiring quantification and potential removal or increased sequencing depth depending on the biological question and target abundance.
  • Negative controls are indispensable for identifying contamination introduced during sample processing, and their taxonomic composition must be compared to biological samples to distinguish true biological signal from reagent or environmental artifacts.

Shotgun metagenomic sequencing generates complex datasets that require systematic quality assessment before any biological interpretation can be trusted. This article defines the core quality control metrics that researchers, laboratory professionals, and bioinformatics practitioners must evaluate after sequencing, explains how to interpret those metrics specifically for metagenomic data, and provides practical thresholds and decision criteria for when to proceed, when to reprocess, and when to reject a sequencing run. The guidance applies to shotgun metagenomic sequencing of microbial communities from environmental, clinical, agricultural, and host-associated samples, with emphasis on the distinct challenges that arise when sequencing DNA from mixed communities instead of from a single isolate or amplicon.

Why Shotgun Metagenomic Data Requires Different Quality Control Than Other Sequencing Data

Shotgun metagenomic sequencing differs fundamentally from whole-genome sequencing of a single organism or from amplicon sequencing of a single marker gene. In a shotgun metagenomic experiment, the sequencing instrument fragments and sequences all DNA present in a sample, which means the resulting reads represent a mixture of genomes from bacteria, archaea, viruses, fungi, and the host organism. This mixed composition creates quality control challenges that do not arise in simpler sequencing projects.

The first challenge is that no single reference genome can serve as the expected outcome. With a bacterial isolate, a researcher knows approximately what genome content should be present and can assess coverage and completeness against that expectation. With metagenomic data, the expected composition is unknown before analysis, so quality metrics must assess the technical properties of the sequencing run itself instead of the biological content. The NCBI Data Resources provide access to reference databases and sequence archives that support these assessments, but the interpretation of quality metrics remains the responsibility of the analyst.

The second challenge is dynamic range. A microbial community typically contains a few abundant taxa and many rare taxa. Sequencing depth that is adequate to characterize the abundant members may be entirely insufficient to detect or quantify the rare members. Quality control metrics must therefore be evaluated in the context of the biological question. A dataset with excellent per-base quality and low duplication might still be inadequate for detecting low-abundance pathogens or rare functional genes.

The third challenge is contamination risk. Because shotgun metagenomic sequencing captures everything in a sample, it also captures contaminants introduced during collection, DNA extraction, library preparation, or sequencing. Distinguishing true biological signal from contamination requires careful attention to negative controls, reagent blanks, and the taxonomic composition of expected versus unexpected organisms. The Galaxy Training Network offers accessible workflows for metagenomic analysis that include contamination assessment steps, and these workflows demonstrate how quality control integrates into the broader analytical pipeline.

The fourth challenge is that metagenomic data quality affects every downstream analysis. Taxonomic profiling, functional annotation, genome assembly, and comparative analysis all depend on the quality of the input reads. Errors introduced during sequencing or artifacts introduced during library preparation propagate through the analysis and can produce false biological conclusions. A study of respiratory samples using metagenomic next-generation sequencing demonstrated that phage community characterization required careful host sequence removal and quality filtering before alignment against viral reference databases, illustrating how quality decisions shape the interpretable output 12.

At a Glance: Core Quality Control Metrics and Decision Criteria

The table below summarizes the primary quality control metrics that should be assessed for every shotgun metagenomic sequencing dataset, the typical information each metric provides, and the general decision criteria for proceeding with analysis.

MetricWhat It AssessesTypical Decision CriteriaMetagenomic-Specific Consideration
Per-base quality scoresSequencing accuracy across read positionsMedian Q30 or higher for the read length, inspect for quality drops at read endsQuality distribution may vary by GC content in mixed communities, check quality across all reads, beyond the mean
Read duplication rateLibrary complexity and PCR amplification artifactsLow duplication for metagenomic libraries, high duplication suggests over-amplificationDuplication inflates apparent abundance of dominant taxa and can obscure rare taxa
Adapter contaminationPresence of sequencing adapter sequences in readsMinimal adapter content after trimming, high adapter content indicates insert size problemsShort inserts in metagenomic libraries produce adapter reads that must be removed before analysis
GC content distributionLibrary composition and potential biasDistribution should reflect the expected community composition, extreme deviation suggests biasDifferent taxa have different GC contents, a narrow GC distribution may indicate contamination or amplification bias
Insert size distributionFragment length of the sequencing libraryConsistent with library preparation protocol, inspect for anomaliesMetagenomic libraries often have variable insert sizes, extreme outliers may indicate chimeric fragments
Host DNA contentProportion of reads originating from the host organismVariable by sample type, must be quantified and either removed or accounted forHigh host content reduces effective microbial sequencing depth and may require additional sequencing
Negative control assessmentContamination introduced during processingNegative controls should show minimal microbial content and no overlap with sample taxaBatch-level negative controls are essential for distinguishing true signal from reagent or environmental contamination
Sequencing depthTotal number of reads and estimated genome coverageSufficient for the biological question, rare taxa require greater depthDepth requirements depend on community complexity and the target of interest

Per-Base Quality Scores and Their Interpretation in Metagenomic Contexts

Per-base quality scores are the most fundamental quality metric produced by sequencing instruments. These scores, typically reported as Phred scores, represent the probability that a base call is incorrect. A Phred score of Q30 corresponds to an error probability of one in one thousand, while Q20 corresponds to one in one hundred. The EMBL-EBI Training resources provide foundational instruction on how quality scores are calculated and how they should be interpreted in sequencing analysis.

For shotgun metagenomic data, per-base quality assessment requires attention to several metagenomic-specific factors. The first factor is that quality scores are typically lower at the ends of reads, and this quality drop can be more pronounced in metagenomic libraries that contain fragments of varying length and composition. Trimming low-quality bases from read ends is a standard preprocessing step, but the trimming threshold must balance error removal against the loss of useful sequence. The SOAPnuke tool implements a quality control and preprocessing workflow that integrates trimming decisions with other filtering steps, and its design illustrates how quality assessment and read processing are linked in practice.

The second factor is that average quality scores can mask important variation. A dataset with a high mean quality score might still contain a substantial fraction of reads with poor quality in specific regions. For metagenomic analysis, where reads are assigned to taxa based on sequence similarity, even a small number of erroneous bases can cause misclassification. The practical approach is to examine the distribution of quality scores across all reads and across read positions, beyond the summary statistics.

The third factor is that quality score interpretation depends on the downstream analysis. Taxonomic classification tools that use exact matching or near-exact matching are more sensitive to base errors than tools that use probabilistic alignment. Functional annotation, which relies on translated protein sequences, can tolerate some nucleotide errors because the genetic code is degenerate. The analyst should understand the error tolerance of the specific tools being used and set quality thresholds accordingly.

A practical assessment workflow for per-base quality includes the following steps. First, generate quality summary statistics for the raw data using a quality assessment tool. Second, examine the per-base quality plot to identify any systematic quality drops or anomalies. Third, assess the fraction of reads that meet the quality threshold for the intended analysis. Fourth, apply trimming or filtering as needed and regenerate the quality summary to confirm improvement. Fifth, document the quality metrics and processing decisions in the analysis record.

Read Duplication Rate and Library Complexity

The read duplication rate measures the fraction of sequencing reads that are identical or nearly identical to other reads in the dataset. In an ideal sequencing library, each DNA fragment is sequenced once, and the duplication rate reflects the complexity of the library. High duplication rates indicate that the library was over-amplified during PCR or that the sequencing depth exceeded the library complexity.

For shotgun metagenomic data, duplication assessment requires careful interpretation. Unlike whole-genome sequencing of a single organism, where duplication can be assessed against a known genome, metagenomic libraries contain fragments from many organisms at varying abundances. Highly abundant taxa will naturally produce more identical reads than rare taxa, and this biological duplication is not an artifact. The challenge is distinguishing biological duplication from PCR duplication.

The practical implication is that duplication metrics for metagenomic data should be interpreted in the context of community composition. A dataset with an overall duplication rate of twenty percent might be perfectly acceptable if the community is dominated by a few abundant taxa, while the same duplication rate in a highly diverse community might indicate a problem. The nf-core Documentation describes community standards for sequencing pipeline configuration, and many metagenomic pipelines include duplication assessment as a standard quality module.

High duplication rates have several consequences for metagenomic analysis. Duplicated reads inflate the apparent abundance of the taxa they represent, which distorts relative abundance estimates. Duplication also reduces the effective sequencing depth by wasting sequencing capacity on redundant reads. For rare taxa, duplication can create false signals of presence when the duplicated reads are actually artifacts of amplification.

The assessment of duplication should also consider whether the library preparation used PCR amplification. PCR-free library preparation methods produce lower duplication rates but require more input DNA. PCR-based methods are often necessary for low-biomass samples, such as clinical specimens or environmental samples with limited microbial content, but they introduce amplification bias and duplication. The Bamdam toolkit for ancient metagenomics computes authentication metrics including k-mer duplicity, demonstrating that duplication assessment is relevant even in challenging metagenomic applications where DNA is damaged and limited.

Adapter Contamination and Read Trimming Decisions

Adapter contamination occurs when sequencing reads contain sequences derived from the adapter oligonucleotides used in library preparation instead of from the sample DNA. This happens when the DNA fragment inserted into the sequencing library is shorter than the read length, causing the sequencer to read through the insert into the adapter sequence.

For shotgun metagenomic libraries, adapter contamination is a common issue because metagenomic DNA is often fragmented to short lengths during library preparation. The fragment size distribution depends on the DNA extraction method and the fragmentation protocol, and samples with degraded DNA, such as clinical specimens or environmental samples, tend to produce shorter fragments. The SOAPnuke workflow includes adapter removal as a core preprocessing function, and its design reflects the recognition that adapter contamination must be addressed before any downstream analysis.

The consequences of unremoved adapter contamination are significant. Adapter sequences do not match any biological reference, so they create alignment failures and reduce the fraction of usable reads. In taxonomic classification, adapter sequences may match adapter sequences present in reference databases, causing false positive assignments. In assembly, adapter sequences create chimeric contigs and reduce assembly quality.

The decision of when to trim adapters requires balancing several factors. Aggressive adapter trimming removes more sequence but may also remove legitimate biological sequence that happens to resemble adapter sequence. Conservative trimming preserves more sequence but may leave adapter contamination in the dataset. The standard approach is to detect adapter sequences using alignment-based methods and trim only the adapter portion, leaving the biological insert intact.

Adapter contamination assessment should be performed both before and after trimming. The pre-trimming assessment quantifies the extent of the problem and informs the trimming strategy. The post-trimming assessment confirms that the trimming was effective and that no adapter sequence remains. This before-and-after assessment is a core component of the quality control workflow described in the Galaxy Training Network metagenomic tutorials.

GC Content Distribution and Its Implications for Metagenomic Analysis

The GC content distribution describes the proportion of guanine and cytosine bases across the reads in a dataset. For a single organism, the GC content distribution should be narrow and centered on the organism's genomic GC content. For a metagenomic sample, the GC content distribution reflects the mixture of organisms present, each with its own characteristic GC content.

Bacterial genomes vary widely in GC content, from approximately twenty-five percent in some species to over seventy-five percent in others. A healthy metagenomic sample from a diverse community should therefore show a broad GC content distribution. A narrow distribution may indicate that the sample is dominated by a single taxon, that the DNA extraction method introduced bias against certain GC contents, or that the sample is contaminated with a single organism.

The deep culturing study of poultry fecal microbiota provides a relevant example of how different methods can yield different views of community composition. That study found that metagenomic sequencing detected DNA from 545 bacterial species while deep culturing produced isolates for 128 species, and some bacterial families were detected only via culturing. This discrepancy illustrates that metagenomic analysis does not provide a complete taxonomic census, and GC content bias is one factor that can contribute to such gaps.

GC content bias in metagenomic sequencing can arise from several sources. DNA extraction methods differ in their efficiency for different cell types, and some methods lyse certain bacteria more effectively than others. PCR amplification during library preparation can introduce bias because GC-rich and GC-poor templates amplify with different efficiencies. Sequencing platforms themselves can have systematic biases related to GC content.

The practical assessment of GC content involves comparing the observed distribution to the expected distribution for the sample type. For well-characterized sample types, such as human gut or soil, reference data on typical community composition can inform this comparison. For novel sample types, the assessment is more challenging, and the analyst must rely on consistency across replicates and comparison to negative controls.

Insert Size Distribution and Library Quality

The insert size distribution describes the length of the DNA fragments that were sequenced. This distribution is determined by the library preparation protocol, specifically the fragmentation method and the size selection steps. The insert size distribution affects several aspects of metagenomic analysis, including read overlap, assembly contiguity, and the ability to detect structural variation.

For shotgun metagenomic libraries, the insert size distribution is typically broader than for whole-genome sequencing libraries because metagenomic DNA is often fragmented to a range of sizes. The distribution should be consistent with the library preparation protocol and should not contain extreme outliers. A large fraction of very short inserts increases the likelihood of adapter contamination, while very long inserts may indicate incomplete fragmentation.

The insert size distribution also affects the assembly process. Metagenomic assembly relies on overlapping reads to build contigs, and the insert size determines the expected overlap between paired reads. If the insert size distribution is inconsistent with the assembly parameters, the assembler may fail to connect reads that should overlap or may incorrectly connect reads that should not overlap.

The phage characterization study in respiratory samples provides an example of how insert size and library quality affect downstream analysis. That study performed viral sequence assembly using multiple tools, and the quality of the assembly depended on the quality of the input reads. Libraries with appropriate insert sizes and minimal adapter contamination produced better assemblies than libraries with suboptimal characteristics.

Assessment of insert size distribution should include examination of the distribution plot, calculation of summary statistics such as mean and median insert size, and comparison to the expected values for the library preparation protocol. Anomalies in the insert size distribution should trigger investigation of the library preparation process and may warrant library re-preparation if the anomalies are severe.

Host DNA Content and Its Management in Metagenomic Analysis

Host DNA contamination is a pervasive challenge in shotgun metagenomic sequencing of host-associated samples. Clinical specimens, animal tissue samples, and plant-associated samples all contain host DNA that can constitute the majority of the sequenced reads. The host DNA dilutes the microbial signal and reduces the effective sequencing depth for microbial analysis.

The proportion of host DNA varies widely by sample type. Blood samples may contain very little microbial DNA relative to host DNA, while fecal samples typically contain a higher proportion of microbial DNA. The respiratory sequencing case series illustrates the interpretive challenges that arise when sequencing yields are low and host or contaminating sequences dominate the dataset. That study described pediatric respiratory samples where retained pathogen-associated contigs were sparse and mapped read support was low, making confident interpretation difficult.

Management of host DNA can occur at several stages. During sample processing, methods such as differential centrifugation, filtration, or enzymatic digestion can enrich for microbial cells before DNA extraction. During library preparation, methods that selectively deplete host DNA can reduce the host fraction. During bioinformatics analysis, computational subtraction of host reads can remove host sequences from the dataset.

The NCBI Data Resources provide reference genomes for many host organisms, which are used for computational host read removal. The process involves aligning all reads to the host reference genome and removing reads that match. The effectiveness of this approach depends on the quality of the host reference genome and the sequencing read length.

The decision of how much host DNA is acceptable depends on the biological question and the sequencing budget. For a study focused on abundant microbial taxa, a host content of ninety percent might be acceptable because the remaining ten percent still provides sufficient microbial reads. For a study focused on rare pathogens or low-abundance community members, even a host content of fifty percent might be problematic because the effective microbial depth is reduced by half.

The pig microbiome standardization review emphasized that variability in DNA extraction methods and sequencing approaches affects reproducibility and data comparability across studies. Host DNA content is one of the variables that differs across extraction methods, and the review proposed standardized metadata reporting to improve cross-study comparisons. Researchers should document host DNA content and the methods used to manage it as part of their metadata.

Negative Controls and Contamination Assessment

Negative controls are essential for interpreting shotgun metagenomic sequencing results, particularly for low-biomass samples where contamination can constitute a substantial fraction of the sequenced reads. A negative control is a sample that undergoes the same processing as the biological samples but contains no biological material, such as extraction blanks or reagent blanks. Sequencing the negative control reveals the contaminants introduced during processing.

The respiratory sequencing case series highlighted the critical role of negative controls in clinical metagenomic interpretation. The study noted that no case had a specimen-matched negative control, and the absence of such controls prevented confident distinction between active infection, carriage or colonization, transient detection, coinfection of uncertain relevance, and contamination. The study emphasized the use of batch-level negative controls as a standard practice.

The assessment of negative controls involves comparing the taxonomic composition of the negative control to the taxonomic composition of the biological samples. Taxa that appear in both the negative control and the biological samples are potential contaminants, and their abundance in the biological samples must be interpreted cautiously. Taxa that appear only in the biological samples are more likely to represent true biological signal.

The interpretation of negative control results requires attention to the relative abundance of contaminants. A contaminant that constitutes a small fraction of the negative control reads might constitute a large fraction of a low-biomass sample. The Bamdam toolkit for ancient metagenomics computes authentication metrics that help distinguish authentic ancient DNA from contamination, and similar principles apply to modern metagenomic samples.

The practical approach to negative controls includes the following steps. First, include negative controls in every sequencing batch. Second, sequence the negative controls to the same depth as the biological samples. Third, compare the taxonomic composition of negative controls to biological samples. Fourth, flag taxa that appear in both for cautious interpretation. Fifth, document the negative control results in the analysis record.

Sequencing Depth and Coverage Considerations

Sequencing depth refers to the total number of reads generated for a sample, while coverage refers to the proportion of the target community or genome that is represented by those reads. For shotgun metagenomic sequencing, the relationship between depth and coverage is complex because the reads are distributed across many genomes at varying abundances.

The sequencing depth required for a metagenomic experiment depends on the biological question. Taxonomic profiling at the phylum level requires less depth than species-level profiling. Detection of rare taxa requires more depth than characterization of abundant taxa. Functional annotation requires more depth than taxonomic classification because functional genes must be covered by sufficient reads to be identified and annotated.

The contributional diversity study demonstrated that metagenomic datasets with taxa linked to the functions they encode are becoming more common as sequencing approaches improve. These datasets require sufficient depth to jointly characterize taxa and functions, and the complexity of these analyses increases the depth requirements.

The assessment of sequencing depth adequacy involves several considerations. The total number of reads provides a basic measure of depth, but the effective depth for microbial analysis depends on the host DNA content and the fraction of reads that pass quality filtering. The number of reads that map to microbial reference sequences provides a more meaningful measure of effective depth. The estimated genome coverage for abundant taxa provides an indication of whether the depth is sufficient for assembly or detailed analysis.

The systematic benchmark of integrative strategies for microbiome and metabolome data noted that no standard currently exists for jointly integrating microbiome and metabolome datasets within statistical models. This lack of standards extends to sequencing depth requirements, which vary by study design and biological question. Researchers should justify their sequencing depth choices based on the expected community complexity and the target of interest.

Quality Control Workflow for Shotgun Metagenomic Data

A systematic quality control workflow ensures that all relevant metrics are assessed consistently across samples and that quality decisions are documented. The workflow should be established before sequencing begins and applied uniformly to all samples in a study.

The first stage of the workflow is raw data assessment. This stage examines the sequencing output files directly, before any processing. Metrics assessed at this stage include per-base quality scores, read length distribution, adapter contamination, and GC content. The SOAPnuke tool implements a quality control and preprocessing workflow that integrates these assessments, and its design demonstrates how raw data assessment feeds into processing decisions.

The second stage is read processing. This stage applies trimming, filtering, and other corrections to the raw reads. The specific processing steps depend on the raw data assessment results and the requirements of the downstream analysis. Common steps include adapter trimming, quality trimming, length filtering, and complexity filtering.

The third stage is post-processing assessment. This stage re-examines the processed reads to confirm that the processing was effective and that no new artifacts were introduced. The post-processing assessment should include the same metrics as the raw data assessment, allowing direct comparison of before and after processing.

The fourth stage is host read removal and contamination assessment. This stage removes host sequences and evaluates the remaining reads for potential contamination. The NCBI Data Resources provide reference databases for host read removal, and the negative control comparison provides the basis for contamination assessment.

The fifth stage is downstream analysis readiness assessment. This stage evaluates whether the processed reads are suitable for the intended downstream analysis. Metrics assessed at this stage include the fraction of reads that map to reference sequences, the estimated community composition, and the coverage of target taxa or functions.

The nf-core Documentation describes community standards for reproducible workflow configuration, and many metagenomic pipelines implement quality control as an integrated module. The use of standardized pipelines facilitates consistent quality control across samples and studies, which is essential for reproducibility.

Common Failure Patterns in Shotgun Metagenomic Quality Control

Several failure patterns recur in shotgun metagenomic sequencing projects. Recognizing these patterns early can prevent wasted sequencing capacity and avoid downstream analysis errors.

The first failure pattern is low library complexity leading to high duplication rates. This pattern occurs when the input DNA quantity is insufficient or when the library preparation over-amplifies the sample. The symptom is a high fraction of duplicated reads, which reduces the effective sequencing depth and distorts abundance estimates. The remedy is to prepare libraries with adequate input DNA and to minimize PCR cycles.

The second failure pattern is excessive adapter contamination. This pattern occurs when the DNA fragments are shorter than the read length, causing the sequencer to read into the adapter sequence. The symptom is a high fraction of reads containing adapter sequence, particularly at the read ends. The remedy is to optimize the fragmentation and size selection steps to produce longer inserts or to use shorter read lengths.

The third failure pattern is high host DNA content that overwhelms the microbial signal. This pattern occurs in host-associated samples where the host genome is large relative to the microbial genomes. The symptom is a low fraction of reads that map to microbial reference sequences. The remedy is to enrich for microbial cells during sample processing or to increase sequencing depth to compensate for the host dilution.

The fourth failure pattern is contamination that mimics biological signal. This pattern occurs when contaminants from reagents, laboratory equipment, or the environment are sequenced alongside the biological sample. The symptom is the presence of unexpected taxa that also appear in negative controls. The remedy is to identify the contamination source and to interpret affected taxa cautiously.

The fifth failure pattern is batch effects that confound biological comparisons. This pattern occurs when samples from different experimental groups are processed or sequenced in different batches, introducing systematic differences that are unrelated to biology. The symptom is clustering of samples by batch instead of by experimental group. The remedy is to randomize samples across batches and to include batch as a covariate in statistical analyses.

The pig microbiome standardization review documented how variability in DNA extraction methods, sequencing platforms, and bioinformatics pipelines affects reproducibility and data comparability. The review proposed a standardized metadata template to address these issues, and the same principles apply to quality control documentation.

Records and Documentation for Quality Control

Documentation of quality control decisions is essential for reproducible metagenomic analysis. The documentation should include the raw quality metrics, the processing steps applied, the post-processing quality metrics, and the rationale for any quality-based decisions.

The quality control record should include the following elements. First, the raw data quality metrics for each sample, including per-base quality, duplication rate, adapter content, GC content, and insert size. Second, the processing parameters used, including trimming thresholds, filtering criteria, and host read removal settings. Third, the post-processing quality metrics, allowing comparison to the raw metrics. Fourth, the negative control results and any contamination flags. Fifth, the final assessment of whether each sample meets the quality criteria for the intended analysis.

The Galaxy Training Network emphasizes reproducibility in bioinformatics analysis, and its tutorials demonstrate how to document analysis steps and parameters. The The Carpentries Lessons provide foundational training in data management and reproducible analysis practices that apply to quality control documentation.

The documentation should also include the versions of all software tools used for quality control and processing. Tool versions affect the results, and version documentation is essential for reproducing the analysis. The Bioconductor project provides versioned packages for genomic analysis, and its documentation practices demonstrate the importance of version tracking.

The quality control record should be stored with the sequencing data and analysis results, not in a separate location. This ensures that the quality information is available when the data are analyzed, shared, or published. Many journals now require quality control documentation as part of data availability statements.

Limitations of Quality Control Metrics

Quality control metrics provide essential information about sequencing data, but they have important limitations that researchers must understand. The metrics assess technical properties of the sequencing run, but they cannot directly assess biological accuracy. A dataset with excellent quality metrics can still produce incorrect biological conclusions if the sample was collected improperly, if the DNA extraction introduced bias, or if the reference databases are incomplete.

The deep culturing study of poultry fecal microbiota demonstrated that metagenomic sequencing does not provide a complete taxonomic census. The study found that some bacterial families were detected only via culturing and not by metagenomic sequencing, indicating that metagenomic analysis has blind spots. Quality control metrics cannot reveal these blind spots because they assess the sequencing process, not the completeness of the biological representation.

Quality control metrics also cannot distinguish between biologically meaningful variation and technical artifacts. A difference in community composition between two samples could reflect true biological differences or could reflect differences in DNA extraction efficiency, sequencing depth, or batch effects. The systematic benchmark of integrative strategies for microbiome and metabolome data noted that no standard currently exists for jointly integrating microbiome and metabolome datasets, and the same lack of standards applies to quality control thresholds.

The thresholds used for quality control decisions are often arbitrary and may not be appropriate for all sample types and research questions. A quality threshold that is appropriate for high-biomass samples may be too stringent for low-biomass samples, and a threshold that is appropriate for taxonomic profiling may be insufficient for functional annotation. Researchers should justify their quality thresholds based on their specific context.

The respiratory sequencing case series illustrated the interpretive challenges that arise when sequencing findings do not align with routine laboratory reports. The study found that low-support hits required cautious interpretation because limited sequencing evidence and the absence of orthogonal confirmation prevented confident distinction between active infection, carriage, colonization, transient detection, coinfection of uncertain relevance, and contamination. Quality control metrics cannot resolve these interpretive challenges.

Professional Escalation Criteria for Quality Control Issues

Some quality control issues require escalation to specialized expertise or additional investigation. The following criteria indicate when a quality control issue exceeds routine handling and requires professional consultation.

The first escalation criterion is the presence of unexpected pathogens or high-consequence organisms in clinical or agricultural samples. If a metagenomic analysis identifies an organism that could pose a health risk, the finding should be escalated to the appropriate clinical or veterinary authority for confirmation and action. The respiratory sequencing case series described the importance of orthogonal confirmation for clinically significant findings, and the absence of such confirmation should trigger escalation.

The second escalation criterion is the presence of contamination that cannot be traced to a known source. If negative controls show contamination that persists across batches or that cannot be eliminated by standard decontamination procedures, the issue should be escalated to laboratory management for investigation of reagents, equipment, and facilities.

The third escalation criterion is systematic quality failure across multiple samples or batches. If a sequencing run produces poor quality metrics for a substantial fraction of samples, the issue may indicate a problem with the sequencing instrument, the reagents, or the library preparation protocol. This situation warrants escalation to the sequencing facility or the instrument manufacturer.

The fourth escalation criterion is the need for orthogonal confirmation of sequencing findings. If a metagenomic finding has clinical, agricultural, or regulatory implications, the finding should be confirmed using an independent method, such as PCR, culture, or serology. The respiratory sequencing case series noted that no case in their series had documented orthogonal confirmation, and this gap limited the interpretability of the findings.

The fifth escalation criterion is the need for regulatory reporting. Some organisms detected by metagenomic sequencing may be subject to regulatory reporting requirements, such as notifiable diseases in livestock or reportable pathogens in clinical settings. Researchers should be aware of the reporting requirements for their sample types and jurisdiction.

Safety and Regulatory Context for Metagenomic Quality Control

Metagenomic sequencing of clinical, agricultural, or environmental samples may be subject to safety and regulatory requirements that affect quality control practices. Researchers should be aware of these requirements and should design their quality control workflows accordingly.

For clinical samples, metagenomic sequencing results may inform patient management decisions, and the quality of the sequencing data directly affects the reliability of those decisions. The respiratory sequencing case series described a reporting framework for low-yield sequencing results, emphasizing the need for cautious interpretation when analytical support is sparse. Clinical laboratories should have protocols for confirming metagenomic findings using orthogonal methods before acting on them.

For agricultural samples, metagenomic sequencing may be used to assess animal health, detect pathogens, or evaluate the effectiveness of interventions. The deep culturing study of poultry fecal microbiota noted that understanding the structure and function of the poultry microbiota could improve health and productivity, including feed conversion ratio, greenhouse gas emissions, and competitive exclusion against pathogens. Quality control is essential for ensuring that these assessments are reliable.

For environmental samples, metagenomic sequencing may be used for monitoring purposes, and the results may inform regulatory decisions. The quality of the sequencing data affects the confidence that can be placed in monitoring conclusions, and inadequate quality control could lead to false positives or false negatives with regulatory consequences.

The EMBL-EBI Training resources provide instruction on responsible data management and analysis practices, and the The Carpentries Lessons provide foundational training in data handling that supports responsible research conduct. Researchers should ensure that their quality control practices align with the ethical and regulatory requirements of their field.

Frequently Asked Questions

What is the minimum sequencing depth for shotgun metagenomic analysis?

The minimum sequencing depth depends on the biological question and the complexity of the microbial community. Taxonomic profiling of abundant taxa can be performed with relatively low depth, while detection of rare taxa, functional annotation, and genome assembly require substantially greater depth. Researchers should estimate the expected community complexity and the abundance of the target organisms to determine the required depth. For low-biomass samples with high host DNA content, the effective microbial depth is reduced, and additional sequencing may be required to compensate.

How do I distinguish biological duplication from PCR duplication in metagenomic data?

Biological duplication occurs when multiple reads originate from the same DNA fragment in the original sample, which is expected for abundant taxa in a mixed community. PCR duplication occurs when the same fragment is amplified multiple times during library preparation, creating identical reads that do not reflect true abundance. Distinguishing the two requires examining the duplication pattern across the dataset. PCR duplication tends to produce clusters of identical reads with the same start and end positions, while biological duplication produces reads that overlap but have different start positions. Tools that assess duplication by comparing read start positions can help distinguish the two types.

What should I do if my negative control shows contamination?

If a negative control shows contamination, the first step is to identify the contaminating taxa and assess whether they also appear in the biological samples. Taxa that appear in both the negative control and the biological samples are potential contaminants, and their abundance in the biological samples should be interpreted cautiously. The contamination source should be investigated, including reagents, laboratory equipment, and the environment. If the contamination is consistent across batches, the source may be a reagent or a laboratory practice that needs to be changed. The contamination assessment should be documented in the quality control record.

How much host DNA is acceptable in a shotgun metagenomic dataset?

The acceptable host DNA content depends on the biological question and the sequencing budget. For studies focused on abundant microbial taxa, a host content of up to ninety percent may be acceptable because the remaining reads still provide sufficient microbial coverage. For studies focused on rare taxa or functional annotation, lower host content is required because the effective microbial depth is reduced by the host reads. Researchers should estimate the required microbial depth and calculate the total sequencing needed based on the expected host content. Host read removal during bioinformatics analysis can reduce the impact of host DNA, but it cannot recover the sequencing capacity that was consumed by host reads.

What quality metrics should I report in a metagenomic publication?

Publications should report the quality metrics that support the reliability of the findings. Essential metrics include the total number of reads, the fraction of reads passing quality filtering, the per-base quality scores, the duplication rate, the adapter contamination level, the host DNA content, and the negative control results. The versions of all software tools used for quality control and processing should also be reported. The pig microbiome standardization review proposed a standardized metadata template that includes DNA extraction protocols and bioinformatic workflows, and similar standards are emerging for quality control reporting.

Can I use quality control metrics to compare samples across different sequencing runs?

Quality control metrics can be used to compare samples across sequencing runs, but the comparison must account for technical differences between runs. Sequencing platforms, reagent lots, and instrument calibration can all affect quality metrics, and these effects can confound biological comparisons. The use of standardized control samples, such as mock communities or reference standards, can help normalize across runs. Batch effects should be assessed and, if present, included as covariates in statistical analyses. The systematic benchmark of integrative strategies for microbiome and metabolome data noted that no standard currently exists for many aspects of metagenomic analysis, and cross-run comparison remains an area of active development.

What is the role of reference databases in metagenomic quality control?

Reference databases are used for several quality control purposes, including host read removal, taxonomic classification, and contamination assessment. The NCBI Data Resources provide reference genomes and sequence databases that support these analyses. The quality of the reference databases affects the reliability of the quality control assessments. Incomplete or inaccurate reference databases can lead to false negative classifications, where true organisms are not detected, or false positive classifications, where reads are assigned to incorrect taxa. Researchers should use the most current reference databases and should document the database versions in their quality control records.

When should I reject a sequencing run and request resequencing?

A sequencing run should be rejected when the quality issues cannot be corrected by bioinformatics processing and when the issues compromise the ability to answer the biological question. Specific criteria include extremely low per-base quality that persists after trimming, very high duplication rates that indicate library failure, excessive adapter contamination that indicates insert size problems, and contamination that cannot be distinguished from biological signal. The decision to reject a run should be documented, and the reasons for rejection should be recorded. In clinical or agricultural contexts, the decision to reject may have implications for sample collection and may require additional sample acquisition.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.