Identifying Carbohydrate-Active Enzymes (CAZymes) in Metagenomes: A Comprehensive Guide to dbCAN and Other Tools
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- CAZyme identification in metagenomes relies on a multi-pronged approach, integrating HMMER searches against profile databases for distant homology, DIAMOND searches against pre-annotated sequences for close homologs and substrate inference, and Hotpep searches for conserved peptide motifs, particularly useful for fragmentary sequences.
- The dbCAN2 meta-server employs a majority-voting rule, requiring at least two of these three search strategies to identify a gene as a high-confidence CAZyme, significantly improving annotation accuracy and reducing false positives compared to single-method approaches.
- CAZyme gene cluster (CGC) prediction, including polysaccharide utilization loci (PULs) in Bacteroidetes, provides functional context by identifying physically linked CAZyme genes and associated proteins that cooperate in degrading specific carbohydrates, offering stronger evidence of metabolic capability than individual enzyme predictions.
- Substrate prediction, while biologically informative, carries substantial uncertainty and is approached via dbCAN-sub subfamily HMMs or homology searches against experimentally characterized PULs in dbCAN-PUL; conflicting predictions between methods are common and require careful reporting of both hypotheses.
- Reproducibility in multi-sample CAZyme annotation projects necessitates meticulous documentation of software versions, database versions, parameter settings, and the chosen assembly strategy (individual, co-assembly, or assembly-free), with normalization strategies like CAZyme proportion or relative abundance being critical for cross-sample comparability.
- Computational limitations, such as the significant runtime for assembly and annotation pipelines, necessitate strategic decisions regarding resource allocation (sequential vs. parallel processing) and the selection of assembly strategies tailored to project scale and research questions.
Researchers studying microbiome carbohydrate metabolism face a specific analytical problem: metagenomic datasets contain thousands of predicted genes, and distinguishing genuine carbohydrate-active enzymes (CAZymes) from other protein families requires specialized databases and search tools. The dbCAN suite, maintained by the Yin lab at the University of Nebraska-Lincoln, provides the most widely used framework for this task, with both web-based and standalone implementations. This article explains how to select and apply these tools within a reproducible metagenomics workflow, how to interpret their outputs, and where their limitations require additional validation.
The Role of CAZymes in Metagenomic Research
CAZymes are enzymes that synthesize, modify, or break down complex carbohydrates. They include glycoside hydrolases (GHs), glycosyltransferases (GTs), polysaccharide lyases (PLs), carbohydrate esterases (CEs), and auxiliary activities (AAs), along with carbohydrate-binding modules (CBMs) that assist in substrate recognition. These enzymes are produced by diverse organisms and are central to complex carbohydrate metabolism in environments such as animal guts, agricultural soils, forest floors, and ocean sediments [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].
For microbiome researchers, CAZyme annotation answers questions about how microbial communities utilize dietary fiber, plant biomass, or host glycans. In animal production systems, for example, understanding which CAZymes are present in rumen or hindgut metagenomes can inform feed formulation and probiotic development. In environmental microbiology, CAZyme profiles reveal how soil or marine communities process organic carbon. The analytical challenge is consistent across these applications: metagenomic assemblies produce tens of thousands of gene predictions, and only a fraction encode CAZymes. Accurate identification requires searching against curated hidden Markov model (HMM) profiles, sequence databases, and conserved peptide motifs, then reconciling results across methods [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
Core Principles of CAZyme Annotation
CAZyme annotation rests on three complementary search strategies, each with distinct strengths and failure modes. The dbCAN2 meta server integrates all three approaches and requires that a candidate gene be identified by at least two of them to be considered a high-confidence CAZyme [<a href="#ref-3">3</a>].
HMMER Search Against Profile Databases
Profile hidden Markov models capture position-specific conservation across a protein family. The dbCAN HMM database contains profiles for CAZyme families curated from the CAZy database. HMMER searches are sensitive to distant homologs because they model conserved positions and allowed substitutions across the full alignment. This approach works well for identifying family membership even when sequence identity to known CAZymes is low. The primary limitation is that HMM profiles describe families, not substrates, so a hit to a GH family does not by itself reveal which glycan the enzyme acts upon [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].
DIAMOND Search Against Pre-Annotated Sequences
DIAMOND performs fast protein sequence alignment against a database of CAZyme sequences that have been pre-annotated from the CAZy database. This approach is computationally efficient and works well for close homologs. A strong DIAMOND hit to a characterized CAZyme provides both family assignment and, in many cases, a putative substrate based on the known biochemistry of the matched sequence. The limitation is sensitivity: highly divergent CAZymes may not produce significant DIAMOND hits even when they are true family members [<a href="#ref-3">3</a>].
Hotpep Search Against Conserved Peptides
Hotpep searches for short conserved peptide motifs that are diagnostic of CAZyme families. This method is independent of both HMM profiles and full-length sequence alignment, providing a third line of evidence. Hotpep is particularly useful for fragmentary sequences from metagenomic assemblies, where full-length alignments may fail due to incomplete gene models. The tradeoff is that short peptide matches can produce false positives when conserved motifs occur in non-CAZyme proteins by chance [<a href="#ref-3">3</a>].
Majority Voting for Confidence
The dbCAN2 meta server combines these three outputs and removes CAZymes identified by only one tool. This majority-voting approach significantly improves annotation accuracy compared to any single method. For practical purposes, a gene predicted by two or three tools is considered a high-confidence CAZyme, while a gene predicted by only one tool should be treated as a candidate requiring manual inspection or additional evidence [<a href="#ref-3">3</a>].
At a Glance: Tool Selection for CAZyme Annotation
| Analysis Need | Recommended Tool | Input Type | Output Type | Key Limitation |
|---|---|---|---|---|
| Single genome or small metagenome, web-based analysis | dbCAN2 meta server | Protein or nucleotide FASTA | CAZyme family assignments, CGC predictions | Upload size limits, no local data control |
| Large metagenome, local processing, data privacy required | run_dbcan standalone package | Short reads or assembled contigs | CAZyme predictions, CGC predictions, substrate predictions | Requires Linux command-line familiarity |
| Substrate prediction at subfamily level | dbCAN-sub (within dbCAN3) | Protein FASTA | Subfamily-level assignments with substrate annotations | Subfamily coverage limited to characterized families |
| Substrate prediction at gene cluster level | dbCAN-PUL homology search | Protein FASTA or CGC sequences | Matches to experimentally characterized PULs | Limited to PULs with known substrates |
| Comparison across multiple samples or environments | dbCAN-seq database | Pre-computed MAG data | CAZyme and CGC profiles from 9,421 MAGs | Reference data limited to four ecological environments |
Practical Workflow for Metagenomic CAZyme Annotation
The complete workflow from raw sequencing data to publication-quality CAZyme abundance plots requires several stages. The protocol described in the dbCAN computational guide assumes familiarity with the Linux command line and the ability to run Python scripts, but does not require programming experience [<a href="#ref-4">4</a>].
Step 1: Data Preparation and Quality Control
Raw metagenomic short reads must be processed to remove adapters, low-quality bases, and host contamination before assembly. This step is essential because assembly quality directly affects gene prediction accuracy, which in turn affects CAZyme identification. The Galaxy Training Network provides accessible tutorials for quality control and preprocessing that can be adapted to metagenomic datasets [<a href="#ref-5">5</a>]. For researchers new to command-line work, The Carpentries lessons offer foundational training in shell scripting and data management that supports reproducible analysis [<a href="#ref-6">6</a>].
Step 2: Metagenome Assembly
Three assembly strategies are available, each with tradeoffs. Individual sample assembly preserves sample-specific variation but produces smaller assemblies with more fragmentation. Co-assembly of multiple samples increases assembly continuity and gene discovery but can obscure sample-specific signals. Assembly-free approaches skip assembly entirely and predict genes directly from reads, which is faster but produces fragmentary gene models. The choice of assembly strategy affects downstream CAZyme counts and should be recorded in the analysis metadata [<a href="#ref-4">4</a>].
Step 3: Gene Prediction
Assembled contigs must be processed with gene prediction software to identify open reading frames. The quality of gene prediction depends on assembly completeness. Fragmented assemblies produce partial gene models that may lack the conserved domains required for HMMER or Hotpep detection. For metagenome-assembled genomes (MAGs), gene prediction is typically performed on each MAG separately, and the resulting protein sequences are used as input for CAZyme annotation [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].
Step 4: CAZyme Prediction with run_dbcan
The run_dbcan standalone package accepts either nucleotide sequences or protein sequences as input. For nucleotide input, the software performs gene prediction internally before CAZyme annotation. The software runs HMMER, DIAMOND, and Hotpep searches and combines results using the majority-voting rule. Output files include per-gene CAZyme family assignments, alignment coordinates, and summary tables [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
For a typical individual sample assembly, the complete run_dbcan workflow takes approximately 33 hours on a Linux computer with 40 CPUs. Co-assembly and assembly-free routes are faster. This computational cost should be factored into project planning, particularly for laboratories with limited computing resources [<a href="#ref-4">4</a>].
Step 5: CAZyme Gene Cluster Prediction
CAZyme gene clusters (CGCs) are physically linked groups of CAZyme genes and associated proteins that work together to degrade specific carbohydrates. In Bacteroidetes, these clusters are called polysaccharide utilization loci (PULs) and have been extensively characterized experimentally. The dbCAN2 and dbCAN3 servers include CGC-Finder, a tool that identifies genomic regions enriched in CAZyme genes. CGC prediction is valuable because it provides functional context: a cluster of CAZymes targeting the same substrate is stronger evidence of carbohydrate utilization capability than individual enzyme predictions [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].
Step 6: Glycan Substrate Prediction
Substrate prediction is the most biologically informative but also the most uncertain step in CAZyme annotation. The dbCAN3 update introduced two approaches for substrate prediction. The first uses dbCAN-sub, a profile HMM database for substrate prediction at the CAZyme subfamily level. The second searches query CGCs against dbCAN-PUL, a database of experimentally characterized PULs with known glycan substrates. A third method applies majority voting across all CAZymes in a CGC with substrate predictions from dbCAN-sub [<a href="#ref-1">1</a>][<a href="#ref-7">7</a>].
The dbCAN-PUL database catalogs experimentally verified CAZyme-containing PULs from the literature, with metadata including substrate, experimental characterization method, and protein sequence information. Users can batch download PUL data by target substrate, species, genus, or characterization method, and can query their own sequences against PUL proteins using an integrated BLASTX service [<a href="#ref-7">7</a>].
Step 7: Data Visualization and Comparison
The final stage involves comparing CAZyme, CGC, and substrate abundance across samples. The dbCAN protocol includes commands for generating publication-quality plots showing the occurrence and abundance of CAZymes and glycan substrates across multiple samples. These visualizations support hypothesis generation about community carbohydrate metabolism and enable statistical comparisons between treatment groups [<a href="#ref-4">4</a>].
Options and Tradeoffs: Web Server vs. Standalone Software
The choice between the dbCAN web server and the run_dbcan standalone package depends on dataset size, data sensitivity, and computing infrastructure.
Web Server Advantages and Limitations
The dbCAN2 and dbCAN3 web servers provide an accessible interface for researchers who do not use the command line. Users upload protein or nucleotide sequences and receive CAZyme annotations, CGC predictions, and substrate predictions through a browser interface. The dbCAN3 server adds improved data browsing and visualization of substrate prediction results [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].
The web server has practical limitations. Upload size limits restrict analysis to smaller datasets. Data privacy is a concern for unpublished or proprietary sequences, since data is transmitted to an external server. For large metagenomic projects with hundreds of samples, the web server is impractical due to manual upload requirements and size constraints [<a href="#ref-4">4</a>].
Standalone Software Advantages and Limitations
The run_dbcan package addresses the scalability and privacy limitations of the web server. Running locally allows researchers to process large datasets without upload limits and keeps sequence data within institutional infrastructure. The software is actively maintained and supports the full workflow from short reads to visualization [<a href="#ref-4">4</a>].
The primary barrier is technical: run_dbcan requires a Linux environment, command-line proficiency, and the ability to install and configure dependencies. Researchers without these skills should either collaborate with a bioinformatics specialist or complete foundational training through resources such as The Carpentries lessons [<a href="#ref-6">6</a>] or the Galaxy Training Network [<a href="#ref-5">5</a>].
Reproducibility Considerations
Reproducibility requires documenting the software versions, database versions, and parameter settings used for each analysis. The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including version pinning and containerization, that can be applied to CAZyme annotation workflows [<a href="#ref-8">8</a>]. Bioconductor provides additional guidance on reproducible genomic analysis practices, including package version management and workflow documentation [<a href="#ref-9">9</a>].
Records and Measurements for CAZyme Annotation Projects
Maintaining detailed records is essential for both scientific validity and regulatory compliance in research settings. The following records should be maintained for every CAZyme annotation project.
Assembly and Quality Metrics
Record the assembly strategy (individual, co-assembly, or assembly-free), the assembler software and version, and assembly quality metrics such as N50, total assembled bases, and number of contigs. These metrics affect gene prediction completeness and therefore CAZyme detection sensitivity. For MAG-based analyses, record the binning approach and completeness/contamination estimates for each MAG [<a href="#ref-2">2</a>][<a href="#ref-4">4</a>].
Gene Prediction Records
Document the gene prediction software and version, the input sequences used, and the total number of predicted genes. The proportion of predicted genes that are annotated as CAZymes provides a useful quality check: values that are unexpectedly high or low relative to comparable datasets may indicate annotation errors or assembly problems.
CAZyme Annotation Parameters
Record the dbCAN database version, HMMER version and E-value thresholds, DIAMOND version and parameters, and Hotpep settings. These parameters directly affect the number and confidence of CAZyme predictions. Changes to any parameter can alter results, so version control is essential for reproducibility [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
Substrate Prediction Records
For substrate predictions, record whether the assignment came from dbCAN-sub subfamily analysis, dbCAN-PUL homology search, or CGC-level majority voting. The dbCAN-seq database update reported that two substrate prediction approaches agreed on only 4,183 of 41,447 CGCs with predicted substrates, indicating substantial method-dependent variation. Researchers should report which method was used and acknowledge the uncertainty associated with substrate assignments [<a href="#ref-2">2</a>].
Common Failure Patterns and Troubleshooting
Several recurring problems affect CAZyme annotation projects. Recognizing these patterns early prevents wasted computation and misinterpretation.
Low CAZyme Detection Rates
If a metagenome yields very few CAZyme predictions, the cause is often poor assembly quality or incomplete gene models. Fragmented assemblies produce partial proteins that lack the conserved domains required for HMMER or Hotpep detection. Check assembly statistics and consider co-assembly to improve contiguity. Alternatively, the assembly-free route may detect CAZymes from reads that are absent from assembled contigs [<a href="#ref-4">4</a>].
Disagreement Between Search Tools
The majority-voting rule in dbCAN2 removes genes identified by only one tool. When HMMER, DIAMOND, and Hotpep disagree substantially, the cause may be divergent CAZyme sequences that evade one or more search methods. Genes identified by only one tool should not be automatically discarded, they should be examined manually using domain prediction and phylogenetic analysis to determine whether they are true CAZymes [<a href="#ref-3">3</a>].
Substrate Prediction Conflicts
When dbCAN-sub and dbCAN-PUL predict different substrates for the same CGC, the conflict reflects the different evidence bases of the two methods. dbCAN-sub uses subfamily-level HMM profiles, while dbCAN-PUL uses homology to experimentally characterized PULs. The dbCAN-seq analysis found that the two approaches agreed on only a minority of CGCs, so conflicting predictions are common. Report both predictions and note the disagreement instead of arbitrarily selecting one [<a href="#ref-2">2</a>].
Database Version Drift
CAZyme family classifications change as new enzymes are characterized. Results obtained with an older dbCAN HMM database may differ from results obtained with a newer version. Always record the database version and avoid comparing CAZyme profiles generated with different database versions. For longitudinal studies, re-run all samples with the same database version to ensure comparability.
Limitations and Interpretation Boundaries
CAZyme annotation provides family-level assignments with high confidence, but substrate-level predictions carry substantial uncertainty. Researchers must understand these boundaries to avoid overinterpreting results.
Family Assignment vs. Substrate Specificity
A GH family can contain enzymes with different substrate specificities. Assigning a gene to GH5, for example, does not reveal whether the enzyme acts on cellulose, xylan, or another glycan. Subfamily-level classification through dbCAN-sub improves substrate resolution but is limited to families with sufficient characterized members to define subfamilies [<a href="#ref-1">1</a>]. For many environmental metagenomes, a large proportion of CAZymes will lack confident substrate predictions.
Experimental Validation Requirements
Computational CAZyme annotation is predictive, not confirmatory. Substrate assignments based on homology or HMM profiles should be treated as hypotheses requiring biochemical validation. The dbCAN-PUL database provides experimentally characterized PULs as reference points, but homology to a characterized PUL does not guarantee identical substrate specificity [<a href="#ref-7">7</a>]. Researchers reporting novel CAZyme functions should design biochemical assays to confirm predicted activities.
Metagenome Assembly Bias
Assembly-based CAZyme annotation is biased toward abundant community members. Rare taxa contribute fewer reads and produce more fragmented assemblies, reducing their representation in CAZyme profiles. The assembly-free route partially addresses this bias by predicting genes directly from reads, but produces fragmentary gene models that may be missed by HMMER or Hotpep [<a href="#ref-4">4</a>]. CAZyme abundance profiles should be interpreted as reflecting the dominant community members, not the full diversity of carbohydrate-active enzymes present.
MAG-Based Analyses
The dbCAN-seq database provides CAZyme and CGC data from 9,421 MAGs spanning human gut, human oral, cow rumen, and marine environments [<a href="#ref-2">2</a>]. These reference data are valuable for comparative analyses but are limited to four ecological environments. Researchers studying other environments should not assume that reference CAZyme profiles from these habitats are representative of their study systems.
Quality Controls and Validation Steps
Implementing systematic quality controls improves the reliability of CAZyme annotation results.
Positive and Negative Controls
Include characterized CAZyme sequences as positive controls in each run to verify that the annotation pipeline detects known enzymes. Include non-CAZyme proteins such as housekeeping enzymes as negative controls to check for false positives. These controls should be embedded in the input file and their annotations verified in the output.
Cross-Validation with Independent Methods
For high-priority genes, validate HMMER-based family assignments using independent domain prediction tools and phylogenetic analysis. A gene assigned to a GH family by HMMER should also produce a significant hit to the corresponding CAZy family when searched against the CAZy sequence database with DIAMOND. Disagreement between methods warrants manual inspection [<a href="#ref-3">3</a>].
Comparison with Reference Databases
For metagenomes from environments represented in dbCAN-seq, compare sample CAZyme profiles with reference profiles from the database. Substantial deviations may indicate technical issues or genuine biological differences. The dbCAN-seq statistics page organizes data by substrate and taxonomic phylum, supporting such comparisons [<a href="#ref-2">2</a>].
Replicate Analysis
For critical samples, run the annotation pipeline twice with identical parameters to confirm reproducibility. Differences between runs indicate software or database instability that must be resolved before reporting results.
Professional Escalation Criteria
Certain situations require consultation with a bioinformatics specialist or CAZyme annotation expert.
Persistent Tool Disagreement
If HMMER, DIAMOND, and Hotpep consistently disagree across many genes, the cause may be an unusual community composition or a technical issue with input data. Consult a specialist before proceeding with downstream analysis, since the majority-voting rule may be discarding genuine CAZymes [<a href="#ref-3">3</a>].
Unexpected Substrate Predictions
Substrate predictions that contradict known biology of the study system should be investigated before reporting. For example, a rumen metagenome predicted to contain high abundance of marine algal polysaccharide-degrading enzymes warrants verification, since such substrates are unlikely to be present in the rumen environment.
Large-Scale Comparative Projects
Projects comparing CAZyme profiles across many samples or environments benefit from expert input on experimental design, statistical analysis, and interpretation. The choice of assembly strategy, normalization method, and statistical model can substantially affect conclusions [<a href="#ref-4">4</a>].
Regulatory or Commercial Applications
CAZyme annotation results used for regulatory submissions, commercial product development, or clinical applications require additional validation beyond computational prediction. Consult with appropriate regulatory and scientific experts to determine the evidence standards required for the specific application.
Building a CAZyme Annotation Decision Framework for Multi-Sample Metagenomic Studies
Researchers moving from single-sample CAZyme annotation to multi-sample comparative projects face a distinct set of decisions that the basic dbCAN workflow does not address. These decisions include how to allocate computing resources across samples, how to normalize CAZyme counts for meaningful comparison, how to handle samples with vastly different sequencing depths, and how to structure the analysis so that results remain interpretable when sample numbers grow into the dozens or hundreds. This section provides a practical decision framework for multi-sample CAZyme annotation projects, with specific attention to resource allocation, normalization strategy, and quality control at scale.
The Multi-Sample Problem in CAZyme Annotation
Single-sample CAZyme annotation is straightforward: assemble, predict genes, run dbCAN, interpret results. Multi-sample projects introduce complications that change the analytical approach. The first complication is computational cost. The dbCAN protocol reports that a typical individual sample assembly takes approximately 33 hours on a Linux computer with 40 CPUs [<a href="#ref-4">4</a>]. A project with 50 samples would therefore require over 1,600 CPU-hours for the assembly route alone, before considering gene prediction, CAZyme annotation, and substrate prediction. This cost forces researchers to make deliberate choices about which assembly strategy to use and how to distribute computing resources.
The second complication is comparability. CAZyme counts from different samples are not directly comparable when samples have different sequencing depths, different assembly qualities, or different numbers of predicted genes. A sample with deeper sequencing will produce more contigs, more genes, and therefore more CAZyme predictions, even if the underlying microbial community is identical to a shallower-sequenced sample. Without normalization, researchers cannot distinguish biological differences in CAZyme content from technical differences in sequencing and assembly.
The third complication is interpretation. Multi-sample projects typically ask comparative questions: does CAZyme abundance differ between treatment groups, does CAZyme diversity change across time points, or do certain substrates dominate in specific environments? Answering these questions requires a consistent annotation pipeline applied uniformly across all samples, with documented parameters and version control [<a href="#ref-8">8</a>].
Decision Point 1: Selecting the Assembly Strategy for Multi-Sample Projects
The assembly strategy decision is the first and most consequential choice in a multi-sample CAZyme project. The dbCAN protocol describes three routes: individual sample assembly, co-assembly, and assembly-free analysis [<a href="#ref-4">4</a>]. Each route has distinct implications for multi-sample comparability.
Individual sample assembly produces one assembly per sample. This approach preserves sample-specific variation and allows each sample to be analyzed independently. The tradeoff is reduced contiguity, particularly for low-abundance community members, and higher total computational cost because each sample requires a separate assembly run. For multi-sample projects, individual assemblies have the advantage of straightforward sample-to-sample comparison: each sample has its own gene set, its own CAZyme calls, and its own abundance estimates. The disadvantage is that fragmented assemblies from low-coverage samples may miss CAZymes that a co-assembly would recover.
Co-assembly combines reads from multiple samples into a single assembly. This approach improves contiguity and gene discovery because shared community members accumulate more coverage across samples. The dbCAN protocol notes that co-assembly is faster than individual sample assembly [<a href="#ref-4">4</a>]. For multi-sample projects, co-assembly has a critical limitation: genes present in only one sample may be absent from the co-assembly if their coverage is too low. Additionally, mapping individual sample reads back to the co-assembly is required to estimate per-sample abundance, adding a downstream analysis step. Co-assembly is most appropriate when samples come from similar environments and share substantial community membership, such as replicate rumen samples from the same treatment group.
Assembly-free analysis predicts genes directly from reads without assembly. This route is fastest and detects CAZymes from reads that are not captured in assemblies [<a href="#ref-4">4</a>]. For multi-sample projects, assembly-free analysis provides the most direct per-sample comparability because each sample is processed independently with no cross-sample mixing. The tradeoff is fragmentary gene models that may be missed by HMMER or Hotpep searches, potentially underestimating CAZyme diversity.
The decision framework for assembly strategy depends on the research question. For projects asking whether CAZyme profiles differ between clearly distinct environments, individual sample assembly preserves the biological signal of each environment. For projects with many replicate samples from similar environments, co-assembly increases sensitivity and reduces computational cost. For projects where sequencing depth is limited or highly variable across samples, assembly-free analysis provides a consistent baseline that does not penalize low-coverage samples.
Decision Point 2: Allocating Computing Resources Across Samples
Multi-sample projects require a resource allocation plan before analysis begins. The 33-hour runtime for individual sample assembly on 40 CPUs [<a href="#ref-4">4</a>] provides a baseline for estimating total computational demand. Researchers should calculate the expected runtime for their sample count and available hardware, then decide whether to process samples sequentially or in parallel.
Sequential processing is simplest to manage and requires the least infrastructure. Each sample is processed through the full pipeline in turn, with results recorded before the next sample begins. This approach is appropriate for projects with fewer than 20 samples or when computing resources are limited to a single workstation. The disadvantage is total wall-clock time: 20 samples at 33 hours each requires 660 hours, or about 28 days of continuous processing.
Parallel processing distributes samples across multiple compute nodes or cores. A laboratory with access to a 200-core cluster could process six samples simultaneously at 40 CPUs each, completing 20 samples in approximately 110 hours. The nf-core documentation describes community standards for reproducible pipelines, including configuration for high-performance computing environments, that support parallel sample processing [<a href="#ref-8">8</a>]. Researchers using parallel processing must ensure that each sample run is independent and that results are collected systematically to avoid file mixing.
A third option is to use the web server for a subset of samples. The dbCAN2 and dbCAN3 web servers accept sequence uploads and return annotations through a browser interface [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>]. This approach is practical for pilot analyses or for verifying local results, but upload size limits and manual submission requirements make it impractical for large multi-sample projects [<a href="#ref-4">4</a>].
Decision Point 3: Normalizing CAZyme Counts for Cross-Sample Comparison
Raw CAZyme counts are not directly comparable across samples. A sample with 100,000 predicted genes will naturally contain more CAZyme predictions than a sample with 20,000 predicted genes, even when the CAZyme proportion is identical. Multi-sample comparison requires a normalization strategy that accounts for these differences.
The simplest normalization divides CAZyme counts by total predicted genes, producing a CAZyme proportion or percentage. This approach is intuitive and easy to report. A sample with 500 CAZymes out of 50,000 genes has a CAZyme proportion of 1 percent. This normalization controls for gene prediction differences but does not account for assembly quality variation. A fragmented assembly produces partial gene models, and partial genes may be missed by HMMER or Hotpep, reducing the apparent CAZyme proportion [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
A second normalization divides CAZyme counts by total assembled bases or by estimated genome equivalents. This approach approximates CAZyme density per unit of community DNA. It is more robust to gene prediction differences but requires accurate assembly size estimates, which are affected by the same fragmentation issues.
A third approach uses relative abundance of CAZyme families within each sample. Each sample's CAZyme profile is converted to proportions that sum to 100 percent across families. This approach supports comparison of community composition independent of total CAZyme count. It is appropriate for questions about which CAZyme families dominate in different environments, but it discards information about total CAZyme content.
The choice of normalization should be recorded in the analysis metadata and reported in publications. The dbCAN protocol includes commands for generating abundance plots from multiple samples [<a href="#ref-4">4</a>], and these plots should clearly state which normalization was applied. Researchers comparing their results to the dbCAN-seq database should note that the database reports CAZyme and CGC data from 9,421 MAGs across four ecological environments [<a href="#ref-2">2</a>], and that MAG-based counts are not directly comparable to read-based or assembly-based counts from metagenomes.
Decision Point 4: Handling Variable Sequencing Depth Across Samples
Multi-sample projects frequently include samples with different sequencing depths. A deeply sequenced sample may have 100 million reads, while a shallow sample has 10 million. This variation affects assembly quality, gene prediction, and CAZyme detection sensitivity.
The first response to variable sequencing depth is to document it. Record the read count, base count, and estimated coverage for each sample before assembly. These metrics should be included in the project metadata and considered during interpretation.
The second response is to consider depth-dependent assembly strategies. Shallow samples may benefit from co-assembly with deeper samples from the same environment, improving contiguity through shared coverage. Alternatively, assembly-free analysis may be more appropriate for shallow samples because it does not depend on assembly completeness [<a href="#ref-4">4</a>].
The third response is to apply a minimum depth threshold for inclusion in comparative analyses. Samples below a specified read count or estimated coverage should be flagged as low-confidence and either excluded from quantitative comparisons or analyzed separately. The threshold should be determined before analysis and recorded in the methods.
Decision Point 5: Structuring the Analysis Pipeline for Reproducibility
Multi-sample projects require a structured pipeline that applies identical parameters to every sample. Manual processing of each sample through separate commands invites parameter drift and documentation errors. The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including version pinning, containerization, and configuration management [<a href="#ref-8">8</a>]. Bioconductor provides additional guidance on reproducible genomic analysis practices, including package version management and workflow documentation [<a href="#ref-9">9</a>].
A reproducible multi-sample pipeline should include the following components. First, a sample manifest listing each sample identifier, the path to its raw reads, and its metadata such as treatment group, time point, or environment. Second, a parameter file specifying the dbCAN database version, HMMER version and thresholds, DIAMOND version and parameters, and Hotpep settings [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>]. Third, a processing script that iterates over the sample manifest and applies the identical pipeline to each sample. Fourth, an output directory structure that separates raw results, processed results, and summary tables for each sample.
The Galaxy Training Network provides accessible tutorials for metagenomic analysis workflows that can be adapted to multi-sample CAZyme projects [<a href="#ref-5">5</a>]. The Carpentries lessons offer foundational training in shell scripting and data management that supports the development of reproducible pipelines [<a href="#ref-6">6</a>]. The EMBL-EBI Training portal provides bioinformatics learning pathways that cover workflow design and data management [<a href="#ref-10">10</a>].
Decision Point 6: Choosing Between MAG-Based and Read-Based Analysis
Multi-sample projects may choose to analyze CAZymes from metagenome-assembled genomes (MAGs) instead of directly from metagenomic assemblies. MAG-based analysis bins assembled contigs into putative genomes, then annotates CAZymes within each MAG. The dbCAN-seq database provides CAZyme and CGC data from 9,421 MAGs spanning human gut, human oral, cow rumen, and marine environments [<a href="#ref-2">2</a>], demonstrating the utility of this approach for comparative studies.
MAG-based analysis has distinct advantages for multi-sample projects. MAGs provide taxonomic context for CAZyme predictions, allowing researchers to ask which taxa contribute which CAZymes. MAGs also reduce data complexity: instead of thousands of contigs per sample, each sample is represented by a manageable number of MAGs. The dbCAN-seq database organizes data by substrate and taxonomic phylum, supporting such comparisons [<a href="#ref-2">2</a>].
MAG-based analysis has limitations. MAGs are typically recovered only for abundant community members, so rare taxa are underrepresented. MAG completeness and contamination vary, affecting CAZyme counts. The dbCAN-seq database is limited to four ecological environments [<a href="#ref-2">2</a>], so researchers studying other environments cannot assume reference profiles are representative.
The decision between MAG-based and read-based analysis depends on the research question. Projects asking which taxa contribute CAZymes should use MAG-based analysis. Projects asking about total community CAZyme content should use read-based or assembly-based analysis. Projects doing both should run both analyses and report the results separately.
Decision Point 7: Managing Substrate Prediction Uncertainty at Scale
Substrate prediction is the most uncertain step in CAZyme annotation, and this uncertainty compounds in multi-sample projects. The dbCAN3 update introduced three substrate prediction methods: dbCAN-sub for subfamily-level prediction, dbCAN-PUL homology search for CGC-level prediction, and majority voting across CAZymes within a CGC [<a href="#ref-1">1</a>]. The dbCAN-seq analysis found that two substrate prediction approaches agreed on only 4,183 of 41,447 CGCs with predicted substrates [<a href="#ref-2">2</a>], indicating substantial method-dependent variation.
For multi-sample projects, substrate prediction uncertainty should be managed systematically. First, record which substrate prediction method was used for each CGC. Second, report the level of agreement between methods when multiple methods are applied. Third, treat substrate predictions as hypotheses instead of confirmed functions, particularly for CGCs where methods disagree [<a href="#ref-2">2</a>][<a href="#ref-7">7</a>].
The dbCAN-PUL database provides experimentally characterized PULs with known glycan substrates [<a href="#ref-7">7</a>]. For multi-sample projects, matching query CGCs to dbCAN-PUL entries provides the strongest substrate evidence because it is based on experimental characterization. However, dbCAN-PUL coverage is limited to PULs that have been experimentally studied, and many environmental CGCs will lack matches.
Decision Point 8: Establishing Quality Control Thresholds for Multi-Sample Inclusion
Multi-sample projects need explicit quality control thresholds to decide which samples enter comparative analysis. These thresholds should be established before analysis and applied uniformly.
Assembly quality thresholds include minimum N50, minimum number of contigs, and minimum total assembled bases. Samples failing these thresholds should be flagged for reassembly or exclusion. Gene prediction thresholds include minimum number of predicted genes and minimum proportion of complete gene models. CAZyme annotation thresholds include minimum number of CAZyme predictions and minimum proportion of CAZymes identified by multiple tools [<a href="#ref-3">3</a>].
The dbCAN protocol provides a baseline for expected performance: the complete workflow from reads to visualization takes approximately 33 hours per sample on 40 CPUs [<a href="#ref-4">4</a>]. Samples that produce dramatically fewer CAZymes than expected for their environment may indicate assembly or annotation failures. The dbCAN-seq database provides reference CAZyme profiles for human gut, human oral, cow rumen, and marine environments [<a href="#ref-2">2</a>], which can serve as benchmarks for samples from these habitats.
Decision Point 9: Planning for Database Version Consistency
CAZyme family classifications change as new enzymes are characterized. The dbCAN HMM database is updated to reflect these changes, and different versions can produce different annotations for the same input sequences. Multi-sample projects must use a single database version for all samples to ensure comparability.
The database version should be recorded in the project metadata and reported in publications. For longitudinal studies that span database updates, researchers should re-run all samples with the same database version instead of mixing results from different versions. The nf-core documentation describes version pinning practices that support this requirement [<a href="#ref-8">8</a>].
Decision Point 10: Structuring Output for Downstream Statistical Analysis
Multi-sample CAZyme projects typically feed into statistical analysis: comparing CAZyme abundance between treatment groups, testing for differences in CAZyme diversity, or correlating CAZyme profiles with environmental variables. The output structure of the annotation pipeline should support these analyses.
The dbCAN protocol includes commands for generating abundance plots from multiple samples [<a href="#ref-4">4</a>]. For statistical analysis, researchers should also produce a sample-by-CAZyme-family matrix, where rows are samples and columns are CAZyme families, with cell values representing counts or normalized abundances. A second matrix should record CGC counts per sample, and a third matrix should record substrate predictions per sample. These matrices should be exported in a format compatible with statistical software.
The Bioconductor project provides packages and workflows for reproducible genomic analysis, including statistical analysis of count data [<a href="#ref-9">9</a>]. Researchers should plan the statistical analysis before running the annotation pipeline to ensure that the output structure meets the analysis requirements.
Common Failure Patterns in Multi-Sample CAZyme Projects
Several failure patterns recur in multi-sample CAZyme annotation projects. Recognizing these patterns early prevents wasted computation and misinterpretation.
The first failure pattern is batch effects from processing samples at different times or with different software versions. Samples processed with an older dbCAN database version will produce different CAZyme calls than samples processed with a newer version. This pattern is prevented by processing all samples with the same software and database versions, ideally in a single pipeline run [<a href="#ref-8">8</a>].
The second failure pattern is sample mislabeling or file mixing during parallel processing. When multiple samples are processed simultaneously, output files can be assigned to the wrong sample if the pipeline does not enforce strict file isolation. This pattern is prevented by using a sample manifest and output directory structure that prevents cross-sample file mixing.
The third failure pattern is overinterpretation of substrate predictions. The dbCAN-seq analysis found that substrate prediction methods agreed on only a minority of CGCs [<a href="#ref-2">2</a>], yet researchers may report substrate predictions as confirmed functions. This pattern is prevented by reporting the prediction method and the level of method agreement for each substrate assignment.
The fourth failure pattern is comparing CAZyme profiles across samples with very different sequencing depths without normalization. This pattern produces apparent differences that reflect sequencing effort instead of biology. It is prevented by applying the normalization strategy described above and reporting the normalization method.
Records and Measurements for Multi-Sample Projects
Multi-sample CAZyme projects require a record system that tracks samples, parameters, and results systematically. The following records should be maintained for every project.
The sample manifest records each sample identifier, raw read file paths, sequencing depth, assembly strategy, assembly quality metrics, gene prediction counts, and CAZyme counts. This manifest serves as the master record for the project and should be updated as analysis progresses.
The parameter log records the software versions, database versions, and parameter settings for each analysis stage. This log should include the dbCAN database version, HMMER version and thresholds, DIAMOND version and parameters, Hotpep settings, and substrate prediction method [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
The quality control log records which samples passed or failed each quality threshold, with the specific metric values that triggered any failures. This log supports transparent reporting of sample inclusion and exclusion decisions.
The substrate prediction log records the substrate prediction method used for each CGC and the level of agreement between methods when multiple methods were applied [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-7">7</a>].
Professional Escalation Criteria for Multi-Sample Projects
Certain situations in multi-sample CAZyme projects require consultation with a bioinformatics specialist or CAZyme annotation expert.
Persistent batch effects that cannot be traced to software or database version changes warrant expert investigation. If samples processed in the same batch show systematic differences from samples processed in other batches, the cause may be an undocumented parameter change or an environmental factor in the computing infrastructure.
Samples that consistently fail quality thresholds despite adequate sequencing depth may indicate a biological or technical issue that requires expert diagnosis. The cause could be unusual community composition, contamination, or a systematic error in the analysis pipeline.
Projects comparing CAZyme profiles across many environments or treatment groups benefit from expert input on statistical analysis. The choice of normalization method, statistical model, and multiple testing correction can substantially affect conclusions [<a href="#ref-4">4</a>].
Regulatory or commercial applications of multi-sample CAZyme data require additional validation beyond computational prediction. Consult with appropriate regulatory and scientific experts to determine the evidence standards required for the specific application.
Frequently Asked Questions
What is the difference between dbCAN2 and dbCAN3?
dbCAN2 is a meta server that combines HMMER, DIAMOND, and Hotpep searches for CAZyme family annotation and includes CGC-Finder for gene cluster prediction. dbCAN3 inherits all dbCAN2 functions and adds three new substrate prediction methods: dbCAN-sub for subfamily-level substrate prediction, dbCAN-PUL homology search for CGC-level substrate prediction, and a majority-voting method that considers all CAZymes with predicted substrates within a CGC [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].
Should I use the dbCAN web server or the run_dbcan standalone package?
Use the web server for small datasets, exploratory analysis, or when command-line tools are not available. Use the standalone package for large metagenomic projects, sensitive data requiring local processing, or when you need to process many samples reproducibly. The standalone package requires Linux command-line familiarity and takes approximately 33 hours for a typical individual sample assembly on a 40-CPU machine [<a href="#ref-4">4</a>].
How do I choose between individual sample assembly, co-assembly, and assembly-free analysis?
Individual sample assembly preserves sample-specific variation but produces more fragmented assemblies. Co-assembly improves contiguity and gene discovery but can obscure sample-specific signals. Assembly-free analysis is fastest and detects genes from reads not captured in assemblies, but produces fragmentary gene models. The choice depends on your research question and the tradeoff between assembly quality and sample-specific resolution [<a href="#ref-4">4</a>].
What does it mean when a CAZyme is identified by only one tool?
The dbCAN2 majority-voting rule removes genes identified by only one of the three search tools to improve annotation accuracy. A gene identified by only one tool may still be a genuine CAZyme, particularly if it is divergent or fragmentary. Such genes should be examined manually using domain prediction and phylogenetic analysis before being discarded or included in downstream analysis [<a href="#ref-3">3</a>].
How reliable are glycan substrate predictions?
Substrate predictions carry more uncertainty than family-level assignments. dbCAN-sub provides subfamily-level predictions, while dbCAN-PUL matches query CGCs to experimentally characterized PULs. The two methods often disagree, and the dbCAN-seq analysis found agreement on only a minority of CGCs with predicted substrates. Report the prediction method used and acknowledge the uncertainty in substrate assignments [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-7">7</a>].
Can I compare CAZyme profiles generated with different dbCAN versions?
Direct comparison is not recommended. CAZyme family classifications and HMM profiles change as new enzymes are characterized, so different database versions can produce different annotations for the same input sequences. For comparative studies, re-run all samples with the same database version and record the version in your methods [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
What training resources are available for learning metagenomic CAZyme analysis?
The Galaxy Training Network provides accessible workflow tutorials for metagenomic analysis [<a href="#ref-5">5</a>]. The Carpentries offers foundational lessons in shell scripting, data management, and programming that support command-line analysis [<a href="#ref-6">6</a>]. The EMBL-EBI Training portal provides bioinformatics learning pathways and data-resource training [<a href="#ref-10">10</a>]. Bioconductor documents reproducible genomic analysis practices [<a href="#ref-9">9</a>], and nf-core documentation describes community standards for reproducible pipelines [<a href="#ref-8">8</a>].
How should I report CAZyme annotation methods in publications?
Report the software versions, database versions, search tools, and parameter thresholds used for annotation. Specify the assembly strategy and gene prediction method. For substrate predictions, state which method was used and note any disagreements between methods. This information allows other researchers to reproduce your analysis and interpret your results appropriately [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>].
Related Bioinformatics Guides
- Functional Annotation of Metagenomes: A Guide to Databases and Pipelines
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Metagenomics Functional Profiling: Tools and Databases for Pathway Analysis
- Metagenomics Data Analysis: From Raw Reads to Biological Insights
- Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [dbCAN3: automated carbohydrate-active enzyme and substrate annotation.](https://pubmed.ncbi.nlm.nih.gov/37125649). Nucleic acids research, 2023. [2] [dbCAN-seq update: CAZyme gene clusters and substrates in microbiomes.](https://pubmed.ncbi.nlm.nih.gov/36399503). Nucleic acids research, 2023. [3] [dbCAN2: a meta server for automated carbohydrate-active enzyme annotation.](https://pubmed.ncbi.nlm.nih.gov/29771380). Nucleic acids research, 2018. [4] [Carbohydrate-active enzyme annotation in microbiomes using dbCAN.](https://pubmed.ncbi.nlm.nih.gov/38260309). bioRxiv : the preprint server for biology, 2024. [5] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [6] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [7] [dbCAN-PUL: a database of experimentally characterized CAZyme gene clusters and their substrates.](https://pubmed.ncbi.nlm.nih.gov/32941621). Nucleic acids research, 2021. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.