# The Impact of Reference Databases on Taxonomic Profiling Accuracy: A Guide to Choosing and Updating Databases for Kraken2 and MetaPhlAn


## Key Takeaways

- Taxonomic profiling accuracy in Kraken2 and MetaPhlAn is fundamentally dictated by the reference database's completeness, version, and taxonomic resolution; larger, more comprehensive databases like NCBI nt or GTDB-derived sets generally improve sensitivity and precision, especially under stringent confidence score settings (e.g., 0.2-0.4 in Kraken2), though they demand greater computational resources.
- MetaPhlAn's marker gene approach offers computational efficiency but is limited by the absence of marker genes for unrepresented organisms, necessitating regular database updates to incorporate new genomic data and ensure accurate species-level abundance estimation.
- Reproducibility hinges on meticulous documentation of database versions, build dates, and all classification parameters; using tools like Galaxy or nf-core and version control practices is crucial for ensuring analyses can be precisely replicated.
- Strain-level resolution requires specialized databases and tools (e.g., Themisto/mSWEEP) as standard Kraken2 databases exhibit very low accuracy (e.g., 1.4% vs. 84.9% for Themisto/mSWEEP in a *Mycoplasma bovis* study) for distinguishing closely related strains.
- Validation through mock communities or simulated datasets is essential to assess classification rate, precision, recall, and abundance estimation accuracy, allowing for parameter optimization and identification of common failure patterns like outdated databases or mismatches between database and sample type.
- Clinical applications demand high diagnostic accuracy, requiring careful database selection and validation against established standards; for instance, a validated Kraken2 workflow achieved rapid species identification in blood cultures, significantly outperforming routine methods.

---

Taxonomic profiling of shotgun metagenomic data depends on the reference database used for classification. The choice of database and its version can change which organisms are detected, how accurately their abundances are estimated, and whether results from different studies can be compared. This article explains how reference databases affect Kraken2 and MetaPhlAn outputs, compares commonly used database options, and provides practical guidance for selecting, updating, and documenting databases in metagenomic research workflows.

## Why Reference Database Choice Matters in Taxonomic Profiling

Metagenomic classifiers assign sequencing reads to taxonomic groups by comparing them against collections of known genomic sequences. Kraken2 uses k-mer based classification, matching short DNA sequences from reads against a database of k-mers derived from reference genomes. MetaPhlAn uses a different approach, relying on clade-specific marker genes to identify organisms and estimate their relative abundances. Both methods depend entirely on the contents of their reference databases. If an organism is not represented in the database, it cannot be detected. If related but distinct organisms share similar sequences, the classifier may assign reads to the wrong taxon.

A 2024 study in aBIOTECH evaluated how database selection and confidence score settings affect Kraken2 performance using simulated metagenomic datasets. The researchers compared databases ranging from the compact Minikraken v1 to the expansive nt and GTDB r202 databases. They found that larger databases combined with moderate confidence scores of 0.2 or 0.4 significantly improved classification accuracy and sensitivity. Smaller databases such as Minikraken and Standard-16 classified no reads when the confidence score exceeded 0.4, while larger databases maintained robust precision and F1 scores under stringent conditions. This evidence demonstrates that database choice is not a minor technical detail but a primary determinant of classification performance.

For MetaPhlAn, the database consists of marker genes extracted from reference genomes. The marker gene approach is computationally efficient because it only searches for specific genomic regions instead of entire genomes. However, this efficiency comes with a tradeoff. Organisms whose marker genes are not present in the database will be missed entirely, and novel strains with divergent marker sequences may be classified at lower taxonomic resolution or not at all.

## Core Principles of Reference Database Selection

### Completeness Versus Computational Cost

Reference databases exist on a spectrum from compact to comprehensive. Compact databases such as Minikraken contain only a subset of representative genomes and require less memory and disk space. They run faster and are suitable for quick exploratory analyses. Comprehensive databases such as the full NCBI nt database or GTDB-derived databases contain many more genomes and provide broader taxonomic coverage, but they require substantial computational resources.

The tradeoff between completeness and computational cost is central to database selection. A 2024 study in Environmental Microbiome evaluated metagenomic classifiers for soil microbiomes and found that classifiers tailored to the specific taxa present in samples led to fewer errors compared with broader databases that included microbial eukaryotes, protozoa, or human genomes. This finding supports the use of targeted databases when the expected community composition is known, even though such databases are less comprehensive than general-purpose options.

### Database Version and Reproducibility

Reference databases are updated regularly as new genomes are sequenced and deposited in public repositories. The National Center for Biotechnology Information maintains sequence databases that grow continuously as researchers submit new genome assemblies. A database version that was current at the start of a study may be outdated by the time the study is published. This creates reproducibility challenges because other researchers may not be able to replicate the exact database configuration used in the original analysis.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics. Similarly, nf-core documentation describes community pipeline standards that include version pinning and containerization to ensure that analyses can be reproduced exactly. These resources illustrate the importance of documenting database versions and parameters as part of any metagenomic analysis workflow.

### Taxonomic Resolution and Reference Genome Quality

The taxonomic resolution achievable with a given database depends on the quality and completeness of the reference genomes it contains. High-quality complete genomes allow classifiers to distinguish closely related species and strains. Draft genomes with gaps and errors may lead to ambiguous classifications. The NCBI provides official descriptions of its sequence databases and search systems, which include information about genome assembly quality and taxonomic curation.

For strain-level classification, the reference database must contain genomes that are sufficiently similar to the organisms in the sample. A 2026 study in Frontiers in Veterinary Science evaluated targeted enrichment shotgun sequencing for detecting Mycoplasma bovis strains in milk. The researchers downloaded 620 M. bovis whole-genome sequences from NCBI and grouped them into genomically clustered sequence variants. They found that Themisto/mSWEEP achieved an average read classification accuracy of 84.9 percent for strain-level detection, while Kraken2 achieved only 1.4 percent accuracy under the same conditions. This dramatic difference illustrates that Kraken2 with a standard database is not suitable for strain-level classification when closely related strains must be distinguished.

## At a Glance: Database Options for Kraken2 and MetaPhlAn

| Database | Classifier Compatibility | Taxonomic Coverage | Computational Requirements | Best Use Cases | Key Limitations |
|----------|------------------------|-------------------|---------------------------|----------------|-----------------|
| Minikraken v1 | Kraken2 | Bacteria and archaea only, reduced representation | Low memory and disk usage, fast classification | Quick exploratory analysis, teaching, resource-limited environments | Poor sensitivity for rare taxa, fails under high confidence score thresholds |
| NCBI RefSeq complete genomes | Kraken2, MetaPhlAn | Bacteria, archaea, viruses, some eukaryotes | Moderate to high memory depending on build | General microbiome studies, clinical pathogen detection | May miss novel or underrepresented taxa, requires regular updates |
| NCBI nt (nucleotide collection) | Kraken2 | All domains including eukaryotes and viruses | Very high memory and disk usage | Comprehensive surveys, environmental samples with diverse taxa | Slow classification, requires substantial computational infrastructure |
| GTDB-derived databases | Kraken2 | Bacteria and archaea with standardized taxonomy | Moderate to high memory | Studies requiring consistent taxonomic framework, soil and environmental microbiomes | Limited representation of eukaryotes and viruses |
| MetaPhlAn marker gene database | MetaPhlAn | Bacteria, archaea, eukaryotes, viruses with marker genes | Low to moderate memory | Fast profiling of microbial communities, species-level abundance estimation | Misses organisms without marker genes, sensitive to database version |
| Custom targeted databases | Kraken2, MetaPhlAn | Depends on included genomes | Variable, often lower than comprehensive databases | Studies focused on specific taxa, clinical diagnostics, outbreak investigations | Requires careful construction and validation, may miss unexpected organisms |

## Kraken2 Database Construction and Selection

### Standard Database Builds

Kraken2 provides a standard database build script that downloads reference genomes from NCBI and constructs a k-mer database. The standard build includes bacterial, archaeal, and viral genomes from RefSeq, along with a small collection of human sequences for host contamination removal. This database is suitable for general microbiome studies but has limitations. It does not include fungal genomes by default, and it may not represent environmental taxa that are absent from RefSeq.

The choice of k-mer length and confidence score parameters interacts with database selection. The aBIOTECH study demonstrated that higher confidence scores decrease classification rates, with the effect being particularly pronounced for smaller databases. For the Standard database, precision and F1 scores improved significantly with increasing confidence scores, indicating that larger databases are more robust to stringent classification thresholds. Researchers should test multiple confidence score settings when using a new database to find the optimal balance between sensitivity and precision for their specific samples.

### GTDB-Derived Databases

The Genome Taxonomy Database provides a standardized taxonomy for bacteria and archaea based on genome phylogeny. GTDB-derived databases for Kraken2 can be constructed using GTDB-TK genomes, which provide a consistent taxonomic framework that differs from NCBI taxonomy in some classifications. The Environmental Microbiome study used a custom database derived from GTDB-TK genomes for soil microbiome analysis and found that it performed well for soil-associated taxa.

GTDB databases are particularly useful when comparing results across studies because they use a consistent taxonomic framework. However, researchers should be aware that GTDB taxonomy may differ from NCBI taxonomy for some organisms, which can complicate comparisons with studies that used NCBI-based databases. The choice between GTDB and NCBI taxonomy should be documented clearly in methods sections.

### Custom and Targeted Databases

Custom databases allow researchers to tailor the reference set to their specific research questions. A targeted database containing only the genomes of expected organisms can improve classification accuracy by reducing the chance of false-positive assignments to unrelated taxa. The Environmental Microbiome study found that classifiers tailored to the specific taxa present in soil samples led to fewer errors compared with broader databases.

Custom databases are particularly valuable for clinical applications where the range of possible pathogens is known. The Lancet Microbe study developed a direct-from-positive blood culture workflow using Oxford Nanopore sequencing and Kraken2 with a comprehensive standard database. The researchers achieved 97 percent sensitivity and 94 percent specificity for species identification compared with routine culture-based diagnostics. They detected 19 additional infections that were missed by conventional methods, including polymicrobial infections and culture-negative samples. This study demonstrates that a well-chosen database combined with appropriate classification models can deliver clinically useful results.

### Building and Validating Custom Databases

When constructing a custom database, researchers should follow a systematic process. First, define the taxonomic scope based on the research question and expected community composition. Second, download genome assemblies from NCBI or other repositories, ensuring that genome quality is assessed and documented. Third, construct the database using the appropriate Kraken2 build command, specifying the k-mer length and other parameters. Fourth, validate the database using mock communities or simulated reads to assess classification accuracy.

The Carpentries provides foundational computing and data lessons that include training on shell, Git, and programming skills needed for reproducible bioinformatics workflows. These skills are essential for managing custom database construction scripts and documenting the process for reproducibility.

## MetaPhlAn Database Considerations

### Marker Gene Approach

MetaPhlAn uses clade-specific marker genes to identify organisms in metagenomic samples. The database contains collections of marker genes that are unique to specific taxonomic clades. When a read matches a marker gene, the classifier assigns it to the corresponding taxon and uses the number of reads to estimate relative abundance.

The marker gene approach is computationally efficient because it only searches for specific genomic regions instead of all sequences in a genome. This efficiency makes MetaPhlAn suitable for large datasets and rapid analysis. However, the approach has limitations. Organisms whose marker genes are not present in the database cannot be detected, and the resolution of classification depends on the specificity of the marker genes.

### Database Updates and Version Control

MetaPhlAn databases are updated periodically to include new marker genes from newly sequenced genomes. Each version of the database may produce different results for the same input data. Researchers should record the exact database version used in their analysis and avoid mixing results from different versions in the same study.

The Bioconductor project provides official documentation for reproducible genomic analysis, including workflows that emphasize version control and documentation. Similarly, the EMBL-EBI Training portal offers bioinformatics learning pathways that cover data-resource training and practical analysis education. These resources can help researchers develop robust practices for managing database versions.

### Species-Level Resolution and Abundance Estimation

MetaPhlAn is designed to estimate relative abundances of microbial species in a sample. The accuracy of these estimates depends on the completeness of the marker gene database and the evenness of marker gene representation across taxa. If some species have more marker genes in the database than others, their abundances may be overestimated relative to species with fewer marker genes.

A 2025 study in Bioscience Trends compared shallow shotgun metagenomic sequencing with full-length 16S rDNA amplicon sequencing for human gut bacterial microbiota analysis. The study title indicates a comparative analysis of these approaches, and the publication metadata supports the relevance of understanding how different sequencing and analysis methods affect taxonomic profiling results. Researchers should be aware that the choice of sequencing depth and analysis method interacts with database selection to determine the final taxonomic profile.

## Practical Workflow for Database Selection

### Step 1: Define the Research Question and Expected Community

The first step in database selection is to define the research question and consider the expected microbial community composition. A study of human gut microbiota will have different database requirements than a study of soil microbiomes or wastewater viruses. The expected taxa should guide the initial choice between general-purpose and targeted databases.

For clinical applications where rapid pathogen identification is critical, a comprehensive database that includes all known pathogens is appropriate. The Lancet Microbe study used a comprehensive standard database for blood culture analysis and achieved rapid species identification within approximately 3 hours and 20 minutes, about 10 hours earlier than routine diagnostic methods. This speed is essential for critically ill patients with bloodstream infections.

### Step 2: Assess Computational Resources

Database selection must account for available computational resources. Comprehensive databases such as the full NCBI nt database require substantial memory and disk space. Researchers working with limited resources may need to use compact databases or cloud-based analysis platforms.

The Galaxy Training Network provides accessible workflow training that includes guidance on running metagenomic analyses on shared infrastructure. The nf-core documentation describes community pipeline standards that support scalable and reproducible workflows. These resources can help researchers optimize their computational resource usage while maintaining analysis quality.

### Step 3: Select and Build the Database

Once the research question and computational resources are defined, select the appropriate database and build it using the classifier-specific tools. For Kraken2, this involves running the build script with the desired reference genomes. For MetaPhlAn, this involves downloading the appropriate marker gene database version.

Document the exact database version, build date, and parameters used. This documentation is essential for reproducibility and for interpreting results in the context of the database contents.

### Step 4: Validate Classification Performance

Before analyzing real samples, validate the database using mock communities or simulated reads. The Environmental Microbiome study generated a custom in-silico mock community containing microbial genomes commonly observed in soil microbiomes and used it to evaluate classifier performance. This approach allows researchers to assess classification accuracy, sensitivity, and precision under controlled conditions.

The aBIOTECH study used simulated metagenomic datasets to evaluate Kraken2 performance across different databases and confidence score settings. Simulated data provides ground truth for assessing classification accuracy, which is not available for real environmental samples.

### Step 5: Apply and Document

Apply the validated database to real samples and document all parameters in the methods section. Include the database version, build date, confidence score settings, and any filtering thresholds applied. This documentation allows other researchers to reproduce the analysis and interpret the results appropriately.

## Records and Measurements for Database Performance

### Classification Rate

Classification rate is the proportion of reads that are assigned to a taxonomic group. The aBIOTECH study found that higher confidence scores decrease classification rates, particularly for smaller databases. Researchers should record the classification rate for each sample and database configuration to assess whether the analysis is capturing sufficient information from the data.

Low classification rates may indicate that the database does not represent the organisms present in the sample. This can occur with environmental samples containing novel or underrepresented taxa. The Environmental Microbiome study highlighted the dearth of soil-specific reference databases available to classifiers, which limits the ability to classify soil microbial communities accurately.

### Precision and Recall

Precision measures the proportion of classifications that are correct, while recall measures the proportion of true organisms that are detected. These metrics are typically assessed using mock communities or simulated data where the true composition is known. The aBIOTECH study reported that precision and F1 scores improved significantly with increasing confidence scores for larger databases, while recovery rates remained mostly stable.

For real samples, precision and recall cannot be measured directly because the true community composition is unknown. However, researchers can use complementary methods such as amplicon sequencing or quantitative PCR to validate selected findings.

### Abundance Estimation Accuracy

Taxonomic classifiers provide estimates of relative abundance for detected taxa. The accuracy of these estimates depends on the database contents and the classification algorithm. The aBIOTECH study evaluated the accuracy of true versus calculated bacterial abundance estimation and found that comprehensive databases combined with moderate confidence scores improved accuracy.

Researchers should compare abundance estimates from different classifiers and databases to assess robustness. Discrepancies between methods may indicate database limitations or classification errors that warrant further investigation.

### Time-to-Result

For clinical applications, the time required to produce taxonomic classifications is critical. The Lancet Microbe study reported species identification results within 3 hours and 20 minutes, approximately 10 hours earlier than routine diagnostic methods. Researchers developing clinical workflows should measure and document time-to-result for their specific pipeline configuration.

## Common Failure Patterns in Database Selection

### Using Outdated Database Versions

Reference databases are updated regularly as new genomes are sequenced. Using an outdated database version can lead to missed organisms and inaccurate classifications. The NCBI maintains current sequence databases that grow continuously, and researchers should update their databases periodically to reflect new genome submissions.

However, updating databases mid-study can introduce inconsistencies if different samples are analyzed with different database versions. Researchers should either use a single database version for the entire study or document the version used for each sample and account for version differences in the analysis.

### Mismatch Between Database and Sample Type

Databases designed for one sample type may perform poorly on another. The Environmental Microbiome study found that soil-specific databases were lacking, and classifiers tailored to the specific taxa present in soil samples led to fewer errors compared with broader databases. Researchers should select databases that represent the expected community composition of their specific sample type.

A 2025 study in Microbiome addressed challenges in capturing the mycobiome from shotgun metagenome data, highlighting the lack of software and databases for fungal detection. The publication metadata supports the finding that fungal taxa are often underrepresented in reference databases, leading to poor detection of fungi in metagenomic samples. Researchers studying fungal communities should be aware of this limitation and consider complementary approaches such as ITS amplicon sequencing.

### Ignoring Confidence Score Settings

The confidence score parameter in Kraken2 controls the stringency of classification. The aBIOTECH study demonstrated that this parameter significantly affects classification performance, with smaller databases failing to classify any reads at confidence scores above 0.4. Researchers should test multiple confidence score settings and select the value that optimizes the balance between sensitivity and precision for their specific database and sample type.

### Overlooking Host Contamination

Metagenomic samples often contain host DNA that can interfere with taxonomic classification. The 2025 study in Journal of Translational Medicine examined dermatological implications of alignment-based de-hosting and bioinformatics pipelines on shotgun microbiome analysis. The publication metadata indicates that de-hosting approaches can affect downstream taxonomic profiling results. Researchers should include host genome sequences in their databases or remove host reads before classification to avoid false-positive assignments to host-associated taxa.

### Using Inappropriate Databases for Strain-Level Analysis

Standard databases such as those used by Kraken2 are not suitable for strain-level classification when closely related strains must be distinguished. The Frontiers in Veterinary Science study found that Kraken2 achieved only 1.4 percent accuracy for strain-level classification of Mycoplasma bovis, while Themisto/mSWEEP achieved 84.9 percent accuracy. Researchers requiring strain-level resolution should use specialized tools and databases designed for this purpose.

## Quality Controls and Validation Approaches

### Mock Community Validation

Mock communities with known composition provide the most direct validation of taxonomic classification accuracy. The Environmental Microbiome study generated a custom in-silico mock community containing microbial genomes commonly observed in soil microbiomes and used it to evaluate classifier performance. Researchers can also use commercially available mock communities containing defined mixtures of microbial strains.

When using mock communities, researchers should assess classification rate, precision, recall, and abundance estimation accuracy. These metrics provide a baseline for evaluating database performance and can guide parameter optimization.

### Cross-Validation with Multiple Classifiers

Using multiple classifiers with different databases can provide complementary information and help identify classification errors. The Microbiome study on vaginal microbiota used shotgun metagenomic sequencing with VIRGO, MetaPhlAn, and Kraken2 to validate taxonomic assignments. Agreement between classifiers increases confidence in the results, while discrepancies highlight areas of uncertainty.

The Water Research study on wastewater viruses compared four virus identification tools including Diamond blast, Kraken2, VirSorter2, and geNomad. The researchers found that different tools detected different viruses, with some viruses unique to specific methods. This finding underscores the importance of using multiple tools for comprehensive viral detection.

### Relative Abundance Thresholds

The Environmental Microbiome study found that optimal classifier performance was achieved when applying a relative abundance threshold of 0.001 percent or 0.005 percent. Low-abundance taxa are more likely to be misclassified due to sequencing errors and database limitations. Applying an abundance threshold can reduce false-positive identifications while retaining biologically meaningful taxa.

Researchers should select abundance thresholds based on their research question and the expected community composition. Thresholds that are too high may miss rare but important taxa, while thresholds that are too low may include spurious classifications.

### Ancient DNA-Specific Filtering

Ancient metagenomic data presents unique challenges for taxonomic classification due to DNA damage and contamination. A 2026 study in Frontiers in Microbiology evaluated filtering strategies for Kraken family tools using simulated microbial and environmental ancient metagenomic data. The study proposed an optimal thresholding strategy tailored to specific sequencing depths in ancient metagenomic datasets.

Researchers working with ancient DNA should use filtering approaches designed for degraded DNA and should validate their results using appropriate controls. The study emphasizes the balance between sensitivity and specificity in ground truth reconstruction, measured by F1-score.

## Limitations of Reference Database Approaches

### Incomplete Representation of Microbial Diversity

Reference databases represent only a fraction of the microbial diversity present in natural environments. Many microorganisms have not been cultured or sequenced, and their genomes are absent from public databases. This limitation is particularly pronounced for environmental samples such as soil and wastewater, where the majority of microbial diversity remains uncharacterized.

The Environmental Microbiome study highlighted the dearth of soil-specific reference databases available to classifiers. Similarly, the Microbiome study on the mycobiome noted the lack of software and databases for capturing fungi from shotgun metagenome data. Researchers should acknowledge these limitations when interpreting their results and should avoid overinterpreting the absence of specific taxa in their samples.

### Database Bias Toward Well-Studied Organisms

Reference databases are biased toward organisms that have been studied extensively, such as human pathogens and model organisms. This bias can lead to overrepresentation of these organisms in taxonomic profiles and underrepresentation of less-studied taxa. The NCBI provides official descriptions of its sequence databases, which include information about the taxonomic distribution of reference genomes.

Researchers should be aware that relative abundance estimates may be inflated for well-represented taxa and deflated for poorly represented taxa. This bias should be considered when interpreting differences in abundance between taxa.

### Taxonomic Framework Inconsistencies

Different databases use different taxonomic frameworks, which can complicate comparisons between studies. NCBI taxonomy and GTDB taxonomy differ for some organisms, and these differences can affect classification results. Researchers should document the taxonomic framework used in their analysis and should be cautious when comparing results across studies that used different frameworks.

The metaFun pipeline described in a 2026 study in Gut Microbes addresses the lack of unified taxonomic criteria that limits cross-study comparability. The pipeline integrates taxonomic profiling with functional profiling and other analyses into a unified framework, providing standardized data interpretation.

### Computational Resource Constraints

Comprehensive databases require substantial computational resources, which may not be available to all researchers. The full NCBI nt database requires hundreds of gigabytes of memory for Kraken2 classification, which exceeds the capacity of many desktop computers. Researchers with limited resources may need to use compact databases or cloud-based analysis platforms.

The Galaxy Training Network provides accessible workflow training that includes guidance on running metagenomic analyses on shared infrastructure. The nf-core documentation describes community pipeline standards that support scalable workflows. These resources can help researchers optimize their computational resource usage while maintaining analysis quality.

## Safety and Regulatory Context for Clinical Applications

### Diagnostic Accuracy Requirements

Clinical applications of metagenomic sequencing require high diagnostic accuracy to avoid misdiagnosis and inappropriate treatment. The Lancet Microbe study achieved 97 percent sensitivity and 94 percent specificity for species identification compared with routine culture-based diagnostics. After adjudication of plausible additional infections, both sensitivity and specificity increased to 100 percent.

Researchers developing clinical metagenomic workflows should validate their methods against established diagnostic standards and should document the performance characteristics of their specific pipeline configuration. The choice of reference database is a critical determinant of diagnostic accuracy and should be carefully evaluated.

### Antimicrobial Resistance Detection

Metagenomic sequencing can enable antimicrobial resistance prediction in addition to species identification. The Lancet Microbe study benchmarked AMR classification tools and databases, including ResFinder, CARD, and NCBI AMRFinderPlus. The NCBI provides official descriptions of its analysis services, including AMRFinderPlus for antimicrobial resistance gene detection.

Researchers should validate AMR detection methods against phenotypic susceptibility testing and should be aware of the limitations of genotypic resistance prediction. The presence of a resistance gene does not always correlate with phenotypic resistance, and the absence of known resistance genes does not rule out resistance.

### Reporting and Escalation Criteria

Clinical metagenomic results should be reported with appropriate caveats about the limitations of the methods used. Results should be confirmed by orthogonal methods when possible, and clinically significant findings should be escalated to appropriate healthcare providers.

The Lancet Microbe study detected 19 additional infections that were missed by conventional methods, including polymicrobial infections, previously unidentifiable infections, and one infection in a culture-negative sample. These findings demonstrate the potential of metagenomic sequencing to improve diagnostic yield, but they also highlight the need for careful interpretation and clinical correlation.

## Professional Escalation Criteria

Researchers should escalate database-related issues to appropriate professionals or resources in the following situations:

### When Classification Rates Are Unexpectedly Low

If classification rates are substantially lower than expected for the sample type, the database may not represent the organisms present in the sample. Consider whether a more comprehensive database or a custom database tailored to the expected community is needed. The Environmental Microbiome study provides guidance on optimizing taxonomic classification parameters and database selection for soil microbiomes.

### When Results Conflict Between Classifiers

If different classifiers or databases produce conflicting results, investigate the source of the discrepancy. The Water Research study found that different virus identification tools detected different viruses, with some viruses unique to specific methods. Consider whether the discrepancy reflects database limitations, algorithmic differences, or actual biological variation.

### When Strain-Level Resolution Is Required

If the research question requires strain-level resolution, standard databases and classifiers may not be sufficient. The Frontiers in Veterinary Science study demonstrated that specialized tools such as Themisto/mSWEEP outperform Kraken2 for strain-level classification. Consult the relevant literature and consider using specialized tools for strain-level analysis.

### When Clinical Decisions Depend on Results

If taxonomic profiling results will inform clinical decisions, validate the results using orthogonal methods and consult with clinical microbiology experts. The Lancet Microbe study provides an example of a clinically validated metagenomic workflow, but individual laboratories should validate their own methods before clinical use.

### When Database Updates Cause Result Changes

If updating a database version changes results for previously analyzed samples, document the changes and consider whether reanalysis is needed. The Bioconductor project provides official documentation for reproducible genomic analysis that emphasizes version control and documentation. Similarly, the EMBL-EBI Training portal offers bioinformatics learning pathways that cover data-resource training.

## Frequently Asked Questions

### How often should I update my Kraken2 or MetaPhlAn reference database?

Reference databases should be updated regularly to include newly sequenced genomes. The NCBI sequence databases grow continuously as researchers submit new genome assemblies. For active research projects, consider updating databases every 6 to 12 months or when beginning a new study. Document the database version and build date for each analysis to ensure reproducibility. If you update a database mid-study, reanalyze all samples with the same database version to maintain consistency.

### What is the difference between NCBI RefSeq and the full NCBI nt database for Kraken2?

RefSeq is a curated collection of reference genomes that provides a representative set of sequences for well-studied organisms. The full nt database includes all nucleotide sequences submitted to NCBI, including many redundant and uncurated sequences. The aBIOTECH study compared databases ranging from Minikraken to the expansive nt and GTDB r202 databases and found that larger databases combined with moderate confidence scores improved classification accuracy. The nt database provides broader coverage but requires substantially more computational resources.

### Can I use the same database for Kraken2 and MetaPhlAn?

No, Kraken2 and MetaPhlAn use different database formats and classification approaches. Kraken2 uses k-mer based classification against a database of k-mers derived from reference genomes. MetaPhlAn uses clade-specific marker genes and requires a marker gene database. You must build or download separate databases for each tool. The choice of reference genomes for each database should be guided by the same considerations of taxonomic coverage and relevance to your sample type.

### How do I build a custom database for Kraken2?

To build a custom Kraken2 database, download the desired reference genomes from NCBI or other repositories, then run the Kraken2 build command with the appropriate options. Assess genome quality before including sequences in the database and document the genomes included. Validate the custom database using mock communities or simulated reads to assess classification accuracy. The Environmental Microbiome study provides an example of building a custom database derived from GTDB-TK genomes for soil microbiome analysis.

### Why does MetaPhlAn miss fungi in my shotgun metagenomic samples?

MetaPhlAn relies on marker genes that are present in its database. Fungal marker genes may be underrepresented or absent from the database, leading to poor detection of fungi. A 2025 study in Microbiome addressed challenges in capturing the mycobiome from shotgun metagenome data, highlighting the lack of software and databases for fungal detection. The Microbiome study on vaginal microbiota found that fungi were identified in 39 of 50 samples with ITS sequencing, while in the metagenome data fungi largely remained undetected due to their low abundance and database issues. Consider using ITS amplicon sequencing for fungal community analysis.

### What confidence score should I use for Kraken2?

The optimal confidence score depends on your database and sample type. The aBIOTECH study found that moderate confidence scores of 0.2 or 0.4 combined with comprehensive databases significantly improved classification accuracy and sensitivity. Smaller databases failed to classify any reads at confidence scores above 0.4. Test multiple confidence score settings using mock communities or simulated data to find the optimal value for your specific configuration.

### How do I document database versions for reproducible research?

Record the exact database version, build date, and parameters used for each analysis. Include this information in the methods section of publications and in analysis documentation. The Bioconductor project provides official documentation for reproducible genomic analysis that emphasizes version control. The nf-core documentation describes community pipeline standards that include version pinning and containerization. The Carpentries provides foundational computing lessons that cover shell, Git, and programming skills needed for reproducible workflows.

### When should I use a targeted database instead of a comprehensive database?

Use a targeted database when the expected community composition is known and the research question focuses on specific taxa. The Environmental Microbiome study found that classifiers tailored to the specific taxa present in soil samples led to fewer errors compared with broader databases. Targeted databases are particularly useful for clinical applications where the range of possible pathogens is known. However, targeted databases may miss unexpected organisms, so consider using a comprehensive database for discovery-oriented research.

## Related Bioinformatics Guides

- [Metagenomics Functional Profiling: Tools and Databases for Pathway Analysis](/knowledge/bioinformatics/metagenomics-functional-profiling-tools-and-databases-for-pathway-analysis)
- [RNA-Seq Databases: Accessing and Using Public RNA-Seq Data](/knowledge/bioinformatics/rna-seq-databases-accessing-and-using-public-rna-seq-data)
- [Metagenomics vs Metatranscriptomics: Choosing the Right Approach for Functional Profiling](/knowledge/bioinformatics/metagenomics-vs-metatranscriptomics-choosing-the-right-approach-for-functional-profiling)
- [Single-Cell Sequencing Databases: Resources for Data Sharing and Exploration](/knowledge/bioinformatics/single-cell-sequencing-databases-resources-for-data-sharing-and-exploration)
- [RNA-Seq vs Microarray: Choosing the Right Gene Expression Profiling Platform](/knowledge/bioinformatics/rna-seq-vs-microarray-choosing-the-right-gene-expression-profiling-platform)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Impact of database choice and confidence score on the performance of taxonomic classification using Kraken2.](https://pubmed.ncbi.nlm.nih.gov/39650139). aBIOTECH, 2024.
- [An in-depth evaluation of metagenomic classifiers for soil microbiomes.](https://pubmed.ncbi.nlm.nih.gov/38549112). Environmental microbiome, 2024.
- [In-depth comparison of untargeted and targeted sequencing for detecting virus diversity in wastewater.](https://pubmed.ncbi.nlm.nih.gov/40373374). Water research, 2025.
- [Rapid diagnosis of common, undetected, and uncultivable bloodstream infections from positive blood cultures using Oxford Nanopore sequencing: a metagenomic pipeline analysis.](https://pubmed.ncbi.nlm.nih.gov/42134371). The Lancet. Microbe, 2026.
- [Unfermented β-fructan Fibers Fuel Inflammation in Select Inflammatory Bowel Disease Patients.](https://pubmed.ncbi.nlm.nih.gov/36183751). Gastroenterology, 2023.
- [The Metagenomic Analysis of Viral Diversity in Colorado Potato Beetle Public NGS Data.](https://pubmed.ncbi.nlm.nih.gov/36851611). Viruses, 2023.
- [Metagenome-validated combined amplicon sequencing and text mining-based annotations for simultaneous profiling of bacteria and fungi: vaginal microbiota and mycobiota in healthy women.](https://pubmed.ncbi.nlm.nih.gov/39731160). Microbiome, 2024.
- [Searching for vertically transmitted endosymbionts in over 4,000 wild strains of three self-fertilizing &lt,i&gt,Caenorhabditis&lt,/i&gt, species.](https://doi.org/10.17912/micropub.biology.002151). 2026.
- [TIPP-SD: A new method for species detection in microbiomes.](https://doi.org/10.1371/journal.pcbi.1014347). 2026.
- [Shotgun metagenomic dataset of leaf endophytic microbiome of the garden sage (Salvia officinalis L.).](https://doi.org/10.1186/s12863-026-01428-4). 2026.
- [metaFun: An analysis pipeline for metagenomic big data with fast and unified functional searches.](https://doi.org/10.1080/19490976.2025.2611544). 2026.
- [&lt,i&gt,In silico&lt,/i&gt, performance of a targeted enriched metagenomics approach to infer &lt,i&gt,Mycoplasma bovis&lt,/i&gt, strains in milk.](https://doi.org/10.3389/fvets.2026.1770245). 2026.
- [Refining filtering criteria of Kraken family of tools for accurate taxonomic profiling of ancient metagenomic data.](https://doi.org/10.3389/fmicb.2026.1603339). 2026.
- [Metrics for evaluating database selection techniques](https://doi.org/10.1023/A:1019241915635). Proceedings. Tenth International Workshop on Database and Expert Systems Applications. DEXA 99, 1999.
- [Cluster-Based Database Selection Techniques for Routing Bibliographic Queries](https://doi.org/10.1007/3-540-48309-8_9). International Conference on Database and Expert Systems Applications, 1999.
- [SWIRL: Selection of Workload-aware Indexes using Reinforcement Learning](https://doi.org/10.48786/edbt.2022.06). International Conference on Extending Database Technology, 2022.
- [Challenges in capturing the mycobiome from shotgun metagenome data: lack of software and databases](https://doi.org/10.1186/s40168-025-02048-3). Microbiome, 2025.
- [Comparative analysis of human gut bacterial microbiota between shallow shotgun metagenomic sequencing and full-length 16S rDNA amplicon sequencing](https://doi.org/10.5582/bst.2024.01393). Bioscience Trends, 2025.
- [Dermatological implications of alignment-based de-hosting and bioinformatics pipelines on shotgun microbiome analysis](https://doi.org/10.1186/s12967-025-07246-z). Journal of Translational Medicine, 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.