Designing a Reproducible Metagenomic Analysis Pipeline: A Template Using Nextflow and Containers
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Reproducibility in metagenomic analysis is critically undermined by unrecorded routine decisions regarding tool versions and parameters; this template mandates explicit recording of workflow logic, software environments via container digests (e.g., Docker/Singularity), and all tool parameters within version-controlled configuration files.
- The pipeline employs modular design, treating each stage (quality control, assembly, binning, annotation) as an independent process with defined inputs and outputs, facilitating testing, replacement, and incremental development.
- Containerization, utilizing pinned image digests (e.g.,
quay.io/biocontainers/fastp:0.23.4--h5f740d0_3), ensures software isolation, guaranteeing identical execution environments across different computational platforms and preventing issues arising from dependency conflicts or OS variations. - Key metagenomic analysis stages, including quality control with
fastp, taxonomic classification withkraken2, assembly withmegahit, binning withmetabat2, and annotation withprokka, are implemented as distinct Nextflow processes, each with specific container requirements and reproducibility controls. - Comprehensive documentation, including README files and inline comments, is integrated as part of the pipeline itself, alongside automated testing with a small, version-controlled test dataset, to ensure clarity, maintainability, and validation of the analysis workflow.
Scope and Reader Context
This article provides a working template for researchers, biology students, and laboratory professionals who need to build a reproducible shotgun metagenomic analysis pipeline. The template uses Nextflow for workflow orchestration and container technologies such as Docker or Singularity for software isolation. The practical outcome is a modular pipeline structure that covers quality control, assembly, binning, and annotation, with configuration files and documentation that allow another researcher to rerun the exact same analysis on the same inputs and obtain the same outputs.
Metagenomic analysis presents a distinct reproducibility problem. Unlike single-genome analysis, metagenomic datasets are large, the software ecosystem is fragmented, and small differences in tool versions or parameters can change biological conclusions. A 2026 review of 438 pig microbiome publications from 2003 to 2023 found that variability introduced by different DNA extraction methods, sequencing platforms, and bioinformatics pipelines had measurable impacts on reproducibility and cross-study comparability, and that metadata reporting for bioinformatic workflows was frequently incomplete. The template described here addresses the bioinformatics portion of that problem by making every step explicit, versioned, and executable.
The intended reader is assumed to have basic familiarity with the Linux command line, FASTQ files, and common metagenomic concepts such as operational taxonomic units and metagenome-assembled genomes. No prior Nextflow experience is required, but the reader should be prepared to install software and edit text files.
Why Reproducibility Fails in Metagenomic Analysis
Reproducibility failures in metagenomic pipelines rarely come from a single dramatic error. They accumulate from routine decisions that are not recorded. A researcher runs a quality trimming step with one tool version, later updates the tool, and the updated version changes default parameters. A second researcher uses the same script but a different operating system, and a dependency compiles differently. A third researcher receives the raw data but not the exact parameter file, and must guess at the trimming thresholds. Each of these situations produces results that differ in ways that are difficult to trace.
The core principle of a reproducible pipeline is that the analysis is defined entirely by files that are under version control, not by the state of a particular computer. The pipeline should specify three things explicitly: the workflow logic, the software environment for each step, and the parameters for each tool. Nextflow provides the workflow logic, containers provide the software environments, and a configuration file provides the parameters.
The nf-core documentation describes community standards for building such pipelines, including requirements for versioned containers, configuration files, and testing. These standards exist because the community has learned that informal approaches do not scale to the complexity of modern metagenomic analysis. A pipeline that works on one machine and fails on another is not reproducible, regardless of how carefully the biological steps were designed.
Containerization solves the software environment problem. A container image contains a complete filesystem with a specific operating system, specific tool versions, and all dependencies. When a process runs inside a container, it sees exactly the software that was present when the image was built. Docker is the most widely used container system, but many high-performance computing clusters do not allow Docker for security reasons. Singularity, now often called Apptainer, is designed for shared computing environments and can run Docker images without requiring root privileges. A well-designed Nextflow pipeline supports both, allowing the same workflow to run on a laptop and on a cluster.
At a Glance
| Pipeline Stage | Primary Tools | Container Requirement | Key Output | Reproducibility Control |
|---|---|---|---|---|
| Quality control and read preprocessing | FastQC, MultiQC, Cutadapt, fastp | Yes | Trimmed and filtered FASTQ files, QC report | Pin tool versions in container image, record trimming parameters in config file |
| Taxonomic classification | Kraken2, Bracken, MetaPhlAn | Yes | Taxonomic abundance table | Use fixed reference database version, record database path in config |
| Metagenome assembly | MEGAHIT, metaSPAdes | Yes | Assembled contigs in FASTA format | Record k-mer parameters and memory limits, use same assembler version |
| Binning and genome recovery | MetaBAT2, MaxBin2, CheckM | Yes | Metagenome-assembled genomes (MAGs) | Record completeness and contamination thresholds, use same CheckM database |
| Functional annotation | Prokka, eggNOG-mapper, DRAM | Yes | Annotated genome files and functional tables | Pin annotation database versions, record all database paths |
The table above shows the five core stages of a shotgun metagenomic pipeline. Each stage requires a containerized environment, produces a specific output, and has a reproducibility control that must be recorded. The template described in this article implements all five stages as modular Nextflow processes.
Core Principles of Pipeline Design
Modularity
A modular pipeline treats each analysis step as an independent process with defined inputs and outputs. The quality control process does not know or care how the assembly process works. It receives raw FASTQ files and produces trimmed FASTQ files. The assembly process receives trimmed FASTQ files and produces contigs. This separation allows individual processes to be tested, replaced, or rerun without disturbing the rest of the pipeline.
Modularity also supports incremental development. A researcher can start with only the quality control and taxonomic classification stages, verify that those work on their data, and then add assembly and binning later. The Bioconductor project applies a similar philosophy to R packages, where each package handles a specific analysis task and packages are combined into workflows. The same principle applies at the pipeline level.
Explicit Parameter Recording
Every parameter that affects the analysis must be recorded in a configuration file. This includes trimming quality thresholds, minimum read lengths, k-mer sizes for assembly, completeness thresholds for binning, and database paths for taxonomic classification. The configuration file is version controlled alongside the pipeline code.
The Galaxy Training Network emphasizes that reproducibility requires more than saving the final results. The complete analysis history, including every tool version and parameter setting, must be preserved. In a Nextflow pipeline, the configuration file serves this role. When the pipeline runs, Nextflow records the exact parameters used in the execution log, so even if the configuration file is later edited, the original run can be reconstructed.
Software Isolation
Each process in the pipeline runs inside a container that specifies the exact software environment. The container image is identified by a digest, which is a cryptographic hash of the image contents. Pinning the digest ensures that the same image is used every time, even if a newer version of the tool has been published.
The nf-core documentation requires that all pipeline processes use containers with pinned versions. This is a standard for community pipelines. The reason is practical: unpinned software versions are a common cause of irreproducible results in bioinformatics.
Documentation as Part of the Pipeline
Documentation is not a separate deliverable produced after the pipeline works. It is part of the pipeline itself. The repository should contain a README that explains how to install dependencies, how to configure the pipeline for new data, and how to interpret the outputs. Each process should have a comment block explaining what it does and why the parameters were chosen.
The EMBL-EBI Training resources emphasize that bioinformatics analysis is a learned skill that requires both conceptual understanding and practical experience. Documentation serves the same role for pipelines. A well-documented pipeline can be understood, modified, and trusted by researchers who did not write it.
The Nextflow Workflow Template
Pipeline Structure
The template pipeline is organized as a Nextflow project with the following directory structure:
metagenome-pipeline/
├── main.nf
├── nextflow.config
├── conf/
│ ├── base.config
│ ├── container.config
│ └── params.config
├── modules/
│ ├── qc.nf
│ ├── taxonomy.nf
│ ├── assembly.nf
│ ├── binning.nf
│ └── annotation.nf
├── bin/
│ └── helper_scripts.py
├── docs/
│ ├── installation.md
│ ├── usage.md
│ └── parameters.md
└── test_data/
└── sample_sheet.csv
The main.nf file defines the workflow, which is the sequence of processes and how their outputs are connected. The nextflow.config file specifies global settings such as the executor, the container engine, and default parameters. The conf directory contains environment-specific configuration files. The modules directory contains one file per pipeline stage, each defining one or more processes. The bin directory contains helper scripts. The docs directory contains the documentation. The test_data directory contains a small sample sheet for testing.
The Sample Sheet
The sample sheet is a CSV file that tells the pipeline which samples to process and where their raw data are located. A minimal sample sheet has three columns: sample ID, path to forward reads, and path to reverse reads. For single-end data, the reverse reads column is left empty.
sample_id,forward_reads,reverse_reads
sample_01,/data/raw/sample_01_R1.fastq.gz,/data/raw/sample_01_R2.fastq.gz
sample_02,/data/raw/sample_02_R1.fastq.gz,/data/raw/sample_02_R2.fastq.gz
The sample sheet is the primary input to the pipeline. It should be committed to version control along with the raw data, or at minimum stored in a location that is backed up and documented. The NCBI Sequence Read Archive provides a standard way to archive raw sequencing data, and the sample sheet should reference the SRA accession numbers when available.
Defining Processes
Each process in Nextflow is defined with a process block that specifies the input, output, script, and container. A minimal quality control process looks like this:
process FASTP {
container 'quay.io/biocontainers/fastp:0.23.4--h5f740d0_3'
input:
tuple val(sample_id), path(forward_reads), path(reverse_reads)
output:
tuple val(sample_id), path("${sample_id}_trimmed_R1.fastq.gz"), path("${sample_id}_trimmed_R2.fastq.gz")
path("${sample_id}_fastp.html")
script:
"""
fastp -i ${forward_reads} -I ${reverse_reads} \\
-o ${sample_id}_trimmed_R1.fastq.gz \\
-O ${sample_id}_trimmed_R2.fastq.gz \\
--html ${sample_id}_fastp.html \\
--qualified_quality_phred ${params.phred_quality} \\
--length_required ${params.min_length}
"""
}
The container directive pins the exact software version. The input and output blocks define the data flow. The script block contains the actual command. Parameters such as params.phred_quality and params.min_length are defined in the configuration file, not hardcoded in the process.
Connecting Processes in the Workflow
The workflow block in main.nf defines how processes are connected:
workflow {
read_pairs = Channel.fromPath(params.sample_sheet)
.splitCsv(header: true)
.map { row -> [row.sample_id, file(row.forward_reads), file(row.reverse_reads)] }
FASTP(read_pairs)
FASTP.out.trimmed_reads.set { trimmed_reads }
KRAKEN2(trimmed_reads)
MEGAHIT(trimmed_reads)
MEGAHIT.out.contigs.set { contigs }
METABAT2(contigs)
CHECKM(METABAT2.out.bins)
PROKKA(METABAT2.out.bins)
}
This workflow runs quality control on all samples, then runs taxonomic classification and assembly in parallel. Assembly is followed by binning, and binning is followed by quality assessment and annotation. The parallel execution of taxonomy and assembly is a deliberate design choice that uses computing resources efficiently.
Quality Control and Read Preprocessing
Why Quality Control Matters
Raw sequencing reads contain adapter sequences, low-quality bases, and technical artifacts. These artifacts can cause spurious taxonomic classifications, fragmented assemblies, and incorrect binning results. Quality control is the first and most important step in the pipeline because every downstream analysis depends on the quality of the input reads.
The Bioinformatics Strategy for 16S and 23S rRNA Metabarcoding Data describes a pipeline that integrates read merging, primer and adapter trimming, quality filtering, dereplication, and chimera removal for amplicon data. Shotgun metagenomic data require a similar but not identical approach. Adapter trimming and quality filtering are always needed. Dereplication and chimera removal are less critical for shotgun data but may be useful for certain analyses.
Implementing Quality Control in Nextflow
The quality control stage in the template pipeline uses fastp for read trimming and filtering, FastQC for per-sample quality reports, and MultiQC to aggregate the FastQC reports into a single HTML file. The fastp process is shown above. The FastQC and MultiQC processes follow the same pattern.
The key parameters for quality control are:
| Parameter | Default Value | Description |
|---|---|---|
params.phred_quality | 20 | Minimum Phred quality score for a base to be retained |
params.min_length | 50 | Minimum read length after trimming |
params.adapter_sequence | auto-detected | Adapter sequence to trim, auto-detection is usually reliable |
params.dedup | false | Whether to remove duplicate reads |
These parameters should be adjusted based on the sequencing platform and the specific research question. A researcher studying very short reads from degraded environmental DNA might use a lower minimum length threshold. A researcher studying full-length 16S rRNA genes would use different parameters than one studying shotgun metagenomes.
Recording Quality Control Decisions
The quality control parameters must be recorded in the configuration file and in the documentation. When the pipeline runs, Nextflow writes a log that includes the exact command executed for each process. This log is the definitive record of what was done.
The Galaxy Training Network provides tutorials that emphasize the importance of understanding quality metrics before proceeding with analysis. A researcher should inspect the MultiQC report before running the assembly stage. If the quality is poor, the assembly will be poor, and no amount of downstream analysis will fix it.
Taxonomic Classification
Reference-Based Classification
Taxonomic classification assigns sequencing reads to taxonomic groups by comparing them to a reference database. The most widely used tools for this purpose are Kraken2 and MetaPhlAn. Kraken2 uses exact k-mer matches against a database of reference genomes. MetaPhlAn uses clade-specific marker genes.
The choice between Kraken2 and MetaPhlAn depends on the research question. Kraken2 provides read-level classification and can detect organisms that are not well represented in the database. MetaPhlAn provides species-level abundance estimates and is more robust to false positives from closely related species. Many pipelines run both and compare the results.
The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study used shotgun metagenomics to complement targeted qPCR for pathogen detection. The metagenomic data provided broad taxonomic screening that could contextualize the qPCR results. This is a good example of how taxonomic classification is used in practice: as a complementary approach that provides broader information than targeted methods alone.
Database Management
The reference database is a critical component of taxonomic classification. Different versions of the same database can produce different results. The database version must be recorded in the configuration file, and the database itself should be stored in a fixed location that is accessible to the pipeline.
The NCBI provides reference databases for taxonomic classification, including RefSeq and GenBank. These databases are updated regularly, and the update schedule should be documented. A pipeline that uses a database from January 2025 will produce different results than one that uses a database from January 2026, even with identical reads and parameters.
Implementing Taxonomic Classification in Nextflow
The Kraken2 process in the template pipeline looks like this:
process KRAKEN2 {
container 'quay.io/biocontainers/kraken2:2.1.3--pl5321h7d875b9_4'
input:
tuple val(sample_id), path(forward_reads), path(reverse_reads)
output:
tuple val(sample_id), path("${sample_id}_kraken2_report.txt"), path("${sample_id}_kraken2_output.txt")
script:
"""
kraken2 --db ${params.kraken2_db} \\
--paired ${forward_reads} ${reverse_reads} \\
--report ${sample_id}_kraken2_report.txt \\
--output ${sample_id}_kraken2_output.txt \\
--confidence ${params.kraken2_confidence} \\
--threads ${task.cpus}
"""
}
The params.kraken2_db parameter points to the database directory. The params.kraken2_confidence parameter controls the classification confidence threshold. A higher confidence threshold reduces false positives but also reduces sensitivity.
Metagenome Assembly
Assembly Strategies
Metagenome assembly reconstructs genomic sequences from the short reads. Unlike single-genome assembly, metagenome assembly must handle a mixture of organisms with different abundances, genome sizes, and sequence compositions. This makes assembly more difficult and more sensitive to parameter choices.
The two most widely used metagenome assemblers are MEGAHIT and metaSPAdes. MEGAHIT is faster and uses less memory, making it suitable for large datasets. metaSPAdes produces more contiguous assemblies but requires more memory and computing time. The choice between them depends on the available computing resources and the research question.
The MINUUR pipeline used metagenome assembly followed by binning to recover metagenome-assembled genomes from host-associated microbiome reads. The pipeline recovered 42 high-quality MAGs with greater than 90 percent completeness and less than 5 percent contamination from 62 mosquito samples. This demonstrates that assembly and binning can recover useful biological information from metagenomic data, but the quality of the results depends on the sequencing depth and the assembly parameters.
Assembly Parameters
The key parameters for metagenome assembly are:
| Parameter | MEGAHIT Default | metaSPAdes Default | Description |
|---|---|---|---|
| K-mer range | 21-141 | auto | Range of k-mer sizes used for assembly |
| Minimum contig length | 200 | 200 | Contigs shorter than this are discarded |
| Memory limit | auto | auto | Maximum memory usage |
| Threads | all available | all available | Number of CPU threads |
The k-mer range is particularly important. Smaller k-mers are better for detecting low-abundance organisms, while larger k-mers produce more contiguous assemblies for high-abundance organisms. MEGAHIT uses a range of k-mers and integrates the results. metaSPAdes automatically selects k-mer sizes based on the read length and coverage.
Implementing Assembly in Nextflow
The MEGAHIT process in the template pipeline looks like this:
process MEGAHIT {
container 'quay.io/biocontainers/megahit:1.2.9--h5b5514e_2'
input:
tuple val(sample_id), path(forward_reads), path(reverse_reads)
output:
tuple val(sample_id), path("${sample_id}_contigs.fa")
script:
"""
megahit -1 ${forward_reads} -2 ${reverse_reads} \\
-o ${sample_id}_megahit \\
--min-contig-len ${params.min_contig_len} \\
--k-min ${params.k_min} \\
--k-max ${params.k_max} \\
--k-step ${params.k_step} \\
--memory ${params.megahit_memory} \\
--num-cpu-threads ${task.cpus}
mv ${sample_id}_megahit/final.contigs.fa ${sample_id}_contigs.fa
"""
}
The assembly process produces a FASTA file of contigs. This file is the input to the binning stage.
Binning and Genome Recovery
What Binning Does
Binning groups contigs into bins that represent putative genomes. The goal is to recover metagenome-assembled genomes (MAGs) from the assembled contigs. Binning algorithms use sequence composition and coverage information to cluster contigs that are likely to come from the same organism.
The most widely used binning tools are MetaBAT2, MaxBin2, and CONCOCT. Each tool uses a different algorithm, and they often produce different results. A common strategy is to run multiple binning tools and then use a bin refinement tool such as DAS Tool or MetaWRAP to combine the results.
The IMP pipeline demonstrated that automated binning could be integrated into a reproducible pipeline for metagenomic and metatranscriptomic data. The pipeline used iterative co-assembly and automated binning to enhance data usage and output quality. This integration is important because binning quality depends on assembly quality, and assembly quality depends on the input reads.
Assessing Bin Quality
Bin quality is assessed using completeness and contamination estimates. Completeness is the percentage of the genome that is present in the bin. Contamination is the percentage of the bin that comes from other organisms. CheckM is the standard tool for these estimates.
The MINUUR pipeline used a threshold of greater than 90 percent completeness and less than 5 percent contamination to define high-quality MAGs. These thresholds are commonly used in the field, but the appropriate thresholds depend on the research question. A study of metabolic potential might accept lower completeness, while a study that aims to describe new species would require higher quality.
Implementing Binning in Nextflow
The MetaBAT2 process in the template pipeline looks like this:
process METABAT2 {
container 'quay.io/biocontainers/metabat2:2.15--h985aee2_4'
input:
tuple val(sample_id), path(contigs)
output:
tuple val(sample_id), path("${sample_id}_bins/*.fa")
script:
"""
mkdir -p ${sample_id}_bins
metabat2 -i ${contigs} \\
-o ${sample_id}_bins/bin \\
--minContig ${params.min_contig_bin} \\
--minSamples ${params.min_samples} \\
--maxP ${params.max_p} \\
--seed ${params.seed}
"""
}
The params.seed parameter is important for reproducibility. MetaBAT2 uses a random seed for some of its calculations. Setting a fixed seed ensures that the same input produces the same output.
Functional Annotation
What Annotation Provides
Functional annotation assigns biological functions to genes in the assembled contigs or MAGs. This is the stage where the pipeline produces biologically interpretable results. Annotation tools identify open reading frames, compare them to functional databases, and assign Gene Ontology terms, Enzyme Commission numbers, or KEGG orthology groups.
The most widely used annotation tools are Prokka, eggNOG-mapper, and DRAM. Prokka is fast and works well for bacterial genomes. eggNOG-mapper uses the eggNOG database and provides more detailed functional annotations. DRAM is designed specifically for metagenomic data and provides curated metabolic pathway information.
The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study used shotgun metagenomics to provide broader microbial information than qPCR alone. Functional annotation would be the next step in such a study, providing information about what the microbial community is doing, beyond which organisms are present.
Annotation Databases
Functional annotation depends on reference databases, and database versions must be recorded. The eggNOG database is updated regularly, and different versions contain different numbers of orthologous groups. The KEGG database has a specific license that may restrict its use in some settings.
The Bioconductor project provides R packages for functional annotation and analysis of annotation results. These packages can be used downstream of the pipeline to perform enrichment analysis, pathway analysis, and visualization.
Implementing Annotation in Nextflow
The Prokka process in the template pipeline looks like this:
process PROKKA {
container 'quay.io/biocontainers/prokka:1.14.6--pl5321h9c3d4e4_5'
input:
tuple val(sample_id), path(bins)
output:
tuple val(sample_id), path("${sample_id}_prokka/*.gff"), path("${sample_id}_prokka/*.faa")
script:
"""
prokka ${bins} \\
--outdir ${sample_id}_prokka \\
--prefix ${sample_id} \\
--kingdom ${params.prokka_kingdom} \\
--cpus ${task.cpus} \\
--force
"""
}
The params.prokka_kingdom parameter specifies the taxonomic kingdom for the annotation. The default is Bacteria, but this can be changed to Archaea or Viruses depending on the sample.
Configuration and Environment Management
The Nextflow Configuration File
The nextflow.config file is the central configuration point for the pipeline. It defines the executor, the container engine, and all parameters. A minimal configuration file looks like this:
params {
sample_sheet = "test_data/sample_sheet.csv"
phred_quality = 20
min_length = 50
kraken2_db = "/data/databases/kraken2_standard_20250101"
kraken2_confidence = 0.1
min_contig_len = 200
k_min = 21
k_max = 141
k_step = 12
megahit_memory = "0.8"
min_contig_bin = 1500
min_samples = 1
max_p = 0.95
seed = 42
prokka_kingdom = "Bacteria"
}
process {
executor = 'local'
cpus = 4
memory = '8 GB'
}
docker {
enabled = true
}
The parameters are organized by pipeline stage. Each parameter has a default value that can be overridden on the command line or in an environment-specific configuration file.
Container Configuration
The container configuration specifies which container engine to use and where to find the container images. The container.config file in the conf directory provides environment-specific settings:
// conf/container.config
docker {
enabled = true
runOptions = '-u $(id -u):$(id -g)'
}
singularity {
enabled = false
}
For a high-performance computing cluster that does not allow Docker, the user would create a conf/cluster.config file that enables Singularity and specifies the cache directory:
// conf/cluster.config
singularity {
enabled = true
cacheDir = '/data/containers'
}
process {
executor = 'slurm'
queue = 'normal'
}
The user selects the appropriate configuration file when running the pipeline:
nextflow run main.nf -profile cluster
Environment-Specific Profiles
Nextflow profiles allow the same pipeline to run in different environments without changing the pipeline code. The nextflow.config file defines profiles:
profiles {
local {
includeConfig 'conf/base.config'
includeConfig 'conf/container.config'
}
cluster {
includeConfig 'conf/base.config'
includeConfig 'conf/cluster.config'
}
test {
includeConfig 'conf/base.config'
includeConfig 'conf/container.config'
params.sample_sheet = 'test_data/sample_sheet.csv'
params.kraken2_db = '/data/databases/kraken2_test'
}
}
The test profile is particularly important. It uses a small test dataset and a small test database, allowing the pipeline to be verified quickly before running on real data.
Testing and Validation
The Test Dataset
Every reproducible pipeline must include a test dataset. The test dataset should be small enough to run in a few minutes but large enough to exercise all pipeline stages. A minimal test dataset for a shotgun metagenomic pipeline might contain two samples with 100,000 read pairs each.
The test dataset should be committed to the pipeline repository. This ensures that the test is always available and that the expected outputs are known. The nf-core documentation requires that all community pipelines include test data and automated tests.
Automated Testing
Nextflow provides a -stub-run option that simulates the pipeline without actually running the tools. This is useful for testing the workflow logic, but it does not test the actual analysis. A more thorough test runs the pipeline on the test dataset and compares the outputs to expected results.
The test can be automated using a continuous integration system such as GitHub Actions. The pipeline repository can be configured to run the test whenever the code is changed. This catches errors early, before they affect real analyses.
Validation Against Known Results
The test dataset should include samples with known biological content. For example, a test dataset might include reads from a mock community with known species composition. The pipeline should recover the expected species in approximately the expected abundances.
The Bioinformatics Strategy for 16S and 23S rRNA Metabarcoding Data pipeline was validated using curated reference databases and produced reliable biological interpretations. A similar validation approach should be used for the shotgun metagenomic pipeline. The test dataset provides a benchmark for evaluating pipeline performance.
Records and Measurements
What to Record
The following records should be maintained for every pipeline run:
| Record | Content | Purpose |
|---|---|---|
| Pipeline version | Git commit hash of the pipeline code | Identifies the exact code version |
| Configuration file | The complete nextflow.config and profile files | Records all parameters |
| Container digests | The cryptographic digests of all container images | Identifies the exact software environments |
| Database versions | Version and download date of all reference databases | Identifies the exact reference data |
| Execution log | The Nextflow execution log | Records the exact commands and resource usage |
| Sample sheet | The CSV file with sample IDs and read paths | Identifies the input data |
| Quality report | The MultiQC HTML report | Records the quality of the input and trimmed reads |
These records should be stored alongside the pipeline outputs. A common practice is to create a directory for each run with the date and a short description:
runs/
├── 2026-01-15_salmon_farm/
│ ├── pipeline_version.txt
│ ├── nextflow.config
│ ├── container_digests.txt
│ ├── database_versions.txt
│ ├── execution_log.txt
│ ├── sample_sheet.csv
│ └── results/
Measuring Reproducibility
Reproducibility can be measured by running the same pipeline twice on the same input and comparing the outputs. The outputs should be identical if the pipeline is fully reproducible. Differences can arise from:
| Source of Variation | How to Control |
|---|---|
| Random number generator seeds | Set fixed seeds for all tools that use randomness |
| Software version differences | Pin container image digests |
| Database version differences | Use fixed database paths and record versions |
| Parallel execution order | Use deterministic algorithms or sort outputs |
| Floating point arithmetic | Rarely causes problems, but can be checked |
The IMP pipeline was designed to be reproducible and was encapsulated in Docker containers. The authors demonstrated that the pipeline could be run by other researchers with the same results. This is the standard that the template pipeline aims to meet.
Common Failure Patterns
Container Pull Failures
The most common failure when running a containerized pipeline is the inability to pull a container image. This can happen when the network is unavailable, the container registry is down, or the image does not exist for the specified architecture.
Diagnosis: The error message will indicate that the container image could not be pulled. Check the image name and digest in the process definition.
Resolution: Verify that the container registry is accessible. If using Singularity, check that the cache directory is writable. If the image is not available for the current architecture, use a different image or build the image locally.
Database Path Errors
The pipeline will fail if the reference database paths in the configuration file are incorrect. This is a common error when moving the pipeline to a new machine.
Diagnosis: The error message will indicate that a database file or directory was not found.
Resolution: Verify that the database paths in the configuration file are correct for the current machine. Update the paths in the environment-specific configuration file, not in the pipeline code.
Memory Exhaustion
Metagenome assembly and binning are memory-intensive. The pipeline will fail if the requested memory exceeds the available memory on the machine or cluster.
Diagnosis: The error message will indicate that the process was killed or that memory allocation failed.
Resolution: Reduce the memory request in the configuration file, or use a machine with more memory. For assembly, consider using MEGAHIT instead of metaSPAdes, as MEGAHIT uses less memory.
Parameter Mismatches
The pipeline will fail if a parameter value is invalid for a particular tool. For example, a k-mer size that is larger than the read length will cause the assembler to fail.
Diagnosis: The error message will indicate that a parameter value is invalid.
Resolution: Check the tool documentation for valid parameter ranges. Adjust the parameter values in the configuration file.
Sample Sheet Format Errors
The pipeline will fail if the sample sheet is not formatted correctly. Common errors include missing columns, incorrect file paths, and duplicate sample IDs.
Diagnosis: The error message will indicate that the sample sheet could not be parsed.
Resolution: Verify that the sample sheet has the correct columns and that all file paths are correct. Check for duplicate sample IDs.
Limitations and Interpretation
What the Pipeline Cannot Do
The template pipeline provides a starting point for metagenomic analysis, but it has limitations that must be understood before interpreting results.
Reference database bias: Taxonomic classification depends on the reference database. Organisms that are not in the database will not be detected. The NCBI databases are comprehensive but not complete, and novel organisms will be missed.
Assembly limitations: Metagenome assembly cannot reconstruct all genomes in a complex community. Low-abundance organisms, organisms with high sequence diversity, and organisms with repetitive genomes are difficult to assemble. The MINUUR pipeline found that samples with more assigned reads produced more MAGs, indicating that sequencing depth is a limiting factor.
Binning uncertainty: Binning is a computational prediction, not a biological fact. Bins can be incomplete, contaminated, or chimeric. The completeness and contamination estimates from CheckM are useful but not definitive.
Functional annotation gaps: Functional annotation depends on reference databases. Genes that are novel or highly divergent may not be annotated. The annotation is a prediction of function, not experimental evidence.
Interpreting Results with Caution
The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study provides a good example of cautious interpretation. The authors found subtle changes in microbial community composition associated with pathogen presence, but they noted that the patterns should be interpreted cautiously given the low pathogen loads. This caution is appropriate for all metagenomic results.
Metagenomic data provide correlative evidence, not causal evidence. If a particular organism or gene is associated with a condition, this does not prove that the organism or gene causes the condition. Follow-up experiments are needed to establish causation.
When to Escalate to Professional Support
The following situations warrant escalation to a bioinformatics specialist or the tool developers:
| Situation | Reason for Escalation |
|---|---|
| The pipeline produces different results on different machines | Indicates a reproducibility problem that requires expert diagnosis |
| The assembly produces very few or very fragmented contigs | May indicate a problem with the input data or assembly parameters |
| The binning produces no high-quality MAGs | May indicate insufficient sequencing depth or a problem with the assembly |
| The taxonomic classification produces unexpected results | May indicate a database problem or a biological phenomenon that requires expert interpretation |
| The pipeline fails with an error that cannot be resolved | The error may be a bug in the pipeline or a tool that requires expert attention |
The EMBL-EBI Training resources and the Galaxy Training Network provide additional learning materials for researchers who need to develop their bioinformatics skills. The The Carpentries lessons provide foundational training in shell, Git, and programming that is useful for working with pipelines.
Safety and Regulatory Context
Data Management
Metagenomic data may include human reads if the samples contain human DNA. This raises privacy concerns, and researchers must comply with applicable regulations for human data protection. The pipeline should include a step to remove human reads if this is a concern.
The NCBI provides guidance on data submission and data management. Researchers should be aware of the data sharing requirements for their funding sources and journals.
Computational Resources
Metagenomic analysis requires significant computational resources. Researchers should be aware of the resource usage of each pipeline stage and plan accordingly. The configuration file should specify resource limits to prevent a single process from consuming all available memory or disk space.
Reproducibility as a Professional Standard
Reproducibility is a professional standard in bioinformatics. The nf-core documentation describes the community standards for reproducible pipelines, and many journals now require that analysis code and data be made available. The template pipeline described in this article provides a way to meet these standards.
Frequently Asked Questions
What is the difference between Nextflow and Snakemake for metagenomic pipelines?
Nextflow and Snakemake are both workflow management systems that support containerization and reproducible analysis. Nextflow uses a Groovy-based syntax and has strong support for cloud computing and high-performance computing clusters. Snakemake uses a Python-based syntax and is often considered easier to learn for researchers with Python experience. The MINUUR pipeline used Snakemake, while the IMP pipeline used Python and Docker. Both tools can produce reproducible pipelines. The choice between them is largely a matter of preference and the computing environment.
How much sequencing depth is needed for metagenome assembly and binning?
The required sequencing depth depends on the complexity of the microbial community and the research question. The MINUUR pipeline found that samples with more assigned reads produced more MAGs, indicating that higher depth improves binning results. For simple communities, 5 to 10 million read pairs per sample may be sufficient. For complex communities such as soil, 50 million or more read pairs per sample may be needed. The appropriate depth should be determined based on the expected community complexity and the desired resolution.
Should I use Kraken2 or MetaPhlAn for taxonomic classification?
The choice depends on the research question. Kraken2 provides read-level classification and can detect organisms that are not well represented in the database. MetaPhlAn provides species-level abundance estimates and is more robust to false positives. Many pipelines run both and compare the results. The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study used shotgun metagenomics for broad taxonomic screening, which is a good use case for Kraken2. For detailed species-level abundance analysis, MetaPhlAn may be more appropriate.
How do I choose between MEGAHIT and metaSPAdes for assembly?
MEGAHIT is faster and uses less memory, making it suitable for large datasets and machines with limited memory. metaSPAdes produces more contiguous assemblies but requires more memory and computing time. The choice depends on the available resources and the research question. If the goal is to recover complete genomes, metaSPAdes may be worth the additional resources. If the goal is a broad survey of the community, MEGAHIT is often sufficient.
What completeness and contamination thresholds should I use for MAGs?
The commonly used thresholds for high-quality MAGs are greater than 90 percent completeness and less than 5 percent contamination, as used in the MINUUR pipeline. However, the appropriate thresholds depend on the research question. A study of metabolic potential might accept lower completeness, while a study that aims to describe new species would require higher quality. The thresholds should be recorded in the configuration file and reported in the results.
How do I handle samples with very low sequencing depth?
Samples with very low sequencing depth will produce poor assembly and binning results. The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study found that metagenomic results from samples with low pathogen loads should be interpreted cautiously. Options for low-depth samples include combining samples for co-assembly, using a reference-based approach instead of assembly, or excluding the samples from assembly and binning analyses.
What should I do if my pipeline produces different results on different machines?
This indicates a reproducibility problem. Check that the container images are the same by verifying the image digests. Check that the reference database versions are the same. Check that all random seeds are fixed. If the problem persists, escalate to a bioinformatics specialist. The nf-core documentation provides guidance on troubleshooting reproducibility issues.
How do I share my pipeline with other researchers?
The pipeline should be committed to a public repository such as GitHub. The repository should include the pipeline code, the configuration files, the documentation, and the test dataset. The README should explain how to install the dependencies and run the pipeline. The nf-core documentation provides standards for community pipelines that can serve as a model for sharing.
Related Bioinformatics Guides
- Metabolomics Data Analysis in R: A Practical Workflow
- Metagenomic Contamination Control: Best Practices for Clean Data
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Longitudinal Microbiome Data Analysis: Methods and Best Practices
- Spatial Transcriptomics Neighborhood Analysis: Tools and Best Practices
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Towards standardization in pig microbiome research based on a comprehensive twenty-year review.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.