Designing a Reproducible Metagenomic Analysis Pipeline: A Template Using Nextflow and Containers

By Dr. Zubair Khalid, DVM, MS, PhD ·

Designing a Reproducible Metagenomic Analysis Pipeline: A Template Using Nextflow and Containers

Key Takeaways

  • Reproducibility in metagenomic analysis is critically undermined by unrecorded routine decisions regarding tool versions and parameters; this template mandates explicit recording of workflow logic, software environments via container digests (e.g., Docker/Singularity), and all tool parameters within version-controlled configuration files.
  • The pipeline employs modular design, treating each stage (quality control, assembly, binning, annotation) as an independent process with defined inputs and outputs, facilitating testing, replacement, and incremental development.
  • Containerization, utilizing pinned image digests (e.g., quay.io/biocontainers/fastp:0.23.4--h5f740d0_3), ensures software isolation, guaranteeing identical execution environments across different computational platforms and preventing issues arising from dependency conflicts or OS variations.
  • Key metagenomic analysis stages, including quality control with fastp, taxonomic classification with kraken2, assembly with megahit, binning with metabat2, and annotation with prokka, are implemented as distinct Nextflow processes, each with specific container requirements and reproducibility controls.
  • Comprehensive documentation, including README files and inline comments, is integrated as part of the pipeline itself, alongside automated testing with a small, version-controlled test dataset, to ensure clarity, maintainability, and validation of the analysis workflow.

Scope and Reader Context

This article provides a working template for researchers, biology students, and laboratory professionals who need to build a reproducible shotgun metagenomic analysis pipeline. The template uses Nextflow for workflow orchestration and container technologies such as Docker or Singularity for software isolation. The practical outcome is a modular pipeline structure that covers quality control, assembly, binning, and annotation, with configuration files and documentation that allow another researcher to rerun the exact same analysis on the same inputs and obtain the same outputs.

Metagenomic analysis presents a distinct reproducibility problem. Unlike single-genome analysis, metagenomic datasets are large, the software ecosystem is fragmented, and small differences in tool versions or parameters can change biological conclusions. A 2026 review of 438 pig microbiome publications from 2003 to 2023 found that variability introduced by different DNA extraction methods, sequencing platforms, and bioinformatics pipelines had measurable impacts on reproducibility and cross-study comparability, and that metadata reporting for bioinformatic workflows was frequently incomplete. The template described here addresses the bioinformatics portion of that problem by making every step explicit, versioned, and executable.

The intended reader is assumed to have basic familiarity with the Linux command line, FASTQ files, and common metagenomic concepts such as operational taxonomic units and metagenome-assembled genomes. No prior Nextflow experience is required, but the reader should be prepared to install software and edit text files.

Why Reproducibility Fails in Metagenomic Analysis

Reproducibility failures in metagenomic pipelines rarely come from a single dramatic error. They accumulate from routine decisions that are not recorded. A researcher runs a quality trimming step with one tool version, later updates the tool, and the updated version changes default parameters. A second researcher uses the same script but a different operating system, and a dependency compiles differently. A third researcher receives the raw data but not the exact parameter file, and must guess at the trimming thresholds. Each of these situations produces results that differ in ways that are difficult to trace.

The core principle of a reproducible pipeline is that the analysis is defined entirely by files that are under version control, not by the state of a particular computer. The pipeline should specify three things explicitly: the workflow logic, the software environment for each step, and the parameters for each tool. Nextflow provides the workflow logic, containers provide the software environments, and a configuration file provides the parameters.

The nf-core documentation describes community standards for building such pipelines, including requirements for versioned containers, configuration files, and testing. These standards exist because the community has learned that informal approaches do not scale to the complexity of modern metagenomic analysis. A pipeline that works on one machine and fails on another is not reproducible, regardless of how carefully the biological steps were designed.

Containerization solves the software environment problem. A container image contains a complete filesystem with a specific operating system, specific tool versions, and all dependencies. When a process runs inside a container, it sees exactly the software that was present when the image was built. Docker is the most widely used container system, but many high-performance computing clusters do not allow Docker for security reasons. Singularity, now often called Apptainer, is designed for shared computing environments and can run Docker images without requiring root privileges. A well-designed Nextflow pipeline supports both, allowing the same workflow to run on a laptop and on a cluster.

At a Glance

Pipeline StagePrimary ToolsContainer RequirementKey OutputReproducibility Control
Quality control and read preprocessingFastQC, MultiQC, Cutadapt, fastpYesTrimmed and filtered FASTQ files, QC reportPin tool versions in container image, record trimming parameters in config file
Taxonomic classificationKraken2, Bracken, MetaPhlAnYesTaxonomic abundance tableUse fixed reference database version, record database path in config
Metagenome assemblyMEGAHIT, metaSPAdesYesAssembled contigs in FASTA formatRecord k-mer parameters and memory limits, use same assembler version
Binning and genome recoveryMetaBAT2, MaxBin2, CheckMYesMetagenome-assembled genomes (MAGs)Record completeness and contamination thresholds, use same CheckM database
Functional annotationProkka, eggNOG-mapper, DRAMYesAnnotated genome files and functional tablesPin annotation database versions, record all database paths

The table above shows the five core stages of a shotgun metagenomic pipeline. Each stage requires a containerized environment, produces a specific output, and has a reproducibility control that must be recorded. The template described in this article implements all five stages as modular Nextflow processes.

Core Principles of Pipeline Design

Modularity

A modular pipeline treats each analysis step as an independent process with defined inputs and outputs. The quality control process does not know or care how the assembly process works. It receives raw FASTQ files and produces trimmed FASTQ files. The assembly process receives trimmed FASTQ files and produces contigs. This separation allows individual processes to be tested, replaced, or rerun without disturbing the rest of the pipeline.

Modularity also supports incremental development. A researcher can start with only the quality control and taxonomic classification stages, verify that those work on their data, and then add assembly and binning later. The Bioconductor project applies a similar philosophy to R packages, where each package handles a specific analysis task and packages are combined into workflows. The same principle applies at the pipeline level.

Explicit Parameter Recording

Every parameter that affects the analysis must be recorded in a configuration file. This includes trimming quality thresholds, minimum read lengths, k-mer sizes for assembly, completeness thresholds for binning, and database paths for taxonomic classification. The configuration file is version controlled alongside the pipeline code.

The Galaxy Training Network emphasizes that reproducibility requires more than saving the final results. The complete analysis history, including every tool version and parameter setting, must be preserved. In a Nextflow pipeline, the configuration file serves this role. When the pipeline runs, Nextflow records the exact parameters used in the execution log, so even if the configuration file is later edited, the original run can be reconstructed.

Software Isolation

Each process in the pipeline runs inside a container that specifies the exact software environment. The container image is identified by a digest, which is a cryptographic hash of the image contents. Pinning the digest ensures that the same image is used every time, even if a newer version of the tool has been published.

The nf-core documentation requires that all pipeline processes use containers with pinned versions. This is a standard for community pipelines. The reason is practical: unpinned software versions are a common cause of irreproducible results in bioinformatics.

Documentation as Part of the Pipeline

Documentation is not a separate deliverable produced after the pipeline works. It is part of the pipeline itself. The repository should contain a README that explains how to install dependencies, how to configure the pipeline for new data, and how to interpret the outputs. Each process should have a comment block explaining what it does and why the parameters were chosen.

The EMBL-EBI Training resources emphasize that bioinformatics analysis is a learned skill that requires both conceptual understanding and practical experience. Documentation serves the same role for pipelines. A well-documented pipeline can be understood, modified, and trusted by researchers who did not write it.

The Nextflow Workflow Template

Pipeline Structure

The template pipeline is organized as a Nextflow project with the following directory structure:

metagenome-pipeline/
├── main.nf
├── nextflow.config
├── conf/
│   ├── base.config
│   ├── container.config
│   └── params.config
├── modules/
│   ├── qc.nf
│   ├── taxonomy.nf
│   ├── assembly.nf
│   ├── binning.nf
│   └── annotation.nf
├── bin/
│   └── helper_scripts.py
├── docs/
│   ├── installation.md
│   ├── usage.md
│   └── parameters.md
└── test_data/
    └── sample_sheet.csv

The main.nf file defines the workflow, which is the sequence of processes and how their outputs are connected. The nextflow.config file specifies global settings such as the executor, the container engine, and default parameters. The conf directory contains environment-specific configuration files. The modules directory contains one file per pipeline stage, each defining one or more processes. The bin directory contains helper scripts. The docs directory contains the documentation. The test_data directory contains a small sample sheet for testing.

The Sample Sheet

The sample sheet is a CSV file that tells the pipeline which samples to process and where their raw data are located. A minimal sample sheet has three columns: sample ID, path to forward reads, and path to reverse reads. For single-end data, the reverse reads column is left empty.

sample_id,forward_reads,reverse_reads
sample_01,/data/raw/sample_01_R1.fastq.gz,/data/raw/sample_01_R2.fastq.gz
sample_02,/data/raw/sample_02_R1.fastq.gz,/data/raw/sample_02_R2.fastq.gz

The sample sheet is the primary input to the pipeline. It should be committed to version control along with the raw data, or at minimum stored in a location that is backed up and documented. The NCBI Sequence Read Archive provides a standard way to archive raw sequencing data, and the sample sheet should reference the SRA accession numbers when available.

Defining Processes

Each process in Nextflow is defined with a process block that specifies the input, output, script, and container. A minimal quality control process looks like this:

process FASTP {
    container 'quay.io/biocontainers/fastp:0.23.4--h5f740d0_3'

    input:
    tuple val(sample_id), path(forward_reads), path(reverse_reads)

    output:
    tuple val(sample_id), path("${sample_id}_trimmed_R1.fastq.gz"), path("${sample_id}_trimmed_R2.fastq.gz")
    path("${sample_id}_fastp.html")

    script:
    """
    fastp -i ${forward_reads} -I ${reverse_reads} \\
          -o ${sample_id}_trimmed_R1.fastq.gz \\
          -O ${sample_id}_trimmed_R2.fastq.gz \\
          --html ${sample_id}_fastp.html \\
          --qualified_quality_phred ${params.phred_quality} \\
          --length_required ${params.min_length}
    """
}

The container directive pins the exact software version. The input and output blocks define the data flow. The script block contains the actual command. Parameters such as params.phred_quality and params.min_length are defined in the configuration file, not hardcoded in the process.

Connecting Processes in the Workflow

The workflow block in main.nf defines how processes are connected:

workflow {
    read_pairs = Channel.fromPath(params.sample_sheet)
        .splitCsv(header: true)
        .map { row -> [row.sample_id, file(row.forward_reads), file(row.reverse_reads)] }

    FASTP(read_pairs)
    FASTP.out.trimmed_reads.set { trimmed_reads }

    KRAKEN2(trimmed_reads)
    MEGAHIT(trimmed_reads)

    MEGAHIT.out.contigs.set { contigs }
    METABAT2(contigs)

    CHECKM(METABAT2.out.bins)
    PROKKA(METABAT2.out.bins)
}

This workflow runs quality control on all samples, then runs taxonomic classification and assembly in parallel. Assembly is followed by binning, and binning is followed by quality assessment and annotation. The parallel execution of taxonomy and assembly is a deliberate design choice that uses computing resources efficiently.

Quality Control and Read Preprocessing

Why Quality Control Matters

Raw sequencing reads contain adapter sequences, low-quality bases, and technical artifacts. These artifacts can cause spurious taxonomic classifications, fragmented assemblies, and incorrect binning results. Quality control is the first and most important step in the pipeline because every downstream analysis depends on the quality of the input reads.

The Bioinformatics Strategy for 16S and 23S rRNA Metabarcoding Data describes a pipeline that integrates read merging, primer and adapter trimming, quality filtering, dereplication, and chimera removal for amplicon data. Shotgun metagenomic data require a similar but not identical approach. Adapter trimming and quality filtering are always needed. Dereplication and chimera removal are less critical for shotgun data but may be useful for certain analyses.

Implementing Quality Control in Nextflow

The quality control stage in the template pipeline uses fastp for read trimming and filtering, FastQC for per-sample quality reports, and MultiQC to aggregate the FastQC reports into a single HTML file. The fastp process is shown above. The FastQC and MultiQC processes follow the same pattern.

The key parameters for quality control are:

ParameterDefault ValueDescription
params.phred_quality20Minimum Phred quality score for a base to be retained
params.min_length50Minimum read length after trimming
params.adapter_sequenceauto-detectedAdapter sequence to trim, auto-detection is usually reliable
params.dedupfalseWhether to remove duplicate reads

These parameters should be adjusted based on the sequencing platform and the specific research question. A researcher studying very short reads from degraded environmental DNA might use a lower minimum length threshold. A researcher studying full-length 16S rRNA genes would use different parameters than one studying shotgun metagenomes.

Recording Quality Control Decisions

The quality control parameters must be recorded in the configuration file and in the documentation. When the pipeline runs, Nextflow writes a log that includes the exact command executed for each process. This log is the definitive record of what was done.

The Galaxy Training Network provides tutorials that emphasize the importance of understanding quality metrics before proceeding with analysis. A researcher should inspect the MultiQC report before running the assembly stage. If the quality is poor, the assembly will be poor, and no amount of downstream analysis will fix it.

Taxonomic Classification

Reference-Based Classification

Taxonomic classification assigns sequencing reads to taxonomic groups by comparing them to a reference database. The most widely used tools for this purpose are Kraken2 and MetaPhlAn. Kraken2 uses exact k-mer matches against a database of reference genomes. MetaPhlAn uses clade-specific marker genes.

The choice between Kraken2 and MetaPhlAn depends on the research question. Kraken2 provides read-level classification and can detect organisms that are not well represented in the database. MetaPhlAn provides species-level abundance estimates and is more robust to false positives from closely related species. Many pipelines run both and compare the results.

The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study used shotgun metagenomics to complement targeted qPCR for pathogen detection. The metagenomic data provided broad taxonomic screening that could contextualize the qPCR results. This is a good example of how taxonomic classification is used in practice: as a complementary approach that provides broader information than targeted methods alone.

Database Management

The reference database is a critical component of taxonomic classification. Different versions of the same database can produce different results. The database version must be recorded in the configuration file, and the database itself should be stored in a fixed location that is accessible to the pipeline.

The NCBI provides reference databases for taxonomic classification, including RefSeq and GenBank. These databases are updated regularly, and the update schedule should be documented. A pipeline that uses a database from January 2025 will produce different results than one that uses a database from January 2026, even with identical reads and parameters.

Implementing Taxonomic Classification in Nextflow

The Kraken2 process in the template pipeline looks like this:

process KRAKEN2 {
    container 'quay.io/biocontainers/kraken2:2.1.3--pl5321h7d875b9_4'

    input:
    tuple val(sample_id), path(forward_reads), path(reverse_reads)

    output:
    tuple val(sample_id), path("${sample_id}_kraken2_report.txt"), path("${sample_id}_kraken2_output.txt")

    script:
    """
    kraken2 --db ${params.kraken2_db} \\
            --paired ${forward_reads} ${reverse_reads} \\
            --report ${sample_id}_kraken2_report.txt \\
            --output ${sample_id}_kraken2_output.txt \\
            --confidence ${params.kraken2_confidence} \\
            --threads ${task.cpus}
    """
}

The params.kraken2_db parameter points to the database directory. The params.kraken2_confidence parameter controls the classification confidence threshold. A higher confidence threshold reduces false positives but also reduces sensitivity.

Metagenome Assembly

Assembly Strategies

Metagenome assembly reconstructs genomic sequences from the short reads. Unlike single-genome assembly, metagenome assembly must handle a mixture of organisms with different abundances, genome sizes, and sequence compositions. This makes assembly more difficult and more sensitive to parameter choices.

The two most widely used metagenome assemblers are MEGAHIT and metaSPAdes. MEGAHIT is faster and uses less memory, making it suitable for large datasets. metaSPAdes produces more contiguous assemblies but requires more memory and computing time. The choice between them depends on the available computing resources and the research question.

The MINUUR pipeline used metagenome assembly followed by binning to recover metagenome-assembled genomes from host-associated microbiome reads. The pipeline recovered 42 high-quality MAGs with greater than 90 percent completeness and less than 5 percent contamination from 62 mosquito samples. This demonstrates that assembly and binning can recover useful biological information from metagenomic data, but the quality of the results depends on the sequencing depth and the assembly parameters.

Assembly Parameters

The key parameters for metagenome assembly are:

ParameterMEGAHIT DefaultmetaSPAdes DefaultDescription
K-mer range21-141autoRange of k-mer sizes used for assembly
Minimum contig length200200Contigs shorter than this are discarded
Memory limitautoautoMaximum memory usage
Threadsall availableall availableNumber of CPU threads

The k-mer range is particularly important. Smaller k-mers are better for detecting low-abundance organisms, while larger k-mers produce more contiguous assemblies for high-abundance organisms. MEGAHIT uses a range of k-mers and integrates the results. metaSPAdes automatically selects k-mer sizes based on the read length and coverage.

Implementing Assembly in Nextflow

The MEGAHIT process in the template pipeline looks like this:

process MEGAHIT {
    container 'quay.io/biocontainers/megahit:1.2.9--h5b5514e_2'

    input:
    tuple val(sample_id), path(forward_reads), path(reverse_reads)

    output:
    tuple val(sample_id), path("${sample_id}_contigs.fa")

    script:
    """
    megahit -1 ${forward_reads} -2 ${reverse_reads} \\
            -o ${sample_id}_megahit \\
            --min-contig-len ${params.min_contig_len} \\
            --k-min ${params.k_min} \\
            --k-max ${params.k_max} \\
            --k-step ${params.k_step} \\
            --memory ${params.megahit_memory} \\
            --num-cpu-threads ${task.cpus}
    mv ${sample_id}_megahit/final.contigs.fa ${sample_id}_contigs.fa
    """
}

The assembly process produces a FASTA file of contigs. This file is the input to the binning stage.

Binning and Genome Recovery

What Binning Does

Binning groups contigs into bins that represent putative genomes. The goal is to recover metagenome-assembled genomes (MAGs) from the assembled contigs. Binning algorithms use sequence composition and coverage information to cluster contigs that are likely to come from the same organism.

The most widely used binning tools are MetaBAT2, MaxBin2, and CONCOCT. Each tool uses a different algorithm, and they often produce different results. A common strategy is to run multiple binning tools and then use a bin refinement tool such as DAS Tool or MetaWRAP to combine the results.

The IMP pipeline demonstrated that automated binning could be integrated into a reproducible pipeline for metagenomic and metatranscriptomic data. The pipeline used iterative co-assembly and automated binning to enhance data usage and output quality. This integration is important because binning quality depends on assembly quality, and assembly quality depends on the input reads.

Assessing Bin Quality

Bin quality is assessed using completeness and contamination estimates. Completeness is the percentage of the genome that is present in the bin. Contamination is the percentage of the bin that comes from other organisms. CheckM is the standard tool for these estimates.

The MINUUR pipeline used a threshold of greater than 90 percent completeness and less than 5 percent contamination to define high-quality MAGs. These thresholds are commonly used in the field, but the appropriate thresholds depend on the research question. A study of metabolic potential might accept lower completeness, while a study that aims to describe new species would require higher quality.

Implementing Binning in Nextflow

The MetaBAT2 process in the template pipeline looks like this:

process METABAT2 {
    container 'quay.io/biocontainers/metabat2:2.15--h985aee2_4'

    input:
    tuple val(sample_id), path(contigs)

    output:
    tuple val(sample_id), path("${sample_id}_bins/*.fa")

    script:
    """
    mkdir -p ${sample_id}_bins
    metabat2 -i ${contigs} \\
             -o ${sample_id}_bins/bin \\
             --minContig ${params.min_contig_bin} \\
             --minSamples ${params.min_samples} \\
             --maxP ${params.max_p} \\
             --seed ${params.seed}
    """
}

The params.seed parameter is important for reproducibility. MetaBAT2 uses a random seed for some of its calculations. Setting a fixed seed ensures that the same input produces the same output.

Functional Annotation

What Annotation Provides

Functional annotation assigns biological functions to genes in the assembled contigs or MAGs. This is the stage where the pipeline produces biologically interpretable results. Annotation tools identify open reading frames, compare them to functional databases, and assign Gene Ontology terms, Enzyme Commission numbers, or KEGG orthology groups.

The most widely used annotation tools are Prokka, eggNOG-mapper, and DRAM. Prokka is fast and works well for bacterial genomes. eggNOG-mapper uses the eggNOG database and provides more detailed functional annotations. DRAM is designed specifically for metagenomic data and provides curated metabolic pathway information.

The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study used shotgun metagenomics to provide broader microbial information than qPCR alone. Functional annotation would be the next step in such a study, providing information about what the microbial community is doing, beyond which organisms are present.

Annotation Databases

Functional annotation depends on reference databases, and database versions must be recorded. The eggNOG database is updated regularly, and different versions contain different numbers of orthologous groups. The KEGG database has a specific license that may restrict its use in some settings.

The Bioconductor project provides R packages for functional annotation and analysis of annotation results. These packages can be used downstream of the pipeline to perform enrichment analysis, pathway analysis, and visualization.

Implementing Annotation in Nextflow

The Prokka process in the template pipeline looks like this:

process PROKKA {
    container 'quay.io/biocontainers/prokka:1.14.6--pl5321h9c3d4e4_5'

    input:
    tuple val(sample_id), path(bins)

    output:
    tuple val(sample_id), path("${sample_id}_prokka/*.gff"), path("${sample_id}_prokka/*.faa")

    script:
    """
    prokka ${bins} \\
           --outdir ${sample_id}_prokka \\
           --prefix ${sample_id} \\
           --kingdom ${params.prokka_kingdom} \\
           --cpus ${task.cpus} \\
           --force
    """
}

The params.prokka_kingdom parameter specifies the taxonomic kingdom for the annotation. The default is Bacteria, but this can be changed to Archaea or Viruses depending on the sample.

Configuration and Environment Management

The Nextflow Configuration File

The nextflow.config file is the central configuration point for the pipeline. It defines the executor, the container engine, and all parameters. A minimal configuration file looks like this:

params {
    sample_sheet = "test_data/sample_sheet.csv"
    phred_quality = 20
    min_length = 50
    kraken2_db = "/data/databases/kraken2_standard_20250101"
    kraken2_confidence = 0.1
    min_contig_len = 200
    k_min = 21
    k_max = 141
    k_step = 12
    megahit_memory = "0.8"
    min_contig_bin = 1500
    min_samples = 1
    max_p = 0.95
    seed = 42
    prokka_kingdom = "Bacteria"
}

process {
    executor = 'local'
    cpus = 4
    memory = '8 GB'
}

docker {
    enabled = true
}

The parameters are organized by pipeline stage. Each parameter has a default value that can be overridden on the command line or in an environment-specific configuration file.

Container Configuration

The container configuration specifies which container engine to use and where to find the container images. The container.config file in the conf directory provides environment-specific settings:

// conf/container.config
docker {
    enabled = true
    runOptions = '-u $(id -u):$(id -g)'
}

singularity {
    enabled = false
}

For a high-performance computing cluster that does not allow Docker, the user would create a conf/cluster.config file that enables Singularity and specifies the cache directory:

// conf/cluster.config
singularity {
    enabled = true
    cacheDir = '/data/containers'
}

process {
    executor = 'slurm'
    queue = 'normal'
}

The user selects the appropriate configuration file when running the pipeline:

nextflow run main.nf -profile cluster

Environment-Specific Profiles

Nextflow profiles allow the same pipeline to run in different environments without changing the pipeline code. The nextflow.config file defines profiles:

profiles {
    local {
        includeConfig 'conf/base.config'
        includeConfig 'conf/container.config'
    }
    cluster {
        includeConfig 'conf/base.config'
        includeConfig 'conf/cluster.config'
    }
    test {
        includeConfig 'conf/base.config'
        includeConfig 'conf/container.config'
        params.sample_sheet = 'test_data/sample_sheet.csv'
        params.kraken2_db = '/data/databases/kraken2_test'
    }
}

The test profile is particularly important. It uses a small test dataset and a small test database, allowing the pipeline to be verified quickly before running on real data.

Testing and Validation

The Test Dataset

Every reproducible pipeline must include a test dataset. The test dataset should be small enough to run in a few minutes but large enough to exercise all pipeline stages. A minimal test dataset for a shotgun metagenomic pipeline might contain two samples with 100,000 read pairs each.

The test dataset should be committed to the pipeline repository. This ensures that the test is always available and that the expected outputs are known. The nf-core documentation requires that all community pipelines include test data and automated tests.

Automated Testing

Nextflow provides a -stub-run option that simulates the pipeline without actually running the tools. This is useful for testing the workflow logic, but it does not test the actual analysis. A more thorough test runs the pipeline on the test dataset and compares the outputs to expected results.

The test can be automated using a continuous integration system such as GitHub Actions. The pipeline repository can be configured to run the test whenever the code is changed. This catches errors early, before they affect real analyses.

Validation Against Known Results

The test dataset should include samples with known biological content. For example, a test dataset might include reads from a mock community with known species composition. The pipeline should recover the expected species in approximately the expected abundances.

The Bioinformatics Strategy for 16S and 23S rRNA Metabarcoding Data pipeline was validated using curated reference databases and produced reliable biological interpretations. A similar validation approach should be used for the shotgun metagenomic pipeline. The test dataset provides a benchmark for evaluating pipeline performance.

Records and Measurements

What to Record

The following records should be maintained for every pipeline run:

RecordContentPurpose
Pipeline versionGit commit hash of the pipeline codeIdentifies the exact code version
Configuration fileThe complete nextflow.config and profile filesRecords all parameters
Container digestsThe cryptographic digests of all container imagesIdentifies the exact software environments
Database versionsVersion and download date of all reference databasesIdentifies the exact reference data
Execution logThe Nextflow execution logRecords the exact commands and resource usage
Sample sheetThe CSV file with sample IDs and read pathsIdentifies the input data
Quality reportThe MultiQC HTML reportRecords the quality of the input and trimmed reads

These records should be stored alongside the pipeline outputs. A common practice is to create a directory for each run with the date and a short description:

runs/
├── 2026-01-15_salmon_farm/
│   ├── pipeline_version.txt
│   ├── nextflow.config
│   ├── container_digests.txt
│   ├── database_versions.txt
│   ├── execution_log.txt
│   ├── sample_sheet.csv
│   └── results/

Measuring Reproducibility

Reproducibility can be measured by running the same pipeline twice on the same input and comparing the outputs. The outputs should be identical if the pipeline is fully reproducible. Differences can arise from:

Source of VariationHow to Control
Random number generator seedsSet fixed seeds for all tools that use randomness
Software version differencesPin container image digests
Database version differencesUse fixed database paths and record versions
Parallel execution orderUse deterministic algorithms or sort outputs
Floating point arithmeticRarely causes problems, but can be checked

The IMP pipeline was designed to be reproducible and was encapsulated in Docker containers. The authors demonstrated that the pipeline could be run by other researchers with the same results. This is the standard that the template pipeline aims to meet.

Common Failure Patterns

Container Pull Failures

The most common failure when running a containerized pipeline is the inability to pull a container image. This can happen when the network is unavailable, the container registry is down, or the image does not exist for the specified architecture.

Diagnosis: The error message will indicate that the container image could not be pulled. Check the image name and digest in the process definition.

Resolution: Verify that the container registry is accessible. If using Singularity, check that the cache directory is writable. If the image is not available for the current architecture, use a different image or build the image locally.

Database Path Errors

The pipeline will fail if the reference database paths in the configuration file are incorrect. This is a common error when moving the pipeline to a new machine.

Diagnosis: The error message will indicate that a database file or directory was not found.

Resolution: Verify that the database paths in the configuration file are correct for the current machine. Update the paths in the environment-specific configuration file, not in the pipeline code.

Memory Exhaustion

Metagenome assembly and binning are memory-intensive. The pipeline will fail if the requested memory exceeds the available memory on the machine or cluster.

Diagnosis: The error message will indicate that the process was killed or that memory allocation failed.

Resolution: Reduce the memory request in the configuration file, or use a machine with more memory. For assembly, consider using MEGAHIT instead of metaSPAdes, as MEGAHIT uses less memory.

Parameter Mismatches

The pipeline will fail if a parameter value is invalid for a particular tool. For example, a k-mer size that is larger than the read length will cause the assembler to fail.

Diagnosis: The error message will indicate that a parameter value is invalid.

Resolution: Check the tool documentation for valid parameter ranges. Adjust the parameter values in the configuration file.

Sample Sheet Format Errors

The pipeline will fail if the sample sheet is not formatted correctly. Common errors include missing columns, incorrect file paths, and duplicate sample IDs.

Diagnosis: The error message will indicate that the sample sheet could not be parsed.

Resolution: Verify that the sample sheet has the correct columns and that all file paths are correct. Check for duplicate sample IDs.

Limitations and Interpretation

What the Pipeline Cannot Do

The template pipeline provides a starting point for metagenomic analysis, but it has limitations that must be understood before interpreting results.

Reference database bias: Taxonomic classification depends on the reference database. Organisms that are not in the database will not be detected. The NCBI databases are comprehensive but not complete, and novel organisms will be missed.

Assembly limitations: Metagenome assembly cannot reconstruct all genomes in a complex community. Low-abundance organisms, organisms with high sequence diversity, and organisms with repetitive genomes are difficult to assemble. The MINUUR pipeline found that samples with more assigned reads produced more MAGs, indicating that sequencing depth is a limiting factor.

Binning uncertainty: Binning is a computational prediction, not a biological fact. Bins can be incomplete, contaminated, or chimeric. The completeness and contamination estimates from CheckM are useful but not definitive.

Functional annotation gaps: Functional annotation depends on reference databases. Genes that are novel or highly divergent may not be annotated. The annotation is a prediction of function, not experimental evidence.

Interpreting Results with Caution

The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study provides a good example of cautious interpretation. The authors found subtle changes in microbial community composition associated with pathogen presence, but they noted that the patterns should be interpreted cautiously given the low pathogen loads. This caution is appropriate for all metagenomic results.

Metagenomic data provide correlative evidence, not causal evidence. If a particular organism or gene is associated with a condition, this does not prove that the organism or gene causes the condition. Follow-up experiments are needed to establish causation.

When to Escalate to Professional Support

The following situations warrant escalation to a bioinformatics specialist or the tool developers:

SituationReason for Escalation
The pipeline produces different results on different machinesIndicates a reproducibility problem that requires expert diagnosis
The assembly produces very few or very fragmented contigsMay indicate a problem with the input data or assembly parameters
The binning produces no high-quality MAGsMay indicate insufficient sequencing depth or a problem with the assembly
The taxonomic classification produces unexpected resultsMay indicate a database problem or a biological phenomenon that requires expert interpretation
The pipeline fails with an error that cannot be resolvedThe error may be a bug in the pipeline or a tool that requires expert attention

The EMBL-EBI Training resources and the Galaxy Training Network provide additional learning materials for researchers who need to develop their bioinformatics skills. The The Carpentries lessons provide foundational training in shell, Git, and programming that is useful for working with pipelines.

Safety and Regulatory Context

Data Management

Metagenomic data may include human reads if the samples contain human DNA. This raises privacy concerns, and researchers must comply with applicable regulations for human data protection. The pipeline should include a step to remove human reads if this is a concern.

The NCBI provides guidance on data submission and data management. Researchers should be aware of the data sharing requirements for their funding sources and journals.

Computational Resources

Metagenomic analysis requires significant computational resources. Researchers should be aware of the resource usage of each pipeline stage and plan accordingly. The configuration file should specify resource limits to prevent a single process from consuming all available memory or disk space.

Reproducibility as a Professional Standard

Reproducibility is a professional standard in bioinformatics. The nf-core documentation describes the community standards for reproducible pipelines, and many journals now require that analysis code and data be made available. The template pipeline described in this article provides a way to meet these standards.

Frequently Asked Questions

What is the difference between Nextflow and Snakemake for metagenomic pipelines?

Nextflow and Snakemake are both workflow management systems that support containerization and reproducible analysis. Nextflow uses a Groovy-based syntax and has strong support for cloud computing and high-performance computing clusters. Snakemake uses a Python-based syntax and is often considered easier to learn for researchers with Python experience. The MINUUR pipeline used Snakemake, while the IMP pipeline used Python and Docker. Both tools can produce reproducible pipelines. The choice between them is largely a matter of preference and the computing environment.

How much sequencing depth is needed for metagenome assembly and binning?

The required sequencing depth depends on the complexity of the microbial community and the research question. The MINUUR pipeline found that samples with more assigned reads produced more MAGs, indicating that higher depth improves binning results. For simple communities, 5 to 10 million read pairs per sample may be sufficient. For complex communities such as soil, 50 million or more read pairs per sample may be needed. The appropriate depth should be determined based on the expected community complexity and the desired resolution.

Should I use Kraken2 or MetaPhlAn for taxonomic classification?

The choice depends on the research question. Kraken2 provides read-level classification and can detect organisms that are not well represented in the database. MetaPhlAn provides species-level abundance estimates and is more robust to false positives. Many pipelines run both and compare the results. The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study used shotgun metagenomics for broad taxonomic screening, which is a good use case for Kraken2. For detailed species-level abundance analysis, MetaPhlAn may be more appropriate.

How do I choose between MEGAHIT and metaSPAdes for assembly?

MEGAHIT is faster and uses less memory, making it suitable for large datasets and machines with limited memory. metaSPAdes produces more contiguous assemblies but requires more memory and computing time. The choice depends on the available resources and the research question. If the goal is to recover complete genomes, metaSPAdes may be worth the additional resources. If the goal is a broad survey of the community, MEGAHIT is often sufficient.

What completeness and contamination thresholds should I use for MAGs?

The commonly used thresholds for high-quality MAGs are greater than 90 percent completeness and less than 5 percent contamination, as used in the MINUUR pipeline. However, the appropriate thresholds depend on the research question. A study of metabolic potential might accept lower completeness, while a study that aims to describe new species would require higher quality. The thresholds should be recorded in the configuration file and reported in the results.

How do I handle samples with very low sequencing depth?

Samples with very low sequencing depth will produce poor assembly and binning results. The Integrated approaches for pathogen monitoring and shotgun metagenomic analysis in Atlantic salmon farming study found that metagenomic results from samples with low pathogen loads should be interpreted cautiously. Options for low-depth samples include combining samples for co-assembly, using a reference-based approach instead of assembly, or excluding the samples from assembly and binning analyses.

What should I do if my pipeline produces different results on different machines?

This indicates a reproducibility problem. Check that the container images are the same by verifying the image digests. Check that the reference database versions are the same. Check that all random seeds are fixed. If the problem persists, escalate to a bioinformatics specialist. The nf-core documentation provides guidance on troubleshooting reproducibility issues.

How do I share my pipeline with other researchers?

The pipeline should be committed to a public repository such as GitHub. The repository should include the pipeline code, the configuration files, the documentation, and the test dataset. The README should explain how to install the dependencies and run the pipeline. The nf-core documentation provides standards for community pipelines that can serve as a model for sharing.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.