Containerization for Reproducible Genome Assembly: A Practical Guide to Docker and Singularity
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Containerization, using tools like Docker and Singularity, is essential for reproducible genome assembly by packaging assembly software (e.g., Flye, hifiasm) with its precise dependencies and execution environment, mitigating failures caused by version mismatches or missing libraries across different computing platforms.
- Docker is suitable for personal workstations with administrative access due to its simpler build and run workflow, while Singularity (Apptainer) is the standard for high-performance computing clusters, as it operates without requiring root privileges and integrates with job schedulers.
- Workflow managers like Nextflow and Snakemake natively support containerized execution, allowing complex assembly pipelines (e.g., quality control, assembly, polishing) to run consistently by specifying container images for each process, ensuring identical software environments from raw reads to final assembly.
- Verifying container reproducibility involves running the same assembly twice within the container and comparing outputs, and rigorously testing across different host systems; recording container metadata, including the image digest, platform, tool version, and command-line arguments, is critical for documenting and ensuring replicability.
- Common failure patterns in containerized assembly include version mismatches between the container and documentation, insufficient memory or thread allocation for resource-intensive tools, file permission issues in mounted directories, and network access problems during image pulling, all of which require specific troubleshooting steps.
Genome assembly projects fail to reproduce when the software environment changes between the original analysis and a repeat attempt. A pipeline that runs on one laboratory workstation may produce different results or fail outright on a university cluster because of mismatched library versions, operating system differences, or missing dependencies. Containerization solves this problem by packaging the assembly tool, its dependencies, and the execution environment into a single portable unit. This article explains how to use Docker and Singularity containers specifically for assembly tools such as Flye and hifiasm, with concrete steps for building, running, and verifying containers in both local and high-performance computing contexts.
The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has raw sequencing data and needs to produce a reliable genome assembly. The practical outcome is a working container workflow that produces identical results across different computing environments. The scope covers data inputs, workflow choices, controls, quality checks, reproducibility, interpretation limits, reporting, and practical decision criteria.
Why Genome Assembly Reproducibility Fails Without Containers
Genome assembly software depends on a precise combination of system libraries, language runtimes, and tool versions. Flye requires specific versions of Python and C libraries. Hifiasm depends on a particular build of the C++ compiler and the zlib compression library. When any of these components differ between environments, the assembly output can change in ways that are difficult to detect until downstream analysis fails.
The practical problem appears in common scenarios. A researcher develops an assembly pipeline on a personal laptop with a recent Linux distribution. The same pipeline is then transferred to a university high-performance computing cluster that runs an older operating system. The cluster lacks a required library version, or the system administrator has installed a conflicting version of Python. The pipeline fails with an obscure error message, and the researcher spends days troubleshooting environment issues instead of analyzing assembly results.
Containerization addresses this problem by bundling the tool and its environment into a single image. The image contains the operating system libraries, the runtime, the assembly tool, and any configuration files needed for execution. When the image is run, the container provides an isolated environment that matches the original development environment regardless of the host system.
The value of this approach is documented in the bioinformatics literature. The PEMA pipeline for environmental DNA metabarcoding analysis uses containerization to ease the sharing and running of software packages across operating systems, which strongly facilitates pipeline development and usage. The authors note that many steps are required to obtain taxonomically assigned matrices from raw data, and each tool's execution parameters need to be tailored to reflect each experiment's idiosyncrasy. Containerization reduces the burden of managing these complex dependencies.
Large-scale metagenome analysis projects face the same challenge. The TOFU-MAaPO pipeline for metagenomic shotgun sequencing data is designed as a portable, automated single-command Nextflow pipeline that analyzes metagenome files locally or directly from the Sequence Read Archive. Its portability depends on containerized execution to ensure that the same software environment is used regardless of where the pipeline runs.
Container Fundamentals for Assembly Workflows
What a Container Contains
A container image for genome assembly contains several layers. The base layer is typically a minimal operating system such as Ubuntu or Alpine Linux. The next layer includes system libraries and tools needed by the assembly software. The application layer contains the assembly tool itself, such as Flye or hifiasm, along with any Python packages, Perl modules, or other language-specific dependencies. The final layer may include configuration files, environment variables, and entry point scripts that define how the container starts.
The key property of a container is isolation. The container shares the host operating system kernel but has its own filesystem, process space, and network interface. This means the container can run a different Linux distribution than the host, or a different version of a library, without conflicting with host software.
Docker and Singularity Compared
Docker is the most widely used container platform. It provides a command-line interface for building, running, and managing containers. Docker images are built from a Dockerfile, which is a text file containing instructions for each layer of the image. Docker is well suited for development work on personal computers and for services that need to run continuously.
Singularity, now known as Apptainer, is designed for high-performance computing environments. The key difference is that Singularity does not require root privileges to run containers. This is important on shared cluster systems where users do not have administrative access. Singularity also integrates with the job scheduler on most clusters, allowing containers to be run as part of batch jobs.
The choice between Docker and Singularity depends on the target environment. If the assembly will run on a personal workstation or a cloud server where the user has administrative control, Docker is the simpler option. If the assembly will run on a shared university or research cluster, Singularity is usually the required format because of security policies that prohibit Docker's root daemon.
Registry and Image Management
Container images are stored in registries. Docker Hub is the default public registry for Docker images. Many bioinformatics projects publish prebuilt images to Docker Hub or to specialized registries such as Quay.io. The National Center for Biotechnology Information provides access to many sequence analysis tools, and some of these are available as container images through community efforts.
For assembly tools, the practical approach is to check whether a maintained image already exists before building a custom one. The Flye and hifiasm projects may provide official images, or community members may maintain images that are regularly updated. Using a maintained image saves time and ensures that the tool version matches the published documentation.
Building a Docker Container for Flye
Writing the Dockerfile
The Dockerfile for Flye follows a standard pattern. The first line specifies the base image. A common choice is ubuntu:22.04 because it provides a stable set of system libraries. The next lines install the dependencies that Flye requires, which include Python 3, pip, and the build tools needed to compile the software from source.
A minimal Dockerfile for Flye might look like this:
FROM ubuntu:22.04
RUN apt-get update && apt-get install -y \
python3 \
python3-pip \
git \
build-essential \
&& rm -rf /var/lib/apt/lists/*
RUN pip3 install --upgrade pip
RUN git clone https://github.com/fenderglass/Flye.git /opt/Flye \
&& cd /opt/Flye \
&& python3 setup.py install
ENV PATH="/opt/Flye/bin:${PATH}"
WORKDIR /data
ENTRYPOINT ["flye"]
This Dockerfile creates an image with Flye installed and sets the working directory to /data, which is where input files should be mounted when the container is run.
Building the Image
The image is built with the docker build command. The command is run from the directory containing the Dockerfile:
docker build -t flye:latest .
The -t flag assigns a tag to the image, which makes it easier to reference later. The build process downloads the base image, runs each instruction in the Dockerfile, and creates a new image layer for each step.
Running the Flye Container
To assemble a genome with the Flye container, the input reads must be mounted into the container. The -v flag mounts a host directory into the container:
docker run -v /path/to/reads:/data flye:latest \
--pacbio-raw /data/reads.fastq \
--out-dir /data/assembly \
--genome-size 5m
The input reads are expected to be in /path/to/reads/reads.fastq on the host. The output assembly is written to /path/to/reads/assembly on the host, which corresponds to /data/assembly inside the container.
Building a Singularity Container for Hifiasm
Converting a Docker Image to Singularity
The most reliable way to create a Singularity image for hifiasm is to start with a Docker image and convert it. This approach ensures that the software environment is identical to the tested Docker environment.
First, build the Docker image for hifiasm using a Dockerfile similar to the Flye example. Then convert the Docker image to a Singularity image file:
singularity build hifiasm.sif docker-daemon://hifiasm:latest
The docker-daemon:// prefix tells Singularity to pull the image from the local Docker daemon. The resulting .sif file is a single-file container that can be copied to a cluster and run without any additional installation.
Running the Hifiasm Container
Singularity containers are run with the singularity exec command. The --bind flag mounts host directories into the container:
singularity exec --bind /path/to/reads:/data hifiasm.sif \
hifiasm -o /data/assembly -t 16 /data/reads.fastq
The -o flag specifies the output prefix, and -t sets the number of threads. The input reads are mounted from the host directory into the container at /data.
Building Directly from a Singularity Definition File
For environments where Docker is not available, a Singularity definition file can be used to build the image directly. The definition file is similar to a Dockerfile but uses Singularity-specific syntax:
Bootstrap: docker
From: ubuntu:22.04
%post
apt-get update && apt-get install -y \
python3 \
python3-pip \
git \
build-essential
pip3 install --upgrade pip
git clone https://github.com/chhylp123/hifiasm.git /opt/hifiasm
cd /opt/hifiasm && make
%environment
export PATH="/opt/hifiasm:${PATH}"
%runscript
exec hifiasm "$@"
The image is built with:
singularity build hifiasm.sif hifiasm.def
At a Glance: Container Choice for Assembly Tools
| Scenario | Recommended Container | Reason | Example Command |
|---|---|---|---|
| Personal workstation with administrative access | Docker | Simple build and run workflow, no cluster security restrictions | docker run -v /data:/data flye:latest --pacbio-raw /data/reads.fastq --out-dir /data/assembly |
| Shared university cluster without root access | Singularity | Runs without root privileges, integrates with job scheduler | singularity exec --bind /data:/data hifiasm.sif hifiasm -o /data/assembly -t 16 /data/reads.fastq |
| Cloud server or virtual machine | Docker | Consistent with deployment tooling, easy to automate | docker build -t flye:latest . && docker run flye:latest |
| Production pipeline with multiple tools | Singularity with Nextflow or nf-core | Pipeline framework manages container execution across steps | nextflow run pipeline -profile singularity |
Workflow Integration for Assembly Pipelines
Using Containers with Workflow Managers
Genome assembly is rarely a single command. A typical assembly workflow includes read quality control, read filtering, assembly, polishing, and quality assessment. Each step may use a different tool, and each tool may have different dependency requirements.
Workflow managers such as Nextflow and Snakemake support container execution natively. The workflow definition specifies which container image to use for each process. The workflow manager pulls the image, runs the process inside the container, and collects the output files.
The nf-core community documentation provides guidance for configuring Nextflow pipelines to use containers. The documentation covers how to specify container images in the pipeline configuration, how to handle container registries, and how to troubleshoot common container-related issues. For assembly pipelines, this approach ensures that each step runs in a consistent environment.
The TOFU-MAaPO pipeline demonstrates this pattern in practice. It is a portable, automated single-command Nextflow pipeline for large-scale metagenomic analysis. The pipeline uses containers to ensure that the same software versions are used regardless of where the pipeline runs, whether on a local machine or a high-performance cluster.
Integrating Flye and Hifiasm into a Containerized Workflow
A containerized assembly workflow for long-read data might include the following steps:
- Read quality assessment with a tool such as NanoPlot
- Read filtering with a tool such as Filtlong
- Initial assembly with Flye or hifiasm
- Assembly polishing with a tool such as Medaka or Racon
- Assembly quality assessment with a tool such as QUAST
Each step runs in its own container. The workflow manager handles the data flow between steps, passing the output of one step as input to the next.
The container images for each tool are specified in the workflow configuration. For example, a Nextflow configuration might specify:
process {
withName: 'ASSEMBLY' {
container = 'quay.io/biocontainers/flye:4.2.1--py39h5a725a7_0'
}
withName: 'POLISHING' {
container = 'quay.io/biocontainers/medaka:1.11.1--py39h5a725a7_0'
}
}
The workflow manager pulls the specified images and runs each process in the appropriate container.
Handling Input and Output Data
Containerized workflows require careful attention to data mounting. The container has its own filesystem, which is separate from the host filesystem. Input files must be mounted into the container, and output files must be written to a mounted directory so they are accessible on the host.
For Docker, the -v flag mounts a host directory into the container. For Singularity, the --bind flag serves the same purpose. When using a workflow manager, the data mounting is handled automatically based on the workflow configuration.
A common failure pattern is mounting the wrong directory or forgetting to mount the input directory. The container starts successfully but cannot find the input files, resulting in an error message that is confusing because the files are visible on the host.
Reproducibility Verification and Quality Controls
Verifying Container Reproducibility
The primary purpose of containerization is reproducibility. To verify that a container produces consistent results, run the same assembly twice in the same container and compare the outputs. The assembly graphs and contig sequences should be identical.
A more rigorous test is to run the same container on different host systems. If the container is properly isolated, the assembly output should be identical regardless of the host operating system or hardware. Differences in output indicate that the container is not fully isolated, possibly because a dependency is being pulled from the host system instead of the container.
The Galaxy Training Network provides accessible workflow training that covers reproducibility concepts. The training materials explain how to document and verify that a workflow produces the same results when run at different times or on different systems. These principles apply directly to containerized assembly workflows.
Recording Container Metadata
Reproducibility requires recording the exact container image used for each assembly run. The image digest is a unique identifier for the image content. The image digest ensures that the same image is used even if a tag such as latest is updated to point to a different image.
The container metadata to record includes:
- The container platform (Docker or Singularity)
- The image name and tag
- The image digest
- The container build date
- The assembly tool version inside the container
- The command used to run the container
This information should be recorded in the assembly report or in a separate reproducibility file that accompanies the assembly data.
Quality Checks for Assembly Output
Containerization ensures that the software environment is consistent, but it does not guarantee that the assembly is biologically correct. Quality checks must be performed on the assembly output regardless of the container used.
Common quality checks for genome assemblies include:
- Assembly size compared to the expected genome size
- Number of contigs and the N50 statistic
- Completeness assessment using BUSCO or similar tools
- Mapping rate of reads back to the assembly
- Comparison with a reference genome if one is available
The NCBI data resources include tools and documentation for evaluating genome assemblies, and the NCBI assembly database provides a framework for comparing assembly quality across projects.
Common Failure Patterns in Containerized Assembly
Version Mismatch Between Container and Documentation
A frequent failure occurs when the container image contains a different version of the assembly tool than the version documented in the software manual. The command-line options may differ between versions, causing the container to fail with an unrecognized argument error.
The solution is to verify the tool version inside the container before running the assembly. For Docker, run:
docker run flye:latest --version
For Singularity, run:
singularity exec hifiasm.sif hifiasm --version
The version output should match the version expected by the workflow documentation.
Insufficient Memory or Thread Allocation
Assembly tools such as Flye and hifiasm are memory-intensive. The container does not limit memory usage by default, but the host system may have limits imposed by the job scheduler or the operating system.
A common failure pattern is running the container with too few threads or too little memory allocated. The assembly process is killed by the operating system when it exceeds the memory limit, producing an error message that does not clearly indicate the cause.
The solution is to check the memory requirements of the assembly tool and allocate sufficient resources. For hifiasm, the memory requirement depends on the genome size and the depth of coverage. For a mammalian genome, 100 gigabytes of memory may be required. The container should be run with a memory allocation that matches the tool requirements.
File Permission Issues in Mounted Directories
Containerized workflows often fail because of file permission problems. The container runs as a specific user, and that user may not have write permission to the mounted host directory. The assembly process fails when it tries to write output files.
The solution is to ensure that the mounted directory is writable by the container user. For Docker, the -u flag can specify the user ID. For Singularity, the --writable flag or the --fakeroot option may be needed depending on the cluster configuration.
Network Access for Image Pulling
Building and pulling container images requires network access. On a cluster with restricted network access, the image pull may fail. The solution is to build the image on a machine with network access and transfer the image file to the cluster.
For Singularity, the .sif file is a single file that can be copied with scp or transferred through the cluster's data transfer system. For Docker, the image can be saved to a tar file with docker save and loaded on the target system with docker load.
Performance Considerations for Containerized Assembly
Overhead of Container Execution
Containers add minimal overhead compared to running software directly on the host. The container shares the host kernel, so system calls are not virtualized. The main performance impact comes from filesystem isolation, which can slow down file access if the container uses a different filesystem layer.
For assembly workloads that are compute-intensive, the container overhead is negligible. The assembly time is dominated by the computational work of the assembly algorithm, not by the container runtime.
Parallel Execution and Thread Management
Assembly tools use multiple threads to speed up computation. The container must be configured to allow the tool to access the required number of CPU cores. For Docker, the --cpus flag limits the number of CPUs available to the container. For Singularity, the container inherits the CPU allocation from the job scheduler.
A common mistake is running the container without specifying the thread count, causing the tool to use the default number of threads, which may be too low or too high for the available resources. The thread count should be specified explicitly in the container run command.
Storage Requirements for Assembly Output
Genome assembly produces large output files. The assembly graph, the contig sequences, and the intermediate files can require tens or hundreds of gigabytes of storage. The mounted output directory must have sufficient free space.
The storage requirement should be checked before running the assembly. The assembly tool documentation usually provides estimates of the output size based on the genome size and coverage. The container run command should include a check for available disk space, or the workflow should include a storage monitoring step.
Limitations of Containerized Assembly
Container Images Are Not a Substitute for Tool Validation
A container image ensures that the software environment is consistent, but it does not validate that the assembly tool produces correct results. The assembly algorithm may have bugs, or the tool may produce suboptimal results for certain types of data. Containerization does not address these issues.
The LEMMIv2 benchmarking framework addresses this problem by providing impartial benchmarks for metagenomic profilers. The framework evaluates tools across several reference databases and provides a catalogue of evaluated tools. Similar benchmarking approaches are needed for assembly tools, and the results should inform tool selection regardless of containerization.
Container Images Can Become Outdated
Container images are static snapshots of a software environment. When the assembly tool is updated with bug fixes or new features, the container image must be rebuilt. A container image that is not updated may contain a version of the tool with known bugs or security vulnerabilities.
The solution is to maintain a container update process. The container image should be rebuilt when the upstream tool releases a new version, and the image digest should be recorded in the workflow documentation. The nf-core documentation provides guidance on container versioning and updates for community pipelines.
Containerization Does Not Solve Data Management Problems
Containerization addresses software environment reproducibility, but it does not address data management. The input reads, the assembly output, and the quality assessment results must be stored and organized according to data management best practices. The NCBI data resources provide data submission and storage options for sequence data, and these should be used to archive the assembly data and metadata.
The EMBL-EBI training resources cover data management for bioinformatics projects, including data organization, documentation, and submission to public databases. These practices are complementary to containerization and are necessary for full reproducibility.
Professional Escalation Criteria
When to Seek Expert Assistance
Containerized assembly workflows can fail in ways that are difficult to diagnose. The following situations warrant escalation to a bioinformatics support specialist, a system administrator, or a colleague with container expertise:
- The container fails to build with an error that is not resolved by following the tool documentation
- The assembly produces different results when run twice in the same container
- The container runs successfully but produces an assembly that fails quality checks
- The cluster scheduler rejects the container job with an error that is not understood
- The container image cannot be pulled from the registry because of network or authentication issues
Documentation for Escalation
When escalating a container issue, provide the following information:
- The exact container image name, tag, and digest
- The container platform and version (Docker or Singularity)
- The host operating system and version
- The exact command used to run the container
- The full error message or log output
- The input data description and file sizes
- The expected output and the actual output
This information allows the support specialist to reproduce the issue and identify the cause.
A Practical Decision Framework for Container Selection in Assembly Projects
Choosing between Docker and Singularity for a genome assembly project is not a one-time decision that applies to every situation. The correct choice depends on the specific computing environment, the scale of the assembly, the collaboration requirements, and the long-term data management plan. This section provides a structured decision framework that researchers can apply when setting up containerized assembly workflows for Flye, hifiasm, or other assembly tools.
Step 1: Assess the Primary Execution Environment
The first decision point is where the assembly will actually run. This is not necessarily where the pipeline is developed. Many researchers develop on a personal workstation and then execute on a shared cluster. The execution environment determines the container format that will be accepted by the system.
For a personal workstation where the user has administrative rights, Docker is the practical choice. The Docker daemon runs with root privileges, and the user can build, pull, and run images without restriction. The workflow is straightforward: write a Dockerfile, build the image, and run the container with the appropriate volume mounts.
For a shared university or research cluster, Singularity is typically the only viable option. Cluster security policies generally prohibit users from running a Docker daemon because it requires root privileges and can pose a security risk to other users. Singularity runs containers as the invoking user without root privileges, which aligns with cluster security requirements. The nf-core documentation provides detailed guidance on configuring Singularity for cluster environments, including how to set the container library path and how to handle image caching.
The decision is not always binary. Some researchers use Docker for development and then convert the image to Singularity for cluster execution. This hybrid approach is common and works well when the Docker image is built on a machine with Docker access and then converted using the singularity build command with the docker-daemon:// source.
Step 2: Evaluate the Collaboration and Sharing Requirements
The second decision point is whether the assembly workflow will be shared with collaborators or published as part of a study. Containerization is a reproducibility tool, and the choice of container format affects how easily others can reproduce the assembly.
Docker images are stored in registries such as Docker Hub and Quay.io. These registries provide a centralized location where images can be published and pulled by other researchers. The image digest provides a unique identifier that can be cited in publications. This makes Docker well suited for sharing workflows with a broad audience.
Singularity images are single files with a .sif extension. These files can be transferred directly to collaborators or stored in a shared filesystem. The single-file format is convenient for cluster environments because it does not require a registry or a daemon. However, sharing a Singularity image requires transferring the file itself, which can be large for assembly tools with many dependencies.
The PEMA pipeline for environmental DNA metabarcoding analysis demonstrates the sharing benefits of containerization. The authors note that software containerization technologies ease the sharing and running of software packages across operating systems, which strongly facilitates pipeline development and usage. When choosing a container format, consider how the workflow will be shared and whether the recipients are likely to have Docker or Singularity available.
Step 3: Consider the Workflow Manager Integration
The third decision point is whether the assembly will be run as a standalone command or as part of a multi-step workflow managed by a tool such as Nextflow or Snakemake. Workflow managers have native support for both Docker and Singularity, but the configuration differs.
Nextflow supports both container platforms through the docker.enabled and singularity.enabled configuration options. The workflow definition specifies which container image to use for each process, and the workflow manager handles the container execution. The nf-core documentation provides comprehensive guidance on configuring Nextflow for containerized execution, including how to specify container images in the process definitions and how to handle container registries.
The TOFU-MAaPO pipeline for metagenomic shotgun sequencing data is an example of a containerized Nextflow pipeline that runs on both local machines and high-performance clusters. The pipeline is designed as a portable, automated single-command workflow that analyzes metagenome files locally or directly from the Sequence Read Archive. Its portability depends on containerized execution to ensure that the same software environment is used regardless of where the pipeline runs.
When using a workflow manager, the container format should match the execution environment. If the workflow will run on a cluster, configure Nextflow or Snakemake to use Singularity. If the workflow will run on a personal workstation or cloud server, Docker is the simpler option.
Step 4: Assess the Image Maintenance Burden
The fourth decision point is the long-term maintenance of the container image. Assembly tools are updated regularly with bug fixes, performance improvements, and new features. The container image must be rebuilt when the upstream tool releases a new version.
Docker images are built from a Dockerfile, which is a text file that can be version-controlled with Git. The Dockerfile documents the exact steps used to create the image, making it easy to rebuild the image when the tool version changes. The build process is reproducible, and the image digest changes when any layer changes.
Singularity images can be built from a definition file, which is also a text file that can be version-controlled. The definition file specifies the base image, the installation steps, and the environment configuration. However, building a Singularity image requires either a local installation of Singularity or access to a remote builder service. On a cluster without Singularity installed, the image must be built elsewhere and transferred.
The maintenance burden also includes updating the workflow configuration to reference the new image digest. The nf-core documentation provides guidance on container versioning and updates for community pipelines. The documentation emphasizes the importance of recording the image digest to ensure that the same image is used for all runs.
Step 5: Document the Decision and Rationale
The final step in the decision framework is to document the container choice and the rationale. This documentation should be included in the assembly report or in a separate reproducibility file that accompanies the assembly data.
The documentation should include:
- The container platform selected and the reason for the selection
- The execution environment where the container will run
- The collaboration and sharing requirements that influenced the choice
- The workflow manager integration details
- The image maintenance plan and update schedule
This documentation allows other researchers to understand why a particular container format was chosen and to make informed decisions if they need to adapt the workflow to a different environment.
A Record System for Containerized Assembly Runs
Reproducibility requires more than choosing the right container format. It requires a systematic record of every assembly run, including the container image, the input data, the parameters, and the output. This section describes a practical record system that researchers can implement with minimal overhead.
The Container Run Log
The container run log is a structured record of each assembly execution. The log should be created for every assembly run, whether it is a test run on a small dataset or a production run on a full genome.
The minimum fields for the container run log are:
- Run identifier: a unique identifier for the assembly run
- Date and time: when the run was executed
- Container platform: Docker or Singularity
- Image name and tag: the full image reference
- Image digest: the unique identifier for the image content
- Tool version: the version of the assembly tool inside the container
- Input data: the file paths and checksums of the input reads
- Parameters: the exact command-line arguments used
- Output location: where the assembly output was written
- Host system: the operating system and hardware where the container ran
- Resource allocation: the CPU and memory allocated to the container
- Run status: completed successfully, failed, or interrupted
- Error messages: any error output from the container
The log can be maintained as a spreadsheet, a plain text file, or a structured format such as JSON or YAML. The key requirement is that the log is updated consistently for every run.
Recording Image Digests
The image digest is the most important piece of metadata for reproducibility. The digest is a cryptographic hash of the image content, and it uniquely identifies the exact software environment. A tag such as latest can change to point to a different image, but the digest always refers to the same image content.
For Docker, the digest can be obtained with:
docker inspect --format='{{index .RepoDigests 0}}' flye:latest
For Singularity, the digest can be obtained with:
singularity inspect hifiasm.sif
The digest should be recorded in the container run log and in the assembly report. When the assembly is published, the digest should be included in the methods section so that other researchers can reproduce the exact environment.
Recording Input Data Checksums
The input data must be recorded with checksums to ensure that the same reads are used in a repeat run. The checksum is a cryptographic hash of the file content, and it changes if the file is modified or corrupted.
The checksum can be calculated with:
sha256sum reads.fastq
The checksum should be recorded in the container run log alongside the file path. This allows a repeat run to verify that the input data has not changed.
The Assembly Report Template
The assembly report should include a reproducibility section that summarizes the container information. The report template should include:
- The container platform and version
- The image name, tag, and digest
- The tool version inside the container
- The exact run command
- The input data checksums
- The output file checksums
- The date and time of the run
- The host system information
This information allows another researcher to reproduce the assembly environment exactly. The Galaxy Training Network provides accessible workflow training that covers reproducibility concepts, including how to document and verify that a workflow produces the same results when run at different times or on different systems. These principles apply directly to containerized assembly workflows.
Troubleshooting Containerized Assembly Failures
Containerized assembly workflows fail in predictable patterns. This section describes a systematic troubleshooting method that addresses the most common failure modes.
Step 1: Verify the Container Image
The first troubleshooting step is to verify that the container image is correct. This includes checking the image digest, the tool version, and the container configuration.
Run the version command inside the container to verify the tool version:
docker run flye:latest --version
For Singularity:
singularity exec hifiasm.sif hifiasm --version
The version output should match the version expected by the workflow documentation. If the version is different, the image may be outdated or the wrong image may have been pulled.
Step 2: Verify the Input Data Mounting
The second troubleshooting step is to verify that the input data is accessible inside the container. The container has its own filesystem, and input files must be mounted from the host.
For Docker, verify the mount with:
docker run -v /path/to/reads:/data flye:latest ls /data
For Singularity:
singularity exec --bind /path/to/reads:/data hifiasm.sif ls /data
The ls command should list the input files. If the directory is empty or the files are missing, the mount is incorrect.
Step 3: Verify Resource Allocation
The third troubleshooting step is to verify that the container has sufficient resources. Assembly tools such as Flye and hifiasm are memory-intensive, and the container may be killed by the operating system if it exceeds the memory limit.
Check the memory allocation for the container. For Docker, the --memory flag limits the memory available to the container. For Singularity, the container inherits the memory limits from the job scheduler.
The memory requirement depends on the genome size and the depth of coverage. For a mammalian genome, 100 gigabytes of memory may be required. The container should be run with a memory allocation that matches the tool requirements.
Step 4: Check the Error Log
The fourth troubleshooting step is to examine the error log. The container run produces output that includes error messages from the assembly tool and from the container runtime.
The error log should be examined for:
- Unrecognized command-line arguments
- Missing input files
- Permission denied errors
- Out-of-memory errors
- Segmentation faults
- Library version mismatches
The error message often indicates the cause of the failure. If the error is not clear, the log should be saved and included in the escalation documentation.
Step 5: Test with a Small Dataset
The fifth troubleshooting step is to test the container with a small dataset. A small test dataset can be generated by taking a subset of the reads or by using a published test dataset from the assembly tool documentation.
The test run should use the same container image and the same parameters as the production run. If the test run succeeds, the problem is likely related to the data size or the resource allocation. If the test run fails, the problem is likely in the container configuration.
Step 6: Compare with a Direct Installation
The sixth troubleshooting step is to compare the containerized run with a direct installation of the assembly tool. If the tool is installed directly on the host system, run the same assembly command without the container.
If the direct installation succeeds but the containerized run fails, the problem is in the container configuration. If both fail, the problem is in the assembly tool or the input data.
This comparison can be time-consuming, but it is the most reliable way to isolate the cause of a container failure.
Common Failure Patterns and Their Resolutions
Pattern 1: The Container Runs but Produces No Output
This failure pattern occurs when the container starts successfully but the assembly tool does not produce the expected output files. The container may exit with a success status, but the output directory is empty.
The likely cause is that the output directory is not mounted correctly, or the assembly tool is writing to a different location than expected. The solution is to verify the output directory mount and to check the assembly tool documentation for the default output location.
Pattern 2: The Container Fails with a Permission Denied Error
This failure pattern occurs when the container user does not have write permission to the mounted output directory. The container runs as a specific user, and that user may not have write access to the host directory.
The solution is to ensure that the mounted directory is writable by the container user. For Docker, the -u flag can specify the user ID. For Singularity, the --writable flag or the --fakeroot option may be needed depending on the cluster configuration.
Pattern 3: The Container Fails with a Library Version Error
This failure pattern occurs when the assembly tool inside the container cannot find a required library. The library may be missing from the container image, or the version may be incompatible with the tool.
The solution is to rebuild the container image with the correct library dependencies. The assembly tool documentation should specify the required libraries and versions.
Pattern 4: The Container Produces Different Results on Different Hosts
This failure pattern occurs when the container is not fully isolated from the host system. The assembly tool may be using a library or configuration file from the host instead of the container.
The solution is to verify that the container image includes all dependencies and that the container is run with the appropriate isolation flags. The LEMMIv2 benchmarking framework provides a platform for evaluating tool performance across different environments, and the results can help identify whether a tool is sensitive to environmental differences.
Professional Escalation Criteria for Container Issues
When to Escalate
Containerized assembly workflows can fail in ways that are difficult to diagnose. The following situations warrant escalation to a bioinformatics support specialist, a system administrator, or a colleague with container expertise:
- The container fails to build with an error that is not resolved by following the tool documentation
- The assembly produces different results when run twice in the same container
- The container runs successfully but produces an assembly that fails quality checks
- The cluster scheduler rejects the container job with an error that is not understood
- The container image cannot be pulled from the registry because of network or authentication issues
- The container fails with a segmentation fault or other low-level error that is not explained by the tool documentation
Documentation for Escalation
When escalating a container issue, provide the following information:
- The exact container image name, tag, and digest
- The container platform and version (Docker or Singularity)
- The host operating system and version
- The exact command used to run the container
- The full error message or log output
- The input data description and file sizes
- The expected output and the actual output
- The container run log entry for the failed run
This information allows the support specialist to reproduce the issue and identify the cause. The EMBL-EBI training resources provide foundational bioinformatics training that includes troubleshooting and problem-solving skills that are relevant to containerized workflows.
The Escalation Process
The escalation process should follow a structured path. First, consult the assembly tool documentation and the container platform documentation. Second, search for the error message in the tool issue tracker or community forums. Third, consult a colleague with container expertise. Fourth, escalate to the institutional bioinformatics support team or the cluster system administrators.
The escalation should include the documentation listed above and a clear description of the troubleshooting steps already attempted. This avoids repeating the same troubleshooting steps and allows the support specialist to focus on the unresolved issue.
Frequently Asked Questions
What is the difference between Docker and Singularity for genome assembly?
Docker requires root privileges to run and is designed for environments where the user has administrative control. Singularity runs without root privileges and is designed for shared high-performance computing clusters. For genome assembly on a personal workstation, Docker is simpler. For assembly on a university cluster, Singularity is usually required because of cluster security policies.
How do I know which container image to use for Flye or hifiasm?
Check whether the tool project provides an official container image. If not, search for community-maintained images on Docker Hub or Quay.io. Verify that the image contains the expected tool version by running the version command inside the container. Record the image digest to ensure that the same image is used for all assembly runs.
Can I use a Docker container on a cluster that only supports Singularity?
Yes. Build the Docker image on a machine with Docker installed, then convert it to a Singularity image file using the singularity build command with the docker-daemon:// source. The resulting .sif file can be transferred to the cluster and run without Docker.
How do I mount my sequencing reads into a container?
For Docker, use the -v flag to mount a host directory into the container. For Singularity, use the --bind flag. The mounted directory appears inside the container at the specified path, and the assembly tool reads the input files from that path.
Why does my containerized assembly fail with an out-of-memory error?
The container inherits the memory limits of the host system or the job scheduler. Assembly tools such as Flye and hifiasm require substantial memory, especially for large genomes. Check the memory requirements for your genome size and coverage, and allocate sufficient memory in the container run command or the job submission script.
How do I verify that my containerized assembly is reproducible?
Run the same assembly twice in the same container and compare the outputs. The contig sequences and assembly statistics should be identical. For a stronger test, run the same container on different host systems and compare the outputs. Record the container image digest and the run command to document the reproducibility.
What should I do if my container image is outdated?
Rebuild the container image using the latest version of the assembly tool. Update the workflow configuration to reference the new image digest. Test the new image with a small test dataset before running the full assembly to ensure that the tool version produces expected results.
How do I record container information in my assembly report?
Include the container platform, image name, tag, digest, build date, tool version, and the exact run command in the assembly report or a separate reproducibility file. This information allows others to reproduce the assembly environment exactly.
Related Bioinformatics Guides
- Evaluating Genome Assembly Quality: Metrics and Tools
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Metagenomics Tools: A Practical Guide to Software and Pipelines
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Docker and Containerization in Reproducible Research
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- PEMA: a flexible Pipeline for Environmental DNA Metabarcoding Analysis of the 16S/18S ribosomal RNA, ITS, and COI marker genes.. GigaScience, 2020.
- TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive.. 2026.
- LEMMIv2: benchmarking framework for metagenomic and 16S amplicon profilers with a catalogue of evaluated tools.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.