From Reads to Reproducible Assembly: A Complete Workflow with Snakemake, Conda, and Git

By Dr. Zubair Khalid, DVM, MS, PhD ·

From Reads to Reproducible Assembly: A Complete Workflow with Snakemake, Conda, and Git

Key Takeaways

  • Genome assembly reproducibility hinges on systematically managing software versions, parameters, and input data provenance, which is critical for validating results and troubleshooting discrepancies.
  • Snakemake orchestrates the assembly pipeline as a directed acyclic graph, defining dependencies and execution order, while Conda isolates software dependencies in reproducible environments, preventing version conflicts.
  • Git version control tracks changes to workflow files (Snakefile, configuration, environment specifications) and parameter adjustments, providing an auditable history essential for reconstructing specific assembly outcomes.
  • Quality control of raw sequencing reads (e.g., adapter trimming, quality filtering) and subsequent assembly assessment using metrics like N50, number of contigs, and BUSCO completeness are vital for evaluating assembly quality.
  • A structured project directory separates raw data, processed files, results, and workflow code, facilitating effective Git management and preventing accidental commitment of large or regenerable files.
  • Common failure patterns include version conflicts (resolved by Conda environments), insufficient memory (requiring resource adjustment or parameter tuning), poor assembly quality (indicating input data issues or inappropriate parameters), and contamination (requiring read filtering).

Genome assembly projects fail in reproducible ways when researchers cannot reconstruct which software versions produced a given assembly, which parameters were used, or how the input reads were filtered. This article provides an end-to-end workflow for small genome assembly projects that integrates Snakemake for workflow management, Conda for environment management, and Git for version control. The workflow is designed for biology students, researchers, and laboratory professionals who have sequencing reads and need a defensible assembly with documented provenance. The practical outcome is a fully annotated pipeline that can be adapted to your own data, with explicit controls for quality assessment, reproducibility, and professional escalation when assembly results fall outside expected parameters.

The Reproducibility Problem in Genome Assembly

Genome assembly is a computational process that transforms raw sequencing reads into contiguous consensus sequences representing an organism's genome. The process involves multiple interdependent steps: quality control of raw reads, error correction, contig assembly, scaffolding, polishing, and quality assessment. Each step typically uses different software tools, and each tool has its own dependencies, version requirements, and parameter sensitivities.

The central problem is that assembly results are highly sensitive to software versions and parameter choices. A change in a single tool version can alter the final assembly in ways that are difficult to predict. Without systematic management of these variables, two researchers running the same analysis on the same data may obtain different assemblies, and neither may be able to reconstruct exactly how their result was produced.

This problem is well recognized in the bioinformatics community. Published workflows increasingly emphasize reproducibility as a primary design goal. For example, the Colora workflow for chromosome-scale de novo genome assembly was developed specifically to provide an automated, portable assembly process that avoids duplication of efforts and variable data quality across laboratories. Similarly, the RASflow RNA-Seq workflow was designed to address library dependencies and version conflicts through the combination of Snakemake and Conda, making the workflow usable by researchers with limited programming skills. The Seq2science workflow for functional genomics analysis was built on the Snakemake workflow language specifically to enable standardized and reproducible analysis across computing infrastructures.

The workflow described here applies the same principles to a small genome assembly project. It is not a substitute for large-scale production pipelines, but it provides a complete, adaptable template that demonstrates the integration of the three core reproducibility tools.

Core Principles of Reproducible Assembly Workflows

Workflow Management with Snakemake

Snakemake is a workflow management system that organizes analysis steps into a directed acyclic graph. Each step, or rule, defines its input files, output files, and the shell command or script that produces the outputs from the inputs. Snakemake determines the order of execution based on file dependencies, so you do not need to specify the order manually.

The key advantage of workflow management is that it makes the analysis explicit and auditable. Every step is defined in a single file, the Snakefile, which can be version-controlled and shared. When you run the workflow, Snakemake checks which outputs already exist and only runs the steps that are needed. This means you can add new samples or change parameters without rerunning the entire pipeline.

Snakemake also supports the specification of software environments for each rule. This is critical for assembly, where different tools may require incompatible versions of shared libraries. The Colora workflow, for example, uses Snakemake to produce chromosome-scale assemblies from Pacific Biosciences HiFi and Hi-C reads, with optional Oxford Nanopore Technologies input. The MosaiCatcher v2 framework for single-cell structural variation detection is implemented using Snakemake and is compatible with both container and conda environments to ensure reproducibility.

Environment Management with Conda

Conda is a package and environment management system that installs software and its dependencies in isolated environments. Each environment can contain specific versions of tools and libraries, and environments do not interfere with each other or with the system Python installation.

For genome assembly, Conda environments solve the dependency problem in a practical way. Assembly tools often require specific versions of C libraries, Python packages, or other tools. If you install everything in a single environment, you risk version conflicts that are difficult to resolve. With Conda, you can create a separate environment for each major analysis stage, or even for each rule in your Snakemake workflow.

The RASflow workflow demonstrates this approach. It uses Snakemake and Conda to alleviate challenges with library dependencies and version conflicts, supporting reproducibility while remaining usable by researchers without programming skills. The ARPEGGIO workflow for epigenetic data analysis in polyploids similarly ensures reproducibility through package management and containerization.

Version Control with Git

Git is a distributed version control system that tracks changes to files over time. For assembly projects, Git serves two purposes. First, it tracks changes to your workflow files, so you can see exactly what changed between versions of your pipeline. Second, it provides a record of when and why changes were made, which is essential for reconstructing the provenance of a specific assembly.

Git is most useful when combined with a hosting platform such as GitHub or GitLab. The published workflows referenced in this article all make their source code available on GitHub, which allows other researchers to inspect, reuse, and adapt the pipelines. The Colora workflow is available on GitHub and has been deposited in Zenodo with a persistent DOI, providing a citable version of the software.

For your own project, Git provides the same benefits at a smaller scale. You can commit your Snakefile, configuration files, and environment specifications at each stage of development. If an assembly produces unexpected results, you can identify exactly which version of the workflow produced it.

Designing the Assembly Workflow

Defining Input Data and Project Scope

The first decision in any assembly project is the scope of the assembly and the type of input data. This workflow is designed for small genome assemblies, which we define as bacterial genomes, small eukaryotic genomes, or targeted regions of larger genomes. These projects typically involve a single sample or a small number of samples, and the computational requirements are modest enough to run on a workstation or small server.

The input data for assembly typically comes from one of two sources. You may have generated your own sequencing data, in which case you need to manage the raw reads from your sequencing facility. Alternatively, you may be using publicly available data from repositories such as the NCBI Sequence Read Archive. The NCBI provides search systems and sequence resources that allow you to locate and download public sequencing data for reanalysis. The Seq2science workflow facilitates the downloading of sequencing data from all major databases, including NCBI SRA, EBI ENA, DDBJ, GSA, and ENCODE, and automates the retrieval of genome assemblies from Ensembl, NCBI, and UCSC.

For this workflow, we assume you have raw sequencing reads in FASTQ format. The reads may be short reads from Illumina platforms or long reads from Pacific Biosciences or Oxford Nanopore Technologies platforms. The choice of assembly tools will depend on the read type, as discussed in the next section.

Choosing Assembly Tools

The choice of assembly software depends primarily on the type of sequencing data you have. Short-read assemblers are designed for Illumina data and typically use de Bruijn graph algorithms. Long-read assemblers are designed for Pacific Biosciences or Oxford Nanopore Technologies data and typically use overlap-layout-consensus algorithms or string graph approaches.

For small genomes with short reads, common choices include SPAdes and Velvet. For long reads, common choices include Flye, Canu, and Raven. Hybrid assemblers can combine short and long reads to improve assembly quality.

The Colora workflow provides a useful reference for the structure of a modern assembly pipeline. It produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. While this workflow is designed for larger genomes, its structure demonstrates the steps that a complete assembly pipeline should include: read filtering, assembly, scaffolding, and quality assessment.

For your small genome assembly, you should select tools that are appropriate for your read type and genome size. The key principle is to document your choice and the rationale behind it in your workflow files.

Structuring the Project Directory

A reproducible assembly project requires a consistent directory structure. This structure separates raw data, intermediate files, final results, and workflow code, so that each component can be managed and version-controlled appropriately.

A recommended structure is as follows:

project/
├── Snakefile
├── config.yaml
├── envs/
│   ├── qc.yaml
│   ├── assembly.yaml
│   └── polishing.yaml
├── scripts/
│   ├── assess_assembly.py
│   └── summarize_results.py
├── data/
│   ├── raw/
│   └── processed/
├── results/
│   ├── qc/
│   ├── assembly/
│   └── assessment/
└── docs/
    └── README.md

The Snakefile contains the workflow rules. The config.yaml file contains parameters that may change between runs, such as sample names, read file paths, and assembly parameters. The envs directory contains Conda environment specifications for each analysis stage. The scripts directory contains any custom Python or shell scripts used by the workflow. The data directory contains raw and processed data, and the results directory contains workflow outputs organized by analysis stage.

This structure keeps workflow code separate from data and results, which is important for Git version control. You should commit the Snakefile, config.yaml, environment specifications, and scripts to Git. You should not commit raw data or large intermediate files, as these are typically managed separately or stored in data repositories.

Implementing the Snakemake Pipeline

The Snakefile Structure

The Snakefile is the central file of your workflow. It defines the rules that transform inputs to outputs, and it specifies the order of execution through file dependencies.

A minimal Snakefile for a small genome assembly might look like this:

configfile: "config.yaml"

rule all:
    input:
        "results/assessment/assembly_summary.txt"

rule quality_control:
    input:
        "data/raw/{sample}_R1.fastq.gz",
        "data/raw/{sample}_R2.fastq.gz"
    output:
        "results/qc/{sample}_R1_trimmed.fastq.gz",
        "results/qc/{sample}_R2_trimmed.fastq.gz",
        "results/qc/{sample}_qc_report.html"
    conda:
        "envs/qc.yaml"
    shell:
        """
        fastp -i {input} -I {input[1]} \
            -o {output} -O {output[1]} \
            -h {output[2]}
        """

rule assemble:
    input:
        "results/qc/{sample}_R1_trimmed.fastq.gz",
        "results/qc/{sample}_R2_trimmed.fastq.gz"
    output:
        "results/assembly/{sample}_contigs.fasta"
    conda:
        "envs/assembly.yaml"
    shell:
        """
        spades.py -1 {input} -2 {input[1]} \
            -o results/assembly/{wildcards.sample} \
            --isolate
        cp results/assembly/{wildcards.sample}/contigs.fasta {output}
        """

This example shows the three essential components of a Snakemake rule: input files, output files, and the command that produces the outputs. The configfile directive loads parameters from a YAML file, and the conda directive specifies the environment in which the rule should run.

The rule all at the top defines the final target of the workflow. Snakemake works backward from this target to determine which rules need to run.

Configuration File

The configuration file separates parameters from workflow logic. This allows you to change sample names, file paths, or assembly parameters without editing the Snakefile.

A sample configuration file might look like this:

samples:
  sample1: data/raw/sample1
  sample2: data/raw/sample2

assembly:
  assembler: "spades"
  mode: "isolate"
  kmer_sizes: 

quality:
  trim_quality: 20
  min_length: 50

The configuration file should be committed to Git, so that you can track changes to parameters over time. When you run the workflow, Snakemake reads the configuration and uses the values in the rules.

Conda Environment Specifications

Each Conda environment is defined in a YAML file that specifies the channels and packages to install. Pinning versions is essential for reproducibility. A minimal environment specification for quality control might look like this:

channels:
  - conda-forge
  - bioconda
dependencies:
  - fastp =0.23.4
  - multiqc =1.14

The assembly environment would include the assembler and its dependencies:

channels:
  - conda-forge
  - bioconda
dependencies:
  - spades =3.15.5
  - quast =5.2.0

When Snakemake encounters a rule with a conda directive, it creates the environment if it does not already exist and runs the rule within that environment. This ensures that each rule uses exactly the software versions specified in the environment file.

Running the Workflow

To run the workflow, you execute Snakemake from the project directory:

snakemake --use-conda --cores 8

The --use-conda flag tells Snakemake to create and use the Conda environments specified in the rules. The --cores flag specifies the number of CPU cores to use.

Snakemake provides several options that are useful for assembly projects. The --dry-run option shows what would be done without executing anything, which is useful for checking that the workflow is correctly specified. The --printshellcmds option shows the actual shell commands being executed, which is useful for debugging. The --report option generates an HTML report that documents the workflow execution, including the commands run and the files produced.

For large assemblies, you may want to run Snakemake on a cluster or cloud infrastructure. Snakemake supports execution on various computing infrastructures, as demonstrated by the Seq2science workflow, which can be run on a range of computing infrastructures. The configuration for cluster execution is beyond the scope of this workflow, but the same Snakefile can be adapted for different execution environments.

Quality Control and Assessment

Read Quality Control

The first analysis step in any assembly project is quality control of the raw reads. This step removes adapter sequences, trims low-quality bases, and filters reads that are too short or of poor quality.

The specific quality control steps depend on your sequencing platform. For Illumina short reads, adapter trimming and quality filtering are standard. For Pacific Biosciences HiFi reads, quality control may involve filtering by read length and predicted accuracy. For Oxford Nanopore Technologies reads, quality control may involve read length filtering and quality score assessment.

The output of quality control should include both the trimmed reads and a quality report. The quality report documents the number of reads before and after trimming, the distribution of read lengths, and the quality scores. This information is essential for interpreting the assembly results and for troubleshooting problems.

Assembly Quality Metrics

After assembly, you need to assess the quality of the resulting contigs. The standard metrics for assembly quality include:

  • N50: The length of the shortest contig such that contigs of this length or longer account for at least half of the total assembly length. A higher N50 indicates a more contiguous assembly.
  • Number of contigs: The total number of contigs in the assembly. Fewer contigs generally indicate a more complete assembly.
  • Total assembly length: The sum of all contig lengths. This should be close to the expected genome size.
  • GC content: The proportion of G and C bases in the assembly. This should be consistent with the expected GC content for the organism.
  • BUSCO completeness: The proportion of conserved single-copy genes that are present in the assembly. This measures the completeness of the assembly in terms of gene content.

The QUAST tool is commonly used to compute assembly quality metrics. It produces a report that includes all of the metrics listed above, along with additional statistics.

For a small genome assembly, you should compare your assembly metrics to the expected values for your organism. If the assembly is for a well-studied organism, you can compare to the reference genome. If the assembly is for a novel organism, you can compare to related species or to the expected genome size estimated from other methods.

Polishing and Iterative Improvement

Assembly polishing is the process of correcting errors in the assembled contigs using additional information. For short-read assemblies, polishing may involve mapping reads back to the contigs and correcting base errors. For long-read assemblies, polishing may involve using the raw signal data or additional sequencing data to correct errors.

The Colora workflow includes polishing as part of its assembly process, producing chromosome-scale assemblies that are complete with organelles. For smaller genomes, polishing is typically simpler but still important for producing a high-quality assembly.

Polishing should be treated as an iterative process. After polishing, you should reassess the assembly quality metrics to determine whether the polishing improved the assembly. If the metrics improve, you may continue polishing. If the metrics do not improve or worsen, you should stop and investigate the cause.

Integrating Git for Version Control

Initializing the Repository

The first step in integrating Git is to initialize a repository in your project directory:

git init

This creates a .git directory that stores the version history. You should then create a .gitignore file that specifies which files should not be tracked by Git. For an assembly project, the .gitignore file should exclude raw data, large intermediate files, and Conda environments:

data/raw/
results/
envs/*/
*.log

The .gitignore file ensures that you do not accidentally commit large files or files that can be regenerated by the workflow.

Committing Workflow Files

You should commit the workflow files at each stage of development. The initial commit should include the Snakefile, configuration file, environment specifications, and any custom scripts:

git add Snakefile config.yaml envs/ scripts/
git commit -m "Initial workflow for genome assembly"

As you modify the workflow, you should commit each change with a descriptive message. This creates a history of changes that documents the development of the pipeline.

Tracking Parameter Changes

One of the most important uses of Git in assembly projects is tracking parameter changes. When you change an assembly parameter, such as the k-mer size or the quality threshold, you should commit the change with a message that explains why the change was made.

This practice is essential for reproducibility. If you produce an assembly and later need to reconstruct exactly how it was produced, the Git history provides the record. You can identify the exact version of the workflow and the exact parameters that were used.

Tagging Releases

For major milestones in your workflow development, you should create Git tags. A tag is a named reference to a specific commit. For example, you might tag the version of the workflow that produced your final assembly:

git tag -a v1.0 -m "Workflow version used for final assembly"

Tags provide a convenient way to reference specific versions of your workflow. If you need to reproduce an assembly, you can check out the tagged version and run it with the appropriate configuration.

At a Glance: Workflow Decision Table

Decision PointShort-Read Data (Illumina)Long-Read Data (PacBio HiFi or ONT)Hybrid Data
Recommended assembler categoryDe Bruijn graph assemblers such as SPAdesOverlap-layout-consensus or string graph assemblers such as Flye or CanuHybrid assemblers that combine short and long reads
Primary quality control focusAdapter trimming and quality filteringRead length filtering and accuracy assessmentBoth short-read trimming and long-read filtering
Expected N50 range for bacterial genomes50 kb to 500 kb depending on coverage and genome complexity1 Mb to 5 Mb for complete or near-complete assemblies500 kb to 5 Mb depending on read combination
Polishing approachMap reads back to contigs and correct base errorsUse raw signal data or additional sequencing dataUse long reads for scaffolding and short reads for base correction
Computational requirementsModerate memory and CPUHigher memory and CPU for overlap stepsHighest memory and CPU due to combined processing

Practical Implementation Steps

Step 1: Set Up the Project Structure

Create the project directory structure described in the previous section. Initialize a Git repository and create the .gitignore file.

Step 2: Define the Configuration

Create the config.yaml file with your sample names, read file paths, and assembly parameters. Start with conservative parameters and document the rationale for each choice.

Step 3: Create Conda Environment Specifications

Create the environment specification files for each analysis stage. Pin the versions of all tools to ensure reproducibility. Test each environment by creating it and running a simple command.

Step 4: Write the Snakefile

Write the Snakefile with rules for quality control, assembly, and quality assessment. Start with a minimal workflow and add steps incrementally. Test each rule individually before running the full workflow.

Step 5: Run the Workflow

Run the workflow with the --use-conda flag. Monitor the output for errors and warnings. If a rule fails, investigate the cause and fix the issue before rerunning.

Step 6: Assess Assembly Quality

Run the quality assessment step and examine the metrics. Compare the metrics to expected values for your organism. If the assembly quality is poor, investigate the cause and adjust parameters or tools.

Step 7: Commit and Tag

Commit the workflow files and any changes to the configuration. Tag the version of the workflow that produced your final assembly.

Records and Measurements

What to Record

For each assembly run, you should record the following information:

  • The version of the workflow (Git commit hash or tag)
  • The configuration file contents
  • The versions of all software tools
  • The input data files and their checksums
  • The date and time of the run
  • The assembly quality metrics
  • Any warnings or errors that occurred during the run

This information should be stored in a run log or in the workflow documentation. The Git history provides some of this information automatically, but a separate run log is useful for recording information that is not captured by Git.

How to Record

The simplest approach is to maintain a run log file in the docs directory. Each entry should include the date, the Git commit hash, the configuration used, and the assembly quality metrics.

You can also use Snakemake's built-in reporting features. The --report option generates an HTML report that includes the workflow structure, the commands run, and the files produced. This report can be archived with the assembly results.

Interpreting Measurements

Assembly quality metrics should be interpreted in the context of your organism and sequencing data. A bacterial genome assembly with an N50 of 100 kb may be excellent, while a eukaryotic genome assembly with the same N50 may be poor. You should establish expected values for your organism before running the assembly, and you should compare your results to these expectations.

The BUSCO completeness score is particularly useful for assessing assembly completeness. A BUSCO score above 95 percent is generally considered good for a small genome assembly. Scores below 90 percent may indicate problems with the assembly or with the input data.

Common Failure Patterns

Failure Pattern 1: Version Conflicts

The most common failure in assembly workflows is version conflicts between tools. A tool may require a specific version of a library that is incompatible with another tool in the same environment.

Observation: The workflow fails with an error message about a missing or incompatible library.

Action: Use separate Conda environments for each analysis stage. If the conflict persists, check the environment specifications and adjust the package versions.

Failure Pattern 2: Insufficient Memory

Assembly tools can require substantial memory, especially for larger genomes or when using certain algorithms.

Observation: The workflow fails with an out-of-memory error.

Action: Increase the memory available to the workflow, or reduce the complexity of the assembly by adjusting parameters. For example, you may reduce the k-mer size or use a simpler assembly mode.

Failure Pattern 3: Poor Assembly Quality

The assembly completes but produces poor quality metrics, such as a low N50 or a low BUSCO score.

Observation: The assembly quality metrics are below expected values.

Action: Investigate the input data quality. Check the quality control report for issues such as adapter contamination or low-quality reads. Consider whether the assembly parameters are appropriate for your data type.

Failure Pattern 4: Contamination

The assembly contains sequences from multiple organisms, indicating contamination in the input data.

Observation: The assembly contains contigs with unexpected GC content or taxonomic assignments.

Action: Check the input data for contamination using a tool such as Kraken or BlobTools. If contamination is detected, filter the reads and rerun the assembly.

Failure Pattern 5: Workflow Non-Reproducibility

The workflow produces different results when run at different times or on different machines.

Observation: Running the same workflow on the same data produces different assemblies.

Action: Check that all software versions are pinned in the Conda environment specifications. Check that the workflow does not depend on environment variables or system-specific paths. Consider using containers for additional reproducibility.

Limitations and Interpretation Boundaries

Computational Limitations

The workflow described here is designed for small genome assemblies. It may not be appropriate for large eukaryotic genomes, which require substantially more computational resources and more complex assembly strategies. The Colora workflow provides a reference for chromosome-scale assemblies, but it requires specialized input data and significant computational infrastructure.

Data Quality Limitations

The quality of the assembly is fundamentally limited by the quality of the input data. Poor quality reads, low coverage, or contaminated samples will produce poor assemblies regardless of the workflow used. You should assess input data quality before running the assembly and interpret assembly results in the context of data quality.

Interpretation Limitations

Assembly quality metrics provide a quantitative assessment of assembly quality, but they do not capture all aspects of assembly correctness. A high N50 does not guarantee that the assembly is free of misassemblies or structural errors. You should interpret assembly metrics in the context of your biological question and consider additional validation if the assembly will be used for downstream analyses.

Reproducibility Limitations

While the combination of Snakemake, Conda, and Git provides strong reproducibility guarantees, it does not guarantee bit-for-bit identical results across all platforms. Differences in hardware, operating systems, or compiler versions can affect the results of some tools. For critical applications, you may need to use containers to achieve full reproducibility.

Safety and Data Management Context

Data Storage and Backup

Genome assembly projects involve large data files that require careful management. Raw sequencing data should be stored in a secure location with regular backups. Processed data and results should be stored separately from raw data to prevent accidental modification.

Public sequencing data should be downloaded from official repositories such as the NCBI. The NCBI provides search systems and sequence resources that allow you to locate and download public data. When using public data, you should record the accession numbers and the date of download.

Data Sharing and Publication

When publishing assembly results, you should make your workflow publicly available. The published workflows referenced in this article all provide their source code on GitHub, and the Colora workflow has been deposited in Zenodo with a persistent DOI. You should follow similar practices to enable others to reproduce your analysis.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help you develop the skills needed for reproducible analysis. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that are relevant to assembly workflows.

Professional Escalation Criteria

You should escalate to a professional bioinformatician or computational biologist when:

  • The assembly consistently produces poor quality metrics despite parameter adjustments
  • The assembly requires computational resources beyond what is available on your workstation
  • The assembly is for a large or complex genome that requires specialized assembly strategies
  • The assembly results will be used for clinical or regulatory decisions
  • You encounter errors that you cannot resolve through the troubleshooting steps described in this article

A Decision Framework for Selecting Assembly Parameters and Tools

Choosing the wrong assembler or parameter set is a common cause of failed assembly projects, yet many researchers make this choice without a structured evaluation process. This section provides a practical decision framework that you can apply before running your workflow, along with a record system for documenting your choices and a troubleshooting method for parameter-related failures.

The Parameter Selection Problem

Assembly tools expose many parameters that interact in complex ways. K-mer size in short-read assemblers affects graph complexity and repeat resolution. Read length thresholds in long-read assemblers affect coverage estimates and overlap detection. Quality trimming thresholds affect how much data survives to the assembly step. Changing one parameter often changes the optimal value of another, which makes trial-and-error approaches time-consuming and difficult to document.

The published workflows referenced in this article address this problem through careful design. The Colora workflow for chromosome-scale assembly specifies a particular combination of input data types and assembly steps that have been tested together. The RASflow workflow supports multiple mapping approaches to accommodate different data types and analysis goals. The Seq2science workflow automates the retrieval of reference data and provides tested workflows for multiple assay types. These workflows succeed because their authors made deliberate, documented choices about parameters and tools.

For your small genome assembly, you need a similar level of deliberation. The framework below helps you make those choices systematically and record them in a way that supports reproducibility.

Step 1: Characterize Your Input Data

Before selecting tools or parameters, you need to know what your data looks like. Run a quick assessment of your raw reads using a tool such as FastQC or a lightweight summary script. Record the following characteristics:

  • Read length distribution and mean read length
  • Total number of reads and estimated coverage
  • Quality score distribution across read positions
  • Presence of adapter sequences or other contaminants
  • GC content distribution and any anomalies

This characterization serves two purposes. First, it tells you whether your data is suitable for assembly at all. Very low coverage, extremely short reads, or heavy contamination may require additional sequencing instead of parameter adjustment. Second, it provides the baseline against which you interpret assembly results. If your assembly produces a GC content distribution that differs dramatically from your raw reads, you may have a contamination or bias problem.

Step 2: Estimate Genome Characteristics

The expected genome size and complexity of your organism should drive your parameter choices. For a bacterial genome, you can often estimate genome size from related species or from the total amount of sequencing data. For a small eukaryotic genome, you may need to consider ploidy, repeat content, and heterozygosity.

Record your estimates for:

  • Expected genome size in base pairs
  • Expected GC content
  • Expected repeat content or genome complexity
  • Ploidy and heterozygosity if known

These estimates help you interpret assembly metrics. A total assembly length that is substantially larger than the expected genome size may indicate contamination or a misassembly. A length that is substantially smaller may indicate incomplete assembly or data quality problems.

Step 3: Select Tools Based on Data Type

Use the At a Glance table in this article as a starting point for tool selection. For short-read data, choose a de Bruijn graph assembler such as SPAdes. For long-read data, choose an overlap-layout-consensus assembler such as Flye or Canu. For hybrid data, choose a tool that explicitly supports combining read types.

Your decision should also consider the maturity of the tool and its maintenance status. Tools that are actively maintained and widely used are more likely to have documented parameter recommendations and community support. The nf-core documentation provides standards for community pipelines that can inform your tool choices, and the Galaxy Training Network offers accessible tutorials for many common assembly tools.

Step 4: Define a Parameter Testing Protocol

instead of guessing at parameters, define a small set of candidate configurations and test them systematically. For a short-read assembly with SPAdes, you might test:

  • Default parameters with the isolate mode
  • A range of k-mer sizes appropriate for your read length
  • With and without read error correction

For a long-read assembly with Flye, you might test:

  • Default parameters
  • Different read length thresholds
  • With and without the optional polishing step

Run each candidate configuration on a subset of your data or on the full dataset if computational resources allow. Record the assembly quality metrics for each configuration. Compare the results using the metrics described in the Records and Measurements section of this article.

Step 5: Select Parameters Based on Evidence

Choose the parameter set that produces the best assembly quality metrics for your organism. The best configuration is not always the one with the highest N50. Consider the trade-offs between contiguity, completeness, and correctness. A configuration that produces a slightly lower N50 but a substantially higher BUSCO completeness score may be preferable for downstream analyses.

Document your selection rationale in your run log. This documentation is essential for reproducibility because it explains why a particular parameter set was chosen. Without this record, a collaborator or reviewer cannot understand whether the parameters were selected based on evidence or by arbitrary choice.

Record System for Assembly Decisions

Create a decision log in your project documentation that records the following for each assembly run:

  • The date and purpose of the run
  • The input data characteristics from Step 1
  • The genome estimates from Step 2
  • The tools and versions used
  • The candidate parameter sets tested
  • The assembly quality metrics for each candidate
  • The selected parameter set and the rationale for the selection
  • Any deviations from the standard workflow and the reasons for those deviations

This decision log complements the Git history. Git records what changed in your workflow files, while the decision log records why those changes were made. Both are necessary for a complete reproducibility record.

Troubleshooting Parameter-Related Failures

When an assembly produces poor results, the cause is often in the parameter selection instead of in the workflow implementation. Use this troubleshooting method to isolate the problem:

Observation 1: The assembly is fragmented with a low N50.

Check whether your k-mer size or read length threshold is appropriate for your data. For short reads, a k-mer size that is too large for your read length will produce a fragmented graph. For long reads, a read length threshold that is too high may discard useful data. Test a range of values and compare the results.

Observation 2: The assembly is too large relative to the expected genome size.

This pattern often indicates contamination or a misassembly. Check the GC content distribution of your assembly and compare it to your raw reads. If you see multiple distinct GC peaks, contamination is likely. Consider filtering your reads with a taxonomic classification tool before assembly.

Observation 3: The assembly is too small relative to the expected genome size.

This pattern may indicate that your quality trimming was too aggressive or that your coverage is insufficient. Check the quality control report to see how many reads were discarded. If coverage is low, consider whether additional sequencing is needed instead of adjusting parameters.

Observation 4: The BUSCO completeness score is low despite a reasonable N50.

This pattern suggests that the assembly is missing conserved genes, which may indicate a problem with the assembly graph or with the input data. Check whether your assembler is appropriate for your data type and whether you have sufficient coverage of the genome.

When to Escalate

If you have tested multiple parameter configurations and the assembly quality remains poor, escalate to a professional bioinformatician. Bring your decision log, the assembly quality metrics for each configuration, and the raw data characteristics. This information allows the specialist to diagnose the problem more quickly than if you present only the final assembly.

The decision framework described here is not a guarantee of a successful assembly, but it provides a structured approach that reduces the likelihood of arbitrary parameter choices and improves the reproducibility of your results. The published workflows referenced throughout this article demonstrate that careful parameter selection is a core component of reproducible assembly practice.

Frequently Asked Questions

What is the difference between Snakemake and a shell script for running an assembly?

A shell script runs commands in a fixed order and does not track dependencies between steps. Snakemake defines the analysis as a set of rules with input and output files, and it determines the order of execution based on file dependencies. This means Snakemake only runs the steps that are needed, can resume interrupted runs, and provides a structured way to specify software environments for each step. The RASflow workflow demonstrates how Snakemake and Conda together alleviate challenges with library dependencies and version conflicts.

Do I need to use Conda if I already have the assembly tools installed?

You do not need to use Conda if your tools are already installed and you can guarantee that the versions will not change. However, Conda provides a way to specify and reproduce the exact software environment used for an assembly. This is important for reproducibility, because tool versions can change when you update your system or install new software. The ARPEGGIO workflow ensures reproducibility by including both package management and containerization, and Conda is the simpler of these two approaches.

How do I choose between short-read and long-read assembly tools?

The choice depends on the type of sequencing data you have. Short-read assemblers such as SPAdes are designed for Illumina data and are appropriate for small genomes with sufficient coverage. Long-read assemblers such as Flye or Canu are designed for Pacific Biosciences or Oxford Nanopore Technologies data and can produce more contiguous assemblies. The Colora workflow uses Pacific Biosciences HiFi reads as its primary input, with optional Oxford Nanopore Technologies reads, demonstrating the use of long-read data for high-quality assemblies.

What is the minimum coverage needed for a good assembly?

The minimum coverage depends on the genome size, the read length, and the assembly algorithm. For bacterial genomes with short reads, coverage of 30 to 50 times is typically sufficient. For larger genomes or for long-read assemblies, higher coverage may be needed. You should assess the coverage of your data before assembly and interpret the assembly quality metrics in the context of the coverage.

How do I know if my assembly is complete?

Assembly completeness is typically assessed using BUSCO, which checks for the presence of conserved single-copy genes. A BUSCO completeness score above 95 percent is generally considered good for a small genome assembly. You should also compare the total assembly length to the expected genome size for your organism. If the assembly is substantially shorter than expected, it may be incomplete.

Can I use this workflow for large eukaryotic genomes?

This workflow is designed for small genome assemblies and may not be appropriate for large eukaryotic genomes. Large genomes require substantially more computational resources and more complex assembly strategies, including scaffolding and possibly Hi-C data for chromosome-scale assembly. The Colora workflow provides a reference for chromosome-scale assemblies using HiFi and Hi-C reads.

How do I share my workflow with collaborators or reviewers?

You should commit your workflow files to a Git repository and push the repository to a hosting platform such as GitHub. You should include a README file that describes the workflow, the input data requirements, and the expected outputs. For published work, you should deposit a version of the workflow in a repository such as Zenodo to obtain a persistent DOI, as was done for the Colora workflow.

What should I do if my assembly produces unexpected results?

First, check the quality control report for the input data to identify any issues with read quality or contamination. Second, check the assembly log for warnings or errors. Third, compare the assembly metrics to expected values for your organism. If the assembly is still unexpected, consider whether the assembly parameters are appropriate for your data type and whether you need to use a different assembler. If the problem persists, escalate to a professional bioinformatician.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [2] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.