Organizing Your Assembly Project: A Directory Structure and Naming Convention Template for Reproducibility
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Explicit Provenance via File Paths: The article advocates for encoding sample identifiers, analysis stages (e.g.,
raw,trimmed,assembly,polished,qc), tool names, and version numbers directly into file names and directory structures. This eliminates ambiguity, preventing confusion between different assembly runs or parameter sets, which is critical for reproducibility. - Hierarchical Data Segregation: A fundamental principle is the strict separation of immutable
raw_datafrom allderived_data(e.g., trimmed reads, assemblies, quality reports) within distinct top-level directories. This ensures that original sequencing reads are never modified, allowing for complete regeneration of downstream results. - Metadata Integration for Auditability: Essential metadata, including sample origins, sequencing platform details, and precise assembly parameters (e.g., k-mer size for de Bruijn graph assemblers like SPAdes, or read type for long-read assemblers like Flye), must be meticulously documented in
READMEfiles andsample_sheet.csvor parameter files. This documentation is crucial for understanding and replicating analyses. - Version Control for Reproducible Scripts: All analysis scripts and configuration files driving the assembly pipeline should be managed under a version control system like Git. This practice allows for precise tracking of script modifications, ensuring that the exact code used for a specific assembly run can be identified and re-executed.
- Standardized Logging for Diagnostics: Every pipeline execution must be logged, capturing standard output and error streams in a dedicated
logsdirectory. Log files, named consistently with sample and tool identifiers, are indispensable for diagnosing pipeline failures and documenting the exact commands executed.
Genome assembly projects generate large volumes of raw sequencing reads, intermediate files, polished contigs, quality reports, and annotation outputs. Without a deliberate organizational scheme, these files accumulate in ambiguous locations with names like final2.fasta or assembly_clean_v3_renamed.fasta, making it impossible to determine which file corresponds to which sample, which pipeline version produced it, or which parameters were used. This article provides a ready-to-use directory template and naming convention system for managing raw reads, assemblies, and QC files across multiple samples and versions. The template is designed for biology students, researchers, laboratory professionals, and life-science practitioners who need a practical, low-overhead method for keeping assembly projects reproducible and auditable.
The core problem is information management instead of technical skill. Assembly software produces dozens of output files per sample, and when you process multiple samples through multiple assembly tools with multiple parameter sets, the number of files multiplies rapidly. A standardized folder structure and file naming scheme solves this problem by making file provenance explicit in the file path itself. This approach follows the reproducibility principles emphasized in bioinformatics training programs, where consistent data organization is treated as a foundational skill for computational biology work. The Galaxy Training Network and The Carpentries both teach file organization and project management as core competencies for researchers working with biological data.
Why Assembly Projects Fail Without Structure
Assembly projects fail in predictable ways when directory structure and naming conventions are not established at the outset. The most common failure is version confusion, where a researcher cannot determine which assembly file corresponds to which polishing round or which parameter adjustment. This confusion leads to wasted computational time, incorrect biological conclusions, and difficulty reproducing results for publication or collaboration.
A second common failure is data loss through overwriting. When multiple assembly tools write output to the same default directory, or when a researcher reruns an assembler without changing the output path, previous results are silently destroyed. This is particularly damaging for long-read assemblies, where a single assembly run can consume significant computational resources and wall-clock time.
A third failure mode is the inability to trace quality issues back to their source. When quality reports are stored separately from the assemblies they describe, or when file names do not encode the sample identifier and assembly tool, diagnosing a low-quality assembly becomes a forensic exercise instead of a straightforward lookup. The EMBL-EBI Training program emphasizes that good data management practices, including consistent naming and directory organization, are essential for making analysis results interpretable and reusable.
The solution does not require complex software or elaborate laboratory information management systems. A simple, consistent directory template combined with a naming convention that encodes critical metadata in every file name solves most assembly project organization problems. This approach works for individual researchers managing a single bacterial genome and for small groups processing dozens of samples.
Core Principles of Assembly Project Organization
Principle 1: Separate Raw Data from Derived Data
Raw sequencing reads are immutable inputs. They should never be modified, moved, or overwritten. Derived data, including trimmed reads, assemblies, polished contigs, and quality reports, can be regenerated from raw data and pipeline parameters. This distinction is fundamental to reproducible assembly projects.
The directory structure must enforce this separation at the top level. A project directory should contain a raw_data directory that is read-only after initial placement, and separate directories for each stage of the analysis pipeline. This separation ensures that a researcher can always return to the original sequencing data and regenerate any downstream result, provided the pipeline parameters are documented.
Principle 2: Encode Metadata in File Names
File names should be self-describing. A researcher should be able to determine the sample, the analysis stage, the tool, and the version from the file name alone, without opening the file or consulting a separate log. This principle is particularly important for assembly projects because assembly tools generate files with generic names like contigs.fasta or scaffolds.fasta that do not identify the sample or the tool version.
The naming convention should include, at minimum, the sample identifier, the analysis stage, the tool name, and a version or date stamp. Additional metadata, such as the sequencing platform or the k-mer size, can be included when relevant. The goal is to make every file name unique within the project and to make the file name informative without being unwieldy.
Principle 3: Document Everything in a README
A README file serves as the project's memory. It records the purpose of the project, the source of the raw data, the pipeline steps, the software versions, the parameters used, and any deviations from the standard workflow. The README should be updated whenever the project changes, and it should be written so that a new researcher, or the original researcher six months later, can understand what was done and why.
The nf-core documentation emphasizes that reproducibility requires saving the analysis scripts along with the context around those scripts, including the input data locations, the software environment, and the parameter choices. A README template is provided later in this article.
Principle 4: Use Version Control for Scripts and Configuration
While raw data and large assembly files are not suitable for version control systems like Git, the scripts, configuration files, and parameter files that drive the assembly pipeline should be version controlled. This practice allows a researcher to track changes to the analysis over time and to reproduce exactly which script version generated which output.
The Carpentries lessons on version control with Git provide practical guidance on tracking changes to scripts and configuration files. For assembly projects, the key practice is to keep all scripts in a scripts directory within the project and to commit changes to those scripts with descriptive commit messages that reference the project stage and the changes made.
The Directory Template
The following directory structure provides a starting point for any assembly project. It separates raw data from derived data, organizes analysis stages into distinct directories, and provides dedicated locations for scripts, logs, and documentation.
project_name/
├── README.md
├── raw_data/
│ ├── sample1_R1.fastq.gz
│ ├── sample1_R2.fastq.gz
│ └── sample2_R1.fastq.gz
├── scripts/
│ ├── 01_trim.sh
│ ├── 02_assemble.sh
│ └── 03_polish.sh
├── analysis/
│ ├── 01_trimmed/
│ │ ├── sample1_R1_trimmed.fastq.gz
│ │ └── sample1_R2_trimmed.fastq.gz
│ ├── 02_assembly/
│ │ ├── sample1_tool1_v1/
│ │ │ ├── contigs.fasta
│ │ │ └── assembly_report.txt
│ │ └── sample1_tool2_v1/
│ │ ├── contigs.fasta
│ │ └── assembly_report.txt
│ ├── 03_polished/
│ │ └── sample1_tool1_v1_polished.fasta
│ └── 04_quality/
│ ├── sample1_trimming_report.html
│ └── sample1_assembly_metrics.txt
├── logs/
│ ├── 01_trim_sample1.log
│ ├── 02_assemble_sample1_tool1.log
│ └── 03_polish_sample1_tool1.log
└── metadata/
├── sample_sheet.csv
└── assembly_parameters.md
Top-Level Directories
The raw_data directory contains the original sequencing reads exactly as received from the sequencing facility. These files should be read-only. The scripts directory contains all analysis scripts, organized by pipeline stage. The analysis directory contains all derived data, organized by analysis stage. The logs directory contains all standard output and error logs from pipeline runs. The metadata directory contains sample information and parameter documentation.
This structure follows the principle that raw data and derived data should be separated, and that analysis stages should be organized in the order they are executed. The numeric prefixes on the analysis subdirectories ensure that the stages appear in the correct order when directories are listed alphabetically.
Analysis Stage Directories
The analysis directory is organized by pipeline stage. Each stage directory contains the outputs of that stage, organized by sample and tool. The stage directories are numbered to reflect the pipeline order: trimming, assembly, polishing, and quality assessment.
Within each stage directory, subdirectories are named using the sample identifier and the tool name. This organization allows multiple assembly tools to be run on the same sample without file collisions, and it allows the results of different tools to be compared directly.
Logs and Metadata
The logs directory stores the standard output and standard error from every pipeline run. Log files are named using the pipeline stage, the sample identifier, and the tool name. These logs are essential for diagnosing pipeline failures and for documenting the exact commands that were executed.
The metadata directory contains the sample sheet and the parameter documentation. The sample sheet records the sample identifier, the raw data file names, the sequencing platform, and any relevant biological information. The parameter documentation records the exact parameters used for each assembly tool, including software versions.
Naming Convention Template
The naming convention encodes critical metadata in every file name. The convention uses underscore-separated fields, with the sample identifier always first. The general format is:
{sample_id}_{stage}_{tool}_{version}.{extension}
Sample Identifiers
Sample identifiers should be short, unique, and consistent across all files for a given sample. A sample identifier might be S1, sample1, or a more descriptive identifier like Ecoli_K12_MG1655. The identifier should not contain spaces or special characters that could cause problems in command-line operations.
The sample identifier should match the identifier used in the sample sheet in the metadata directory. This linkage ensures that a researcher can look up the biological details of any sample from the file name alone.
Stage Identifiers
The stage identifier indicates where in the pipeline the file was generated. Common stage identifiers include raw, trimmed, assembly, polished, and qc. The stage identifier should match the directory name in the analysis directory, so that a file can be located by its name alone.
Tool and Version Identifiers
The tool identifier indicates which software generated the file. Common tool identifiers include fastp, trimmomatic, flye, canu, raven, medaka, and polca. The version identifier indicates the software version, such as v2.9 or 20240115.
Including the tool and version in the file name is critical for reproducibility because assembly results can vary substantially between software versions. A file named S1_assembly_flye_v2.9.fasta is immediately distinguishable from S1_assembly_flye_v2.8.fasta, and a researcher can determine which version produced which result without consulting a log file.
Example File Names
The following examples illustrate the naming convention for a typical assembly project:
S1_raw_R1.fastq.gzandS1_raw_R2.fastq.gzfor raw paired-end readsS1_trimmed_R1.fastq.gzandS1_trimmed_R2.fastq.gzfor trimmed readsS1_assembly_flye_v2.9.fastafor a Flye assemblyS1_assembly_canu_v2.2.fastafor a Canu assemblyS1_polished_medaka_v1.11.fastafor a Medaka-polished assemblyS1_qc_quast_v5.2.txtfor a QUAST quality report
These file names are self-describing. A researcher can determine the sample, the pipeline stage, the tool, and the version from the file name alone.
At a Glance
| Project Element | Recommended Practice | Common Mistake | Consequence of Mistake |
|---|---|---|---|
| Raw data storage | Store in raw_data directory, read-only, never modified | Editing or re-formatting raw reads in place | Loss of original data, inability to regenerate results |
| File naming | Encode sample, stage, tool, and version in every file name | Using generic names like contigs.fasta or final.fasta | Version confusion, inability to trace results to parameters |
| Analysis stage organization | Separate directories for each pipeline stage, numbered in execution order | Mixing all outputs in a single directory | File collisions, difficulty locating specific outputs |
| Pipeline documentation | Maintain README with software versions, parameters, and run dates | Relying on memory or informal notes | Inability to reproduce results, difficulty troubleshooting |
| Script version control | Track scripts and configuration files in Git | Storing only final scripts without history | Inability to determine which script version produced which output |
| Log retention | Save standard output and error for every run in logs directory | Discarding logs after successful runs | Inability to diagnose failures or document exact commands |
Practical Implementation Steps
Step 1: Create the Directory Structure
Create the top-level project directory and all subdirectories before beginning any analysis. The directory structure can be created with a single command or with a file manager. The key is to create the complete structure at the outset, so that all outputs have a designated location from the start.
The directory names should be consistent across projects. Using the same directory names for every assembly project allows a researcher to navigate any project without consulting documentation. The numeric prefixes on the analysis stage directories ensure that the stages appear in the correct order.
Step 2: Place Raw Data and Set Permissions
Copy the raw sequencing reads into the raw_data directory. Verify that the file names match the sample sheet entries. After verification, change the permissions on the raw data files to read-only. This step prevents accidental modification or overwriting of the original sequencing data.
The raw data files should be named using the sample identifier and the read pair designation. For paired-end reads, the _R1 and _R2 suffixes distinguish the forward and reverse reads. For long-read data, a single file per sample is typical.
Step 3: Create the Sample Sheet
Create a sample sheet in the metadata directory. The sample sheet should include the sample identifier, the raw data file names, the sequencing platform, the sequencing depth, and any relevant biological information such as the species or strain. The sample sheet serves as the authoritative record of what data exists for each sample.
The sample sheet format can be a simple CSV file with columns for each metadata field. The sample identifier column must match the sample identifiers used in the file naming convention.
Step 4: Write the README Template
Create the README file in the project root directory. The README should include the project title, the project purpose, the raw data source, the pipeline overview, the software versions, the parameter files, and the run dates. The README should be updated whenever the project changes.
A README template is provided in the next section. The template can be adapted to the specific needs of any assembly project.
Step 5: Initialize Version Control for Scripts
Initialize a Git repository in the scripts directory. Commit the initial versions of all analysis scripts. As the pipeline evolves, commit changes with descriptive messages that reference the pipeline stage and the nature of the change.
The Carpentries lessons on version control provide practical guidance on using Git for tracking changes to scripts and configuration files. The key practice is to commit early and often, with messages that describe what changed and why.
Step 6: Run the Pipeline with Logging
Run each pipeline stage with output redirection to the logs directory. The log file name should match the naming convention, such as 01_trim_S1.log for the trimming stage of sample S1. This practice ensures that every run is documented and that the exact commands and outputs are preserved.
The log files are essential for troubleshooting. When an assembly fails or produces unexpected results, the log file provides the first evidence of what went wrong.
README Template
The following README template provides a starting point for documenting an assembly project. The template includes sections for project overview, data sources, pipeline steps, software versions, and parameter documentation.
## Project Title
## Project Overview
Brief description of the project purpose and the biological question being addressed.
## Data Sources
- Sequencing facility:
- Sequencing platform:
- Library preparation:
- Raw data location: raw_data/
## Sample Information
See metadata/sample_sheet.csv for sample identifiers and biological details.
## Pipeline Overview
1. Quality trimming: fastp v0.23.4
2. Genome assembly: Flye v2.9
3. Assembly polishing: Medaka v1.11
4. Quality assessment: QUAST v5.2
## Software Versions
- fastp: v0.23.4
- Flye: v2.9
- Medaka: v1.11
- QUAST: v5.2
## Parameter Files
- Trimming parameters: scripts/01_trim.sh
- Assembly parameters: metadata/assembly_parameters.md
- Polishing parameters: scripts/03_polish.sh
## Run Logs
- Trimming logs: logs/01_trim_*.log
- Assembly logs: logs/02_assemble_*.log
- Polishing logs: logs/03_polish_*.log
## Notes
Any deviations from the standard workflow, known issues, or observations.
The README should be updated whenever the pipeline changes or when new samples are added. The goal is to make the README sufficient for a new researcher to understand the project and reproduce the analysis.
Records and Measurements
What to Record
The assembly project should maintain records of the following measurements and observations for each sample:
- Raw read count and total bases before trimming
- Trimmed read count and total bases after trimming
- Assembly statistics including number of contigs, N50, total assembly length, and GC content
- Polishing statistics including the number of rounds and the changes in assembly metrics
- Quality assessment metrics including completeness estimates and contamination estimates
These measurements should be recorded in a consistent format, either in a spreadsheet or in text files within the analysis/04_quality directory. The measurements provide the basis for comparing assembly results across tools and parameter sets.
How to Record
Quality metrics should be recorded in a tabular format with one row per sample and one column per metric. The table should include the sample identifier, the assembly tool, the assembly version, and the date of the analysis. This format allows direct comparison of metrics across samples and tools.
The EMBL-EBI Training resources emphasize the importance of recording analysis parameters and results in a structured format that supports reproducibility. A simple spreadsheet or CSV file in the metadata directory serves this purpose.
When to Record
Quality metrics should be recorded at each pipeline stage. The trimming report should be recorded immediately after trimming. The assembly metrics should be recorded immediately after assembly. The polishing metrics should be recorded after each polishing round. This practice ensures that the measurements are captured while the relevant files are still available and the context is fresh.
Common Failure Patterns
Failure Pattern 1: Generic File Names
The most common failure pattern is the use of generic file names that do not identify the sample or the tool. Assembly tools typically write output to files named contigs.fasta or scaffolds.fasta. When multiple samples are assembled in the same directory, these files overwrite each other, and the results are lost.
The solution is to run each assembly in a separate directory named with the sample identifier and the tool name, as shown in the directory template. The output files can then retain their generic names within the sample-specific directory, and the directory name provides the context.
Failure Pattern 2: Mixing Raw and Derived Data
A second common failure pattern is storing raw reads and derived data in the same directory. This practice makes it difficult to distinguish original data from processed data and increases the risk of accidentally modifying raw reads.
The solution is to maintain the separation between raw_data and analysis directories. Raw data should be read-only, and all derived data should be stored in the analysis directories.
Failure Pattern 3: Incomplete Documentation
A third common failure pattern is incomplete documentation of software versions and parameters. A researcher may remember the assembly tool but not the version, or may remember the parameters but not the exact command. This incomplete documentation makes reproduction difficult or impossible.
The solution is to record software versions and parameters in the README and in the parameter files at the time the analysis is run. The log files provide a backup record of the exact commands executed.
Failure Pattern 4: Overwriting Previous Versions
A fourth common failure pattern is overwriting previous versions of assemblies or polished contigs. When a researcher reruns an assembly with different parameters, the previous output is destroyed if the output path is the same.
The solution is to include the version or date in the output directory name, such as S1_assembly_flye_v2.9 and S1_assembly_flye_v2.9_v2. This practice preserves the history of the analysis and allows comparison of results across versions.
Quality Controls and Assessment
Assembly Quality Metrics
Assembly quality is assessed using metrics that describe the contiguity, completeness, and accuracy of the assembly. The N50 statistic describes the contiguity of the assembly, representing the length of the contig at which half of the total assembly length is contained in contigs of that length or longer. The number of contigs and the total assembly length provide additional context for interpreting N50.
Completeness estimates, such as those provided by BUSCO, assess whether the assembly contains expected single-copy orthologs. These estimates provide a measure of how much of the genome is represented in the assembly. Contamination estimates assess whether sequences from other organisms are present in the assembly.
The Galaxy Training Network provides tutorials on genome assembly quality assessment that describe these metrics and their interpretation. The tutorials emphasize that no single metric is sufficient to assess assembly quality and that multiple metrics should be considered together.
When to Assess Quality
Quality should be assessed at each pipeline stage. The raw reads should be assessed for quality before trimming. The trimmed reads should be assessed to confirm that trimming improved quality. The assembly should be assessed immediately after assembly and again after polishing. This staged assessment allows a researcher to identify the stage at which problems arise.
Interpreting Quality Metrics
Quality metrics should be interpreted in the context of the expected genome size and complexity. A bacterial genome assembly with an N50 of 100 kb may be excellent, while a eukaryotic genome assembly with the same N50 may be fragmented. The expected genome size and the sequencing platform should be considered when interpreting quality metrics.
The EMBL-EBI Training resources provide guidance on interpreting assembly quality metrics in the context of different genome types and sequencing platforms. The key principle is that quality metrics are comparative and should be evaluated relative to expectations for the organism and the data type.
Workflow Choices and Tradeoffs
Assembly Tool Selection
The choice of assembly tool depends on the sequencing platform and the genome characteristics. Long-read assemblers such as Flye, Canu, and Raven are designed for Oxford Nanopore and PacBio data. Short-read assemblers such as SPAdes are designed for Illumina data. Hybrid assemblers combine both data types.
The directory template and naming convention accommodate multiple assembly tools by including the tool name in the directory and file names. This design allows a researcher to run multiple assemblers on the same sample and compare the results directly.
The nf-core documentation describes community-developed pipelines that integrate multiple assembly tools and quality assessment steps. These pipelines provide a standardized workflow that can be adapted to individual projects.
Polishing Strategy
Polishing is the process of correcting errors in the assembly using additional sequencing data. Long-read assemblies are typically polished using the raw reads or using short-read data. The polishing strategy depends on the error profile of the sequencing platform and the availability of additional data.
The naming convention should include the polishing tool and the polishing round in the file name, such as S1_polished_medaka_v1.11_round2.fasta. This practice allows a researcher to track the polishing history and to compare the assembly before and after each polishing round.
Quality Assessment Tools
Quality assessment tools include QUAST for assembly statistics, BUSCO for completeness estimation, and CheckM for contamination estimation. These tools provide complementary information about assembly quality.
The output files from quality assessment tools should be stored in the analysis/04_quality directory with names that identify the sample and the tool, such as S1_qc_quast_v5.2.txt and S1_qc_busco_v5.4.txt. This organization allows a researcher to locate all quality reports for a sample in a single directory.
Limitations and Interpretation Constraints
Assembly Quality Is Relative
Assembly quality metrics are relative to the expected genome size, the sequencing platform, and the assembly tool. A metric that indicates a good assembly for one organism may indicate a poor assembly for another. Quality metrics should be interpreted in the context of the specific project.
Completeness Estimates Have Limitations
Completeness estimates based on single-copy orthologs assume that the genome contains the expected set of orthologs. Genomes with unusual gene content, such as highly reduced genomes or genomes with extensive gene duplication, may produce misleading completeness estimates. The Computational Analysis of Telomerase RNA Evolution in Caenorhabditis Species study demonstrates that even well-studied genomes can contain unexpected features that affect computational analysis, highlighting the need for careful interpretation of automated results.
Assembly Errors Persist After Polishing
Polishing reduces but does not eliminate assembly errors. Repetitive regions, structural variants, and regions with extreme GC content may remain problematic even after multiple polishing rounds. The assembly should be considered a draft unless validated by additional evidence.
Software Versions Affect Results
Assembly results can vary substantially between software versions. A file named with the tool name but without the version does not provide sufficient information for reproduction. The naming convention must include the version to ensure that the exact software environment can be reconstructed.
Professional Escalation Criteria
When to Seek Assistance
A researcher should seek assistance from a bioinformatics core facility, a collaborator with assembly expertise, or an online community when the following situations arise:
- The assembly quality metrics are consistently poor across multiple tools and parameter sets
- The assembly fails to complete or crashes repeatedly
- The quality assessment indicates contamination that cannot be explained by the expected biology
- The assembly results are inconsistent with the expected genome size or complexity
The Bioconductor project provides a community of developers and users who can assist with genomic analysis problems. The Galaxy Training Network provides tutorials and a user community for troubleshooting analysis workflows.
What to Prepare Before Seeking Assistance
Before seeking assistance, a researcher should prepare the following materials:
- The sample sheet with sample identifiers and biological details
- The README with software versions and parameters
- The log files from the pipeline runs
- The quality assessment reports
- A description of the specific problem and the steps already taken
These materials allow an expert to diagnose the problem without requiring the researcher to repeat the analysis or to provide missing context.
When to Consider Alternative Approaches
If the assembly quality does not improve after multiple attempts with different tools and parameters, the researcher should consider alternative approaches. These alternatives may include:
- Generating additional sequencing data, such as higher coverage or a different sequencing platform
- Using a hybrid assembly approach that combines long-read and short-read data
- Using a reference-guided assembly approach if a closely related reference genome is available
- Seeking assistance from a specialized assembly service or core facility
The decision to pursue alternative approaches should be based on the quality metrics and the biological question being addressed. The StrainCascade workflow demonstrates that automated, modular pipelines can integrate assembly, annotation, and functional profiling into a single reproducible framework, providing a model for comprehensive analysis that extends beyond basic assembly.
Safety and Data Management Context
Data Storage and Backup
Assembly projects generate large data volumes, particularly for long-read sequencing. The raw data and the final assemblies should be backed up to a separate storage system. The intermediate files can be regenerated from the raw data and the pipeline parameters, so they do not require the same level of backup protection.
The backup strategy should be documented in the README. The backup location and the backup frequency should be recorded so that a researcher can verify that the data is protected.
Data Sharing and Publication
When assembly results are shared or published, the directory structure and naming convention facilitate data transfer and interpretation. A collaborator who receives the project directory can navigate the structure and locate the relevant files without requiring a separate explanation.
The NCBI Data Resources provide repositories for depositing raw sequencing reads and assembled genomes. The directory structure and naming convention described in this article can be adapted to the submission requirements of these repositories.
Long-Term Data Preservation
Assembly projects should be preserved in a form that can be interpreted in the future. The README, the sample sheet, and the parameter files provide the context needed to interpret the assembly files. These documentation files should be preserved alongside the data files.
The EMBL-EBI Training resources emphasize that data preservation requires more than storing the data files. The metadata and the analysis context must be preserved to make the data interpretable and reusable.
Decision Framework for Choosing Between Single-Project and Multi-Project Organization
The directory template presented above works well for a single assembly project, but many researchers manage multiple related projects simultaneously. A common practical question is whether to create one large directory structure containing all samples or to maintain separate project directories. The answer depends on the relationship between the samples and the intended use of the assemblies.
Single-Project Structure
A single-project structure places all samples under one top-level directory. This approach is appropriate when all samples are part of the same study, share the same sequencing platform, and will be analyzed with the same pipeline. Examples include a comparative genomics study of multiple strains of the same species or a mutation screening project where all samples were sequenced in the same batch.
The single-project structure simplifies comparative analysis because all assemblies reside in the same directory tree. Quality metrics can be compared directly across samples, and a single README documents the entire study. The main limitation is that the directory can become unwieldy when the project contains dozens of samples or when different samples require different analysis approaches.
Multi-Project Structure
A multi-project structure maintains separate top-level directories for each distinct study or experimental group. This approach is appropriate when samples come from different sequencing runs, use different sequencing platforms, or address different biological questions. Examples include a laboratory that maintains separate projects for different species, or a researcher who has one project for clinical isolates and another for environmental samples.
The multi-project structure keeps each study self-contained, which simplifies data sharing and publication. Each project directory can be archived or transferred independently. The main limitation is that cross-project comparisons require navigating multiple directory trees and reconciling potentially different naming conventions.
Decision Criteria
Use the following criteria to decide between single-project and multi-project organization:
- Sequencing platform consistency: If all samples were sequenced on the same platform with the same library preparation, a single project is appropriate. If platforms differ, separate projects reduce confusion about platform-specific quality expectations.
- Analysis pipeline consistency: If all samples will be processed through the same assembly and polishing pipeline, a single project simplifies pipeline management. If different samples require different tools or parameters, separate projects prevent accidental application of the wrong parameters.
- Biological question alignment: If all samples address the same biological question, a single project supports direct comparison. If samples address different questions, separate projects keep the analysis focused.
- Collaboration and sharing: If different collaborators will receive different subsets of samples, separate projects make data transfer straightforward. A single project requires extracting subsets before sharing.
- Publication timeline: If samples will be published in separate manuscripts, separate projects align with the publication units. A single project covering multiple manuscripts requires careful documentation of which samples belong to which manuscript.
The Galaxy Training Network emphasizes that project organization should reflect the analysis workflow and the intended use of the results. The decision between single-project and multi-project organization should be made before any analysis begins, because reorganizing directories after data generation is error-prone and can break file paths recorded in logs and scripts.
Hybrid Approach
A hybrid approach maintains separate project directories but uses a consistent naming convention across all projects. This approach combines the benefits of both structures. Each project is self-contained, but the consistent naming convention allows a researcher to navigate any project without consulting project-specific documentation.
The hybrid approach is particularly useful for laboratories that maintain multiple ongoing assembly projects. A researcher can move between projects without relearning the directory structure, and cross-project comparisons are possible because file names encode the same metadata fields.
The nf-core documentation describes how community pipelines use consistent directory and file naming conventions across different analysis runs. This consistency is a key factor in making the pipelines reproducible and transferable between research groups.
Record System for Project Organization Decisions
The decision between single-project and multi-project organization should be documented in the README file. The following record system captures the decision and its rationale:
Project Organization Record
Create a section in the README titled Project Organization that records:
- The organization type chosen (single-project or multi-project)
- The rationale for the choice, referencing the decision criteria above
- The date the decision was made
- Any changes to the organization structure and the date of each change
This record ensures that a researcher returning to the project after a gap understands why the directory structure looks the way it does. It also provides context for collaborators who receive the project directory.
Sample-to-Project Mapping
For multi-project structures, maintain a mapping of samples to projects. This mapping can be a simple table in each project README or a central spreadsheet in the laboratory. The mapping should include the sample identifier, the project directory name, and the date the sample was added to the project.
The EMBL-EBI Training resources emphasize that sample tracking is a critical component of reproducible research. A sample-to-project mapping prevents samples from being analyzed in the wrong project context and ensures that all samples are accounted for.
Pipeline Version Tracking Across Projects
When the same pipeline is used across multiple projects, track the pipeline version in each project README. This tracking allows a researcher to determine whether differences in assembly results between projects are due to biological differences or pipeline version differences.
The Bioconductor project provides guidance on managing software versions in reproducible analyses. The key practice is to record the exact software versions and parameters in each project, even when the same pipeline is used across projects.
Troubleshooting Method for Directory Structure Problems
Even with a well-designed directory template, problems can arise. The following troubleshooting method addresses common directory structure issues.
Symptom: Files Appear in the Wrong Directory
When assembly output files appear in unexpected locations, the first step is to check the pipeline script for hard-coded output paths. Many assembly tools write output to the current working directory or to a default output directory. If the script does not explicitly set the output directory, files will land wherever the script was executed.
The solution is to modify the pipeline script to set the output directory explicitly using the naming convention. The log files in the logs directory provide evidence of where the script was executed and where the output was written.
Symptom: Duplicate Files with Different Names
Duplicate files with different names typically arise when the same analysis is run multiple times with slightly different naming conventions. This situation creates confusion about which file is the authoritative version.
The solution is to compare the file contents using checksums. If the files are identical, delete the duplicate and standardize on the naming convention. If the files differ, determine which version was generated with the correct parameters and archive the other version.
Symptom: Missing Files Referenced in Logs
When log files reference files that no longer exist, the files were either moved, deleted, or overwritten. The first step is to check the backup system. If no backup exists, the analysis must be rerun to regenerate the missing files.
This situation highlights the importance of the read-only raw data directory and the version control system for scripts. Raw data cannot be lost if permissions are set correctly, and scripts can be recovered from version control.
Symptom: Inconsistent Naming Across Samples
Inconsistent naming across samples typically arises when different researchers or different pipeline versions contribute to the same project. This inconsistency makes it difficult to locate files and compare results.
The solution is to establish a naming convention checklist and to verify file names against the checklist when files are added to the project. The sample sheet in the metadata directory provides the authoritative list of sample identifiers, and file names should be verified against this list.
The Carpentries lessons on project organization recommend periodic audits of file naming consistency. A simple script can check that all files in the project follow the naming convention and report any deviations.
Comparison of Organization Approaches
| Organization Approach | Best For | Key Advantage | Key Limitation |
|---|---|---|---|
| Single-project structure | One study, one platform, one pipeline | Direct comparison across samples | Becomes unwieldy with many samples |
| Multi-project structure | Multiple studies, different platforms | Self-contained, easy sharing | Cross-project comparison requires navigation |
| Hybrid approach | Laboratories with ongoing projects | Consistent navigation across projects | Requires discipline to maintain consistency |
The choice between these approaches should be made at the project outset and documented in the README. The decision affects every subsequent analysis step, and changing the organization structure after data generation is costly and error-prone.
Practical Assessment Steps
Step 1: Assess Current Project Organization
Before applying the decision framework, assess the current state of the project. List all samples, their sequencing platforms, and the analysis pipelines applied. Determine whether the samples share a common biological question and whether they will be published together.
Step 2: Apply the Decision Criteria
Apply the decision criteria to determine whether a single-project or multi-project structure is appropriate. Document the rationale for the decision in the README. If the project is already underway, assess whether the current structure matches the recommended structure.
Step 3: Implement the Chosen Structure
Create the directory structure for the chosen approach. For a single-project structure, create one top-level directory with the template described above. For a multi-project structure, create separate top-level directories for each project, each with the same internal template.
Step 4: Verify File Placement
After implementing the structure, verify that all existing files are in the correct directories. Move any misplaced files and update the README to reflect the current state. Verify that the log files reference the correct file paths.
Step 5: Document the Decision
Record the organization decision and the rationale in the README. Include the date of the decision and any changes made during implementation. This documentation ensures that the organization structure is transparent to collaborators and to future researchers.
The Galaxy Training Network provides practical tutorials on project organization that complement the decision framework presented here. These tutorials emphasize that consistent organization is a prerequisite for reproducible analysis and that the effort invested in organization pays dividends throughout the project lifecycle.
Frequently Asked Questions
Why should I use a directory template instead of organizing files as I go?
Organizing files as you go leads to inconsistent naming and directory structures that become difficult to navigate as the project grows. A directory template established at the outset ensures that every file has a designated location and that the structure is consistent across samples and projects. This consistency reduces the cognitive load of managing the project and makes it easier to locate files, compare results, and reproduce analyses.
What is the minimum information that should be in every file name?
Every file name should include the sample identifier, the pipeline stage, the tool name, and the tool version. This minimum information allows a researcher to determine what the file contains, which sample it belongs to, and which software version generated it. Additional metadata, such as the polishing round or the parameter set, should be included when relevant.
How do I handle multiple assembly tools for the same sample?
Run each assembly tool in a separate directory named with the sample identifier and the tool name. This approach prevents file collisions and allows direct comparison of the results. The naming convention should include the tool name in the file names as well, so that files can be identified even if they are moved out of their original directories.
Should I store raw reads in a compressed format?
Raw sequencing reads are typically stored in compressed format to reduce storage requirements. The compression should be applied by the sequencing facility or during the initial data transfer. The raw data directory should contain the files exactly as received, without additional processing or re-formatting.
How do I document software versions and parameters?
Record software versions and parameters in the README file and in the parameter files in the metadata directory. The log files from each pipeline run provide a backup record of the exact commands executed. The README should be updated whenever the software environment or the parameters change.
What should I do if I accidentally overwrite an assembly file?
If an assembly file is overwritten, the original result may be recoverable from the log files or from the version control history if the output was committed. If the file cannot be recovered, the assembly must be regenerated by rerunning the pipeline with the same parameters. This situation highlights the importance of the naming convention and the directory structure in preventing accidental overwrites.
How do I compare assemblies from different tools?
Store the quality assessment reports for each assembly in the analysis/04_quality directory with file names that identify the sample and the tool. The quality metrics can then be compared directly in a spreadsheet or table. The comparison should consider multiple metrics, including N50, number of contigs, completeness, and contamination.
When should I update the README file?
Update the README whenever the pipeline changes, when new samples are added, when software versions change, or when any deviation from the standard workflow occurs. The README should be updated at the time the change is made, not later, to ensure that the documentation is accurate and complete.
Related Bioinformatics Guides
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Evaluating Genome Assembly Quality: Metrics and Tools
- Metagenomics Pipeline: From Raw Reads to Taxonomic and Functional Profiles
- Genomic Data Processing: From Raw Sequencing to Analysis-Ready Files
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- PYEAST - A Computational Toolkit for Saccharomyces cerevisiae Genetic Engineering.. 2026.
- Protocol to predict gene expression from transcriptomic data using PREDICT.. 2026.
- <,i>,StrainCascade<,/i>,: An automated, modular workflow for high-throughput long-read bacterial genome reconstruction and characterization.. 2026.
- Computational Analysis of Telomerase RNA Evolution in <,i>,Caenorhabditis<,/i>, Species.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.