# How to Write a Reproducible Assembly Methods Section: A Checklist for Publications and Database Submissions


## Key Takeaways

- **Input Data Precision is Paramount:** Accurately describe sequencing platforms (e.g., PacBio Sequel II with HiFi chemistry vs. Illumina NovaSeq 2x150 bp), library preparation methods, read length distributions, and estimated coverage depth, as these directly influence read error profiles and assembly strategies. Omitting these details, such as only providing a GenBank accession, prevents reproducibility.

- **Exact Software Versions and Parameters are Non-Negotiable:** Specify the precise version number (or commit hash for GitHub-based tools) for *every* software tool used, including assemblers, polishers, and QC utilities. Report all non-default parameters with their exact values and the rationale for their selection, as default settings can change significantly between software versions, impacting assembly outcomes.

- **Executable Commands are the Gold Standard for Reproducibility:** Instead of prose descriptions, provide the complete, executable command lines used for each assembly step, including all flags and arguments. This ensures that parameters, input files, and execution environments are precisely captured, allowing independent replication of the workflow.

- **Comprehensive Quality Assessment is Essential:** Report standard assembly metrics such as N50, L50, total assembly length, and GC content, alongside the specific tools used for their calculation. For eukaryotic genomes, include completeness assessment using tools like BUSCO, and for bacterial genomes, report expected single-copy gene presence and contamination levels.

- **Detailed Documentation of Quality Control Steps is Crucial:** Clearly outline every quality control step applied to raw reads and the final assembly, including read trimming (tool, version, quality threshold, minimum read length), filtering criteria, and error correction methods. Documenting these steps allows for the assessment of data integrity and its impact on the assembly.

---

Genome assembly methods sections fail when they omit the details that let another researcher repeat the work. Reviewers and database curators need to know what sequencing data went in, which software versions were used, what parameters were set, and how quality was assessed. This article provides a practical checklist for writing assembly methods that meet the standards of peer review and public sequence databases.

The scope here covers genome assembly from short reads, long reads, or hybrid approaches. The guidance applies to small bacterial genomes, large eukaryotic genomes, and metagenomic assemblies. The focus is on the methods text itself, not on how to run the assembly. You will find concrete lists of details to include, example phrasing for common situations, and explanations of why each detail matters for reproducibility.

## At a Glance

The table below summarizes the essential components of a reproducible assembly methods section and the common pitfalls that make each component inadequate.

| Component | What Reviewers Expect | Common Failure |
| --- | --- | --- |
| Input data description | Sequencing platform, read length, coverage depth, library preparation, quality filtering steps | Listing only the GenBank accession without describing the raw data characteristics |
| Software and versions | Exact version numbers for every tool, including assemblers, polishers, and QC utilities | Writing "used SPAdes" without a version number or commit hash |
| Parameters and settings | All non-default parameters, plus the rationale for choosing them | Stating "default parameters were used" when the tool version changed defaults |
| Quality assessment | Metrics such as N50, completeness, and base accuracy, with the tools used to calculate them | Reporting only contig count without BUSCO or read mapping validation |
| Reproducibility aids | Commands, scripts, or workflow files that capture the exact execution | Describing the workflow in prose without providing executable code |

## Why Assembly Methods Sections Fail Reproducibility Checks

Reproducibility in genome assembly means that an independent researcher can take the same input data and the same described procedure and obtain a comparable assembly. The scientific literature contains many assembly papers that cannot be reproduced because the methods text omits critical operational details.

The Critical Assessment of Metagenome Interpretation (CAMI) challenge provides direct evidence that parameter settings markedly affect assembly performance. In the first CAMI benchmark, researchers found that parameter settings had a substantial impact on program performance, underscoring their importance for reproducibility [<a href="#ref-1">1</a>]. The second CAMI challenge, which analyzed 5,002 results from 76 program versions, found that reproducibility remained a concern in clinical pathogen detection [<a href="#ref-2">2</a>]. These findings show that the choice of parameters is not a minor detail. It can change the outcome of the assembly as much as the choice of software itself.

The practical implication is straightforward. A methods section that says "default parameters" without specifying the software version is not reproducible, because defaults change between versions. A methods section that describes the assembly in prose without providing the actual commands forces the reader to guess at the execution details.

## Core Principles for Writing Assembly Methods

### Describe the Input Data With Precision

The assembly methods section must begin with a complete description of the sequencing data. This includes the biological sample, the sequencing platform, the library preparation method, the read length distribution, and the estimated coverage.

For the biological sample, state the species, strain, or isolate identifier. If the sample is a metagenome, describe the source environment and the DNA extraction method. If the sample is a cell line or tissue, state the exact source and any relevant handling steps.

For the sequencing platform, name the instrument model and the chemistry version. A PacBio Sequel II run with HiFi chemistry produces different data than a PacBio RS II run with continuous long-read chemistry. An Illumina MiSeq run with 2x300 bp reads produces different data than a NovaSeq run with 2x150 bp reads. The platform and chemistry determine the error profile of the reads, which directly affects the assembly strategy.

For coverage, report the estimated depth of sequencing. This can be calculated from the read count, the read length, and the estimated genome size. For a genome of unknown size, state the method used to estimate coverage, such as k-mer analysis or flow cytometry.

### Name Every Software Tool With Its Exact Version

Every tool used in the assembly workflow must be named with its exact version. This includes the assembler, the read quality control tool, the error corrector, the polisher, the scaffolding tool, and the quality assessment tool.

The version number alone is often insufficient. Many tools are developed on GitHub or similar platforms, and the version number may not capture the exact state of the code. For tools installed from a package manager such as Bioconductor, the package version and the Bioconductor release version should both be stated [<a href="#ref-3">3</a>]. For tools run through a workflow manager such as nf-core, the pipeline version and the configuration used should be reported [<a href="#ref-4">4</a>].

When the exact version is not known, the methods section should state the date the tool was installed and the source from which it was obtained. This allows a reader to identify the exact code state even if the version number is ambiguous.

### Report Parameters as Executable Commands

The most reproducible way to report assembly parameters is to include the actual commands used. This can be done in the methods text, in a supplementary file, or in a public repository. The commands should include every flag and every argument, beyond the non-default ones.

For example, instead of writing "the genome was assembled with Flye using default parameters," write the full command:

```
flye --nano-raw reads.fastq --genome-size 5m --out-dir flye_output --threads 32
```

This command tells the reader the input file name, the genome size estimate, the output directory, and the thread count. If the reader wants to reproduce the assembly, they can run this exact command.

For tools that use configuration files, such as the workflow managers used in nf-core pipelines, the configuration file should be included or referenced [<a href="#ref-4">4</a>]. For tools that use parameter sets, such as the Bioconductor packages for sequence analysis, the parameter objects should be described [<a href="#ref-3">3</a>].

### Document the Quality Control Steps

The methods section must describe every quality control step applied to the raw reads and to the final assembly. This includes read trimming, read filtering, error correction, and contamination screening.

For read trimming, state the tool, the version, the quality threshold, and the minimum read length retained. For read filtering, state the criteria used to remove reads, such as adapter contamination or low complexity. For error correction, state whether it was performed and with which tool.

For the final assembly, report the quality metrics and the tools used to calculate them. The standard metrics include N50, L50, contig count, total assembly length, and GC content. For eukaryotic genomes, completeness should be assessed with a tool such as BUSCO. For bacterial genomes, completeness can be assessed by checking for the presence of expected single-copy genes.

## The Practical Workflow for Writing the Methods Section

### Step 1: Record Every Command During the Assembly

The methods section cannot be written accurately from memory. The assembly workflow must be documented as it is executed. This means keeping a lab notebook, a shell history, or a workflow file that records every command.

The Carpentries lessons on the Unix shell and version control with Git provide the foundational skills for this kind of documentation [<a href="#ref-5">5</a>]. The shell history captures the commands as they are run. Git captures the state of the scripts and configuration files. Together, they provide a complete record of the assembly workflow.

For researchers who use graphical interfaces or point-and-click tools, the documentation burden is higher. Every menu selection and every parameter change must be recorded manually. This is one reason why command-line tools and workflow managers are preferred for reproducible assembly.

### Step 2: Identify the Essential Details for Each Tool

Each tool in the assembly workflow has a set of essential details that must be reported. The table below lists the details for common categories of tools.

| Tool Category | Essential Details to Report |
| --- | --- |
| Read QC and trimming | Tool name, version, quality threshold, minimum read length, adapter sequences removed |
| Error correction | Tool name, version, whether correction was performed before or after assembly |
| Assembler | Tool name, version, input read type, genome size estimate, k-mer sizes (for k-mer based assemblers), expected coverage |
| Polisher | Tool name, version, input read type, number of polishing rounds |
| Scaffolder | Tool name, version, input data (such as mate pairs or Hi-C), minimum scaffold length |
| Quality assessment | Tool name, version, reference database used, metrics calculated |

### Step 3: Write the Methods Text From the Record

The methods text should be written directly from the recorded commands and parameters. The text should be organized in the order the steps were performed. Each paragraph should describe one stage of the workflow.

The first paragraph describes the input data. The second paragraph describes read quality control. The third paragraph describes the assembly. The fourth paragraph describes polishing and scaffolding. The fifth paragraph describes quality assessment.

This structure matches the expectations of reviewers and database curators. It allows a reader to follow the workflow from raw data to final assembly without jumping back and forth.

### Step 4: Include the Commands in a Supplementary File or Repository

The methods text should summarize the workflow in prose, but the full commands should be available in a supplementary file or a public repository. This is the standard practice for reproducible bioinformatics.

The Galaxy Training Network provides tutorials that demonstrate how to document and share analysis workflows [<a href="#ref-6">6</a>]. The nf-core documentation describes how to package workflows for distribution and reproducibility [<a href="#ref-4">4</a>]. These resources show the expected format for sharing executable workflows.

For researchers who cannot share their data publicly, the commands can still be shared. The commands describe the procedure, not the data. Sharing the commands allows another researcher to run the same procedure on their own data.

## Options and Tradeoffs in Assembly Methods Reporting

### Prose Description Versus Executable Commands

The traditional methods section describes the assembly in prose. The modern standard includes executable commands. The tradeoff is readability versus precision.

Prose descriptions are easier to read but lose precision. A phrase like "the reads were assembled with Flye" does not tell the reader which version was used, what parameters were set, or how the input was prepared. Executable commands are precise but harder to read. A long command with many flags can obscure the biological logic of the workflow.

The solution is to use both. The prose describes the workflow in biological terms. The commands provide the exact execution details. The prose should reference the commands, and the commands should be available in a supplementary file or repository.

### Default Parameters Versus Explicit Parameters

Many assembly tools have default parameters that work well for common cases. The question is whether the methods section should report only the non-default parameters or all parameters.

The CAMI challenge found that parameter settings markedly affected performance [<a href="#ref-1">1</a>]. This means that the default parameters of one version may not be the default parameters of another version. Reporting only the non-default parameters is insufficient, because the reader cannot know what the defaults were.

The safest approach is to report all parameters, including the defaults. This can be done by including the full command or the full configuration file. If the tool has a command that prints the parameters used, such as a log file, that log file should be included.

### Version Numbers Versus Commit Hashes

Version numbers are the standard way to identify software. However, version numbers can be ambiguous. A tool may have multiple releases with the same version number, or the version number may not be updated for every change.

For tools developed on GitHub, the commit hash provides a more precise identifier. The commit hash identifies the exact state of the code. The nf-core documentation recommends reporting the pipeline version and the configuration used [<a href="#ref-4">4</a>]. For tools installed from Bioconductor, the package version and the Bioconductor release version should be reported [<a href="#ref-3">3</a>].

The practical approach is to report both the version number and the source. If the tool was installed from a package manager, state the package manager and the version. If the tool was installed from source, state the commit hash or the download date.

## Observations and Measurements That Support Reproducibility

### The Impact of Long-Read Data on Assembly Quality

The choice of sequencing technology has a direct impact on assembly quality. The CAMI II challenge found substantial improvements in assembly, some due to long-read data [<a href="#ref-2">2</a>]. This finding supports the decision to use long-read sequencing for complex assemblies.

The pig reference genome project provides a concrete example. The draft reference genome Sscrofa10.2, built with older clone-based sequencing methods, was incomplete and contained unresolved redundancies and errors. The improved assemblies Sscrofa11.1 and USMARCv1.0, built with long-read technologies, achieved substantially higher continuity and accuracy [<a href="#ref-7">7</a>]. The methods section for this project would need to describe the long-read platform, the coverage, and the assembly strategy to be reproducible.

For researchers choosing between short-read and long-read assembly, the methods section should justify the choice. The justification can cite the expected genome complexity, the available sequencing budget, and the downstream analysis requirements.

### The Role of High-Fidelity Reads in Assembly Accuracy

High-fidelity long reads, such as PacBio HiFi reads, provide both long read length and high per-base accuracy. The HiCanu assembler was designed to leverage the full potential of HiFi reads through homopolymer compression, overlap-based error correction, and aggressive false overlap filtering [<a href="#ref-8">8</a>].

For a methods section describing a HiFi assembly, the key details include the read length distribution, the per-base accuracy, and the coverage. The HiCanu study reported that for diploid human genomes sequenced to 30x HiFi coverage, the assembler achieved superior accuracy and allele recovery compared to the current state of the art [<a href="#ref-8">8</a>]. The methods section should state the coverage and the expected accuracy to allow the reader to assess the assembly quality.

### The Challenge of Related Strains in Metagenomic Assembly

Metagenomic assembly faces a specific challenge that single-genome assembly does not. The presence of related strains in a metagenome substantially affects assembly and binning performance. The first CAMI challenge found that assembly and genome binning programs performed well for species represented by individual genomes but were substantially affected by the presence of related strains [<a href="#ref-1">1</a>].

A methods section for a metagenomic assembly should describe how the workflow handles strain diversity. This includes the choice of assembler, the parameters for handling coverage variation, and the binning strategy. The methods section should also state the expected limitations, such as the difficulty of resolving closely related strains.

## Records and Measurements for the Methods Section

### What to Record During the Assembly

The following records should be kept during the assembly workflow:

1. The exact version of every software tool, including the operating system and the installation method
2. The full command for every step, including all flags and arguments
3. The input file names and their checksums, to verify data integrity
4. The output file names and their sizes, to track the assembly progress
5. The runtime and memory usage for each step, to identify bottlenecks
6. The log files from each tool, which often contain parameter summaries and warnings

These records serve two purposes. They allow the methods section to be written accurately, and they allow the assembly to be debugged if problems arise.

### What to Measure in the Final Assembly

The final assembly should be measured with standard quality metrics. The choice of metrics depends on the type of assembly.

For all assemblies, report the total assembly length, the contig count, the N50, and the GC content. The N50 is the contig length at which half of the assembly is in contigs of that length or longer. The L50 is the number of contigs needed to reach half of the assembly.

For eukaryotic assemblies, report the completeness with BUSCO or a similar tool. BUSCO assesses the presence of expected single-copy genes. The pig reference genome project used annotation to assess the quality of the assemblies, finding that the improved assemblies had substantially higher continuity and accuracy [<a href="#ref-7">7</a>].

For bacterial assemblies, report the completeness and the contamination. Completeness is the percentage of expected genes present. Contamination is the percentage of sequences that come from other organisms.

For metagenomic assemblies, report the number of bins, the completeness and contamination of each bin, and the taxonomic assignment of each bin. The CAMI challenges provide benchmarks for assessing metagenomic assembly and binning performance [<a href="#ref-2">2</a>][<a href="#ref-1">1</a>].

## Common Failure Patterns in Assembly Methods Sections

### Failure Pattern 1: Missing Version Numbers

The most common failure is the omission of version numbers. A methods section that says "the genome was assembled with SPAdes" without a version number cannot be reproduced. The reader does not know whether the assembly used SPAdes 3.15.0 or SPAdes 4.0.0, and these versions have different defaults and different performance characteristics.

The fix is to record the version number of every tool at the time of installation. This can be done with the tool's version command, such as `spades.py --version`, or with the package manager's query command.

### Failure Pattern 2: Vague Parameter Descriptions

The second most common failure is the vague description of parameters. A methods section that says "the reads were assembled with high sensitivity" does not tell the reader what parameters were used. The phrase "high sensitivity" may correspond to a specific parameter set in one tool but not in another.

The fix is to report the exact parameters. This can be done by including the full command or by describing the parameter values in the text. For example, "the reads were assembled with SPAdes 3.15.0 using the following parameters: --cov-cutoff auto, --k 21,33,55,77."

### Failure Pattern 3: Missing Quality Assessment

The third common failure is the omission of quality assessment. A methods section that describes the assembly but does not report the quality metrics leaves the reader unable to assess the reliability of the assembly.

The fix is to report the standard quality metrics and the tools used to calculate them. The metrics should be reported in the methods section or in a supplementary table.

### Failure Pattern 4: Incomplete Input Data Description

The fourth common failure is the incomplete description of input data. A methods section that says "the genome was assembled from PacBio reads" does not tell the reader the read length, the coverage, or the library preparation method.

The fix is to describe the input data in detail. This includes the sequencing platform, the chemistry, the read length distribution, the coverage, and the library preparation method.

### Failure Pattern 5: No Access to the Commands

The fifth common failure is the absence of executable commands. A methods section that describes the workflow in prose but does not provide the commands forces the reader to reconstruct the workflow from the description.

The fix is to include the commands in a supplementary file or a public repository. The commands should be complete and executable, with all input file names and parameters specified.

## Limitations of Assembly Methods Reporting

### The Limits of Textual Description

A methods section, no matter how detailed, cannot capture every aspect of the assembly workflow. The hardware environment, the operating system, the file system, and the exact state of the input data all affect the assembly outcome. The methods section should describe the software environment, but the hardware environment is often omitted.

The practical implication is that a methods section provides a guide to reproduction, not a guarantee of identical results. Two researchers running the same commands on different hardware may obtain slightly different assemblies. The methods section should acknowledge this limitation and focus on the aspects of the workflow that are under the researcher's control.

### The Limits of Quality Metrics

The standard quality metrics, such as N50 and BUSCO completeness, provide a partial view of assembly quality. They do not capture all errors. The HiCanu study found that gaps and errors still remained within the most challenging regions of the genome, even with high-quality HiFi data [<a href="#ref-8">8</a>]. The methods section should acknowledge the limitations of the quality metrics and describe any additional validation performed.

For example, the pig reference genome project used annotation to assess the quality of the assemblies [<a href="#ref-7">7</a>]. The methods section should describe any such additional validation, such as mapping RNA-seq reads to the assembly or comparing the assembly to a known reference.

### The Limits of Benchmark Comparisons

Benchmark comparisons, such as the CAMI challenges, provide guidance for software selection but do not guarantee performance on a specific dataset. The CAMI II challenge found that related strains were still challenging for assembly and genome recovery through binning [<a href="#ref-2">2</a>]. A methods section that cites a benchmark to justify a software choice should acknowledge that the benchmark may not reflect the specific characteristics of the dataset.

## Quality and Welfare Controls in Assembly Methods

### The Role of Quality Control in Reproducibility

Quality control is not a separate step in the assembly workflow. It is an integral part of the workflow that affects the final assembly. The methods section must describe the quality control steps in enough detail that the reader can repeat them.

The quality control steps include read trimming, read filtering, error correction, and contamination screening. Each step should be described with the tool, the version, and the parameters. The quality metrics of the reads before and after each step should be reported.

### The Role of Documentation in Reproducibility

Documentation is the foundation of reproducibility. The methods section cannot be written accurately without a complete record of the assembly workflow. The Carpentries lessons on the Unix shell and version control provide the skills needed to create this record [<a href="#ref-5">5</a>].

The documentation should include the shell history, the scripts, the configuration files, and the log files. The documentation should be stored in a version control system, such as Git, to track changes over time. The documentation should be shared with the methods section, either as a supplementary file or as a public repository.

### The Role of Training in Reproducibility

Reproducible assembly requires skills that are not always taught in standard biology courses. The EMBL-EBI Training program provides learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-9">9</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-6">6</a>]. The Carpentries provides foundational computing and data lessons [<a href="#ref-5">5</a>].

Researchers who are new to genome assembly should complete these training programs before attempting to write a reproducible methods section. The training provides the vocabulary and the skills needed to document the workflow accurately.

## Safety and Regulatory Context for Assembly Methods

### Data Submission Requirements

Public sequence databases have specific requirements for data submission. The NCBI provides official descriptions of its databases, search systems, sequence resources, and analysis services [<a href="#ref-10">10</a>]. The methods section should be written to meet the submission requirements of the target database.

For NCBI submissions, the methods section should describe the assembly workflow in enough detail that the curators can assess the quality of the assembly. The methods section should also describe the quality metrics, because the curators use these metrics to validate the assembly.

### Data Availability Statements

The methods section should include a data availability statement that describes where the raw data and the assembly can be accessed. The statement should include the accession numbers for the raw reads and the assembly. The statement should also describe any restrictions on data access.

The data availability statement is a critical part of reproducibility. Without access to the raw data, the reader cannot reproduce the assembly. The statement should be written to comply with the data sharing policies of the target journal and the funding agency.

### Professional Escalation Criteria

When the assembly quality is poor or the workflow cannot be reproduced, the researcher should escalate the issue to a professional. The escalation criteria include:

1. The assembly has a very low N50 or a very high contig count, suggesting a problem with the assembly strategy
2. The BUSCO completeness is below the expected range for the organism, suggesting missing genes or contamination
3. The assembly contains sequences from unexpected organisms, suggesting contamination
4. The workflow cannot be reproduced by an independent researcher, suggesting missing documentation
5. The quality metrics cannot be calculated because the tools fail or produce errors

In these cases, the researcher should consult with a bioinformatics specialist or a sequence database curator. The specialist can help identify the source of the problem and recommend corrective actions.

## A Decision Framework for Choosing Assembly Parameters and Documenting Their Justification

The most common reproducibility failure in assembly methods sections is not the absence of parameters but the absence of reasoning behind them. Reviewers and database curators can verify that a command runs, but they cannot assess whether the parameter choices were appropriate for the biological question unless the methods section explains the decision process. This section provides a practical framework for making parameter decisions during the assembly and for documenting those decisions in a way that survives peer review.

### The Three-Tier Parameter Decision Framework

Before running any assembler, classify every parameter decision into one of three tiers. This classification determines how much documentation each parameter requires.

**Tier 1: Data-Derived Parameters**

These parameters are determined directly from the sequencing data or the biological sample. Examples include the estimated genome size, the expected coverage, the read length distribution, and the k-mer size range for k-mer based assemblers. These parameters must be justified with the evidence used to derive them.

For genome size estimation, state whether you used k-mer analysis, flow cytometry, or a related species reference. The pig reference genome project provides a relevant example. The improved assemblies Sscrofa11.1 and USMARCv1.0 were built with a whole-genome shotgun strategy using long-read technologies, and the methods required accurate genome size estimates to guide the assembly [<a href="#ref-7">7</a>]. If you used a related species genome size, name the species and the source of the estimate. If you used k-mer analysis, name the tool and the k-mer value.

For coverage estimation, report the calculation. The formula is the product of read count and read length divided by the estimated genome size. If the genome size is unknown, state the estimation method and its limitations.

**Tier 2: Tool-Specific Parameters**

These parameters are specific to the assembler and control its internal behavior. Examples include the overlap settings in Canu, the error correction mode in HiCanu, the polishing rounds in Racon, and the scaffolding parameters in LINKS. These parameters must be reported with their exact values and the rationale for choosing them.

The CAMI challenge found that parameter settings markedly affected performance, underscoring their importance for program reproducibility [<a href="#ref-1">1</a>]. This finding means that tool-specific parameters cannot be dismissed as implementation details. They change the assembly outcome.

For each tool-specific parameter, document three things. First, the default value for the exact software version used. Second, the value you chose. Third, the reason for the change, such as a pilot assembly result, a benchmark comparison, or a known limitation of the default for your data type.

**Tier 3: Resource and Environment Parameters**

These parameters control the computational environment instead of the assembly logic. Examples include thread count, memory allocation, and temporary directory location. These parameters do not usually change the assembly result, but they affect runtime and resource usage.

The CAMI II challenge analyzed runtime and memory usage across 76 program versions and identified efficient programs, including top performers on other metrics [<a href="#ref-2">2</a>]. This analysis shows that resource parameters matter for practical reproducibility. A reader who cannot run the assembly because the memory requirement is undocumented cannot reproduce the work.

Report resource parameters in a supplementary table or in the data availability statement. State the maximum memory used, the total runtime, and the hardware configuration. This information helps reviewers assess whether the workflow is feasible in their environment.

### A Record System for Parameter Decisions

The decision framework requires a record system that captures also the final parameters but the reasoning behind them. The following record format works for both individual researchers and collaborative projects.

Create a parameter decision log with one entry per parameter. Each entry contains five fields:

1. **Parameter name**: The exact flag or configuration key
2. **Tool and version**: The software and version that interprets this parameter
3. **Default value**: The value used if the parameter is not specified
4. **Chosen value**: The value used in the final assembly
5. **Rationale**: The evidence or reasoning behind the chosen value

The rationale field is the most important for reproducibility. It distinguishes a parameter choice based on data analysis from a parameter choice based on convenience. For example, a rationale that states "k-mer size 77 was chosen because the k-mer analysis showed a peak at 77 for the estimated genome size" is reproducible. A rationale that states "k-mer size 77 was chosen because it worked in a previous project" is not reproducible, because the reader cannot assess whether the previous project is comparable.

The parameter decision log should be maintained during the assembly, not after. The Carpentries lessons on the Unix shell and version control with Git provide the skills for maintaining this kind of record [<a href="#ref-5">5</a>]. The shell history captures the commands, and the parameter decision log captures the reasoning. Together, they provide a complete record of the assembly workflow.

### A Troubleshooting Method for Parameter-Related Assembly Failures

When an assembly fails or produces poor quality results, the parameter decision log becomes the primary troubleshooting tool. The following method uses the log to isolate the cause of the failure.

**Step 1: Check the Data-Derived Parameters**

Verify that the genome size estimate and the coverage estimate are accurate. A genome size estimate that is too large or too small can cause the assembler to make incorrect assumptions about the data. The HiCanu study demonstrated that accurate parameter estimation is critical for assembling complex regions such as segmental duplications, satellites, and allelic variants [<a href="#ref-8">8</a>]. If the assembly fails to resolve these regions, the genome size estimate or the coverage estimate may be wrong.

To check the genome size estimate, run a k-mer analysis on the raw reads and compare the estimated genome size to the value used in the assembly. To check the coverage estimate, calculate the coverage from the read count and read length and compare it to the value used in the assembly.

**Step 2: Check the Tool-Specific Parameters**

Verify that the tool-specific parameters are appropriate for the data type. The CAMI challenge found that assembly programs performed well for species represented by individual genomes but were substantially affected by the presence of related strains [<a href="#ref-1">1</a>]. If the assembly fails to resolve related strains, the parameters that control strain diversity handling may need adjustment.

For k-mer based assemblers, check the k-mer range. A k-mer range that is too narrow may miss the optimal k-mer value. A k-mer range that is too wide may increase runtime without improving the assembly. The CAMI II challenge found that related strains were still challenging for assembly and genome recovery through binning [<a href="#ref-2">2</a>]. This finding suggests that parameter adjustments alone may not resolve strain diversity issues.

**Step 3: Check the Resource Parameters**

Verify that the resource parameters are sufficient for the assembly. If the assembly ran out of memory or was killed by the operating system, the resource parameters may be too low. The CAMI II challenge identified efficient programs with lower runtime and memory usage [<a href="#ref-2">2</a>]. If the assembly fails due to resource exhaustion, consider switching to a more efficient program or increasing the resource allocation.

**Step 4: Escalate to a Professional**

If the parameter decision log does not reveal the cause of the failure, escalate the issue to a bioinformatics specialist. The escalation criteria include:

1. The assembly fails repeatedly with different parameter combinations
2. The assembly produces a very low N50 or a very high contig count despite reasonable parameters
3. The assembly contains sequences from unexpected organisms, suggesting contamination
4. The quality metrics cannot be calculated because the tools fail or produce errors

A bioinformatics specialist can review the parameter decision log and the assembly output to identify the source of the problem. The specialist may recommend a different assembler, a different parameter set, or additional quality control steps.

### Comparing Parameter Documentation Approaches

The decision framework supports three approaches to parameter documentation. Each approach has tradeoffs in precision, readability, and reviewer acceptance.

**Approach 1: Full Command Listing**

The methods section includes the complete command for every step. This approach is the most precise and the most reproducible. The reader can run the exact command without interpretation.

The tradeoff is readability. A long command with many flags can obscure the biological logic of the workflow. The methods text must explain the command in prose while the command provides the execution details.

The nf-core documentation describes how community pipelines package workflows for distribution and reproducibility [<a href="#ref-4">4</a>]. These pipelines include the full commands and the configuration files. The methods section can reference the pipeline version and the configuration used, instead of listing every command.

**Approach 2: Parameter Table**

The methods section includes a table that lists every parameter, its default value, and the chosen value. This approach is more readable than a full command listing but less precise. The reader must reconstruct the command from the table.

The parameter table works well for tools with a small number of parameters. For tools with many parameters, the table becomes unwieldy. The parameter decision log can be included as a supplementary table.

**Approach 3: Workflow Repository**

The methods section references a public repository that contains the full workflow, including the commands, the configuration files, and the parameter decision log. This approach is the most complete and the most reproducible.

The Galaxy Training Network provides tutorials that demonstrate how to document and share analysis workflows [<a href="#ref-6">6</a>]. The nf-core documentation describes how to package workflows for distribution [<a href="#ref-4">4</a>]. These resources show the expected format for sharing executable workflows.

The tradeoff is the requirement for a public repository. Some researchers cannot share their workflows publicly due to institutional or funding restrictions. In these cases, the workflow can be shared as a supplementary file with the manuscript.

### Implementing the Decision Framework in Practice

The following steps integrate the decision framework into the assembly workflow.

**Step 1: Create the Parameter Decision Log**

Before running the assembly, create a parameter decision log with the five fields described above. This log will be maintained throughout the assembly.

**Step 2: Classify Each Parameter**

For each parameter in the assembler and the associated tools, classify it as data-derived, tool-specific, or resource and environment. This classification determines the documentation requirements.

**Step 3: Document the Rationale for Each Parameter**

For each parameter, record the rationale in the parameter decision log. The rationale should reference the evidence used to choose the value, such as a k-mer analysis result, a pilot assembly, or a benchmark comparison.

**Step 4: Run the Assembly and Record the Results**

Run the assembly and record the results, including the quality metrics and the runtime and memory usage. Update the parameter decision log with any parameter changes made during the assembly.

**Step 5: Write the Methods Section From the Log**

Write the methods section from the parameter decision log. The methods text should describe the workflow in biological terms and reference the parameter decision log for the exact values and rationales.

**Step 6: Share the Log and the Workflow**

Share the parameter decision log and the workflow in a supplementary file or a public repository. The methods section should reference the location of these files.

### The Role of Training in Parameter Documentation

The decision framework requires skills in data analysis, software configuration, and documentation. The EMBL-EBI Training program provides learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-9">9</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-6">6</a>]. The Carpentries provides foundational computing and data lessons [<a href="#ref-5">5</a>].

Researchers who are new to genome assembly should complete these training programs before attempting to implement the decision framework. The training provides the vocabulary and the skills needed to document parameter decisions accurately.

The Bioconductor project provides official documentation for reproducible genomic-analysis workflows [<a href="#ref-3">3</a>]. This documentation includes examples of parameter documentation for Bioconductor packages. Researchers who use Bioconductor packages for assembly or quality assessment should follow the documentation standards described in the project materials.

The NCBI provides official descriptions of its databases, search systems, sequence resources, and analysis services [<a href="#ref-10">10</a>]. Researchers who submit assemblies to NCBI should review the submission requirements before writing the methods section. The parameter decision log can be included in the submission to demonstrate the reproducibility of the assembly.

## Frequently Asked Questions

### What is the minimum information needed for a reproducible assembly methods section?

The minimum information includes the input data description, the software versions, the parameters, and the quality metrics. The input data description must state the sequencing platform, the read length, the coverage, and the library preparation method. The software versions must be stated for every tool. The parameters must be stated for every non-default setting. The quality metrics must include the standard assembly statistics and the completeness assessment.

### How do I report parameters when I used default settings?

Report the defaults explicitly. State the software version and the default parameters for that version. If the tool has a command that prints the parameters, include the output of that command. Do not write "default parameters were used" without specifying the version, because defaults change between versions.

### Should I include the actual commands in the methods section?

Yes. The actual commands are the most precise way to describe the assembly workflow. Include the commands in the methods text, in a supplementary file, or in a public repository. The commands should include all flags and arguments, beyond the non-default ones.

### How do I describe a hybrid assembly using both short and long reads?

Describe each data type separately, then describe how they were combined. State the platform, read length, and coverage for each data type. Describe the assembly strategy, such as using the long reads for the initial assembly and the short reads for polishing. State the software and parameters for each step.

### What quality metrics should I report for a bacterial genome assembly?

Report the total assembly length, the contig count, the N50, the GC content, and the completeness. For bacterial genomes, completeness is often assessed by checking for the presence of expected single-copy genes. Report the contamination, which is the percentage of sequences from other organisms.

### What quality metrics should I report for a eukaryotic genome assembly?

Report the total assembly length, the contig count, the N50, the scaffold N50 if scaffolding was performed, and the GC content. Report the completeness with BUSCO or a similar tool. If the assembly was annotated, report the number of genes and the annotation statistics.

### How do I describe the quality control steps in the methods section?

Describe each quality control step with the tool, the version, and the parameters. State the quality threshold for read trimming and the minimum read length retained. State the criteria for read filtering, such as adapter contamination or low complexity. Report the number of reads before and after each step.

### What should I do if my assembly cannot be reproduced by another researcher?

First, check the documentation. Verify that the software versions, parameters, and commands are complete and accurate. Second, check the input data. Verify that the raw reads are accessible and that the accession numbers are correct. Third, consult with a bioinformatics specialist. The specialist can help identify the source of the problem and recommend corrective actions.

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Metagenomic Assembly Overview: Challenges and Applications](/knowledge/bioinformatics/metagenomic-assembly-overview-challenges-and-applications)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Critical Assessment of Metagenome Interpretation-a benchmark of metagenomics software.](https://pubmed.ncbi.nlm.nih.gov/28967888). Nature methods, 2017.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Critical Assessment of Metagenome Interpretation: the second round of challenges.](https://pubmed.ncbi.nlm.nih.gov/35396482). Nature methods, 2022.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [An improved pig reference genome sequence to enable pig genetics and genomics research.](https://pubmed.ncbi.nlm.nih.gov/32543654). GigaScience, 2020.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads.](https://pubmed.ncbi.nlm.nih.gov/32801147). Genome research, 2020.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.