How to Write a Reproducible RNA-seq Analysis Report: A Template for Methods and Results
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Reproducible RNA-seq analysis requires meticulous documentation of every computational step, including exact software versions, all parameters used (e.g., quality thresholds for trimming, alignment parameters like
--outFilterMultimapScoreCutoffin STAR), and the specific reference genome and annotation versions (e.g., GRCh38 with Ensembl v110). - Complete data provenance is critical, detailing sequencing platform (e.g., Illumina NovaSeq 6000), library preparation method (e.g., SMART-seq2 for single-cell), read length, sequencing depth, and the number of biological replicates per condition, alongside accession numbers for public repositories.
- The computational environment must be precisely specified, including the operating system, all software package versions (e.g., R 4.3.1, Bioconductor release 3.18), and any containerization (e.g., Docker, Singularity) or workflow management tools (e.g., Nextflow, Snakemake) used.
- Every parameter choice, from read trimming (e.g.,
fastp --detect_adapter_for_pe --cut_right --cut_right_window_size 5 --cut_right_mean_quality 20) to differential expression modeling (e.g.,~ condition + batchin DESeq2), must be documented with a clear rationale, citing training resources (e.g., Galaxy Training Network) or established protocols. - Quality control metrics (e.g., percentage of reads mapped, number of genes detected per sample) and the rationale for filtering thresholds must be explicitly recorded, alongside a log of all analysis decisions, including alternatives considered and reasons for rejection.
- Independent verification of results, such as validating differentially expressed genes with RT-qPCR, and clear documentation of analysis limitations (e.g., short-read limitations for isoform resolution) are essential for robust interpretation.
RNA sequencing has become a standard tool for measuring gene expression across biological conditions, yet many published reports lack the detail needed for another laboratory to repeat the analysis. This article provides a structured template for writing the methods and results sections of an RNA-seq analysis report, with placeholders for describing data processing, alignment, quantification, and statistical analysis. The template aligns with reporting standards used by major bioinformatics training resources and community pipeline projects, and it is intended for biology students, researchers, laboratory professionals, and life-science practitioners who need to document their work in a way that supports verification and reuse.
The Reproducibility Problem in RNA-seq Reporting
A reproducible RNA-seq analysis report must contain enough information for an independent analyst to take the raw sequencing data and arrive at the same results. This requirement extends beyond listing software names and versions. It includes the exact parameters used at each step, the reference resources employed, the quality thresholds applied, and the rationale for each decision. Without this level of detail, a report becomes a summary of what was done instead of a specification of how to do it.
The scale of the problem is well recognized across the bioinformatics community. Training programs at the European Bioinformatics Institute emphasize that data analysis skills must be paired with an understanding of how to document workflows properly [<a href="#ref-1">1</a>]. The Galaxy Training Network provides hands-on tutorials that model reproducible analysis practices, showing learners how to record each step in a way that others can follow [<a href="#ref-2">2</a>]. Community pipeline projects such as nf-core have established documentation standards that require clear descriptions of pipeline usage, configuration options, and expected outputs [<a href="#ref-3">3</a>]. These resources all point to the same conclusion: reproducibility is a design principle that must be built into the analysis from the start.
For researchers writing an RNA-seq report, the practical implication is that documentation must happen alongside the analysis, not after it. Every command, every parameter choice, and every quality assessment should be recorded at the time it is made. This approach mirrors the record-keeping practices used in laboratory notebooks, where observations are recorded immediately instead of reconstructed from memory.
Core Principles for Documenting RNA-seq Analysis
Record the Complete Data Provenance
The first requirement for a reproducible report is a complete description of the input data. This includes the sequencing platform, the library preparation method, the read length, the sequencing depth, and the number of biological replicates per condition. For example, a study using total RNA sequencing of extracellular RNA from blood plasma would need to document the SMART cDNA synthesis technology used for library construction, as described in protocols for biofluid exRNA analysis [<a href="#ref-4">4</a>]. A study using targeted single-cell RNA sequencing would need to specify which targeted method was selected and why, since different targeted approaches address different limitations in transcript detection [<a href="#ref-5">5</a>].
The data provenance section should also include information about how the raw data were stored and accessed. If the data were deposited in a public repository such as the NCBI Sequence Read Archive, the accession numbers must be provided [<a href="#ref-6">6</a>]. If the data are subject to controlled access, the report should state the conditions under which other researchers can request access.
Specify the Computational Environment
A reproducible report must describe the computational environment in which the analysis was performed. This includes the operating system, the versions of all software packages, and the hardware specifications if they could affect the results. For example, a report that used Bioconductor packages for differential expression analysis should list the Bioconductor release and the R version, since package behavior can change between releases [<a href="#ref-7">7</a>].
Containerization and workflow management tools offer a practical solution to environment reproducibility. The nf-core documentation describes how Nextflow pipelines can be configured to run in reproducible environments, with each process isolated in its own container [<a href="#ref-3">3</a>]. The Carpentries lessons teach foundational skills in using version control with Git, which allows researchers to track changes to analysis scripts and document the evolution of their workflow [<a href="#ref-8">8</a>]. A report that describes its computational environment in this level of detail gives other researchers a realistic path to repeating the analysis.
Document Every Parameter Choice
Each step in an RNA-seq analysis involves parameter choices that can affect the final results. The report must document these choices explicitly. For read trimming, the report should state the quality threshold, the minimum read length, and whether adapter trimming was performed. For alignment, the report should state the aligner, the reference genome version, the gene annotation version, and any non-default parameters. For quantification, the report should state whether gene-level or transcript-level quantification was performed and which counting method was used.
The rationale for parameter choices should also be documented. For example, if a study used a particular quality threshold because it was recommended in a training tutorial from the Galaxy Training Network, that source should be cited [<a href="#ref-2">2</a>]. If a study used a specific alignment parameter because it was the default in a Bioconductor workflow, that should be stated [<a href="#ref-7">7</a>]. This documentation of rationale helps other researchers understand whether the choices were based on established practice, preliminary analysis, or specific properties of the data.
At a Glance: Key Components of a Reproducible RNA-seq Report
| Report Section | Required Information | Common Documentation Gaps |
|---|---|---|
| Data provenance | Sequencing platform, library method, read length, depth, replicate numbers, repository accessions | Missing accession numbers, no description of library preparation details |
| Computational environment | Operating system, software versions, container or workflow system, hardware | No version numbers, no description of environment isolation method |
| Quality control | Trimming parameters, quality metrics, filtering thresholds, number of reads removed at each step | No record of reads removed, no justification for thresholds |
| Alignment and quantification | Aligner, reference genome version, annotation version, counting method, non-default parameters | No reference version, no annotation version, no parameter documentation |
| Statistical analysis | Model formula, normalization method, contrast definitions, multiple testing correction, significance thresholds | No model formula, no description of how contrasts were defined |
| Results reporting | Number of genes tested, number of differentially expressed genes, effect size reporting, visualization parameters | No effect sizes, no description of how thresholds were applied |
Practical Workflow for Writing the Methods Section
Step 1: Describe the Experimental Design
The methods section should begin with a clear statement of the experimental design. This includes the number of biological replicates per condition, the total number of samples sequenced, and the sequencing strategy. For example, a study comparing gene expression between two conditions with three biological replicates per condition would state this explicitly. The report should also describe any pooling or multiplexing strategies used during library preparation.
The experimental design description should include the rationale for the number of replicates. While the specific power calculation may be described in the results or supplementary materials, the methods section should state whether the replicate number was based on a formal power analysis, on previous experience with similar experiments, or on practical constraints such as cost or sample availability.
Step 2: Document the Quality Control Process
Quality control is a critical step in RNA-seq analysis, and the methods section must describe it in detail. This includes the software used for quality assessment, the metrics examined, and the thresholds applied for read filtering. The report should state how many reads passed each quality filter and how many were removed.
The quality control description should also address the assessment of sequencing artifacts. Research on host-virus chimeric events in RNA-seq data has shown that a substantial fraction of reads can be artifactually chimeric due to template switching during reverse transcription [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. A reproducible report should describe whether chimeric reads were assessed and how they were handled in the analysis. This is particularly important for studies of viral infection or studies using spike-in controls, where chimeric artifacts could be misinterpreted as biological events.
Step 3: Specify the Alignment and Quantification Strategy
The alignment section must specify the reference genome version, the gene annotation version, and the alignment software with its version and parameters. For example, a report might state that reads were aligned to the human genome assembly GRCh38 with Ensembl annotation version 110 using a splice-aware aligner with default parameters. The report should also describe how multi-mapping reads were handled, since this decision can affect quantification results.
For quantification, the report should specify whether gene-level counts, transcript-level estimates, or both were generated. The choice of quantification method should be justified. For example, a study focused on differential isoform usage would need transcript-level quantification, while a study focused on gene-level expression changes could use gene-level counting. The report should also describe how the quantification was integrated with the alignment step, whether through a combined pipeline or through separate steps.
Step 4: Describe the Statistical Analysis
The statistical analysis section must describe the normalization method, the statistical model, and the tests applied. For differential expression analysis, the report should state the model formula, including any covariates or batch effects included in the model. The report should describe how contrasts were defined and which multiple testing correction method was applied.
The statistical analysis description should also address the handling of low-count genes. Many differential expression methods require a minimum count threshold, and the report should state what threshold was applied and how many genes were filtered out. The report should also describe how the significance threshold was chosen and whether it was adjusted for multiple testing.
Step 5: Document the Results Reporting Format
The methods section should describe how results will be reported, including the criteria for defining differentially expressed genes and the format for presenting results. This includes the fold change threshold, the adjusted p-value threshold, and whether results are reported for all genes or only for those passing the significance thresholds.
The results reporting description should also address the presentation of quality control metrics. A reproducible report should include the key quality metrics for each sample, such as the number of reads sequenced, the percentage of reads aligned, and the number of genes detected. These metrics allow readers to assess the overall quality of the dataset and to identify any samples that may be problematic.
Options and Tradeoffs in RNA-seq Analysis Workflows
Read Trimming and Quality Filtering
The decision to trim reads before alignment involves a tradeoff between removing low-quality bases and preserving useful sequence information. Aggressive trimming can remove bases that would otherwise align, while minimal trimming may leave low-quality bases that cause alignment errors. The report should document the trimming parameters and the rationale for the chosen approach.
Some analysis pipelines recommend skipping read trimming entirely when using aligners that are tolerant of mismatches. Other pipelines recommend trimming to a specific quality threshold. The choice depends on the sequencing platform, the read length, and the downstream analysis goals. A reproducible report should state the trimming strategy and provide the metrics that justify the choice.
Alignment Strategy
The choice of alignment strategy depends on whether the analysis requires splice-aware alignment, whether the reference genome is well annotated, and whether the study involves non-model organisms. Splice-aware aligners are required for RNA-seq data because reads span exon-exon junctions. The report should specify the aligner and its version, the reference genome and annotation versions, and any parameters that affect splice junction detection.
For studies involving organisms without a well-annotated reference genome, the analysis may require de novo transcriptome assembly. This approach is more complex and requires additional documentation of the assembly parameters and the quality assessment of the assembled transcripts. The report should describe the assembly process and the methods used to evaluate assembly completeness.
Quantification Approach
The choice between alignment-based quantification and alignment-free quantification involves tradeoffs in speed, accuracy, and the ability to detect novel transcripts. Alignment-based methods provide more information about splice junctions and can detect novel isoforms, while alignment-free methods are faster and require less computational resources. The report should specify the quantification method and justify the choice based on the analysis goals.
For studies that require transcript-level quantification, the report should describe how transcript abundances were estimated and how isoform-level results were summarized. For studies that focus on gene-level analysis, the report should describe how transcript-level estimates were aggregated to the gene level.
Differential Expression Analysis
The choice of differential expression method involves tradeoffs between sensitivity, specificity, and the ability to handle small sample sizes. Different methods make different assumptions about the distribution of count data and the relationship between mean and variance. The report should specify the method used and the version, and it should describe the model formula and the normalization approach.
The report should also describe how the results were validated. This may include visual inspection of the data, examination of specific genes of interest, or comparison with results from an alternative method. The validation approach should be documented so that readers can assess the robustness of the findings.
Records and Measurements for a Reproducible Report
Quality Metrics to Record
A reproducible RNA-seq report should include a table of quality metrics for each sample. These metrics include the total number of reads, the number of reads passing quality filters, the percentage of reads aligned to the reference genome, the percentage of reads assigned to genes, and the number of genes detected. These metrics provide a snapshot of data quality and allow readers to identify potential issues.
The report should also include metrics that assess the overall quality of the sequencing run, such as the percentage of bases with a quality score above a threshold and the GC content distribution. These metrics can reveal systematic biases in the library preparation or sequencing process.
Analysis Decision Log
A reproducible report should include a log of analysis decisions, documenting when each decision was made and the rationale behind it. This log serves as a record of the analysis process and helps other researchers understand why specific choices were made. The log should include the date of each decision, the person making the decision, and the information that informed the decision.
The decision log is particularly important for decisions that involve subjective judgment, such as the choice of quality thresholds or the handling of outlier samples. By documenting these decisions, the report provides transparency about the analysis process and allows readers to assess the potential impact of these choices on the results.
Version Control Records
The report should document the version control history of the analysis scripts and workflows. This includes the repository location, the commit identifiers for each analysis step, and the dates of each commit. Version control records allow other researchers to access the exact scripts used in the analysis and to track changes over time.
The Carpentries lessons provide foundational training in using Git for version control, and the nf-core documentation describes how version control is integrated into pipeline development [<a href="#ref-3">3</a>][<a href="#ref-8">8</a>]. A reproducible report should follow these practices and document the version control history as part of the methods.
Common Failure Patterns in RNA-seq Reporting
Missing Version Information
One of the most common failures in RNA-seq reporting is the omission of software version numbers. A report that states "reads were aligned using STAR" without specifying the version is not reproducible, because different versions of STAR can produce different results. The same applies to reference genome versions and annotation versions. A reproducible report must specify the exact version of every software tool and reference resource used.
Undocumented Parameter Changes
Another common failure is the use of non-default parameters without documentation. If a researcher changes a parameter from its default value, the report must state this change and the rationale behind it. Undocumented parameter changes make it impossible for other researchers to repeat the analysis, because they will not know which parameters were modified.
Incomplete Quality Control Description
Many reports describe quality control in vague terms, such as "reads were trimmed for quality" without specifying the trimming parameters or the quality thresholds. This level of detail is insufficient for reproducibility. The report must specify the exact quality control steps, the parameters used, and the number of reads removed at each step.
Failure to Describe the Statistical Model
Differential expression analysis requires a statistical model, and the report must describe this model in detail. This includes the model formula, the normalization method, and the multiple testing correction approach. A report that states "differential expression analysis was performed using DESeq2" without describing the model is not reproducible, because the model specification affects the results.
Omission of Data Access Information
A reproducible report must provide information about how to access the raw data. If the data are deposited in a public repository, the accession numbers must be provided. If the data are not publicly available, the report should state the conditions under which they can be accessed. Without this information, other researchers cannot verify the results or repeat the analysis.
Quality Controls and Verification Steps
Independent Verification of Results
A reproducible report should describe any independent verification of the results. This may include validation of differentially expressed genes using an alternative method such as quantitative PCR, as described in studies that combine RNA-seq with qRT-PCR validation [<a href="#ref-11">11</a>][<a href="#ref-12">12</a>]. The report should describe the validation approach and the results of the validation.
For studies that use RNA-seq to identify novel transcripts or splicing events, the report should describe how these findings were verified. This may include examination of the raw sequencing reads, comparison with independent datasets, or experimental validation. The verification approach should be documented so that readers can assess the confidence in the findings.
Assessment of Technical Variability
The report should describe how technical variability was assessed. This may include examination of the correlation between replicate samples, principal component analysis of the expression data, or assessment of the relationship between library size and expression levels. The assessment of technical variability helps readers understand the reliability of the results.
For studies that involve complex experimental designs, the report should describe how batch effects were assessed and handled. This may include the inclusion of batch as a covariate in the statistical model or the use of batch correction methods. The report should document the approach used and the rationale for the choice.
Documentation of Analysis Limitations
A reproducible report should acknowledge the limitations of the analysis. This includes limitations imposed by the sequencing depth, the number of replicates, the quality of the reference genome, and the statistical power. The report should describe how these limitations might affect the interpretation of the results.
For example, a study using short-read RNA-seq may not detect all transcript isoforms, as noted in reviews of single-cell long-read sequencing that highlight the limitations of short-read platforms for resolving transcriptome complexity [<a href="#ref-13">13</a>]. A reproducible report should acknowledge such limitations and describe their potential impact on the findings.
Limitations and Interpretation Boundaries
Technical Limitations of RNA-seq
RNA-seq analysis has inherent technical limitations that should be acknowledged in the report. These include the potential for artifacts introduced during library preparation, such as the template switching that can generate chimeric reads [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. The report should describe how these potential artifacts were assessed and handled.
The report should also acknowledge the limitations of the reference resources used. If the reference genome or annotation is incomplete, the report should describe how this might affect the results. For example, a study using a reference transcriptome assembled from full-length sequencing data would need to describe the completeness of the assembly and its potential limitations [<a href="#ref-11">11</a>].
Biological Interpretation Boundaries
The report should clearly distinguish between the computational results and the biological interpretation. The computational analysis identifies genes that are differentially expressed, but the biological significance of these changes requires additional context. The report should describe the boundaries of the interpretation and the evidence needed to support biological conclusions.
For example, a study that identifies differentially expressed genes in a disease model should acknowledge that the computational results alone do not establish causation. The report should describe what additional experiments or analyses would be needed to support the biological interpretation.
Generalizability of Results
The report should acknowledge the limits of generalizability of the results. Findings from a specific cell type, tissue, or experimental condition may not apply to other contexts. The report should describe the scope of the findings and the conditions under which they are likely to be generalizable.
For example, a study of gene expression in hibernating retinas would need to acknowledge that the findings may be specific to the hibernation model and may not generalize to other conditions [<a href="#ref-14">14</a>]. The report should describe the boundaries of the generalizability and the evidence needed to extend the findings to other contexts.
Safety and Regulatory Context
Data Management and Privacy
RNA-seq data from human subjects are subject to privacy and regulatory requirements. The report should describe how data privacy was protected, including any de-identification procedures and data access controls. If the data are deposited in a public repository, the report should describe the data access conditions and any restrictions on data use.
The NCBI provides resources for data deposition and access, including controlled-access databases for sensitive data [<a href="#ref-6">6</a>]. The report should describe how these resources were used and how data access was managed.
Compliance with Reporting Standards
The report should describe how the analysis complies with relevant reporting standards. This includes standards for data deposition, analysis documentation, and results reporting. The report should reference the specific standards that were followed and describe how compliance was achieved.
For studies that involve clinical samples or that have regulatory implications, the report should describe the regulatory context and how the analysis supports compliance with regulatory requirements. For example, studies that support investigational new drug applications would need to document the analysis in a way that meets regulatory standards [<a href="#ref-15">15</a>].
Ethical Considerations
The report should address ethical considerations related to the data and the analysis. This includes the ethical approval for the study, the informed consent for sample collection, and the responsible use of the data. The report should describe how these ethical considerations were addressed.
For studies involving human subjects, the report should describe the ethical approval process and the measures taken to protect participant privacy. For studies involving animal subjects, the report should describe the ethical approval and the measures taken to ensure animal welfare.
Professional Escalation Criteria
When to Seek Additional Expertise
Researchers should seek additional expertise when the analysis involves complex decisions that are beyond their current skill level. This includes decisions about statistical modeling, the handling of confounding factors, and the interpretation of complex results. The report should describe when additional expertise was sought and how it informed the analysis.
For example, a researcher who is uncertain about the appropriate statistical model for a complex experimental design should consult with a biostatistician. A researcher who is uncertain about the interpretation of results involving novel transcripts or splicing events should consult with a molecular biologist. The report should document these consultations and how they influenced the analysis.
When to Repeat the Analysis
Researchers should repeat the analysis when there is evidence of technical problems or when the results are unexpectedly different from previous findings. The report should describe the criteria for deciding when to repeat the analysis and the process for doing so.
For example, if quality control metrics indicate a problem with a particular sample, the analysis may need to be repeated with the problematic sample removed. If the results are inconsistent with previous findings, the analysis may need to be repeated to verify the findings. The report should describe these decision criteria and the process for repeating the analysis.
When to Escalate to Supervisors or Collaborators
Researchers should escalate concerns to supervisors or collaborators when the analysis reveals potential problems that could affect the validity of the findings. This includes concerns about data quality, analysis methodology, or interpretation of results. The report should describe the escalation process and the criteria for escalation.
For example, if the analysis reveals a systematic bias in the data that cannot be explained, the researcher should escalate the concern to the study team. If the analysis produces results that contradict established knowledge, the researcher should escalate the concern to the appropriate experts. The report should document these escalations and their outcomes.
Decision Framework for Selecting Analysis Components and Recording Justifications
A reproducible RNA-seq report requires more than documenting what was done. It also requires a structured method for deciding which analysis components to use and for recording why each choice was made. Without a decision framework, researchers often select tools and parameters based on habit or convenience, then struggle to reconstruct the rationale when writing the report. This section provides a practical decision framework that integrates with the reporting template described throughout this article, along with a record system for capturing the reasoning behind each analysis choice.
The Four-Stage Decision Framework
The framework organizes analysis decisions into four stages that correspond to the major sections of an RNA-seq report: data preparation, alignment and quantification, statistical analysis, and results interpretation. At each stage, the researcher answers a set of structured questions and records the answers in a decision log that becomes part of the final report.
Stage 1: Data Preparation Decisions
The first stage addresses questions about read trimming, quality filtering, and artifact assessment. The researcher must decide whether to trim reads, which quality threshold to apply, and whether to assess chimeric reads. The decision framework requires answering three questions before making these choices.
The first question asks what sequencing platform generated the data and what error profile is expected. Different platforms produce different error patterns, and the trimming strategy should match the platform characteristics. The second question asks what downstream analysis steps will be performed. If the analysis includes fusion detection or viral integration assessment, chimeric read evaluation becomes critical, since research has shown that approximately 1 percent of RNA-seq reads can be artifactually chimeric due to template switching during reverse transcription [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. The third question asks what reference resources are available for the organism under study. A well-annotated reference genome supports different quality control decisions than a poorly annotated one.
For each decision, the framework requires the researcher to state the alternative options considered and the reason for rejecting them. For example, if the researcher decides to skip read trimming, the decision log must record that trimming was considered and explain why it was not performed, such as the use of a mismatch-tolerant aligner or the short read length of the dataset.
Stage 2: Alignment and Quantification Decisions
The second stage addresses the choice of aligner, reference genome version, annotation version, and quantification method. The decision framework requires the researcher to document the version of every resource used and to justify each choice in relation to the study goals.
The first question in this stage asks whether the study requires detection of novel transcripts, splice variants, or fusion genes. Studies that require full-length transcript information may benefit from long-read sequencing approaches, which can resolve transcriptome complexity that short-read platforms miss [<a href="#ref-13">13</a>]. The second question asks whether the organism has a well-annotated reference genome. For organisms without complete annotations, the analysis may require de novo transcriptome assembly, which introduces additional documentation requirements. The third question asks what computational resources are available, since some alignment and quantification methods require substantially more memory and processing time than others.
The decision log for this stage must record the specific versions of the reference genome and annotation used. For example, a report might state that reads were aligned to the human genome assembly GRCh38 with Ensembl annotation version 110. The log must also record whether the researcher considered alternative reference versions and why the chosen version was selected.
Stage 3: Statistical Analysis Decisions
The third stage addresses the choice of normalization method, statistical model, and multiple testing correction approach. The decision framework requires the researcher to specify the model formula, the handling of low-count genes, and the definition of contrasts between conditions.
The first question asks what experimental design was used, including the number of biological replicates per condition and whether batch effects are expected. The second question asks whether the analysis requires gene-level or transcript-level statistics, since this choice affects both the quantification method and the statistical model. The third question asks what significance thresholds will be applied and whether they will be adjusted for multiple testing.
The decision log for this stage must record the exact model formula used, including any covariates. For example, a report might state that the model included condition and batch as factors, with the formula expression ~ condition + batch. The log must also record how many genes were filtered out due to low counts and what threshold was applied.
Stage 4: Results Interpretation Decisions
The fourth stage addresses how results will be summarized, validated, and presented. The decision framework requires the researcher to specify the criteria for defining differentially expressed genes, the validation approach, and the format for presenting results.
The first question asks what fold change and adjusted p-value thresholds will define significance. The second question asks whether results will be validated using an independent method, such as quantitative PCR, which is commonly used to confirm RNA-seq findings [<a href="#ref-11">11</a>][<a href="#ref-12">12</a>]. The third question asks how the results will be visualized and what information will be included in the results tables.
The decision log for this stage must record the validation approach and the outcome of the validation. If validation was not performed, the log must state this explicitly and explain why.
The Decision Log Record System
The decision log is a structured record that captures the reasoning behind each analysis choice. It should be maintained throughout the analysis, not created after the analysis is complete. The log can be maintained as a spreadsheet, a text document, or a version-controlled file in the analysis repository.
Each entry in the decision log should include the date of the decision, the person making the decision, the analysis step being decided, the options considered, the option selected, and the rationale for the selection. The log should also record any information that informed the decision, such as results from preliminary analyses or recommendations from training resources.
The decision log serves multiple purposes. It provides the raw material for writing the methods section of the report. It allows other researchers to understand why specific choices were made. And it provides a record that can be reviewed if the analysis needs to be repeated or if questions arise about the validity of the results.
Integrating the Decision Framework with the Report Template
The decision framework and decision log should be integrated with the report template described in the previous sections. Each section of the methods should reference the corresponding decision log entries. For example, the quality control section should reference the decision log entries that document the trimming parameters and the rationale for the chosen thresholds.
The integration also works in the reverse direction. When writing the methods section, the researcher should consult the decision log to ensure that every decision is documented and that the rationale is stated. This process helps identify gaps in the documentation before the report is submitted.
Common Failure Patterns in Decision Documentation
Several common failure patterns emerge when researchers do not use a structured decision framework. The first is the failure to document alternatives that were considered and rejected. A report that states "reads were aligned using STAR" without explaining why STAR was chosen over other aligners is incomplete. The decision log should record that other aligners were considered and the reasons for rejecting them.
The second failure pattern is the retrospective reconstruction of rationale. Researchers who do not maintain a decision log during the analysis often invent plausible-sounding rationales after the fact. This practice undermines the credibility of the report and can lead to inaccurate documentation. The decision log should be maintained in real time to avoid this problem.
The third failure pattern is the omission of decisions that were made implicitly. Many analysis choices are made without conscious deliberation, such as accepting default parameters without considering alternatives. The decision framework requires the researcher to explicitly consider whether each default parameter is appropriate for the study and to document the consideration in the decision log.
Practical Implementation Steps
To implement the decision framework, researchers should follow these steps at the start of each analysis project. First, create a decision log file in the analysis repository and record the date and project name. Second, work through the four stages of the framework, answering each question and recording the answers in the log. Third, update the log whenever a decision is made or changed during the analysis. Fourth, use the log as the basis for writing the methods section of the report.
The decision log should be stored alongside the analysis scripts and data in the version-controlled repository. This ensures that the log is preserved with the analysis and can be accessed by other researchers who want to understand the analysis process. The Carpentries lessons provide training in using Git for version control, which supports this practice [<a href="#ref-8">8</a>].
Relationship to Community Pipeline Standards
The decision framework complements the documentation standards used by community pipeline projects. The nf-core documentation describes standards for pipeline usage, configuration, and expected outputs [<a href="#ref-3">3</a>]. These standards ensure that pipelines are documented consistently, but they do not replace the need for individual researchers to document their own decisions about which pipeline to use and how to configure it for their specific study.
The Galaxy Training Network provides tutorials that model reproducible analysis practices [<a href="#ref-2">2</a>]. These tutorials often include explanations of why specific tools and parameters are recommended, which can inform the decision log. Researchers should cite these sources in their decision logs when they inform the choices made.
The decision framework also aligns with the training provided by the European Bioinformatics Institute, which emphasizes the importance of documenting workflows properly [<a href="#ref-1">1</a>]. By integrating the decision framework with the report template, researchers can meet the documentation standards promoted by these training resources.
Records and Measurements for the Decision Log
The decision log should include specific measurements and records that support the analysis decisions. These include the results of any preliminary analyses that informed parameter choices, the quality metrics that justified filtering thresholds, and the output of any diagnostic plots that guided the statistical analysis.
For example, if the researcher decided to set a particular quality threshold based on examination of quality score distributions, the decision log should reference the diagnostic plot and describe what was observed. If the researcher decided to include a batch effect in the statistical model based on principal component analysis, the decision log should reference the PCA plot and describe the clustering pattern that justified the decision.
These records provide evidence that the decisions were based on data instead of arbitrary choices. They also allow other researchers to review the evidence and assess whether they would have made the same decisions.
Professional Escalation Criteria for Analysis Decisions
The decision framework should include criteria for escalating decisions to supervisors, collaborators, or bioinformatics experts. Researchers should escalate a decision when they are uncertain about the appropriate choice, when the decision could substantially affect the results, or when the decision involves statistical or computational methods beyond their expertise.
For example, a researcher who is uncertain about the appropriate statistical model for a complex experimental design should consult with a biostatistician before finalizing the analysis. A researcher who is uncertain about the handling of chimeric reads in a viral infection study should consult with a bioinformatics expert who has experience with this type of analysis. The decision log should record these consultations and how they influenced the final decisions.
The escalation criteria should also include situations where the analysis produces unexpected results that may indicate technical problems. In these cases, the researcher should escalate the concern before proceeding with the analysis, since the problem may affect the validity of all downstream results.
Frequently Asked Questions
What is the minimum information needed for a reproducible RNA-seq methods section?
The minimum information includes the sequencing platform and library preparation method, the number and type of samples, the quality control steps with parameters, the alignment software and reference resources with versions, the quantification method, and the statistical model with normalization and multiple testing correction details. Each software tool must be listed with its exact version, and each reference resource must be listed with its version or accession. The report should also include the data repository accession numbers and the computational environment description.
How should I document the reference genome and annotation versions?
The reference genome should be identified by its assembly name and version, such as GRCh38 for human or GRCm39 for mouse. The annotation should be identified by its source and version, such as Ensembl release 110 or GENCODE version 44. The report should state where these resources were obtained and how they were accessed. If a custom reference was used, the report should describe how it was constructed and provide access to the custom resource.
What quality control metrics should be reported for each sample?
The report should include the total number of reads sequenced, the number of reads passing quality filters, the percentage of reads aligned to the reference genome, the percentage of reads assigned to genes, and the number of genes detected. Additional metrics may include the percentage of reads mapping to exonic, intronic, and intergenic regions, the read duplication rate, and the distribution of insert sizes. These metrics should be presented in a table that allows comparison across samples.
How should I describe the statistical model for differential expression analysis?
The statistical model should be described in enough detail that another researcher could implement it. This includes the model formula, the normalization method, the handling of low-count genes, the multiple testing correction method, and the significance thresholds. The report should state whether the model included covariates or batch effects and how these were incorporated. The report should also describe how contrasts were defined for comparisons between conditions.
What should I do if my results cannot be reproduced by another researcher?
If another researcher cannot reproduce the results, the first step is to compare the analysis documentation with the actual analysis. Check whether all software versions, parameters, and reference resources were documented correctly. Check whether the computational environment was described accurately. If the documentation is correct, the next step is to examine whether differences in the computational environment could explain the discrepancy. If the discrepancy cannot be resolved, the issue should be escalated to the study team or to a bioinformatics expert.
How should I handle chimeric reads in my RNA-seq analysis?
Chimeric reads can arise from template switching during reverse transcription and are a known artifact in RNA-seq data [<a href="#ref-9">9</a>][<a href="#ref-10">10</a>]. The report should describe whether chimeric reads were assessed and how they were handled. For studies where chimeric reads could be misinterpreted as biological events, such as studies of viral integration or gene fusions, the report should describe the methods used to distinguish true biological events from artifacts. The report should also acknowledge the limitations of the methods used to detect chimeric reads.
What are the common mistakes in writing a reproducible RNA-seq report?
Common mistakes include omitting software version numbers, failing to document non-default parameters, describing quality control in vague terms, failing to describe the statistical model in detail, and omitting data access information. Another common mistake is describing the analysis steps without providing the rationale for the choices made. A reproducible report should document both what was done and why it was done.
How should I document the computational environment for my analysis?
The computational environment should be described in enough detail that another researcher could recreate it. This includes the operating system and version, the versions of all software packages, and the hardware specifications if they could affect the results. If containers or workflow management systems were used, the report should describe how these were configured and how the environment was isolated. The report should also describe how the environment was validated to ensure that it produced consistent results.
Related Bioinformatics Guides
- RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- RNA-Seq Normalization Methods: TPM, RPKM, and Beyond
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
- RNA-Seq vs qPCR: Validation and Comparison
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [2] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [3] [nf-core Documentation](https://nf-co.re/docs). nf-core. [4] [Protocol for total RNA sequencing analysis of extracellular RNA from biofluids.](https://doi.org/10.1016/j.xpro.2026.104540). 2026. [5] [A practical guide to targeted single-cell RNA sequencing technologies.](https://doi.org/10.1038/s42003-026-09675-y). 2026. [6] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [7] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Host-Virus Chimeric Events in SARS-CoV-2-Infected Cells Are Infrequent and Artifactual.](https://pubmed.ncbi.nlm.nih.gov/33980601). Journal of virology, 2021. [10] [Host-virus chimeric events in SARS-CoV2 infected cells are infrequent and artifactual.](https://pubmed.ncbi.nlm.nih.gov/33619483). bioRxiv : the preprint server for biology, 2021. [11] [Unveiling the molecular mechanism of sepal curvature in Dendrobium Section Spatulata through full-length transcriptome and RNA-seq analysis](https://doi.org/10.3389/fpls.2024.1497230). Frontiers in Plant Science, 2024. [12] [RNA-seq analysis of shrimp tropomyosin-induced allergic reactions through PI3K/Akt pathway](https://doi.org/10.3389/fnut.2025.1623971). Frontiers in Nutrition, 2025. [13] [Beyond counting: how single-cell long-read sequencing turns transcriptome complexity into precision targets.](https://doi.org/10.3389/fonc.2026.1800370). 2026. [14] [Integrated transcriptomic and metabolomic analysis reveals adaptive changes of hibernating retinas.](https://pubmed.ncbi.nlm.nih.gov/28542832). Journal of cellular physiology, 2018. [15] [High-fidelity PAMless base editing of hematopoietic stem cells to treat chronic granulomatous disease.](https://pubmed.ncbi.nlm.nih.gov/39413163). Science translational medicine, 2024.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.