MINSEQE Compliance for RNA-seq: A Practical Checklist for Transparent Reporting
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- MINSEQE compliance for RNA-seq is critical for experimental reproducibility and data reanalysis, requiring detailed reporting of raw data, processed data, protocols, annotation information, and experimental design parameters.
- Raw data must be deposited in its original format (e.g., FASTQ, BCL) without prior quality filtering or trimming to allow for independent downstream analysis and validation of computational choices.
- Processed data files, such as count matrices and differential expression results, must be accompanied by explicit descriptions of the reference genome/transcriptome version, alignment/quantification tools and their versions, and all parameters used.
- Comprehensive documentation of both wet-laboratory (e.g., RNA extraction, enrichment strategy, strandedness) and computational (e.g., operating system, tool versions, reference files) protocols is essential, as experimental factors like mRNA enrichment and strandedness significantly influence gene expression measurements.
- Accurate reporting of annotation information, including the specific genome build and gene model annotation version, is vital as it directly impacts alignment and quantification accuracy, particularly for novel transcript discovery.
- Detailed experimental design parameters, encompassing sample characteristics, biological replicates, conditions, and batch structure, are necessary to interpret findings and distinguish biological variation from technical or methodological artifacts.
RNA sequencing has become a standard method for measuring gene expression across biological systems, but the value of any experiment depends on whether another laboratory can interpret what was done and reproduce the analysis. The Minimum Information About a Sequencing Experiment (MINSEQE) framework establishes the reporting requirements that journals and public repositories expect authors to meet. This article provides a practical checklist for researchers who need to ensure their RNA-seq experiments and analyses meet MINSEQE reporting standards for publication and data sharing. The checklist covers experimental design, library preparation details, sequencing parameters, raw data deposition, processed data files, and the metadata that makes those files interpretable.
What MINSEQE Requires and Why It Matters
MINSEQE was developed to address a persistent problem in genomics: datasets deposited in public archives that lack the information needed to understand how they were generated or to reanalyze them independently. The core principle is that a reader should be able to take the deposited data and the accompanying description and understand exactly what biological question was asked, what material was sequenced, how libraries were prepared, what platform produced the reads, and how the raw data were converted into expression measurements.
The standard applies to any high-throughput sequencing experiment that measures gene expression, including bulk RNA-seq, single-cell RNA-seq, and long-read transcriptome sequencing. The requirements are organized around five elements: the raw data files, the final processed data files, the essential experimental and data-processing protocols, the annotation information for the genome or transcriptome used, and the experimental design parameters including sample characteristics and replicates.
For researchers working with public repositories, the practical consequence of MINSEQE is that submission to databases such as the NCBI Sequence Read Archive, Gene Expression Omnibus, or the European Nucleotide Archive requires structured metadata that maps to these elements. The NCBI Data Resources provide the infrastructure for sequence data storage, search, and analysis, and its submission portals enforce metadata collection at the point of deposit. Understanding what those portals ask for and why helps researchers prepare the necessary files before submission instead of reconstructing them afterward.
The standard matters beyond compliance. A 2024 benchmarking study across 45 laboratories using the Quartet and MAQC reference materials found that experimental factors including mRNA enrichment and strandedness, along with each bioinformatics step, were primary sources of variation in gene expression measurements. The study demonstrated greater inter-laboratory variation in detecting subtle differential expression among reference samples. This finding means that the details MINSEQE asks researchers to report are not bureaucratic formalities. They are the variables that determine whether a result will replicate in another laboratory.
The Five Core Reporting Elements
MINSEQE organizes reporting requirements into five categories. Each category maps to specific files and descriptions that must be present in a submission.
Raw Data Files
The raw data files are the sequence reads produced by the instrument, typically in FASTQ format or the native binary format of the sequencing platform. For Illumina platforms this is often BCL format before conversion to FASTQ. The submission must include the raw reads as they came off the sequencer, before any quality filtering, adapter trimming, or alignment. This requirement exists because downstream processing choices vary, and a reviewer or secondary analyst should be able to start from the original reads and apply their own pipeline.
For long-read platforms such as Oxford Nanopore or Pacific Biosciences, the raw data may include signal-level data in addition to base-called reads. The same principle applies: deposit the primary output so that reanalysis is possible. A 2024 assessment by the Long-read RNA-Seq Genome Annotation Assessment Project generated over 427 million long-read sequences from complementary DNA and direct RNA datasets across human, mouse, and manatee species. The consortium found that libraries with longer, more accurate sequences produced more accurate transcripts than those with increased read depth, while greater read depth improved quantification accuracy. These findings underscore that raw data characteristics directly influence what can be detected, and reporting those characteristics is essential for interpreting results.
Processed Data Files
Processed data files are the derived measurements that support the conclusions in a paper. For a differential expression analysis, this includes the gene-level or transcript-level count matrix, normalized expression values, and the results tables with statistics for each gene. For isoform-level analyses, the processed data include transcript quantification files. For single-cell experiments, the processed data include the cell-by-gene expression matrix and the metadata describing cell annotations.
The processed data must be accompanied by a description of how they were derived. This includes the reference genome or transcriptome version, the alignment or quantification tool and its version, the parameters used, and the filtering thresholds applied. A 2024 study mapping RNA isoform diversity in the aged human frontal cortex with deep long-read RNA-seq identified 1,917 medically relevant genes expressing multiple isoforms, with 1,018 having multiple isoforms with different protein-coding sequences. The study also uncovered 53 new RNA isoforms in medically relevant genes. These findings depended on specific alignment and quantification approaches, and the processed data files alone would be uninterpretable without the pipeline description.
Experimental and Data Processing Protocols
The protocols section must describe both the wet-laboratory methods and the computational methods. Wet-laboratory protocols include RNA extraction method, rRNA depletion or poly-A enrichment strategy, fragmentation conditions, cDNA synthesis details, adapter ligation, PCR amplification cycles, and library size selection. Computational protocols include the operating system and environment, tool names and versions, reference files and their versions, and parameter settings for each step.
The level of detail matters. A 2024 multi-center benchmarking study found that mRNA enrichment and strandedness were among the experimental factors that most influenced gene expression measurements. Reporting that libraries were prepared with a particular kit is insufficient. The description should state whether poly-A selection or rRNA depletion was used, whether the library was stranded or unstranded, and how these choices were validated.
Annotation Information
Annotation information refers to the genome assembly and gene model annotation used for alignment and quantification. This includes the genome build, the annotation source and version, and any modifications made to the annotation. For organisms with multiple annotation releases, the specific version must be identified. For experiments using transcriptome references instead of genome references, the transcript set and its source must be specified.
The choice of annotation is not neutral. The long-read benchmarking study found that in well-annotated genomes, tools based on reference sequences demonstrated the best performance. The study advised incorporating additional orthogonal data and replicate samples when aiming to detect rare and novel transcripts or using reference-free approaches. These findings indicate that annotation choice interacts with the biological question, and the annotation details must be reported for the analysis to be evaluated.
Experimental Design Parameters
Experimental design parameters include the biological question, the sample characteristics, the number of biological replicates, the conditions compared, and any batch structure. Sample characteristics include the organism, tissue or cell type, developmental stage, treatment conditions, and any relevant clinical or environmental variables. The design description should make clear which samples are compared and what statistical model was used.
A systematic review of coding and non-coding RNA expression in cumulus cells from women with polycystic ovary syndrome included 33 studies and found that the cumulus cell transcriptome was strongly modified by oocyte maturation stage, stimulation protocols, and clinical phenotypes such as insulin resistance and obesity. The review noted substantial clinical and methodological heterogeneity among studies. This heterogeneity is exactly what MINSEQE reporting is designed to make visible. Without detailed sample characteristics, readers cannot determine whether differences between studies reflect biology or methodology.
At a Glance: MINSEQE Submission Checklist
| Reporting Element | Required Files and Information | Common Submission Location |
|---|---|---|
| Raw data | FASTQ or native platform format files for all samples, including technical replicates | NCBI Sequence Read Archive or equivalent repository |
| Processed data | Count matrices, normalized expression tables, differential expression results, isoform quantifications | NCBI Gene Expression Omnibus or equivalent repository |
| Protocols | Wet-laboratory library preparation details and computational pipeline with tool versions and parameters | Methods section of manuscript and submission metadata |
| Annotation | Genome build, annotation source and version, reference transcript set, any custom modifications | Submission metadata and methods section |
| Experimental design | Sample characteristics, biological replicates, conditions, batch structure, statistical model | Submission metadata and methods section |
Practical Implementation Steps for MINSEQE Compliance
Meeting MINSEQE requirements is easier when the necessary information is collected during the experiment instead of reconstructed at submission time. The following steps provide a practical workflow.
Step 1: Define the Experimental Design Before Library Preparation
Write a study plan that specifies the biological question, the sample types, the number of biological replicates per condition, and the expected effect size. Determine whether the experiment requires poly-A selection or rRNA depletion based on the RNA species of interest. Decide whether stranded or unstranded library preparation is appropriate. Record these decisions in a laboratory notebook or electronic lab notebook with dates and personnel.
The number of replicates should be based on the expected variability and the magnitude of expression changes of interest. The multi-center benchmarking study using Quartet and MAQC reference materials found greater inter-laboratory variation in detecting subtle differential expressions. This finding implies that experiments designed to detect small expression changes require more replicates and more careful batch management than experiments detecting large changes.
Step 2: Document Library Preparation Details
For each library, record the RNA extraction method, the input RNA quantity and quality metrics such as RIN values, the library preparation kit and protocol version, the adapter sequences, the number of PCR cycles, and the final library concentration and fragment size distribution. Record whether the library is stranded or unstranded and whether poly-A enrichment or rRNA depletion was used.
These details are not optional. The benchmarking study identified mRNA enrichment and strandedness as primary sources of variation in gene expression. A reviewer cannot evaluate whether these factors were controlled without the documentation.
Step 3: Record Sequencing Parameters
For each sequencing run, record the instrument model, the flow cell type, the read length and whether the reads are single-end or paired-end, the sequencing depth per sample, and the base calling software version. For long-read platforms, record the library preparation protocol, the sequencing chemistry version, and the base calling model.
The long-read benchmarking study found that libraries with longer, more accurate sequences produced more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. These tradeoffs mean that sequencing parameters directly affect the biological conclusions that can be drawn, and they must be reported.
Step 4: Establish the Analysis Pipeline and Version Control
Choose the tools for quality control, adapter trimming, alignment or quantification, and differential expression analysis. Record the exact versions of each tool and the parameters used. Use a workflow management system that captures the pipeline configuration. The nf-core Documentation describes community standards for pipeline usage and configuration that support reproducible workflow execution. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility. Bioconductor provides official package and workflow documentation for reproducible genomic analysis.
Version control is essential. Tools change frequently, and results can differ between versions. The Carpentries Lessons provide foundational training in Git and shell computing that supports version-controlled analysis. A pipeline that cannot be reconstructed from its documentation does not meet MINSEQE standards.
Step 5: Perform Quality Control and Record the Results
Run quality control at multiple stages: raw reads, after trimming, after alignment, and after quantification. Record the metrics at each stage, including read counts, duplication rates, alignment rates, gene detection rates, and any samples that failed quality thresholds. Record the decisions made in response to quality issues, such as removing a sample or changing a filtering threshold.
The multi-center benchmarking study provided best practice recommendations for experimental designs, strategies for filtering low-expression genes, and the optimal gene annotation and analysis pipelines. These recommendations are only actionable if the quality metrics that drive filtering decisions are recorded.
Step 6: Prepare Submission Files During the Analysis
Create the processed data files in the formats required by the repository. For the Gene Expression Omnibus, this includes the processed data matrix and the metadata table. For the Sequence Read Archive, this includes the raw read files and the run metadata. Prepare these files while the analysis is fresh instead of months later when details have been forgotten.
Step 7: Submit to a Public Repository and Obtain Accession Numbers
Submit the raw data to the Sequence Read Archive or an equivalent repository and the processed data to the Gene Expression Omnibus or an equivalent repository. Obtain accession numbers for each dataset. Include these accession numbers in the manuscript and in the submission metadata.
The NCBI Data Resources provide the infrastructure for sequence data storage, search, and analysis. Its submission portals guide researchers through the metadata collection process. Using these portals ensures that the submission meets the repository requirements, which align with MINSEQE.
Records and Measurements to Maintain
The following records should be maintained for each RNA-seq experiment. These records serve both the submission process and the internal quality assurance process.
Sample Tracking Records
Maintain a sample manifest that links each biological sample to its RNA extraction date, library preparation date, sequencing run, and any quality metrics. The manifest should include unique identifiers for each sample and each library. This manifest is the foundation for the metadata table required by repositories.
Quality Control Metrics
Record the following metrics for each sample: RNA quantity and quality, library concentration, fragment size distribution, raw read count, read count after trimming, alignment rate, gene detection rate, and any other metrics generated by the quality control tools. Record the thresholds used and any samples that failed thresholds.
Pipeline Configuration Files
Maintain the exact configuration files used for each analysis step. This includes the tool versions, reference files, and parameters. The nf-core Documentation describes how pipeline configuration supports reproducible workflow execution. The Bioconductor documentation describes how package versions and session information support reproducible analysis.
Analysis Decision Log
Maintain a log of decisions made during the analysis, including filtering thresholds, normalization methods, batch correction approaches, and statistical models. Record the rationale for each decision. This log is essential for responding to reviewer questions and for reanalysis.
Common Failure Patterns in MINSEQE Compliance
Several recurring problems appear in submissions that fail to meet MINSEQE standards. Recognizing these patterns helps researchers avoid them.
Missing Raw Data
Some submissions include only processed data or only summary statistics. This pattern prevents reanalysis and fails the most basic MINSEQE requirement. The raw data must be deposited in a public repository. The NCBI Sequence Read Archive is the standard destination for raw sequencing data.
Incomplete Protocol Descriptions
Many submissions describe the library preparation kit but omit the specific protocol version, the input quantity, or the number of PCR cycles. Others describe the analysis tools but omit the versions or parameters. These omissions make the experiment impossible to reproduce. The multi-center benchmarking study found that each bioinformatics step was a primary source of variation in gene expression, which means that incomplete pipeline descriptions hide the variables that most affect results.
Ambiguous Sample Annotations
Some submissions use sample names that do not clearly indicate the condition, tissue, or replicate. Others omit batch information or clinical variables. A systematic review of cumulus cell transcriptomics in polycystic ovary syndrome found that the transcriptome was strongly modified by oocyte maturation stage, stimulation protocols, and clinical phenotypes. Without these sample characteristics in the metadata, the data cannot be interpreted or compared across studies.
Unspecified Annotation Versions
Some submissions state that reads were aligned to the human genome but do not specify the genome build or annotation version. This omission is critical because gene models change between annotation releases. The long-read benchmarking study found that tools based on reference sequences demonstrated the best performance in well-annotated genomes, but this performance depends on the specific reference version used.
Processed Data Without Pipeline Description
Some submissions include processed data files but do not describe how they were generated. The count matrix alone is insufficient. The description must include the quantification tool, the reference used, and the filtering thresholds. Without this information, the processed data cannot be evaluated or compared to other datasets.
Limitations of MINSEQE and Related Reporting Standards
MINSEQE establishes a minimum standard, but it does not cover every aspect of experimental reporting. Researchers should be aware of the limitations and of complementary standards.
MINSEQE Does Not Mandate Specific Analysis Methods
The standard requires that analysis methods be described, but it does not prescribe which methods should be used. This flexibility is appropriate because different biological questions require different approaches. However, it means that two studies addressing the same question may use different pipelines, and the results may not be directly comparable. The multi-center benchmarking study found that each bioinformatics step contributed to variation in gene expression, which means that pipeline differences can produce different conclusions.
MINSEQE Does Not Address Statistical Rigor
The standard requires that the experimental design be described, but it does not require that the design be statistically adequate. A study with one replicate per condition can meet MINSEQE requirements if the design is described accurately. Researchers should apply appropriate statistical standards in addition to MINSEQE compliance.
Single-Cell Experiments Require Additional Metadata
The MINSEQE standard was developed primarily for bulk sequencing experiments. Single-cell RNA-seq experiments require additional metadata, including cell capture methods, cell viability metrics, and cell annotation criteria. Guidelines for reporting single-cell RNA-seq experiments describe a minimum set of metadata to sufficiently describe these experiments and ensure reproducibility of data analyses. Researchers conducting single-cell experiments should consult these guidelines in addition to MINSEQE.
Long-Read Experiments Have Specific Reporting Needs
Long-read RNA-seq experiments require reporting of library preparation protocols, sequencing chemistry, and base calling models that differ from short-read experiments. The long-read benchmarking study provided a benchmark for current practices and direction for future method development in transcriptome analysis. Researchers using long-read platforms should report the platform-specific parameters that affect transcript identification and quantification.
Clinical Applications Require Additional Validation
RNA-seq is increasingly used in clinical contexts, and the reporting requirements extend beyond MINSEQE. A review of RNA-seq data analysis through visualization techniques and tools found that effective visualization is important for helping clinicians and biomedical researchers understand complex patterns of gene expression associated with health and disease. A study of RNA-seq in acute lymphoblastic leukemia demonstrated that rapid detection and accurate reporting of clinically relevant alterations requires specialized tools and validation approaches. Researchers working toward clinical applications should consult the relevant regulatory and validation frameworks.
Quality Controls and Their Role in Transparent Reporting
Quality control is not separate from MINSEQE compliance. The quality metrics generated during an analysis are part of the experimental record, and they should be reported alongside the processed data.
Raw Read Quality Assessment
Assess raw read quality using metrics such as per-base quality scores, GC content, adapter contamination, and duplication rates. Record these metrics for each sample. The Galaxy Training Network provides accessible workflow training that includes quality control tutorials. The Bioconductor documentation describes packages for quality assessment and visualization.
Alignment and Quantification Quality Assessment
Assess alignment rates, read distribution across genomic features, and gene detection rates. Record these metrics and any samples that fail thresholds. The multi-center benchmarking study provided recommendations for filtering low-expression genes, and these recommendations depend on the quality metrics that reveal which genes are reliably measured.
Reproducibility Assessment
Assess the reproducibility of the analysis by running the pipeline on a subset of samples twice or by using a reference sample. The Quartet and MAQC reference materials were developed for this purpose. The benchmarking study using these materials found that experimental execution had a profound influence on gene expression measurements, which means that reproducibility assessment should be part of the experimental workflow.
Common Analysis Workflows and Their Reporting Requirements
Different RNA-seq analysis workflows require different reporting details. Understanding the workflow-specific requirements helps researchers prepare complete submissions.
Bulk RNA-seq Differential Expression Workflow
The standard bulk RNA-seq workflow includes quality control, adapter trimming, alignment or pseudo-alignment, quantification, normalization, and differential expression analysis. Each step must be documented with tool versions and parameters. The Galaxy Training Network provides tutorials that walk through each step of this workflow with reproducible commands. The Bioconductor project provides packages for each step, and its documentation includes session information that captures package versions.
Single-Cell RNA-seq Workflow
Single-cell RNA-seq workflows include additional steps for cell calling, quality filtering, normalization, dimensionality reduction, clustering, and cell type annotation. The guidelines for reporting single-cell RNA-seq experiments describe the minimum metadata needed to reproduce these analyses. The nf-core Documentation describes community pipelines for single-cell analysis that capture the full configuration in a reproducible format.
Long-Read Transcriptome Workflow
Long-read workflows require reporting of the base calling model, the alignment tool and parameters, and the transcript identification approach. The long-read benchmarking study found that libraries with longer, more accurate sequences produced more accurate transcripts than those with increased read depth. This finding means that the read length distribution and accuracy metrics should be reported alongside the analysis parameters.
Visualization and Interpretation
A systematic review of RNA-seq visualization tools found that 56% of studies used visualization techniques for single-cell RNA-seq data, 23% for bulk RNA-seq data, 18% for circular RNA-seq data, and 3% for long non-coding RNA-seq data. The review emphasized that visualization choices affect how clinicians and researchers interpret complex expression patterns. The visualization methods used should be described in the methods section, and the underlying data should be available in the processed data files.
Metadata Standards and Repository Submission
The metadata submitted to repositories must be structured according to the repository requirements. Understanding these requirements helps researchers prepare submissions that meet MINSEQE standards.
NCBI Submission Portals
The NCBI Data Resources provide submission portals for sequence data and processed data. The Sequence Read Archive accepts raw sequencing data, and the Gene Expression Omnibus accepts processed data and metadata. The submission portals guide researchers through the metadata collection process, asking for the information that maps to MINSEQE elements.
Metadata Fields for Raw Data Submission
The raw data submission requires the following metadata fields: sample identifiers, organism, tissue or cell type, library preparation method, sequencing platform, read length, and read orientation. The submission portal also asks for the library construction protocol and the sequencing instrument model.
Metadata Fields for Processed Data Submission
The processed data submission requires the processed data files, the analysis pipeline description, the reference genome and annotation versions, and the normalization method. The submission portal asks for the processed data matrix and the metadata table that describes each sample.
Data Format Requirements
The raw data must be submitted in FASTQ format or the native format of the sequencing platform. The processed data must be submitted in a tabular format that can be read by standard tools. The Bioconductor documentation describes the standard data structures for genomic data, and the Galaxy Training Network provides tutorials on data format conversion.
Professional Escalation Criteria
Researchers should seek additional guidance or escalate concerns in the following situations.
When Repository Requirements Are Unclear
If the submission portal requirements are unclear, consult the repository documentation or contact the repository help desk. The NCBI Data Resources provide documentation and support for its databases. The Galaxy Training Network provides tutorials that include data submission guidance.
When Analysis Results Are Unexpected
If quality metrics indicate problems that cannot be explained by known sample characteristics, consult a bioinformatics specialist or the tool developers. Unexpected patterns in alignment rates, gene detection, or batch effects may indicate technical problems that require investigation before the data are submitted.
When Clinical or Regulatory Use Is Anticipated
If the data will be used for clinical decision-making or regulatory submissions, consult the relevant regulatory framework and seek guidance from specialists in clinical genomics. The multi-center benchmarking study laid the foundation for developing and quality control of RNA-seq for clinical diagnostic purposes, and its recommendations should be considered.
When Novel Transcripts or Isoforms Are Reported
If the analysis identifies novel transcripts or isoforms, additional validation may be required. The long-read benchmarking study advised incorporating additional orthogonal data and replicate samples when aiming to detect rare and novel transcripts. The study of RNA isoform diversity in the aged human frontal cortex identified new RNA isoforms in medically relevant genes, and such findings require careful validation and reporting.
When Single-Cell Data Are Submitted
If the experiment involves single-cell RNA-seq, consult the guidelines for reporting single-cell RNA-seq experiments in addition to MINSEQE. These guidelines describe the minimum metadata needed to ensure reproducibility of single-cell analyses.
A Decision Framework for Choosing MINSEQE Reporting Depth by Experiment Type
MINSEQE compliance is not a single fixed standard applied identically to every experiment. The reporting burden and the specific metadata fields that matter most vary by experiment type, platform, and biological question. Researchers who apply the same checklist to every project either over-report irrelevant details or under-report critical variables. This section provides a practical decision framework for calibrating MINSEQE reporting depth to the experiment at hand, with concrete criteria for determining which metadata fields are essential, which are recommended, and which can be omitted without compromising reproducibility.
Step 1: Classify the Experiment by Primary Question
The first decision point is the biological question the experiment addresses. This classification determines which reporting elements carry the most weight. Three broad categories cover most RNA-seq experiments.
Category A: Differential Expression Between Conditions. These experiments compare gene expression levels across two or more conditions, such as treated versus untreated, diseased versus healthy, or time course samples. The critical reporting elements are the experimental design parameters, the number of biological replicates, the batch structure, and the statistical model. The multi-center benchmarking study using Quartet and MAQC reference materials found greater inter-laboratory variation in detecting subtle differential expressions, which means that experiments designed to detect small changes require more detailed reporting of the statistical approach and filtering decisions. For this category, the processed data files must include the full count matrix and the differential expression results with all statistics, beyond the significant genes.
Category B: Transcript Discovery and Isoform Characterization. These experiments aim to identify novel transcripts, characterize isoform diversity, or annotate genomes. The critical reporting elements are the annotation information, the alignment and quantification approach, and the platform-specific parameters. The Long-read RNA-Seq Genome Annotation Assessment Project found that libraries with longer, more accurate sequences produced more accurate transcripts than those with increased read depth, while greater read depth improved quantification accuracy. For this category, the raw data files and the annotation files used for analysis are the most important submission components. The processed data must include transcript-level quantifications and the specific transcript models identified.
Category C: Biomarker Discovery and Clinical Correlation. These experiments link expression patterns to clinical outcomes, diagnostic categories, or prognostic groups. The critical reporting elements are the sample characteristics, including clinical variables, and the validation approach. A systematic review of coding and non-coding RNA expression in cumulus cells from women with polycystic ovary syndrome found that the transcriptome was strongly modified by oocyte maturation stage, stimulation protocols, and clinical phenotypes such as insulin resistance and obesity. For this category, the sample metadata must include all relevant clinical variables, and the processed data must be accompanied by the full statistical results linking expression to outcomes.
Step 2: Determine Platform-Specific Reporting Requirements
The sequencing platform determines which metadata fields are essential. Short-read and long-read platforms have different failure modes and different parameters that affect results.
Short-read platforms (Illumina and similar). The essential reporting elements are the read length, read orientation (single-end or paired-end), strandedness, and the library preparation method. The multi-center benchmarking study identified mRNA enrichment and strandedness as primary sources of variation in gene expression. These two variables must be reported for every short-read experiment. The specific kit version and the number of PCR cycles are also essential because they affect duplication rates and library complexity.
Long-read platforms (Oxford Nanopore and Pacific Biosciences). The essential reporting elements are the library preparation protocol, the sequencing chemistry version, the base calling model, and the read length distribution. The Long-read RNA-Seq Genome Annotation Assessment Project found that libraries with longer, more accurate sequences produced more accurate transcripts than those with increased read depth. This finding means that the read length distribution and accuracy metrics must be reported alongside the analysis parameters. For direct RNA sequencing, the RNA preparation method and any RNA modifications that affect base calling must also be documented.
Single-cell platforms. Single-cell experiments require additional metadata beyond the MINSEQE baseline. The guidelines for reporting single-cell RNA-seq experiments describe a minimum set of metadata to sufficiently describe these experiments, including cell capture methods, cell viability metrics, and cell annotation criteria. The cell-by-gene expression matrix and the cell metadata must be deposited as processed data files.
Step 3: Assess the Annotation Dependency
The choice of reference genome and annotation version is not neutral, and the reporting depth should reflect how dependent the analysis is on the annotation.
Well-annotated organisms. For human, mouse, and other model organisms with mature annotations, the reporting requirement is to specify the genome build and annotation version precisely. The Long-read RNA-Seq Genome Annotation Assessment Project found that in well-annotated genomes, tools based on reference sequences demonstrated the best performance. The annotation version must be reported because gene models change between releases, and results from different annotation versions are not directly comparable.
Poorly annotated organisms. For non-model organisms or newly sequenced genomes, the annotation is often incomplete, and the analysis may depend on de novo transcript assembly or reference-free approaches. The Long-read RNA-Seq Genome Annotation Assessment Project advised incorporating additional orthogonal data and replicate samples when aiming to detect rare and novel transcripts or using reference-free approaches. For these experiments, the reporting burden is higher. The submission must include the genome assembly version, the evidence used to construct the annotation, and the parameters used for any de novo assembly steps.
Custom annotations. If the analysis uses a modified annotation, such as adding novel isoforms or filtering specific gene models, the modifications must be described and the modified annotation file must be deposited. The processed data are uninterpretable without the exact annotation used for quantification.
Step 4: Determine the Required Metadata Granularity
The level of detail in the sample metadata should match the expected sources of biological and technical variation.
Minimum metadata for all experiments. Every sample must have a unique identifier, the organism, the tissue or cell type, the condition or treatment, and the replicate number. These fields are required for any submission to the NCBI Sequence Read Archive or Gene Expression Omnibus.
Expanded metadata for heterogeneous samples. If the samples come from different batches, different collection sites, or different time points, the batch information and collection details must be included. The multi-center benchmarking study found that experimental execution was a primary source of variation in gene expression. Batch structure is part of that execution, and it must be reported for the statistical model to be evaluated.
Clinical metadata for human studies. For human samples, the metadata must include the relevant clinical variables, such as age, sex, disease subtype, treatment history, and any other variables that could affect gene expression. The systematic review of cumulus cell transcriptomics in polycystic ovary syndrome found that the transcriptome was strongly modified by clinical phenotypes such as insulin resistance and obesity. These variables must be in the metadata for the data to be interpretable or comparable across studies.
Step 5: Match the Pipeline Documentation to the Analysis Complexity
The pipeline description must be detailed enough for another researcher to reproduce the analysis, but the required detail varies with the analysis complexity.
Simple pipelines. A standard differential expression pipeline with quality control, trimming, alignment, quantification, and differential expression analysis requires documentation of each tool, its version, and the parameters used. The Bioconductor documentation describes how package versions and session information support reproducible analysis. The Galaxy Training Network provides tutorials that walk through each step with reproducible commands.
Complex pipelines. Pipelines that include batch correction, multiple normalization approaches, or custom filtering steps require additional documentation. The decision log should record the rationale for each choice. The nf-core Documentation describes how pipeline configuration supports reproducible workflow execution, and its community pipelines capture the full configuration in a reproducible format.
Machine learning and AI-assisted analysis. A review of artificial intelligence in transcriptomics described a shift from human-in-the-loop approaches to agentic AI systems that can autonomously retrieve, pre-process, and analyze transcriptomics data. If the analysis uses AI methods, the reporting requirements extend beyond MINSEQE. The specific model, training data, version, and parameters must be documented, and the reproducibility of AI-assisted analysis must be assessed separately. The review noted that these approaches raise important considerations regarding the advantages and concerns of automated analysis, and researchers should document the human oversight applied to the analysis.
Step 6: Apply the Escalation Criteria for Additional Reporting
Certain experimental features trigger additional reporting requirements beyond the baseline MINSEQE elements.
Rare or novel transcript detection. If the experiment aims to detect rare transcripts or novel isoforms, the reporting must include the read depth, the library complexity, and the validation approach. The Long-read RNA-Seq Genome Annotation Assessment Project advised incorporating additional orthogonal data and replicate samples when aiming to detect rare and novel transcripts. The processed data must include the evidence supporting each novel transcript call.
Clinical or regulatory use. If the data will be used for clinical decision-making or regulatory submissions, the reporting must follow the relevant regulatory framework. The multi-center benchmarking study laid the foundation for developing and quality control of RNA-seq for clinical diagnostic purposes. The study found that experimental execution had a profound influence on gene expression measurements, and its best practice recommendations should be followed.
Multi-omics integration. If the RNA-seq data are integrated with other omics data, such as proteomics or epigenomics, the reporting must describe the integration approach and the data versions used. An integrative systems biology study of colorectal cancer combined transcriptomics, network topology, and immune profiling to identify prognostic biomarkers. The study demonstrated that single-gene markers often overlook the network-based nature of tumorigenesis, and the reporting must capture the integration methods for the results to be evaluated.
A Practical Decision Table for Reporting Depth
| Experiment Feature | Essential Reporting | Recommended Reporting | Optional Reporting |
|---|---|---|---|
| Differential expression in well-annotated organism | Raw data, count matrix, differential expression results, annotation version, library preparation method, strandedness | Batch structure, filtering thresholds, normalization method, quality metrics | Visualization parameters, exploratory analysis details |
| Transcript discovery with long-read platform | Raw signal data, base called reads, base calling model, read length distribution, annotation files | Orthogonal validation data, replicate information, de novo assembly parameters | Cross-platform comparisons |
| Single-cell experiment | Cell-by-gene matrix, cell metadata, cell capture method, cell viability metrics | Cell annotation criteria, clustering parameters, batch correction approach | Marker gene lists, trajectory analysis parameters |
| Biomarker study with clinical samples | Raw data, processed data, full clinical metadata, statistical results | Validation cohort details, model performance metrics, confounder analysis | Exploratory subgroup analyses |
| Poorly annotated organism | Raw data, genome assembly version, annotation evidence, de novo assembly parameters | Orthogonal data for validation, replicate recommendations | Comparative genomics analyses |
Implementing the Decision Framework in Practice
The decision framework should be applied at the experimental design stage, not at the submission stage. When writing the study plan, classify the experiment into one of the three categories and identify the platform-specific reporting requirements. This classification determines which metadata fields to collect during the experiment and which quality metrics to record.
For example, a researcher planning a differential expression experiment in human cell lines should record the library preparation kit and version, the strandedness, the read length, and the number of PCR cycles during library preparation. The researcher should also record the batch information for each sample and the quality metrics at each analysis stage. At submission time, the researcher needs only to compile the records that were already collected.
A researcher planning a transcript discovery experiment in a non-model organism should plan for additional replicates and orthogonal validation data from the start. The Long-read RNA-Seq Genome Annotation Assessment Project found that libraries with longer, more accurate sequences produced more accurate transcripts than those with increased read depth, which means the sequencing strategy should prioritize read length and accuracy over depth. The reporting plan should include the read length distribution and accuracy metrics as essential fields.
A researcher planning a biomarker study with clinical samples should identify the clinical variables that affect gene expression in the relevant disease context. The systematic review of cumulus cell transcriptomics in polycystic ovary syndrome found that the transcriptome was strongly modified by oocyte maturation stage, stimulation protocols, and clinical phenotypes. These variables must be collected and recorded for each sample from the start of the study.
Common Mistakes in Applying the Decision Framework
The most common mistake is applying the same reporting template to every experiment without considering the experiment type. This approach either produces submissions with missing critical fields or submissions with excessive irrelevant detail that obscures the essential information.
A second common mistake is treating the decision framework as a one-time assessment. The reporting requirements should be revisited at each analysis stage. If the analysis reveals unexpected batch effects, the batch structure must be added to the metadata. If the analysis identifies novel transcripts, the validation approach must be documented. The decision framework should be applied iteratively throughout the experiment.
A third mistake is assuming that more reporting is always better. The goal is sufficient reporting for reproducibility, not exhaustive documentation of every parameter. The decision framework helps researchers identify which fields are essential for their experiment type and focus their documentation efforts on those fields.
When to Seek Professional Guidance
The decision framework covers the common experiment types, but some experiments fall outside these categories or have unusual features. Researchers should seek additional guidance in the following situations.
Multi-platform comparisons. If the experiment compares results across different sequencing platforms, the reporting must capture the platform-specific parameters for each platform. The Long-read RNA-Seq Genome Annotation Assessment Project compared different protocols and sequencing platforms, and its findings can guide the reporting requirements for multi-platform studies.
Integration with AI methods. If the analysis uses AI models for gene expression analysis and integration, the reporting requirements extend beyond MINSEQE. The review of artificial intelligence in transcriptomics described the capabilities of AI models to retrieve, analyze, integrate, and generate data recursively. The specific AI methods, training data, and validation approaches must be documented, and the human oversight applied to the analysis must be described.
Regulatory submissions. If the data will be used for regulatory submissions, consult the relevant regulatory framework and seek guidance from specialists in clinical genomics. The multi-center benchmarking study provided best practice recommendations for experimental designs, strategies for filtering low-expression genes, and the optimal gene annotation and analysis pipelines for clinical diagnostic purposes.
The decision framework provides a structured approach to determining MINSEQE reporting depth, but it does not replace the judgment of the researcher. The framework identifies the fields that are likely to matter for each experiment type, and the researcher must verify that the specific experiment has no additional variables that affect the results. The goal is a submission that another laboratory can use to reproduce the analysis and evaluate the conclusions.
Frequently Asked Questions
What is the difference between MINSEQE and MIAME?
MIAME was developed for microarray experiments and requires information about the array design, sample preparation, hybridization protocols, and scanning parameters. MINSEQE adapts these principles for high-throughput sequencing and requires information about library preparation, sequencing platform, read processing, and data analysis. Both standards share the goal of making experiments interpretable and reproducible, but the specific reporting elements differ because the technologies differ.
Which repositories accept MINSEQE-compliant RNA-seq data?
The NCBI Sequence Read Archive accepts raw sequencing data, and the NCBI Gene Expression Omnibus accepts processed data and metadata. Equivalent repositories include the European Nucleotide Archive and the DNA Data Bank of Japan. Most journals require deposition in one of these public repositories as a condition of publication.
How much detail is required for the analysis pipeline description?
The pipeline description must include the tool names, versions, and parameters for each step from raw reads to final results. This includes the quality control tools, trimming tools, alignment or quantification tools, and differential expression tools. The description should be sufficient for another researcher to run the same pipeline and obtain the same results.
What should be done if a sample fails quality control?
The quality control failure should be documented, and the decision to include or exclude the sample should be recorded with the rationale. The raw data for the failed sample should still be deposited, and the exclusion should be noted in the metadata. Excluding samples without documentation is a common failure pattern in submissions.
Are single-cell RNA-seq experiments covered by MINSEQE?
MINSEQE provides the baseline for single-cell experiments, but additional metadata are required. Guidelines for reporting single-cell RNA-seq experiments describe a minimum set of metadata to sufficiently describe these experiments. Researchers should follow both the MINSEQE standard and the single-cell specific guidelines.
How should batch effects be reported?
Batch information should be included in the sample metadata, and the batch correction approach should be described in the analysis pipeline. The experimental design description should indicate whether batches were balanced across conditions. The multi-center benchmarking study found that experimental execution was a primary source of variation, and batch structure is part of that execution.
What is the role of reference materials in RNA-seq reporting?
Reference materials such as the Quartet and MAQC samples provide a common benchmark for evaluating RNA-seq performance across laboratories. The benchmarking study using these materials found greater inter-laboratory variation in detecting subtle differential expressions. Including reference samples in an experiment provides a basis for evaluating whether the results are comparable to other laboratories.
How should novel isoforms be reported?
Novel isoforms should be reported with the evidence supporting their existence, including the read support, the splicing patterns, and any validation data. The long-read benchmarking study found that libraries with longer, more accurate sequences produce more accurate transcripts, and the study advised incorporating additional orthogonal data when aiming to detect rare and novel transcripts. The annotation files used for the analysis should be deposited so that others can evaluate the novel isoform calls.
Related Bioinformatics Guides
- RNA-Seq Databases: Accessing and Using Public RNA-Seq Data
- RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- RNA-Seq vs qPCR: Validation and Comparison
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Gene and non-coding RNA expression in human cumulus cells in polycystic ovary syndrome: A systematic review.. 2026.
- Artificial Intelligence in Transcriptomics: From Human-in-the-Loop to Agentic AI.. 2026.
- Integrative systems biology identifies <,i>,PRELP<,/i>, and <,i>,TAGLN<,/i>, as prognostic biomarkers and immune modulators in colorectal cancer.. 2025.
- Insights from meta-analysis and experimental validation identify exosomal miR-146a-5p as a potential biomarker for sporadic amyotrophic lateral sclerosis.. 2025.
- Decoding the archipelago: single-cell biomarkers rechart the molecular geography of acute myeloid leukemia.. 2026.
- Exploring RNA-Seq Data Analysis Through Visualization Techniques and Tools: A Systematic Review of Opportunities and Limitations for Clinical Applications. Bioengineering, 2025.
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification. Nature Methods, 2024.
- Mapping medically relevant RNA isoform diversity in the aged human frontal cortex with deep long-read RNA-seq. Nature Biotechnology, 2024.
- RaScALL: Rapid (Ra) screening (Sc) of RNA-seq data for prognostically significant genomic alterations in acute lymphoblastic leukaemia (ALL). PLoS Genetics, 2022.
- Guidelines for reporting single-cell RNA-seq experiments. Nature Biotechnology, 2019.
- A real-world multi-center RNA-seq benchmarking study using the Quartet and MAQC reference materials. Nature Communications, 2024.
- The landscape of abiotic and biotic stress-responsive splice variants with deep RNA-seq datasets in hot pepper. Scientific Data, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.