RNA Sequencing Protocol: From Sample to Data
RNA sequencing (RNA-seq) is a laboratory workflow that converts RNA molecules into complementary DNA (cDNA), prepares sequencing libraries, and generates digital readouts of transcript abundance and structure. This protocol covers the complete path from RNA extraction through library preparation and sequencing, with quality control checkpoints at each stage. The intended readers are laboratory students, technicians, researchers, and diagnostic professionals who need a practical, decision-oriented reference for executing RNA-seq experiments and interpreting their results.
RNA-seq has become one of the most widely used technologies in transcriptomics because it can reveal relationships between genetic alterations and complex biological processes, with value in diagnostics, prognostics, and therapeutics of tumors. The protocol described here follows the general structure of RNA extraction, cDNA synthesis, cDNA fragmentation, and library generation, with attention to the choices that determine data quality and interpretability.
At a Glance
The table below summarizes the major workflow stages, the key decisions at each stage, and the quality indicators that should be monitored before proceeding to the next step.
| Workflow Stage | Primary Decisions | Quality Checkpoints |
|---|---|---|
| RNA Extraction | Source tissue type, extraction method, DNase treatment | RNA integrity number (RIN), yield, purity ratios (A260/A280, A260/230), absence of genomic DNA contamination |
| RNA Selection or Depletion | Poly-A enrichment versus ribosomal RNA depletion versus total RNA | Proportion of reads mapping to mRNA, retained representation of non-coding RNAs, input RNA amount compatibility |
| cDNA Synthesis and Library Preparation | Strand-specific versus non-strand-specific, fragmentation time, ligation approach, amplification cycles | Library yield, fragment size distribution, adapter dimer presence, index assignment accuracy |
| Sequencing | Read length, paired-end versus single-end, sequencing depth, platform choice | Cluster density, Q30 scores, read duplication rate, on-target alignment rate |
| Data Analysis | Alignment tool, quantification method, normalization approach, differential expression method | Alignment rate, gene detection count, batch effects, biological replicate concordance |
Scope and Context of RNA Sequencing Protocols
RNA sequencing protocols vary substantially depending on the biological question, the input material, and the available equipment. A protocol that works well for cultured cells with abundant RNA may fail for clinical biopsies with limited material. The choice of protocol affects which transcripts are detected, how accurately expression levels are measured, and whether splicing events and structural variants can be resolved.
The evolution of next-generation sequencing technologies over the past decade has brought tremendous progress in speed, read length, and throughput, along with a sharp reduction in per-base cost. These advances have made RNA-seq accessible for basic science and for translational research areas such as clinical diagnostics, agrigenomics, and forensic science. Library preparation protocols have also improved, with options ranging from traditional methods to full-length double-stranded cDNA approaches and automated miniaturized workflows.
For diagnostic applications, the laboratory must operate within a quality management framework. The World Health Organization Laboratory Quality Management System Handbook describes the components of a quality management system that apply to all stages of the testing process, including pre-examination, examination, and post-examination phases. RNA-seq protocols in diagnostic settings should be validated, documented, and monitored with the same rigor applied to any laboratory-developed test.
Core Principles of RNA Sequencing
RNA-seq measures the abundance and sequence of RNA molecules in a biological sample. The workflow begins with RNA extraction, proceeds through conversion of RNA to cDNA, adds sequencing adapters, and culminates in high-throughput sequencing of the resulting library. Each step introduces potential bias, and the cumulative effect of these biases determines the biological validity of the final data.
RNA as the Starting Material
RNA is chemically less stable than DNA. Ribonucleases are ubiquitous in the environment and on human skin, so sample handling requires strict precautions. The quality of the extracted RNA is the single most important determinant of downstream success. Degraded RNA produces libraries with reduced complexity, biased gene body coverage, and unreliable quantification of low-abundance transcripts.
The RNA extraction method must be matched to the sample type. Solid tissues require homogenization and often benefit from column-based purification. Cells in suspension can be lysed directly. Blood samples may require red blood cell lysis or density gradient separation before RNA extraction. The extraction method should remove genomic DNA, proteins, and chemical inhibitors that could interfere with enzymatic steps.
Conversion of RNA to cDNA
Reverse transcriptase converts RNA into cDNA, which is stable and compatible with DNA sequencing platforms. The efficiency of reverse transcription depends on the enzyme, the priming strategy, and the RNA secondary structure. Random hexamers prime throughout the transcript, while oligo-dT primers selectively prime polyadenylated mRNAs. The choice of priming strategy affects which RNA species are represented in the final library.
For single-cell RNA sequencing, the reverse transcription step often incorporates unique molecular identifiers (UMIs), which are random barcodes that label individual mRNA molecules. UMIs allow amplification noise to be distinguished from true biological variation. Methods that use UMIs, such as CEL-seq2, Drop-seq, MARS-seq, and SCRB-seq, quantify mRNA levels with less amplification noise than methods without UMIs.
Library Construction and Amplification
The cDNA is fragmented, end-repaired, and ligated to sequencing adapters. The adapter sequences provide binding sites for the sequencing instrument and allow multiplexing of multiple samples in a single sequencing run. The library is then amplified by PCR to generate sufficient material for sequencing.
The number of PCR cycles must be minimized to reduce amplification bias and duplicate reads. However, low-input samples require more cycles to produce adequate library yield. The tradeoff between yield and bias is a central decision in library preparation.
Sequencing and Signal Generation
The completed library is loaded onto a sequencing flow cell, where clonal amplification generates clusters of identical templates. The sequencing instrument reads the nucleotide sequence of each cluster, producing millions of short reads. Read length and paired-end versus single-end configuration are chosen based on the experimental goals.
RNA Extraction and Quality Assessment
RNA extraction is the first and most error-prone stage of the RNA-seq workflow. The goal is to obtain intact, pure RNA that accurately represents the transcriptome of the source tissue.
Sample Collection and Stabilization
The time between sample collection and RNA stabilization should be minimized. Fresh tissue should be snap-frozen in liquid nitrogen or placed in an RNA stabilization reagent immediately after collection. Delayed stabilization leads to RNA degradation and changes in gene expression profiles that reflect the stress response instead of the biological state of interest.
For blood samples, the stabilization method depends on whether the goal is to profile circulating cells or cell-free RNA. PAXgene tubes stabilize intracellular RNA at the time of blood collection, while plasma or serum requires separation within a defined time window to prevent RNA release from blood cells.
Extraction Method Selection
Column-based kits are the most common approach for RNA extraction. They use silica membranes to bind RNA in the presence of chaotropic salts, followed by washing and elution. The advantages are consistency, speed, and the removal of enzymatic inhibitors. The disadvantages are cost per sample and the potential for column overload at high RNA concentrations.
Organic extraction with phenol-chloroform is an alternative that can yield higher quantities of RNA from difficult samples. The method is labor-intensive and requires careful phase separation to avoid contamination with the interphase containing DNA and proteins. Organic extraction is often used for samples with high lipid content or fibrous tissue.
Magnetic bead-based extraction is increasingly common in automated workflows. Beads bind RNA reversibly, allowing washing and elution in a format compatible with liquid handlers. Automated extraction reduces hands-on time and improves reproducibility across large batches.
DNase Treatment
Genomic DNA contamination is a serious problem for RNA-seq because it produces reads that align to intronic and intergenic regions, confounding expression quantification. Most RNA extraction kits include an optional on-column DNase treatment. For samples where genomic DNA contamination is expected to be high, such as formalin-fixed paraffin-embedded tissue, an additional DNase digestion after elution may be necessary.
RNA Quality Metrics
RNA integrity is assessed by microfluidic electrophoresis, which separates RNA molecules by size and produces an electropherogram. The RNA integrity number (RIN) is a score from 1 to 10 that reflects the degree of degradation. High-quality RNA from fresh tissue typically has a RIN above 8. Degraded samples have lower RIN values and produce libraries with reduced complexity.
The A260/A280 ratio indicates protein contamination, with values around 2.0 for pure RNA. The A260/230 ratio indicates contamination with chaotropic salts or organic solvents, with values above 1.8 considered acceptable. These spectrophotometric measurements are useful but should be interpreted alongside the electropherogram, because degraded RNA can still produce acceptable absorbance ratios.
Documentation of RNA Quality
Every RNA extraction should be documented with the sample identifier, extraction date, method, yield, RIN value, and absorbance ratios. This documentation supports troubleshooting if downstream steps fail and provides the traceability required for diagnostic applications. The World Health Organization Laboratory Quality Management System Handbook emphasizes the importance of records that allow the reconstruction of the entire testing process.
Library Preparation Strategies
Library preparation converts purified RNA into a sequencing-ready library. The major decisions are whether to enrich for polyadenylated RNA, whether to deplete ribosomal RNA, whether to preserve strand information, and how to fragment the RNA or cDNA.
Poly-A Enrichment
Poly-A enrichment selects for messenger RNA by hybridizing the poly-A tails of mature mRNAs to oligo-dT beads. This approach removes ribosomal RNA, which constitutes the majority of total RNA, and enriches for protein-coding transcripts. Poly-A enrichment is the standard choice for gene expression studies focused on mRNA.
The limitation of poly-A enrichment is that it excludes non-polyadenylated RNAs, including many long non-coding RNAs, circular RNAs, and some viral RNAs. Samples with degraded RNA may have lost their poly-A tails, making poly-A enrichment ineffective.
Ribosomal RNA Depletion
Ribosomal RNA depletion removes rRNA by hybridization to complementary probes, leaving both polyadenylated and non-polyadenylated RNAs in the library. This approach is preferred when the experimental question involves non-coding RNAs, when the sample is degraded, or when the organism of interest has unusual RNA composition.
Ribosomal RNA depletion is more expensive than poly-A enrichment and requires careful optimization of the depletion probes for the target species. Incomplete depletion reduces the proportion of informative reads and increases the sequencing depth required for adequate coverage of the transcriptome.
Comparison of Library Preparation Methods
A comparison of three commercially available RNA-seq library preparation kits found that the traditional TruSeq method detected transcripts and splicing events better than full-length double-stranded cDNA methods using SMARTer and TeloPrime. The expression patterns between TruSeq and SMARTer correlated strongly, while SMARTer and TeloPrime underestimated the expression of relatively long transcripts. Genes with low expression levels were undetected stochastically regardless of the method used.
TeloPrime detected a significantly higher proportion of reads at the transcription start site, but its coverage of the gene body was not uniform. SMARTer was proposed to yield nonspecific genomic DNA amplification. The detected splicing event number was highest with TruSeq, while the percent spliced in index was highly correlated across all three methods.
These findings indicate that the choice of library preparation method affects the biological conclusions that can be drawn from the data. Methods that perform well for gene expression quantification may be suboptimal for splicing analysis, and vice versa.
Strand-Specific Library Preparation
Standard library preparation loses the information about which DNA strand was transcribed. Strand-specific protocols preserve this information by marking the second cDNA strand during synthesis or by ligating adapters in a directional manner. Strand-specific libraries allow the identification of antisense transcription and the accurate quantification of overlapping genes on opposite strands.
A systematic comparison of strand-specific library preparation methods for low input samples found that normalized read counts between treatment groups were in high agreement. The Swift RNA kits offered shorter workflow times enabled by their patented Adaptase technology. The Swift RNA kit produced the fewest number of differentially expressed genes and pathways directly attributable to input mRNA amount.
Input RNA Amount and Miniaturization
Library preparation protocols differ in their input RNA requirements. Standard protocols typically require 100 to 1000 nanograms of total RNA. Low-input protocols can generate libraries from 10 nanograms or less, which is essential for clinical biopsies and sorted cell populations.
A protocol for RNA-seq library preparation from low-volume total RNA by RNA/cDNA hybrid tagmentation, known as SHERRY, profiles polyadenylated RNAs by direct tagging of RNA/DNA hybrids. The protocol describes steps of RNA purification, reverse transcription, hybrid tagmentation, and library generation from 200 nanograms of total RNA.
Automated and miniaturized workflows for RNA library preparation minimize reagent usage and processing time per sample. Reduced-volume libraries show similar behavior to full-scale libraries with comparable numbers of genes detected and reproducible clustering of samples.
Fragment Size and Insert Length
The fragment size of the library determines the read length that can be fully sequenced and the ability to detect structural variants. Longer inserts reduce the overlap of paired reads and preserve information about exon connectivity.
A study of TruSeq Stranded library preparation protocols for acute lymphoblastic leukemia found that poly-A enrichment outperformed ribosomal RNA depletion for the analysis of gene expression and structural aberrations. Ribosomal RNA depletion was more suitable for detection of various classes of RNAs, mutations, or polymorphisms. Reduced RNA fragmentation time, which generates longer inserts, positively affected detection of structural RNA changes without introducing bias into gene expression analysis.
Sequencing Platforms and Run Configuration
The sequencing platform determines the read length, throughput, error profile, and cost per sample. The choice of platform should be guided by the experimental question and the number of samples to be processed.
Short-Read Sequencing
Short-read sequencing platforms generate reads of 50 to 300 base pairs. These platforms offer high throughput and low per-base cost, making them the standard choice for most RNA-seq applications. Paired-end sequencing reads both ends of each fragment, providing information about the distance between the two reads and improving the accuracy of transcript assembly and splice junction detection.
The sequencing depth required depends on the complexity of the transcriptome and the goals of the experiment. Gene expression quantification of abundant transcripts requires fewer reads than the detection of rare splice isoforms or the assembly of full-length transcripts.
Long-Read and Direct RNA Sequencing
Long-read sequencing platforms generate reads of thousands to tens of thousands of base pairs. These platforms can sequence full-length transcripts, resolving isoform structure without computational assembly. Direct RNA sequencing bypasses the cDNA conversion step, allowing the detection of RNA base modifications that are lost during reverse transcription.
A gene-specific RNA enrichment protocol for nanopore direct-RNA sequencing enables targeted analysis of specific transcripts. This approach is valuable when the experimental question focuses on a small number of genes and when RNA modifications are of interest.
Single-Cell RNA Sequencing
Single-cell RNA sequencing (scRNA-seq) profiles the transcriptome of individual cells, revealing heterogeneity that is hidden in bulk analysis. A comparison of six prominent scRNA-seq methods found that Smart-seq2 detected the most genes per cell and across cells, while CEL-seq2, Drop-seq, MARS-seq, and SCRB-seq quantified mRNA levels with less amplification noise due to the use of unique molecular identifiers.
Power simulations at different sequencing depths showed that Drop-seq is more cost-efficient for transcriptome quantification of large numbers of cells, while MARS-seq, SCRB-seq, and Smart-seq2 are more efficient when analyzing fewer cells. The choice among scRNA-seq methods depends on the number of cells to be profiled and the depth of coverage required per cell.
Single-cell RNA sequencing has been applied to breast cancer research to cluster cancer cell populations with different molecular subtypes, identify distinct populations that may correlate with poor prognosis and drug resistance, and explain tumor microenvironment heterogeneity by identifying distinct immune cell subsets. The technology has diverse applications beyond exploring heterogeneity, including the analysis of cell-cell communications, regulatory single-cell states, and immune cell distributions.
Cost Considerations
The cost of RNA-seq includes library preparation reagents, sequencing consumables, instrument depreciation, and bioinformatics analysis. Cost-effective protocols have been developed for large-scale gene expression studies, and tissue multiplexing strategies can reduce costs by 2 to 4 times compared to existing protocols for single-cell RNA sequencing.
A cost-effective protocol for single-cell RNA sequencing of human skin uses a label-free sample multiplexing strategy that enables analysis of paired blood and skin samples. The protocol provides detailed instructions for simultaneous flow cytometry analysis from the same sample, with adaptations for both healthy and inflamed skin specimens. The tissue multiplexing strategy mitigates technical batch effects and reduces costs.
Quality Control Checkpoints
Quality control is not a single step but a continuous process that spans the entire workflow. Each checkpoint has defined acceptance criteria, and samples that fail these criteria should be flagged for review or repeated.
Post-Extraction Quality Control
The extracted RNA should be assessed for yield, purity, and integrity before proceeding to library preparation. The minimum acceptable yield depends on the library preparation protocol. The RIN value should be recorded, and samples with low RIN should be evaluated for whether the experimental question can still be answered with degraded RNA.
Post-Library Quality Control
The completed library should be assessed for concentration, fragment size distribution, and the presence of adapter dimers. Adapter dimers are short fragments that contain only adapter sequences and no insert. They consume sequencing capacity without producing useful data and should be removed by size selection or bead purification.
The library concentration is measured by fluorometric methods that are specific to double-stranded DNA. The fragment size distribution is assessed by microfluidic electrophoresis. The expected size range depends on the library preparation protocol and the intended read length.
Sequencing Run Quality Metrics
The sequencing instrument generates quality metrics for each run, including cluster density, Q30 scores, and read yield. Cluster density that is too low wastes sequencing capacity, while cluster density that is too high causes signal overlap and reduced base calling accuracy. Q30 scores indicate the probability of an incorrect base call, with Q30 corresponding to an error rate of 1 in 1000.
Post-Sequencing Quality Control
After sequencing, the raw reads should be assessed for quality scores, GC content, adapter contamination, and duplication rate. Reads with low quality scores should be trimmed or filtered. Adapter contamination indicates that the library fragments were shorter than the read length. High duplication rates suggest that the library complexity was low, which can result from insufficient input RNA or excessive PCR amplification.
The alignment rate indicates the proportion of reads that map to the reference genome or transcriptome. Low alignment rates may indicate contamination with non-target species, ribosomal RNA carryover, or problems with the reference sequence.
Data Analysis Workflow
The data analysis workflow transforms raw sequencing reads into biological conclusions. The major steps are alignment, quantification, normalization, and differential expression analysis.
Alignment and Quantification
The first step is to align the reads to a reference genome or transcriptome. The choice of alignment tool affects the sensitivity and specificity of read mapping. Spliced aligners are required for RNA-seq data because reads span exon-exon junctions.
After alignment, reads are assigned to genes or transcripts to produce a count matrix. The count matrix tallies the number of reads or fragments within each gene for each sample. This matrix is the starting point for downstream analysis.
Normalization
Normalization adjusts the count data to account for differences in sequencing depth and library composition between samples. The choice of normalization method affects the results of differential expression analysis. A study using Shannon entropy as a benchmark found that TPM, RLE, and TMM normalization, coupled with a threshold of log2 fold change of at least 1 for identifying differentially expressed genes, yielded the best results.
Differential Expression Analysis
Differential expression analysis identifies genes whose expression differs between experimental conditions. Three widely used methods are limma, edgeR, and DESeq2. These methods use different statistical models. Limma uses a linear model, while edgeR and DESeq2 use the negative binomial distribution. Normalized RNA-seq count data is necessary for edgeR and limma but is not necessary for DESeq2.
The results of the three methods are partly overlapping, and each method has its own advantages. The choice of method depends on the data and the experimental design. The detailed protocols for limma, DESeq2, and edgeR are similar but have different steps in the analysis process.
Single-Cell Data Analysis
Single-cell RNA-seq data requires dedicated analysis methods because the technical noise and data complexity differ from bulk RNA-seq. A computational workflow for low-level analyses of scRNA-seq data covers quality control, data exploration, normalization, cell cycle phase assignment, identification of highly variable and correlated genes, clustering into subpopulations, and marker gene detection.
RNA velocity can be directly estimated by distinguishing between unspliced and spliced mRNAs in common single-cell RNA sequencing protocols. RNA velocity is a high-dimensional vector that predicts the future state of individual cells on a timescale of hours. This approach aids the analysis of developmental lineages and cellular dynamics.
Common Failure Patterns and Troubleshooting
RNA-seq experiments fail in characteristic ways, and recognizing the failure pattern is the first step toward correction.
Low RNA Yield
Low RNA yield can result from inefficient cell lysis, RNA degradation during extraction, or loss during purification. The extraction method should be reviewed, and the sample type should be considered. Some tissues, such as bone and cartilage, require specialized extraction protocols. An optimized protocol for meniscus cell extraction for single-cell RNA sequencing addresses the challenges of fibrous connective tissue.
RNA Degradation
RNA degradation produces low RIN values and biased representation of transcripts. The 5-prime ends of long transcripts are lost preferentially, leading to apparent downregulation of long genes. Prevention is the primary strategy: minimize the time between sample collection and stabilization, maintain cold temperatures during processing, and use RNase-free consumables.
Genomic DNA Contamination
Genomic DNA contamination produces reads that map to introns and intergenic regions. The problem is detected by inspecting the alignment metrics and the proportion of reads mapping to exonic regions. DNase treatment should be repeated, and the extraction protocol should be reviewed.
Adapter Dimers
Adapter dimers appear as a sharp peak at approximately 120 to 130 base pairs in the library fragment size distribution. They consume sequencing capacity and reduce the number of informative reads. Adapter dimers are removed by size selection, and the ligation conditions should be optimized to prevent their formation.
PCR Duplicates
PCR duplicates are identical reads that arise from amplification of the same original fragment. High duplication rates reduce the effective sequencing depth and bias expression quantification. The problem is addressed by increasing the input RNA amount, reducing the number of PCR cycles, or using unique molecular identifiers.
Batch Effects
Batch effects are systematic differences between samples processed in different batches. They can obscure biological differences and produce false positives. Batch effects are minimized by randomizing sample processing, using the same reagent lots, and including technical replicates. Computational methods can identify and correct batch effects during data analysis.
Biosafety and Laboratory Practices
RNA-seq involves handling biological samples that may contain infectious agents. The World Health Organization Laboratory Biosafety Manual provides guidance on risk assessment, containment levels, and safe laboratory practices. All work with human samples should follow standard precautions, including the use of gloves, lab coats, and eye protection.
Sample processing should be performed in a biosafety cabinet when the sample is known or suspected to contain infectious agents. Centrifugation steps should use sealed rotors to prevent aerosol generation. Waste disposal should follow institutional guidelines for biohazardous material.
Chemical hazards in the RNA-seq workflow include phenol, chloroform, guanidinium salts, and acrylamide. These reagents should be handled in a fume hood with appropriate personal protective equipment. Material safety data sheets should be reviewed before handling new reagents.
The Laboratory Quality Management System Handbook describes the quality assurance requirements for laboratories, including document control, equipment calibration, internal audits, and proficiency testing. RNA-seq protocols should be written as standard operating procedures, and deviations should be documented and reviewed.
Records and Documentation
Complete documentation is essential for reproducible RNA-seq experiments and for diagnostic applications. The following records should be maintained for each sample:
Sample information including the source, collection date, and storage conditions. Extraction records including the method, reagent lot numbers, and quality metrics. Library preparation records including the protocol version, input amount, and amplification cycles. Sequencing records including the platform, run identifier, and quality metrics. Analysis records including the software versions, parameters, and output files.
The World Health Organization Laboratory Quality Management System Handbook emphasizes that records must be legible, permanent, and retrievable. Electronic records should be backed up and protected from unauthorized modification.
Professional Escalation Criteria
Laboratory personnel should escalate problems to a supervisor or principal investigator when the issue exceeds their training or when the problem persists despite troubleshooting. Specific escalation criteria include:
Repeated failure of quality control checkpoints that cannot be resolved by standard troubleshooting. Unexpected results that suggest contamination, reagent failure, or instrument malfunction. Discrepancies between replicate samples that exceed acceptable variation. Results that will be used for clinical decisions and require confirmation by an independent method. Safety incidents involving biological or chemical hazards.
The escalation process should be documented, and the outcome should be recorded in the laboratory records.
Frequently Asked Questions
What is the minimum RNA quality required for RNA sequencing?
The minimum RNA quality depends on the library preparation method and the experimental question. High-quality RNA with a RIN above 8 is preferred for most applications. Degraded RNA can still be used for some experiments, particularly those using ribosomal RNA depletion instead of poly-A enrichment, but the results will have reduced sensitivity for low-abundance transcripts and may show biased coverage of long genes.
How much RNA is needed for library preparation?
Standard library preparation protocols require 100 to 1000 nanograms of total RNA. Low-input protocols can generate libraries from 10 nanograms or less. The SHERRY protocol generates libraries from 200 nanograms of total RNA. The minimum input amount is determined by the protocol and should be verified with the manufacturer instructions.
What is the difference between poly-A enrichment and ribosomal RNA depletion?
Poly-A enrichment selects for messenger RNA by hybridizing the poly-A tails of mature mRNAs to oligo-dT beads. Ribosomal RNA depletion removes rRNA by hybridization to complementary probes, leaving both polyadenylated and non-polyadenylated RNAs. Poly-A enrichment is standard for gene expression studies, while ribosomal RNA depletion is preferred for non-coding RNAs and degraded samples.
How do I choose between single-end and paired-end sequencing?
Paired-end sequencing reads both ends of each fragment, providing information about the distance between the two reads. Paired-end reads improve the accuracy of transcript assembly and splice junction detection. Single-end sequencing is less expensive and may be sufficient for gene expression quantification of known transcripts. Paired-end sequencing is recommended for novel transcript discovery and isoform analysis.
What sequencing depth is needed for RNA-seq?
The sequencing depth depends on the transcriptome complexity and the experimental goals. Gene expression quantification of abundant transcripts requires fewer reads than the detection of rare splice isoforms. Power simulations for single-cell RNA-seq methods showed that Drop-seq is more cost-efficient for transcriptome quantification of large numbers of cells, while other methods are more efficient when analyzing fewer cells.
How do I know if my library preparation worked correctly?
The library should be assessed for concentration, fragment size distribution, and the presence of adapter dimers. The expected fragment size range depends on the protocol and the intended read length. Adapter dimers appear as a sharp peak at approximately 120 to 130 base pairs and should be removed by size selection.
What causes high duplication rates in RNA-seq data?
High duplication rates result from low library complexity, which can be caused by insufficient input RNA, excessive PCR amplification, or degradation of the starting material. The problem is addressed by increasing the input RNA amount, reducing the number of PCR cycles, or using unique molecular identifiers to distinguish biological duplicates from PCR duplicates.
How do I choose between limma, edgeR, and DESeq2 for differential expression analysis?
The three methods use different statistical models. Limma uses a linear model, while edgeR and DESeq2 use the negative binomial distribution. Normalized RNA-seq count data is necessary for edgeR and limma but is not necessary for DESeq2. The results of the three methods are partly overlapping, and the choice of method depends on the data and the experimental design.
Related Diagnostic Guides
- DNA Shearing for NGS Library Preparation: Methods and Quality Control
- RNA Extraction Using TRIzol Reagent: Protocol, Troubleshooting, and Best Practices
- Northern Blotting Protocol: Step-by-Step Guide for RNA Detection
- RNA Extraction from Plant Tissues: Methods and Troubleshooting
- How to Perform a Gram Stain: Protocol and Quality Control
References and Further Reading
- Laboratory Quality Management System Handbook. World Health Organization.
- Laboratory Biosafety Manual. World Health Organization.
- Assay Guidance Manual. National Center for Advancing Translational Sciences.
- Bioanalytical Method Validation Guidance. U.S. Food and Drug Administration.
- NCBI Literature Resources. National Center for Biotechnology Information.
- Comparative Analysis of Single-Cell RNA Sequencing Methods.. Molecular cell, 2017.
- Three Differential Expression Analysis Methods for RNA Sequencing: limma, EdgeR, DESeq2.. Journal of visualized experiments : JoVE, 2021.
- RNA velocity of single cells.. Nature, 2018.
- Single-cell RNA sequencing in breast cancer: Understanding tumor heterogeneity and paving roads to individualized therapy.. Cancer communications (London, England), 2020.
- Integration of single-cell and bulk RNA sequencing to identify a distinct tumor stem cells and construct a novel prognostic signature for evaluating prognosis and immunotherapy in LUAD.. Journal of translational medicine, 2025.
- RNA sequencing.. Methods in molecular biology (Clifton, N.J.), 2011.
- Ten years of next-generation sequencing technology.. Trends in genetics : TIG, 2014.
- A cost-effective protocol for single-cell RNA sequencing of human skin.. Frontiers in immunology, 2024.
- Protocol for RNA-seq library preparation from low-volume total RNA by RNA/cDNA hybrid tagmentation.. 2025.
- A comparison of mRNA sequencing (RNA-Seq) library preparation methods for transcriptome analysis.. 2022.
- A Comparison of mRNA Sequencing (RNA-Seq) Library Preparation Methods for Transcriptome Analysis. 2021.
- High-throughput Minitaturized RNA-Seq Library Preparation.. 2020.
- Systematic comparative analysis of strand-specific RNA-seq library preparation methods for low input samples.. 2022.
- RNA-seq library preparation for comprehensive transcriptome analysis in cancer cells: The impact of insert size.. 2021.
- RNA-Seq workflow: gene-level exploratory analysis and differential expression. F1000Research, 2015.
- Assessing RNA-Seq Workflow Methodologies Using Shannon Entropy. Biology, 2024.
- Analysis of a Single Cell RNA-seq Workflow by Random Matrix Theory Methods. Bulletin of Mathematical Biology, 2024.
- Enhanced single-cell RNA-seq workflow reveals coronary artery disease cellular cross-talk and candidate drug targets. Atherosclerosis, 2021.
- A step-by-step workflow for low-level analysis of single-cell RNA-seq data with Bioconductor. F1000Research, 2016.
- Plant Single-Cell/Nucleus RNA-seq Workflow.. Methods in molecular biology, 2023.
- Characterization of novel small non-coding RNAs and their modifications in bladder cancer using an updated small RNA-seq workflow. Frontiers in Molecular Biosciences, 2022.
- A gene-specific RNA enrichment protocol for nanopore direct-RNA sequencing. Plos One, 2026.
- An optimized protocol of meniscus cell extraction for single-cell RNA sequencing. Nan Fang Yi Ke Da Xue Xue Bao Journal of Southern Medical University, 2021.
- Universal Library Preparation Protocol for Efficient High-Throughput Sequencing of Double-Stranded RNA Viruses. Methods in Molecular Biology, 2020.
- A cost-effective RNA sequencing protocol for large-scale gene expression studies. Scientific Reports, 2015.
- Deep-sequencing protocols influence the results obtained in small-RNA sequencing. Plos One, 2012.
- RNA Sequencing Protocols for Short-Read Sequencing. Methods in Molecular Biology, 2025.
This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.