Sequence RNA: Methods, Steps, and Applications

By Dr. Zubair Khalid, DVM, MS, PhD ·

Sequence RNA: Methods, Steps, and Applications

Introduction to RNA Sequencing

What is RNA Sequencing?

RNA sequencing (RNA-seq) is a high-throughput technique used to determine the sequence and quantity of RNA molecules in a biological sample at a given moment. Unlike genome sequencing, which reveals the static DNA blueprint of an organism, RNA-seq captures the dynamic snapshot of gene expression—which genes are active, at what levels, and in what alternatively spliced forms. The complete set of RNA transcripts in a cell or tissue is called the transcriptome, and RNA-seq is the primary tool for characterizing it.

The core workflow involves converting RNA into complementary DNA (cDNA), attaching sequencing adapters, and then determining the nucleotide order of millions of cDNA fragments in parallel. The resulting data are aligned to a reference genome or assembled de novo to identify which genes are expressed and quantify their expression levels. Because RNA-seq is not limited to known sequences, it can discover novel transcripts, splice variants, and non-coding RNAs that other methods would miss.

Why Sequence RNA Instead of DNA?

DNA sequencing tells you what is possible; RNA sequencing tells you what is actually happening. The central dogma of molecular biology—DNA to RNA to protein—means that RNA is the intermediate that reflects gene expression. Several reasons make RNA the preferred target for expression studies:

  1. Dynamic range: Gene expression varies across tissues, developmental stages, and conditions. RNA levels can differ by orders of magnitude between genes, and RNA-seq can capture this range quantitatively.
  2. Isoform information: Alternative splicing produces multiple mRNA isoforms from a single gene. RNA-seq reads spanning exon-exon junctions reveal which isoforms are present, something DNA sequencing cannot do.
  3. Post-transcriptional regulation: MicroRNAs, RNA-binding proteins, and RNA degradation pathways all affect mRNA stability. RNA-seq captures the steady-state level of RNA after these regulatory processes have acted.
  4. Non-coding RNAs: Many functional RNAs—such as microRNAs, long non-coding RNAs, and circular RNAs—are transcribed from DNA but never translated into protein. Only RNA sequencing can detect them.

DNA sequencing of genomic DNA cannot reveal which promoters are active, which exons are included in mature mRNA, or how RNA abundance changes in response to a stimulus. For these questions, RNA-seq is the method of choice.

RNA Extraction and Quality Control

RNA Isolation Techniques

The success of any RNA-seq experiment depends on isolating RNA that is pure, intact, and free of contaminants. RNA is chemically less stable than DNA—the 2′ hydroxyl group on the ribose sugar makes it susceptible to alkaline hydrolysis and enzymatic degradation by ribonucleases (RNases), which are ubiquitous in the environment and on human skin.

Two main approaches dominate RNA isolation:

TRIzol (acid guanidinium thiocyanate-phenol-chloroform) extraction is a classic liquid-phase method. Cells are lysed in a monophasic solution containing guanidinium thiocyanate, which denatures proteins including RNases, and phenol, which partitions cellular components. After adding chloroform and centrifuging, the mixture separates into three phases: an aqueous upper phase containing RNA, an interphase containing DNA, and an organic lower phase containing proteins and lipids. The RNA is precipitated from the aqueous phase with isopropanol, washed with 75% ethanol to remove residual salts and guanidinium, and resuspended in RNase-free water or Tris-EDTA buffer. This method yields total RNA including small RNAs (<200 nucleotides) and is inexpensive, but it requires careful technique to avoid contamination with DNA or phenol.

Column-based kits (e.g., from Qiagen, Zymo, or Thermo Fisher) use silica membrane spin columns. Samples are lysed in a chaotropic buffer (typically containing guanidinium thiocyanate) that denatures RNases and promotes RNA binding to the silica membrane. After loading the lysate onto the column, RNA binds to the silica; contaminants are washed away with ethanol-containing buffers; and pure RNA is eluted in water or low-salt buffer. These kits are faster, more reproducible, and safer than TRIzol, but they may not retain small RNAs unless specifically designed for that purpose.

For mRNA-focused experiments, many protocols include a poly(A) enrichment step. Since most eukaryotic mRNAs have a 3′ polyadenylated tail, they can be captured using oligo(dT) beads—magnetic beads coated with thymidine oligonucleotides that hybridize to the poly(A) tail. This removes ribosomal RNA (rRNA), which constitutes 80–90% of total RNA, and enriches for mRNA. Alternatively, rRNA depletion uses probes complementary to rRNA sequences to remove them from total RNA, which is preferable when studying non-polyadenylated RNAs such as bacterial transcripts or certain long non-coding RNAs.

Assessing RNA Quality: RIN and Gel Electrophoresis

RNA quality is paramount. Degraded RNA produces biased results—short fragments align poorly, and expression of long genes is underestimated. Two standard quality checks are used:

Gel electrophoresis on a denaturing agarose gel (containing formaldehyde to prevent secondary structure) can visualize ribosomal RNA bands. In intact eukaryotic RNA, the 28S rRNA band should be approximately twice as intense as the 18S rRNA band. A smeared appearance or loss of the 28S band indicates degradation. This method is simple but qualitative and requires relatively large amounts of RNA.

The RNA Integrity Number (RIN) is a more quantitative measure generated by microfluidic electrophoresis instruments such as the Agilent Bioanalyzer or the TapeStation. These instruments separate RNA by size in a microchannel and produce an electropherogram. The RIN algorithm assigns a score from 1 (completely degraded) to 10 (fully intact) based on the shape of the electropherogram, the ratio of 28S to 18S peaks, and the presence of degradation products. A RIN of 7 or higher is generally acceptable for standard RNA-seq; lower-quality samples may still be usable with specialized protocols but will yield compromised data.

Additional quality metrics include A260/A280 absorbance ratios (measured by spectrophotometry), which should be between 1.8 and 2.1 for pure RNA. A lower ratio suggests protein contamination; a higher ratio may indicate residual chaotropic salts. The A260/A230 ratio (typically >1.5) detects contamination by guanidinium, phenol, or carbohydrates.

Library Preparation for RNA Sequencing

Reverse Transcription to cDNA

RNA cannot be sequenced directly on most platforms because sequencing enzymes (DNA polymerases) use DNA as a template. Therefore, RNA must first be converted to complementary DNA (cDNA) through reverse transcription, catalyzed by the enzyme reverse transcriptase.

The reaction uses a primer to initiate synthesis. For poly(A)-enriched RNA, an oligo(dT) primer (a string of 15–30 thymidines) anneals to the poly(A) tail and primes first-strand synthesis. For total RNA or rRNA-depleted samples, random hexamer primers (six-nucleotide random sequences) anneal throughout the transcript, providing more uniform coverage. The reverse transcriptase extends the primer, synthesizing a cDNA strand complementary to the RNA template.

Key reaction components include:

  • Reverse transcriptase: Commonly used enzymes include Moloney Murine Leukemia Virus (M-MLV) reverse transcriptase and its engineered derivatives (e.g., SuperScript II/III/IV), which have high processivity and reduced RNase H activity.
  • dNTPs: Deoxynucleotide triphosphates at concentrations of 0.5–1 mM each.
  • RNase inhibitor: A protein that blocks RNase activity, protecting the RNA template.
  • Reaction buffer: Typically Tris-HCl (pH 8.3), KCl, and MgCl₂ at concentrations optimized for the specific enzyme.

The reaction is incubated at 42–55°C for 30–60 minutes. After first-strand synthesis, the RNA template is removed by treatment with RNase H (which degrades RNA in RNA-DNA hybrids), and second-strand synthesis is performed using DNA polymerase I. The result is double-stranded cDNA that is stable and compatible with downstream steps.

Fragmentation and Size Selection

For short-read sequencing platforms, cDNA must be fragmented into pieces of 200–500 base pairs. Fragmentation can be performed before reverse transcription (on RNA) or after cDNA synthesis. Common methods include:

  • Enzymatic fragmentation: Endonucleases such as Fragmentase or NEBNext dsDNA Fragmentase cleave cDNA at random sites. This method is gentle and produces uniform fragment sizes.
  • Mechanical fragmentation: Sonication uses high-frequency sound waves to shear cDNA. This is effective but can be less reproducible and may require more input material.
  • Chemical fragmentation: For RNA, incubation with divalent cations (Mg²⁺ or Zn²⁺) at elevated temperatures (94°C for 5–10 minutes) hydrolyzes RNA at random positions.

After fragmentation, size selection removes fragments that are too short or too long. This is typically done using AMPure XP beads—paramagnetic beads that bind DNA in the presence of polyethylene glycol (PEG) and salt. By adjusting the bead-to-sample ratio, fragments within a desired size range are retained while others are washed away. Alternatively, gel extraction or automated size-selection instruments (e.g., BluePippin) can be used for tighter size distributions.

Adapter Ligation and PCR Amplification

Sequencing adapters—short, double-stranded oligonucleotides with known sequences—must be ligated to both ends of the cDNA fragments. These adapters serve several purposes:

  1. They provide binding sites for the sequencing primers.
  2. They contain sequences that anchor the fragment to the flow cell (for Illumina) or the sequencing chip (for Ion Torrent).
  3. They include index sequences (barcodes), which are unique 6–10 nucleotide tags that allow multiple samples to be pooled and sequenced together in a single run (multiplexing).

The ligation reaction uses T4 DNA ligase, which catalyzes the formation of phosphodiester bonds between the 5′ phosphate of the adapter and the 3′ hydroxyl of the cDNA. The reaction is incubated at 20°C for 15–30 minutes. To prevent adapter dimers (adapters ligated to each other without insert), the cDNA ends are often end-repaired (blunt-ended) and A-tailed (a single adenine added to the 3′ ends) before ligation. The adapters have complementary thymidine overhangs, ensuring efficient ligation.

Finally, PCR amplification enriches the library and adds sufficient material for sequencing. The PCR uses primers complementary to the adapter sequences and typically runs for 8–15 cycles. Fewer cycles are preferred to minimize PCR duplicates—identical reads arising from amplification of the same original fragment—which can skew quantification. High-fidelity DNA polymerases (e.g., Phusion or Q5) are used to minimize polymerase errors. The final library is purified, quantified by qPCR or fluorometry, and quality-checked on a Bioanalyzer to confirm the size distribution and absence of adapter dimers.

RNA Sequencing Platforms and Methods

Short-Read Sequencing (Illumina)

Illumina sequencing-by-synthesis is the most widely used RNA-seq platform. The workflow is as follows:

  1. Cluster generation: The library is denatured to single strands and loaded onto a flow cell coated with oligonucleotides complementary to the adapter sequences. Each fragment is amplified in place by bridge amplification, creating a cluster of ~1,000 identical copies.
  2. Sequencing-by-synthesis: Fluorescently labeled reversible terminator nucleotides are added one at a time. After each nucleotide incorporation, the flow cell is imaged to record which base was added to each cluster. The fluorescent label and terminator are then cleaved, allowing the next nucleotide to be added.
  3. Read length and throughput: Typical read lengths are 50–150 base pairs, paired-end (both ends of each fragment are sequenced). A single NovaSeq run can produce 1–6 billion reads, sufficient for multiple samples.

Advantages of Illumina include high accuracy (>99.9% per base), high throughput, and low cost per read. Limitations include short read lengths, which make it difficult to resolve repetitive regions, assemble full-length transcripts, or distinguish highly similar isoforms. PCR amplification during library preparation can also introduce bias.

Long-Read Sequencing (PacBio, Nanopore)

Long-read platforms can sequence single molecules of 10–100 kilobases or more, providing full-length transcript information.

Pacific Biosciences (PacBio) SMRT sequencing uses zero-mode waveguides to observe DNA polymerase incorporating fluorescently labeled nucleotides in real time. The circular consensus sequencing (CCS) mode reads the same molecule multiple times, achieving high accuracy (Q20 or better) for shorter inserts. Iso-Seq is the PacBio method for full-length transcript sequencing: RNA is converted to cDNA, and the entire transcript is sequenced in one read, revealing complete isoform structures.

Oxford Nanopore sequencing passes single-stranded DNA or RNA through a protein nanopore embedded in a membrane. As nucleotides pass through the pore, they cause characteristic disruptions in an electrical current, which are decoded into a sequence. Nanopore can sequence RNA directly (without reverse transcription), avoiding cDNA synthesis artifacts. It produces very long reads (up to megabases) but has lower per-base accuracy (typically 90–99%) compared to Illumina.

Advantages of long-read sequencing include the ability to resolve complex isoforms, detect base modifications (e.g., m6A), and assemble transcriptomes de novo. Limitations include higher cost per base, lower throughput, and higher error rates, though these are improving rapidly.

Single-Cell RNA Sequencing

Bulk RNA-seq measures the average expression across millions of cells, masking cellular heterogeneity. Single-cell RNA sequencing (scRNA-seq) profiles the transcriptome of individual cells, revealing distinct cell types, states, and trajectories.

The most common method is the 10x Genomics Chromium platform, which encapsulates individual cells in nanoliter-scale droplets (GEMs) with barcoded beads. Each bead carries a unique cell barcode (a 16-nucleotide sequence) and a poly(dT) primer with a unique molecular identifier (UMI)—a 10-nucleotide random sequence that tags each mRNA molecule. After cell lysis within the droplet, reverse transcription occurs, producing cDNA labeled with both the cell barcode and the UMI. This allows every read to be traced back to its cell of origin and to its original mRNA molecule, eliminating PCR duplicates computationally.

Other scRNA-seq methods include SMART-seq (full-length, plate-based, higher sensitivity but lower throughput) and MARS-seq. scRNA-seq requires specialized bioinformatics to cluster cells by expression profile, identify marker genes, and reconstruct developmental trajectories. The technology has revolutionized immunology, neurobiology, and cancer research by enabling the identification of rare cell populations.

Bioinformatics Analysis of RNA-Seq Data

Quality Control and Trimming

Raw sequencing data arrive as FASTQ files containing nucleotide sequences and per-base quality scores (Phred scores, where Q30 corresponds to 99.9% accuracy). The first step is quality assessment using tools like FastQC, which reports per-base quality, GC content, adapter contamination, and overrepresented sequences.

Trimming removes low-quality bases and adapter sequences. Common tools include Trimmomatic, cutadapt, and fastp. Typical parameters:

  • Remove bases with Phred quality <20 (i.e., <1% error rate).
  • Remove adapter sequences (e.g., Illumina TruSeq adapters) when they appear at the 3′ end of reads.
  • Discard reads shorter than 36 nucleotides after trimming.

For paired-end reads, trimming must maintain the pairing between forward and reverse reads. After trimming, FastQC should be rerun to confirm improvement.

Read Alignment to Reference Genome

The trimmed reads are aligned to a reference genome or transcriptome. The choice of aligner depends on the application:

  • Splice-aware aligners (STAR, HISAT2) are the standard for RNA-seq because they can align reads spanning exon-exon junctions. STAR first builds a genome index, then performs a two-step search: it finds the longest exact match for each read, then extends alignments across splice junctions. HISAT2 uses a hierarchical indexing strategy for faster performance.
  • Transcriptome aligners (Salmon, Kallisto) use a different approach called pseudoalignment. Instead of aligning reads to the genome, they map reads to a reference transcriptome and quantify abundance directly. These tools are extremely fast and memory-efficient, making them popular for large datasets.

Alignment statistics to check include the percentage of reads mapped (typically >80% for good samples), the percentage uniquely mapped, and the distribution of reads across gene features (exons, introns, intergenic regions).

Quantification and Differential Expression

Quantification counts the number of reads mapping to each gene or transcript. This can be done at the gene level (summing all reads across exons) or the transcript level (assigning reads to specific isoforms, which is more challenging). Tools like featureCounts and HTSeq-count generate gene-level count matrices.

Raw counts are not directly comparable between samples because of differences in sequencing depth and library composition. Normalization methods include:

  • CPM/TPM: Counts per million or transcripts per million, which divide by total reads. TPM additionally normalizes for transcript length.
  • DESeq2's median-of-ratios: Estimates size factors to account for library size and composition bias.
  • edgeR's TMM (trimmed mean of M-values): Computes normalization factors based on the weighted mean of log ratios between samples.

Differential expression analysis identifies genes whose expression changes significantly between conditions (e.g., treated vs. untreated). DESeq2 and edgeR use negative binomial models to account for the overdispersion inherent in RNA-seq count data. Both tools output:

  • Log2 fold change: The magnitude of expression change.
  • Adjusted p-value: The significance after multiple testing correction (e.g., Benjamini-Hochberg FDR). A common threshold is adjusted p < 0.05 and |log2 fold change| > 1.

The results are typically visualized as volcano plots (showing fold change vs. significance) and heatmaps (showing expression of differentially expressed genes across samples). Downstream analyses include gene ontology (GO) enrichment to identify biological processes overrepresented among differentially expressed genes, and pathway analysis using databases like KEGG or Reactome.

Applications of RNA Sequencing

Gene Expression Profiling

The most common application of RNA-seq is comparing gene expression between conditions. Examples include:

  • Identifying genes upregulated in cancer cells compared to normal tissue, revealing potential oncogenes or tumor suppressors.
  • Profiling the response of cells to drug treatment, identifying pathways activated or repressed.
  • Characterizing developmental changes—for instance, comparing embryonic stem cells to differentiated neurons to identify key regulatory genes.
  • Studying host-pathogen interactions by simultaneously sequencing host and pathogen transcripts.

RNA-seq has largely replaced microarrays for these applications because it has a wider dynamic range, detects novel transcripts, and does not require prior knowledge of gene sequences.

Discovery of Novel Transcripts and Splice Variants

Because RNA-seq does not rely on known probes, it can discover previously unannotated transcripts. Examples include:

  • Novel isoforms: Alternative splicing produces multiple mRNA variants from a single gene. RNA-seq reads spanning exon-exon junctions can identify splice variants that differ from reference annotations. For example, the TP53 gene (encoding p53) has multiple isoforms with distinct functions in apoptosis and cell cycle regulation.
  • Long non-coding RNAs (lncRNAs): Thousands of lncRNAs have been discovered by RNA-seq, many of which regulate gene expression through chromatin remodeling, transcriptional interference, or acting as molecular sponges for microRNAs.
  • Fusion genes: Chromosomal rearrangements can create chimeric transcripts, such as the BCR-ABL1 fusion in chronic myeloid leukemia. RNA-seq can detect these fusions by identifying reads that align to two different genes.
  • RNA editing: A-to-I editing by ADAR enzymes changes the sequence of mature mRNA. RNA-seq can detect these editing events by comparing RNA sequences to the genomic DNA sequence.

Clinical Applications

RNA-seq is increasingly used in clinical settings:

  • Cancer diagnostics: RNA-seq can identify fusion genes, splice variants, and expression signatures that guide treatment decisions. For example, detecting ALK fusions in lung adenocarcinoma identifies patients who may benefit from ALK inhibitors like crizotinib.
  • Infectious disease: RNA-seq can identify the causative pathogen in unexplained infections by sequencing total RNA from patient samples and detecting pathogen transcripts.
  • Genetic disorders: RNA-seq can complement DNA sequencing by revealing splicing defects or allele-specific expression that explain disease phenotypes. In rare diseases, RNA-seq of patient tissues can identify aberrant splicing caused by deep intronic variants that are missed by exome sequencing.
  • Pharmacogenomics: Expression profiling can predict drug response. For example, low expression of DPYD (encoding dihydropyrimidine dehydrogenase) predicts severe toxicity to the chemotherapy drug 5-fluorouracil.

The CRISPR Sequence Example for Gene Therapy demonstrates how sequence-level understanding of gene regulation can be applied therapeutically; RNA-seq provides the transcriptome-wide view that complements such targeted approaches.

Common Pitfalls and Troubleshooting in RNA Sequencing

RNA Degradation and Contamination

Pitfall: RNA degrades rapidly after tissue collection. RNases are everywhere—on skin, in dust, and in laboratory glassware. Even brief exposure can fragment RNA, biasing results toward the 5′ ends of transcripts.

Solutions:

  • Work quickly and keep samples on ice.
  • Use RNase-free consumables and treat surfaces with RNase decontamination solutions (e.g., RNaseZap).
  • Snap-freeze tissues in liquid nitrogen immediately after collection and store at −80°C.
  • Add RNase inhibitors to lysis buffers.
  • For clinical samples, use stabilization reagents like RNAlater, which permeabilize tissues and precipitate RNases.

Pitfall: Genomic DNA contamination in RNA preparations. DNA can be amplified during library preparation, producing reads that align to introns or intergenic regions and skewing quantification.

Solutions:

  • Treat RNA with DNase I (RNase-free) after extraction.
  • Verify the absence of DNA by PCR using primers spanning an intron—no product should be amplified from RNA.
  • Check alignment statistics: a high percentage of reads in intronic regions may indicate DNA contamination.

Library Preparation Artifacts

Pitfall: Adapter dimers—adapters ligated to each other without an insert—are a common problem. They waste sequencing capacity and produce low-quality reads.

Solutions:

  • Optimize the adapter-to-insert ratio. Too much adapter favors dimer formation.
  • Include a size-selection step after ligation to remove small fragments.
  • Check the library on a Bioanalyzer; adapter dimers appear as a peak around 120–130 bp.

Pitfall: PCR duplicates—multiple reads from the same original fragment—inflate apparent expression levels. This is especially problematic when starting with low RNA input.

Solutions:

  • Minimize PCR cycles (use 8–12 cycles, not 20).
  • Use UMIs to identify and remove duplicates computationally.
  • Start with sufficient RNA input to reduce the need for extensive amplification.

Pitfall: Index hopping or barcode swapping—reads assigned to the wrong sample due to adapter misidentification during multiplexed sequencing.

Solutions:

  • Use unique dual indexes (different i5 and i7 barcodes for each sample) rather than single indexes.
  • Include negative controls (water or empty wells) to detect contamination.

Bioinformatics Mistakes

Pitfall: Using an incorrect or outdated reference genome. This leads to poor alignment and missed transcripts.

Solutions:

  • Download the latest reference genome and annotation from Ensembl or UCSC.
  • Build the aligner index with the same genome version used for annotation.
  • Document the genome version in your methods.

Pitfall: Ignoring batch effects. Samples processed on different days, by different operators, or on different sequencing runs can have systematic differences unrelated to biology.

Solutions:

  • Randomize sample processing across batches.
  • Include technical replicates.
  • Use batch correction tools like ComBat-seq or include batch as a covariate in DESeq2.

Pitfall: Overinterpreting differential expression results without validation.

Solutions:

  • Validate key findings by RT-qPCR or Western blotting.
  • Check that differentially expressed genes are biologically plausible.
  • Use independent replication to confirm results.

Summary and Best Practices

RNA sequencing is a powerful, versatile technology for studying gene expression and transcriptome structure. The workflow spans experimental and computational domains, each with its own critical steps:

  1. Experimental design: Define biological replicates (at least 3 per condition), choose the appropriate sequencing depth (typically 20–50 million reads per sample for differential expression), and decide between total RNA and poly(A)-enriched RNA.
  2. RNA extraction: Isolate high-quality RNA using appropriate methods, verify integrity (RIN > 7), and eliminate DNA contamination.
  3. Library preparation: Convert RNA to cDNA, fragment, ligate adapters, and amplify with minimal cycles. Include appropriate controls and consider UMIs for quantitative accuracy.
  4. Sequencing: Choose the platform that matches your question—Illumina for high-throughput quantification, long-read platforms for isoform resolution, and scRNA-seq for cellular heterogeneity.
  5. Bioinformatics: Perform quality control, trim adapters, align to a reference, quantify expression, and apply appropriate statistical models for differential expression.
  6. Validation: Confirm key findings with independent methods.

Best practices include documenting every step, using consistent protocols across samples, and depositing raw data in public repositories (GEO, SRA) for reproducibility. RNA-seq is a mature technology, but its success depends on careful attention to quality at every stage.

Frequently Asked Questions

How to sequence RNA?

RNA sequencing involves five main steps: (1) isolate RNA from cells or tissues, (2) convert RNA to cDNA by reverse transcription, (3) fragment the cDNA and ligate sequencing adapters, (4) amplify and sequence the library on a high-throughput platform (typically Illumina), and (5) analyze the resulting data bioinformatically—trimming reads, aligning to a reference genome, and quantifying gene expression. The specific protocols vary by platform and application, but this general workflow applies to all RNA-seq experiments.

What is the difference between RNA-seq and microarray?

Microarrays measure gene expression by hybridizing fluorescently labeled cDNA to known probe sequences attached to a solid surface. RNA-seq sequences all RNA molecules directly. Key differences: RNA-seq has a wider dynamic range (detecting both very low and very high expression), can discover novel transcripts and isoforms, does not require prior genome annotation, and can distinguish closely related sequences. Microarrays are cheaper per sample and have simpler data analysis, but they are limited to known genes and suffer from background hybridization and signal saturation.

Why is RNA sequencing important?

RNA sequencing is important because it provides a comprehensive, quantitative view of the transcriptome—the complete set of RNA molecules in a cell. This reveals which genes are active, at what levels, and in what alternatively spliced forms. RNA-seq has transformed our understanding of gene regulation, development, disease mechanisms, and cellular heterogeneity. It is essential for identifying biomarkers, understanding drug responses, discovering novel transcripts, and characterizing complex biological systems.

What are the steps of RNA sequencing?

The steps of RNA sequencing are: (1) RNA extraction and quality assessment, (2) removal of ribosomal RNA or enrichment of mRNA, (3) fragmentation of RNA or cDNA, (4) reverse transcription to generate cDNA, (5) adapter ligation and PCR amplification to create a sequencing library, (6) sequencing on a high-throughput platform, and (7) bioinformatics analysis including quality control, alignment, quantification, and differential expression testing.

How much RNA is needed for RNA-seq?

Standard RNA-seq library preparation requires 100 ng to 1 μg of total RNA. Low-input protocols can work with as little as 1–10 ng of RNA, and single-cell RNA-seq works with approximately 10 pg of RNA per cell. The required amount depends on the library preparation kit and the sequencing platform. For degraded or low-quality RNA, more input may be needed to compensate for losses during preparation.

What is the difference between RNA-seq and RT-PCR?

RT-PCR (reverse transcription polymerase chain reaction) measures the expression of a small number of specific genes (typically 1–100) using gene-specific primers and fluorescent probes. RNA-seq measures the expression of all genes simultaneously without prior knowledge of their sequences. RT-PCR is cheaper, faster, and more sensitive for targeted measurements, making it ideal for validating RNA-seq results. RNA-seq provides genome-wide coverage and can discover novel transcripts, but it is more expensive and requires substantial bioinformatics expertise.

What is single-cell RNA sequencing?

Single-cell RNA sequencing (scRNA-seq) profiles the transcriptome of individual cells rather than a bulk population. Cells are isolated and encapsulated in droplets or wells, each receiving a unique barcode. After cell lysis, reverse transcription tags each cell's cDNA with its specific barcode, allowing the transcriptomes of thousands to millions of individual cells to be sequenced together. Analysis then clusters cells by similar expression profiles, identifying distinct cell types, rare populations, and developmental trajectories. scRNA-seq has revolutionized the study of complex tissues like the brain and tumors by revealing cellular heterogeneity that bulk RNA-seq averages away.

Key Takeaways

  • RNA sequencing (RNA-seq) quantifies the complete transcriptome, revealing gene expression levels, splice variants, and non-coding RNAs in a single experiment.
  • The workflow involves RNA extraction, quality assessment, cDNA synthesis, library preparation, high-throughput sequencing, and bioinformatics analysis.
  • RNA integrity (measured by RIN) is critical; degraded RNA produces biased, unreliable results.
  • Illumina short-read sequencing is the standard for quantification, while long-read platforms (PacBio, Nanopore) provide full-length isoform information.
  • Single-cell RNA-seq reveals cellular heterogeneity by profiling individual cells with barcoded beads and UMIs.
  • Differential expression analysis uses negative binomial models (DESeq2, edgeR) to identify genes with statistically significant expression changes between conditions.
  • RNA-seq has broad applications in basic research, cancer diagnostics, infectious disease, and rare genetic disorders, but requires careful experimental design and rigorous quality control at every step.

Further Reading

  • Rivas E. Evolutionary conservation of RNA sequence and structure. Wiley interdisciplinary reviews. RNA. 2021. PubMed 33754485
  • Banfalvi G. Origin of Coding RNA from Random-Sequence RNA. DNA and cell biology. 2019. PubMed 30638405
  • Xu L, Seki M. Recent advances in the detection of base modifications using the Nanopore sequencer. Journal of human genetics. 2020. PubMed 31602005
  • Henley RY, Carson S, Wanunu M. Studies of RNA Sequence and Structure Using Nanopores. Progress in molecular biology and translational science. 2016. PubMed 26970191
  • Hamada M et al. Predictions of RNA secondary structure by combining homologous sequence information. Bioinformatics (Oxford, England). 2009. PubMed 19478007
  • Zirbel CL et al. Identifying novel sequence variants of RNA 3D motifs. Nucleic acids research. 2015. PubMed 26130723

Related Topics

Related Clinical & Scientific Guides