Gene Splicing: Mechanisms, Types, and Biological Significance
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Gene Splicing
What is Gene Splicing?
Gene splicing, in its molecular biological sense, is the post-transcriptional process by which introns (non-coding intervening sequences) are removed from a precursor messenger RNA (pre-mRNA) and exons (coding sequences) are joined together to form a mature mRNA molecule. This process is a mandatory step in the expression of most eukaryotic protein-coding genes. The term "gene splicing" is often used interchangeably with "RNA splicing," and the two refer to the same biochemical pathway. The mature mRNA produced by splicing is then exported to the cytoplasm, where it serves as the template for translation by ribosomes.
The discovery of splicing in 1977 by Phillip Sharp and Richard Roberts, working independently on adenovirus, overturned the then-prevailing "one gene, one polypeptide" model. They demonstrated that genes are not contiguous stretches of DNA but are interrupted by non-coding segments that are removed at the RNA level. This fundamental insight reshaped our understanding of genome architecture and gene regulation. Today, we know that in humans, over 95% of multi-exon genes undergo alternative splicing, meaning a single gene can produce multiple distinct mRNA transcripts.
The biological importance of splicing cannot be overstated. It is a critical quality-control checkpoint in gene expression. Errors in splicing can produce non-functional or even toxic proteins, and mis-splicing is implicated in a substantial fraction of human genetic diseases. Moreover, splicing provides a powerful mechanism for generating proteomic diversity from a limited number of genes, allowing a single gene to encode multiple protein isoforms with distinct—and sometimes opposing—functions.
Gene Splicing vs. Genetic Engineering
A common source of confusion is the dual use of the term "gene splicing" in popular media and older literature. In genetic engineering, "gene splicing" refers to the deliberate manipulation of DNA—cutting and pasting genes from one organism into another using restriction enzymes and DNA ligases. This is a laboratory technique, sometimes called recombinant DNA technology, and is fundamentally different from the natural, cellular process of RNA splicing.
The distinction is critical. Natural gene splicing is an intracellular process that occurs on RNA molecules during gene expression. It is catalyzed by a large ribonucleoprotein complex called the spliceosome (or by self-splicing introns). Genetic engineering, by contrast, is an ex vivo manipulation of DNA. When a biotechnologist "splices" a human insulin gene into a bacterial plasmid, they are performing genetic engineering, not RNA splicing. The bacterial host will then transcribe and translate that gene, but bacteria lack the splicing machinery, which is why engineered genes used in bacteria are typically cDNA copies that have already had their introns removed. Throughout this article, "gene splicing" refers exclusively to the natural RNA processing event. For a broader view of how splicing fits into the overall pathway of gene expression, see mRNA Splicing.
The Spliceosome: The Molecular Machinery
snRNPs and Their Roles
The spliceosome is a massive, dynamic ribonucleoprotein complex that catalyzes the removal of introns from pre-mRNA. It is composed of five small nuclear ribonucleoproteins (snRNPs, pronounced "snurps") and numerous non-snRNP protein factors. Each snRNP consists of a small nuclear RNA (snRNA) molecule—U1, U2, U4, U5, or U6—complexed with a set of seven common Sm proteins (B/B′, D1, D2, D3, E, F, and G) and a variable set of particle-specific proteins.
The snRNAs are the key functional components. They base-pair with conserved sequence elements in the pre-mRNA and with each other, providing the specificity for splice site recognition and the structural framework for catalysis. The U1 snRNA contains a sequence complementary to the 5′ splice site consensus sequence. The U2 snRNA base-pairs with the branch point sequence. The U5 snRNA interacts with exon sequences at both the 5′ and 3′ splice sites, aligning the two exons for ligation. U4 and U6 snRNAs are extensively base-paired with each other; U6 is the catalytic core of the spliceosome, and U4 acts as a chaperone that keeps U6 inactive until the correct moment.
The protein components of snRNPs serve structural and regulatory roles. The Sm proteins form a ring around the Sm site of the snRNA, stabilizing the snRNP. Additional proteins, such as the U1-specific U1-70K and U1A, and the U2-specific U2B″, facilitate snRNP–pre-mRNA interactions and protein–protein contacts during assembly. In total, the human spliceosome contains over 150 distinct proteins, making it one of the most complex macromolecular machines in the cell.
Spliceosome Assembly and Catalysis
Spliceosome assembly is a highly ordered, stepwise process that occurs on the pre-mRNA. It proceeds through several discrete complexes, designated E, A, B, Bact, B*, and C, before catalysis and disassembly. The assembly is ATP-dependent; the DEAD/H-box helicases (such as Prp5, Prp28, UAP56, and Brr2) provide the energy to remodel RNA–RNA and RNA–protein interactions, driving the conformational rearrangements required for catalysis.
The process begins with the E (early) complex, in which U1 snRNP binds to the 5′ splice site via base-pairing between the U1 snRNA and the conserved GU dinucleotide at the intron 5′ end. Simultaneously, the splicing factor SF1 (also called BBP, branch point binding protein) binds to the branch point adenine, and U2AF (U2 auxiliary factor) binds to the polypyrimidine tract and the 3′ splice site AG dinucleotide. This "commitment complex" defines the intron boundaries.
Next, U2 snRNP is recruited to the branch point, displacing SF1 in an ATP-dependent reaction. The U2 snRNA base-pairs with the branch point sequence, bulging out the conserved adenosine that will perform the first catalytic step. This forms the A complex (pre-spliceosome). The U4/U6.U5 tri-snRNP then joins, forming the B complex. At this point, the spliceosome undergoes major rearrangements: U1 is released from the 5′ splice site, and U6 base-pairs with the 5′ splice site, displacing U1. U4 is released, allowing U6 to base-pair with U2 and form the catalytic center. This yields the Bact complex, which is then activated to the B* complex, the catalytically competent form.
The B* complex catalyzes the first transesterification reaction, producing the C complex. After the second transesterification, the mature mRNA is released, and the spliceosome disassembles into its component snRNPs, which are recycled for subsequent rounds of splicing. This entire assembly–catalysis–disassembly cycle is the essence of Spliceosome Splicing.
Steps of Gene Splicing
Splice Site Recognition
The accuracy of splicing depends on the recognition of three conserved sequence elements in the intron: the 5′ splice site, the branch point sequence, and the 3′ splice site. In metazoans, these sequences are short and degenerate, which is why the spliceosome relies on both RNA–RNA base-pairing and protein–RNA interactions for recognition.
The 5′ splice site consensus sequence in humans is AG|GURAGU (where | denotes the exon–intron boundary, R is a purine, and N is any nucleotide). The GU dinucleotide at the intron start is nearly invariant. The 3′ splice site consensus is YAG| (where Y is a pyrimidine), preceded by a polypyrimidine tract of 10–20 pyrimidines. The branch point sequence is less conserved; in humans, the consensus is YNYURAC (where the underlined A is the branch point adenosine), typically located 18–40 nucleotides upstream of the 3′ splice site.
Recognition occurs through a combination of mechanisms. U1 snRNA base-pairs with the 5′ splice site. U2AF65 binds the polypyrimidine tract, and U2AF35 binds the 3′ splice site AG. SF1 recognizes the branch point. In higher eukaryotes, additional "exonic splicing enhancers" (ESEs) and "intronic splicing enhancers" (ISEs) are recognized by SR proteins (serine/arginine-rich proteins), which promote splice site recognition. Conversely, exonic splicing silencers (ESSs) and intronic splicing silencers (ISSs) are bound by hnRNP proteins (heterogeneous nuclear ribonucleoproteins), which inhibit splice site recognition. The balance of these positive and negative factors determines whether a given splice site is used.
Two Transesterification Reactions
Splicing proceeds via two sequential transesterification reactions, both catalyzed by the RNA components of the spliceosome. No external energy is required for the chemistry itself; ATP is used only for the conformational rearrangements of the spliceosome.
Step 1: Branch point attack. The 2′-hydroxyl of the branch point adenosine performs a nucleophilic attack on the phosphate at the 5′ splice site. This cleaves the phosphodiester bond between the 5′ exon and the intron, producing two intermediates: the free 5′ exon (with a 3′-hydroxyl group) and the intron–3′ exon lariat intermediate. In this lariat, the 5′ end of the intron is covalently linked to the 2′-hydroxyl of the branch point adenosine via a 2′–5′ phosphodiester bond, forming a branched RNA structure.
Step 2: Exon ligation. The 3′-hydroxyl of the free 5′ exon performs a nucleophilic attack on the phosphate at the 3′ splice site. This cleaves the phosphodiester bond between the intron and the 3′ exon, joining the two exons together and releasing the intron as a lariat structure. The lariat intron is subsequently debranched by the enzyme Dbr1 (debranching enzyme), which hydrolyzes the 2′–5′ bond, and the linear intron is degraded by exonucleases.
The chemistry of both steps is an SN2-like inline attack, resulting in inversion of configuration at the phosphorus atom. The active site is formed by U6 snRNA, specifically the U6 internal stem-loop and the U6–U2 helix Ia. Divalent metal ions, particularly Mg²⁺, coordinate the attacking hydroxyl groups and stabilize the developing negative charge on the leaving groups. This makes the spliceosome a ribozyme—an RNA enzyme—with proteins serving as structural scaffolds and regulatory elements.
Types of Gene Splicing
Constitutive Splicing
Constitutive splicing is the default pathway in which all introns are removed and all exons are joined in the order they appear in the pre-mRNA. This produces a single, predictable mRNA species from a given pre-mRNA. Constitutive splicing is the rule for the vast majority of introns in a given transcript; even genes that undergo alternative splicing typically have some introns that are always removed constitutively.
The term "constitutive" does not imply that these splice sites are weak or unregulated; rather, it means that the splicing pattern is invariant under all physiological conditions. Constitutive splice sites are typically strong—they match the consensus sequences well and are efficiently recognized by the spliceosome. Housekeeping genes, such as those encoding actin (ACTB) or GAPDH, are often constitutively spliced, producing a single protein isoform essential for basic cellular functions.
Alternative Splicing
Alternative splicing is the process by which a single pre-mRNA can be spliced in multiple ways, producing different mRNA isoforms that may encode different protein variants. This is achieved by the differential inclusion or exclusion of exons or parts of exons. The major patterns of alternative splicing are:
- Exon skipping (cassette exon): an entire exon is either included or excluded. This is the most common pattern in humans.
- Alternative 5′ splice site selection: two or more different 5′ splice sites compete for pairing with a given 3′ splice site, leading to extension or truncation of the upstream exon.
- Alternative 3′ splice site selection: two or more 3′ splice sites compete, altering the downstream exon boundary.
- Intron retention: an intron is not removed and remains in the mature mRNA. This is common in plants and lower eukaryotes but rare in mammals.
- Mutually exclusive exons: two or more adjacent exons are included in a mutually exclusive manner; only one is retained in any given mRNA.
Alternative splicing is particularly prevalent in the human nervous system, where genes such as DSCAM (Down syndrome cell adhesion molecule) in Drosophila can theoretically generate over 38,000 isoforms through mutually exclusive exon selection. In humans, the gene encoding neurexin (NRXN1) produces thousands of isoforms through alternative promoter usage and alternative splicing, contributing to the molecular diversity of synapses.
Self-Splicing Introns
Some introns are capable of removing themselves from pre-mRNA without the assistance of the spliceosome. These are called self-splicing introns, and they fall into two main classes: Group I and Group II introns.
Group I introns are found in nuclear rRNA genes, mitochondrial genes, and chloroplast genes of lower eukaryotes, as well as in bacteriophages. They use an exogenous guanosine cofactor (GTP, GDP, GMP, or guanosine itself) as the attacking nucleophile. The 3′-hydroxyl of the guanosine attacks the 5′ splice site, becoming covalently attached to the intron. The free 3′-hydroxyl of the 5′ exon then attacks the 3′ splice site, ligating the exons and releasing the linear intron. Group I introns fold into a conserved secondary structure with several paired regions (P1–P10) that form the catalytic core.
Group II introns are found in bacterial genomes, mitochondrial and chloroplast genes of fungi and plants, and some archaeal genes. They use the 2′-hydroxyl of an internal branch point adenosine as the attacking nucleophile, producing a lariat intermediate exactly like spliceosomal splicing. This mechanistic similarity has led to the widely accepted hypothesis that the spliceosome evolved from a Group II intron ancestor. Group II introns also encode a reverse transcriptase (maturase) that allows them to retrotranspose into new genomic locations. Some Group II introns are true ribozymes that can splice in vitro in the absence of any proteins.
Trans-Splicing
Trans-splicing is a form of splicing that joins exons from two separate pre-mRNA molecules. It is distinct from the cis-splicing described above, where all exons come from a single transcript. Trans-splicing occurs in several organisms, including trypanosomes, nematodes (such as Caenorhabditis elegans), flatworms, and some plants.
The best-characterized example is SL (spliced leader) trans-splicing in trypanosomes and nematodes. A short, non-coding RNA called the SL RNA provides a 5′ exon (the spliced leader) that is joined to the 5′ end of a pre-mRNA encoding a protein. This adds a capped 5′ leader sequence to the mature mRNA, which is required for its stability and translation. In C. elegans, approximately 70% of genes are trans-spliced, either with the SL1 leader (added to the first exon) or SL2 (added to downstream exons in operons). Trans-splicing requires a modified spliceosome that recognizes the SL RNA as a pseudo-substrate.
Alternative Splicing and Proteome Diversity
Mechanisms of Alternative Splicing
The decision to include or exclude an exon is governed by the strength of its splice sites and the regulatory context. Weak splice sites—those that deviate from the consensus—are more likely to be skipped, whereas strong splice sites are constitutively used. The regulation of alternative splicing is achieved by the combinatorial action of trans-acting factors that bind to cis-acting RNA elements.
The core mechanism is the competition between splice sites. For an exon to be included, the spliceosome must recognize the 5′ splice site at its 5′ end and the branch point/3′ splice site at its 3′ end. If either site is weak, the exon may be skipped. SR proteins typically promote exon inclusion by binding to ESEs and recruiting U1 snRNP to the 5′ splice site and U2AF to the 3′ splice site. hnRNP proteins typically promote exon skipping by binding to ESSs or ISSs and sterically blocking the access of the splicing machinery.
A classic example is the alternative splicing of the Drosophila gene doublesex (dsx), which controls sexual differentiation. In females, the SR protein Tra (transformer) and Tra2 bind to an ESE in exon 4, promoting its inclusion. In males, Tra is absent, exon 4 is skipped, and a premature stop codon in exon 5 leads to the production of a truncated, male-specific protein. This simple binary switch illustrates how a single regulatory protein can determine the splicing outcome of an entire gene.
Regulation by Splicing Factors
Splicing factors are proteins that regulate splice site selection. They are classified into two major families: SR proteins and hnRNP proteins.
SR proteins are characterized by one or two N-terminal RNA recognition motifs (RRMs) and a C-terminal arginine/serine-rich (RS) domain. The RS domain is phosphorylated by SR protein kinases (SRPKs) and the Clk/Sty kinase family, which regulates their subcellular localization and activity. SR proteins bind to ESEs and promote spliceosome assembly by recruiting U1 snRNP to the 5′ splice site (via interaction with U1-70K) and U2AF to the 3′ splice site (via interaction with U2AF35). Examples include SF2/ASF (SRSF1), SC35 (SRSF2), and SRp20 (SRSF3).
hnRNP proteins are a diverse family of RNA-binding proteins that generally repress splicing. They bind to ESSs and ISSs and can inhibit splice site recognition by competing with SR proteins, by multimerizing along the RNA to form a repressive complex, or by looping out the exon. Examples include hnRNP A1, hnRNP I (PTB, polypyrimidine tract binding protein), and hnRNP H.
The regulation of alternative splicing is often tissue-specific and developmentally regulated. For example, the gene encoding the neuronal splicing factor Nova-1 is itself alternatively spliced, creating a feedback loop. The PTB protein is expressed in most tissues but is downregulated in neurons, where its paralog nPTB (neural PTB, also called brPTB) takes over. This switch in splicing factor expression drives neuron-specific splicing patterns.
Splicing Errors and Disease
Common Splicing Defects
Splicing errors can arise from mutations in the cis-acting sequence elements (splice sites, branch points, ESEs, ESSs) or in the trans-acting factors (splicing factors). The consequences range from exon skipping to intron retention to the activation of cryptic splice sites.
Mutations at the invariant GU and AG dinucleotides of the 5′ and 3′ splice sites are the most severe; they typically abolish splicing at that site entirely, leading to exon skipping or intron retention. Mutations in the branch point sequence can shift the branch point or abolish splicing. Mutations in ESEs can weaken SR protein binding, causing exon skipping; mutations in ESSs can strengthen hnRNP binding, also causing exon skipping. Mutations that create new splice sites (cryptic splice sites) can lead to the inclusion of intronic sequences or the exclusion of exonic sequences.
It is estimated that up to 15% of all disease-causing point mutations affect splicing, and this is likely an underestimate because many exonic mutations are assumed to be missense or nonsense mutations without being tested for splicing effects.
Diseases Caused by Splicing Errors
Several well-characterized human diseases are caused by splicing defects:
Spinal muscular atrophy (SMA) is caused by mutations in the SMN1 (survival of motor neuron 1) gene. The human genome contains a nearly identical copy, SMN2, which differs by a single C-to-T transition in exon 7. This mutation disrupts an ESE, causing exon 7 to be skipped in the majority of SMN2 transcripts. The resulting protein is truncated and unstable. Because SMN1 is deleted or mutated in SMA patients, they rely on SMN2 for SMN protein production, but the splicing defect means they produce insufficient full-length protein. This is the basis for the FDA-approved drug nusinersen (Spinraza), an antisense oligonucleotide that blocks an ISS in SMN2 intron 7, promoting exon 7 inclusion.
β-thalassemia is caused by mutations in the HBB gene encoding β-globin. Many of these mutations affect splicing. For example, the IVS1-110 G>A mutation creates a new 3′ splice site in intron 1, leading to the inclusion of 19 nucleotides of intronic sequence in the mRNA, which introduces a premature stop codon. Other mutations inactivate the normal splice sites, causing intron retention.
Frontotemporal dementia and Parkinsonism linked to chromosome 17 (FTDP-17) is caused by mutations in the MAPT gene encoding tau protein. Some of these mutations affect the alternative splicing of exon 10, which encodes a microtubule-binding domain. Increased inclusion of exon 10 leads to an overproduction of tau isoforms with four microtubule-binding repeats (4R tau), which aggregate abnormally in neurons.
Duchenne muscular dystrophy (DMD) is caused by mutations in the DMD gene, the largest human gene. Many DMD mutations disrupt the reading frame, leading to a truncated, non-functional dystrophin protein. However, some mutations can be "skipped" by antisense oligonucleotides that promote exon skipping, restoring the reading frame and producing a shorter but partially functional protein. This is the basis for the exon-skipping drugs eteplirsen and golodirsen.
Splicing factors themselves can also be mutated in disease. Mutations in SF3B1, a component of the U2 snRNP, are common in myelodysplastic syndromes and chronic lymphocytic leukemia. Mutations in U2AF1 and SRSF2 are also found in myeloid malignancies. These mutations alter the splicing of many genes, contributing to the pathogenesis of these cancers. The role of splicing dysregulation in cancer is an active area of research, and splicing factors are considered potential therapeutic targets.
Methods to Study Gene Splicing
RT-PCR and Quantitative PCR
Reverse transcription PCR (RT-PCR) is the most direct method to analyze splicing patterns. Total RNA is isolated from cells or tissues, and reverse transcriptase is used to synthesize cDNA. PCR primers are designed to flank the region of interest—typically spanning one or more alternatively spliced exons. The PCR products are then separated by agarose gel electrophoresis, and the different isoforms are visualized as bands of different sizes.
For example, to analyze the alternative splicing of exon 7 in SMN2, primers are designed in exons 6 and 8. The full-length isoform (including exon 7) produces a larger amplicon than the isoform lacking exon 7. The relative intensity of the bands reflects the relative abundance of the isoforms, although this is only semi-quantitative.
Quantitative PCR (qPCR) can provide more precise measurements. Two approaches are common. In the first, isoform-specific primers are designed to amplify only one isoform (e.g., a primer spanning the exon 6–7 junction for the full-length isoform and a primer spanning the exon 6–8 junction for the skipped isoform). In the second, a TaqMan probe is designed to hybridize specifically to the exon junction of interest. Both approaches allow the relative abundance of different isoforms to be quantified with high sensitivity and reproducibility. A typical qPCR reaction uses 10–50 ng of cDNA, 200–400 nM primers, and 100–200 nM probe, with cycling conditions of 95°C for 15 seconds and 60°C for 1 minute for 40 cycles.
RNA Sequencing (RNA-seq)
RNA-seq has revolutionized the study of splicing. In a typical RNA-seq experiment, total RNA is depleted of ribosomal RNA (or poly-A selected), fragmented, converted to cDNA, and sequenced using high-throughput platforms such as Illumina. The resulting short reads (typically 50–150 bp) are aligned to a reference genome or transcriptome.
Splicing analysis in RNA-seq data relies on the detection of reads that span exon–exon junctions. A read that aligns to two non-contiguous genomic regions indicates that those exons are joined in the mature mRNA. The number of reads supporting each junction is proportional to the abundance of that isoform. Tools such as rMATS, MISO, and DEXSeq are used to quantify alternative splicing events and to identify differentially spliced genes between conditions.
RNA-seq has several advantages over RT-PCR: it is unbiased (no prior knowledge of isoforms required), genome-wide, and quantitative. However, it requires substantial bioinformatics expertise and computational resources. For a typical human RNA-seq experiment, 20–50 million paired-end reads per sample are recommended for splicing analysis. The alignment and quantification steps are computationally intensive, and the analysis of splicing events requires careful filtering to avoid false positives from mapping artifacts.
Minigene Splicing Reporters
Minigene reporters are plasmid constructs that contain a splicing reporter gene—typically a constitutively expressed promoter (such as CMV or a β-globin promoter), a portion of the gene of interest containing the alternative exon and flanking intronic sequences, and two fluorescent protein genes (such as GFP and RFP) in different reading frames. The inclusion or exclusion of the alternative exon determines which fluorescent protein is expressed.
For example, a common design places the alternative exon between two exons encoding GFP and RFP. If the alternative exon is included and is in-frame, the GFP–exon–RFP fusion protein is produced, and the cells fluoresce green and red. If the exon is skipped, the GFP and RFP are out of frame, and a premature stop codon leads to the production of only GFP (or no fluorescent protein, depending on the design). The ratio of green to red fluorescence provides a quantitative readout of splicing.
Minigene reporters are widely used to study the cis-acting elements that regulate splicing. By mutating putative ESEs or ESSs in the minigene and measuring the change in splicing, researchers can identify functional regulatory elements. They are also used to test the effects of patient-derived mutations on splicing. The minigene is transfected into cells, and after 24–48 hours, RNA is harvested and analyzed by RT-PCR, or the cells are analyzed by flow cytometry or fluorescence microscopy.
Common Pitfalls and Misconceptions
Splicing vs. Other RNA Processing Steps
A frequent source of confusion is the distinction between splicing and the other two major RNA processing events: 5′ capping and 3′ polyadenylation. These are three separate processes that occur on the same pre-mRNA molecule but are mechanistically distinct.
5′ capping is the addition of a 7-methylguanosine cap to the 5′ end of the pre-mRNA. It occurs co-transcriptionally, when the transcript is only 20–30 nucleotides long, and is catalyzed by the capping enzyme complex (triphosphatase, guanylyltransferase, and methyltransferase). The cap protects the mRNA from 5′→3′ exonucleolytic degradation and is required for translation initiation.
3′ polyadenylation is the cleavage of the pre-mRNA at a site downstream of the AAUAAA polyadenylation signal and the addition of a poly(A) tail of 200–250 adenosines. This is catalyzed by the cleavage and polyadenylation specificity factor (CPSF), cleavage stimulation factor (CstF), and poly(A) polymerase (PAP). The poly(A) tail protects the mRNA from 3′→5′ degradation and is required for translation and mRNA export.
Splicing is the removal of introns and joining of exons. It is catalyzed by the spliceosome and occurs either co-transcriptionally or post-transcriptionally. These three processes are coordinated by the C-terminal domain (CTD) of RNA polymerase II, which serves as a platform for the recruitment of processing factors. However, they are independent reactions with distinct enzymes and substrates.
Prokaryotic vs. Eukaryotic Splicing
Another common misconception is that splicing occurs in all organisms. In fact, splicing is predominantly a eukaryotic process. Prokaryotes (bacteria and archaea) do not have spliceosomes and do not splice their mRNAs. Their genes are typically continuous, with no introns in protein-coding genes. Transcription and translation are coupled in bacteria: ribosomes begin translating the mRNA while it is still being transcribed, leaving no opportunity for splicing.
There are exceptions. Some archaeal genes contain introns, but these are removed by a different mechanism involving a bulge-helix-bulge endonuclease, not a spliceosome. Some bacterial genes contain Group I or Group II introns, but these are self-splicing and are often found in tRNA genes or mobile genetic elements rather than in typical mRNA genes. The presence of self-splicing introns in bacteria is thought to reflect the ancient origin of these elements, which predate the eukaryotic spliceosome.
In eukaryotes, the extent of splicing varies widely. Yeast (Saccharomyces cerevisiae) has only about 300 introns in its entire genome, and most genes have a single intron. In contrast, the human genome has an average of 8 introns per gene, and some genes (such as titin, TTN) have over 300 introns. The complexity of splicing regulation also increases with organismal complexity, with higher eukaryotes having more SR proteins, more hnRNP proteins, and more extensive alternative splicing.
Summary and Key Takeaways
Gene splicing is the removal of introns and joining of exons in pre-mRNA, a mandatory step in the expression of most eukaryotic genes. The spliceosome, a complex of five snRNPs and numerous proteins, catalyzes the two transesterification reactions that accomplish this. Splicing can be constitutive (all introns removed, all exons joined) or alternative (different combinations of exons selected), and it can be catalyzed by the spliceosome, by self-splicing introns, or by trans-splicing between separate RNA molecules. Alternative splicing is a major source of proteomic diversity, regulated by the combinatorial action of SR proteins and hnRNP proteins. Errors in splicing cause numerous human diseases, including spinal muscular atrophy, β-thalassemia, and many cancers. Modern methods for studying splicing include RT-PCR, RNA-seq, and minigene reporters.
Frequently Asked Questions
What are the types of gene splicing?
The main types are constitutive splicing (all introns removed, all exons joined in order), alternative splicing (different exon combinations), self-splicing (Group I and Group II introns that catalyze their own removal), and trans-splicing (joining exons from separate pre-mRNA molecules). Constitutive and alternative splicing are catalyzed by the spliceosome; self-splicing is catalyzed by the intron RNA itself; trans-splicing uses a modified spliceosome.
What are the steps of gene splicing?
Splicing proceeds in two transesterification reactions. First, the 2′-hydroxyl of the branch point adenosine attacks the 5′ splice site, cleaving the RNA and forming a lariat intermediate. Second, the 3′-hydroxyl of the free 5′ exon attacks the 3′ splice site, ligating the exons and releasing the intron as a lariat. These reactions are preceded by spliceosome assembly: U1 binds the 5′ splice site, U2 binds the branch point, and the U4/U6.U5 tri-snRNP joins to form the active site.
What is the difference between gene splicing and alternative splicing?
Gene splicing (or RNA splicing) is the general process of intron removal and exon joining. Alternative splicing is a specific type of splicing in which a single pre-mRNA can be spliced in multiple ways, producing different mRNA isoforms. All alternative splicing is gene splicing, but not all gene splicing is alternative. Constitutive splicing produces a single mRNA; alternative splicing produces multiple mRNAs from the same pre-mRNA.
Where does gene splicing occur in the cell?
Splicing occurs in the nucleus, either co-transcriptionally (while the pre-mRNA is still being synthesized by RNA polymerase II) or post-transcriptionally (after transcription is complete but before the mRNA is exported to the cytoplasm). The spliceosome assembles on the nascent pre-mRNA, and splicing factors are enriched in nuclear speckles, which are storage and assembly sites for splicing components.
Why is gene splicing important?
Splicing is essential for gene expression in eukaryotes because most genes contain introns that must be removed to produce a functional mRNA. Splicing also provides a mechanism for generating protein diversity through alternative splicing, allowing a single gene to encode multiple protein isoforms. Additionally, splicing serves as a quality-control checkpoint; errors in splicing can lead to non-functional proteins and disease.
What happens if gene splicing goes wrong?
Errors in splicing can lead to exon skipping, intron retention, or the use of cryptic splice sites, all of which can alter the reading frame or introduce premature stop codons. The resulting mRNA may be degraded by nonsense-mediated decay, or it may produce a truncated or aberrant protein. Splicing errors cause diseases such as spinal muscular atrophy, β-thalassemia, and many cancers. Mutations in splicing factors themselves can also cause disease by globally altering splicing patterns.
Do prokaryotes undergo gene splicing?
Most prokaryotes do not undergo splicing of their mRNAs. Bacterial and archaeal genes are typically continuous, and transcription and translation are coupled. However, some archaeal genes contain introns that are removed by a bulge-helix-bulge endonuclease, and some bacterial genes contain self-splicing Group I or Group II introns. These are exceptions rather than the rule, and they do not involve a spliceosome.
Key Takeaways
- Gene splicing is the removal of introns and joining of exons in pre-mRNA, catalyzed by the spliceosome via two transesterification reactions.
- The spliceosome is composed of five snRNPs (U1, U2, U4, U5, U6) and over 150 proteins; U6 snRNA forms the catalytic core.
- Splicing occurs in the nucleus, often co-transcriptionally, and is coordinated with capping and polyadenylation.
- Alternative splicing generates multiple mRNA isoforms from a single gene, vastly increasing proteomic diversity; it is regulated by SR proteins (activators) and hnRNP proteins (repressors).
- Self-splicing Group I and Group II introns are ribozymes that catalyze their own removal; Group II introns are likely ancestors of the spliceosome.
- Splicing errors cause numerous human diseases, including spinal muscular atrophy, β-thalassemia, and cancers; antisense oligonucleotide therapies can correct some splicing defects.
- Splicing is studied using RT-PCR, RNA-seq, and minigene reporters, each with distinct advantages and limitations.
Further Reading
- Wang Y et al. Alternative splicing of inner-ear-expressed genes. Frontiers of medicine. 2016. PubMed 27376950
- Xu W et al. Fifty-four novel mutations in the NF1 gene and integrated analyses of the mutations that modulate splicing. International journal of molecular medicine. 2014. PubMed 24789688
- Ratni H et al. Discovery of Risdiplam, a Selective Survival of Motor Neuron-2 ( SMN2) Gene Splicing Modifier for the Treatment of Spinal Muscular Atrophy (SMA). Journal of medicinal chemistry. 2018. PubMed 30044619
- Suslow T. Gene-splicing experiment. Science (New York, N.Y.). 1984. PubMed 6587569
- Barnaby W. Sweden debates gene-splicing. Nature. 1977. PubMed 11661514
- Speakman E, Gunaratne GH. On a kneading theory for gene-splicing. Chaos (Woodbury, N.Y.). 2024. PubMed 38579148