Spliceosome Splicing: Mechanism, Regulation, and Disease
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Spliceosome Splicing
What is RNA Splicing?
RNA splicing is the post-transcriptional process by which introns—non-coding intervening sequences—are removed from a precursor messenger RNA (pre-mRNA) and exons—coding sequences—are joined to form a mature mRNA. In humans, the vast majority of protein-coding genes contain introns; the average human gene has roughly eight introns, and some, like the gene encoding the muscle protein titin (TTN), contain over 360. The removal of these introns is therefore an obligatory step in the expression of most human genes.
The biological significance of splicing extends beyond simple intron removal. Splicing is a prerequisite for mRNA export to the cytoplasm, as the exon junction complex deposited during splicing marks the mRNA as processed. It also determines the open reading frame of the mRNA, and errors in splice site selection can introduce premature stop codons that trigger nonsense-mediated decay. The process is constitutive for some introns—meaning they are always removed in the same way—but for many others, splicing is regulated, allowing a single gene to produce multiple distinct mRNA isoforms. This regulated process, called alternative splicing, is discussed in detail in Section 5.
The Spliceosome: A Ribonucleoprotein Machine
The spliceosome is the large, dynamic ribonucleoprotein complex that catalyzes splicing. It is composed of five small nuclear ribonucleoproteins (snRNPs)—U1, U2, U4, U5, and U6—each containing a small nuclear RNA (snRNA) and a set of associated proteins. In addition to the snRNPs, the spliceosome contains numerous non-snRNP protein factors that are required for assembly, regulation, and catalysis. In total, the assembled spliceosome can contain over 150 distinct proteins, making it one of the most complex macromolecular machines in the cell. For a detailed breakdown of its molecular architecture, see Spliceosome Structure and Spliceosome Composed.
The spliceosome does not pre-exist as a single stable entity. Instead, it assembles de novo on each intron in a stepwise, highly ordered manner. The snRNPs themselves are pre-assembled in the cytoplasm and nucleus: the Sm core proteins (B/B′, D1, D2, D3, E, F, and G) bind to a conserved Sm site on the snRNA, and specific proteins associate with each snRNP. For example, U1 snRNP contains the U1-70K, U1-A, and U1-C proteins, while U2 snRNP contains the U2-A′ and U2-B″ proteins. The assembly of the spliceosome on pre-mRNA is a major point of regulation, and its disruption underlies many splicing-related diseases.
The Chemistry of Splicing: Two Transesterification Reactions
Splicing proceeds via two sequential transesterification reactions. A transesterification is a chemical reaction in which an ester is converted into another ester by exchanging the alkoxy group; in the context of RNA, this means breaking one phosphodiester bond and forming another. Importantly, the overall reaction is isoenergetic—no ATP is consumed in the chemistry itself. ATP is instead used to drive the conformational rearrangements of the spliceosome during assembly and disassembly.
Step 1: Branch Point Nucleophilic Attack
In the first step, the 2′-hydroxyl group of an adenosine residue at the branch point—a conserved sequence located 18–40 nucleotides upstream of the 3′ splice site—attacks the phosphate at the 5′ splice site. This nucleophilic attack breaks the phosphodiester bond between the 5′ exon and the intron, generating two products: a free 5′ exon with a 3′-hydroxyl group, and the intron in the form of a lariat, in which the 5′ end of the intron is covalently linked to the branch point adenosine via a 2′–5′ phosphodiester bond.
The branch point adenosine is not simply any adenosine; its 2′-hydroxyl must be positioned precisely within the catalytic core of the spliceosome. This positioning is achieved by the U2 snRNA, which base-pairs with the branch point sequence, and by the protein SF3b, which binds the branch point and holds it in a bulged conformation. The bulged adenosine is thereby activated for nucleophilic attack. The chemistry of this step is analogous to the first step of group II intron self-splicing, and it is now widely accepted that the spliceosome is a ribozyme—that is, the RNA components of the spliceosome, specifically U6 snRNA, provide the catalytic magnesium ions that coordinate the reaction.
Step 2: Exon Ligation and Lariat Release
In the second step, the 3′-hydroxyl group of the free 5′ exon attacks the phosphate at the 3′ splice site. This reaction joins the 5′ exon to the 3′ exon, releasing the lariat intron. The lariat is subsequently debranched by the enzyme DBR1, which hydrolyzes the 2′–5′ phosphodiester bond, and the linear intron is degraded by exonucleases.
The second step requires that the two exons be brought into close proximity and that the 3′ splice site be correctly positioned. This is achieved by the U5 snRNP, which aligns the two exons: the loop 1 of U5 snRNA base-pairs with the last few nucleotides of the 5′ exon and the first few nucleotides of the 3′ exon, holding them in the correct register for ligation. The transition from step 1 to step 2 involves a conformational rearrangement of the spliceosome, in which the branch point is moved out of the active site and the 3′ splice site is recruited in. This rearrangement is ATP-dependent and is mediated by the RNA helicase Prp16.
Spliceosome Assembly and Disassembly Cycle
The assembly of the spliceosome is a highly ordered, ATP-dependent process that proceeds through a series of discrete complexes, designated E, A, B, B*, C, and P. The nomenclature reflects the order of assembly, and each complex is defined by its snRNP composition. The entire cycle is reviewed in detail in Spliceosome Assembly and Spliceosome Complex.
Early Complexes: E, A, and B
The assembly begins with the recognition of the 5′ splice site by U1 snRNP and the branch point/polypyrimidine tract by the non-snRNP proteins SF1 (also called BBP, branch point binding protein) and U2AF (U2 auxiliary factor). U2AF is a heterodimer of a 65 kDa subunit (U2AF65) that binds the polypyrimidine tract and a 35 kDa subunit (U2AF35) that recognizes the 3′ splice site AG dinucleotide. This initial complex is called the E (early) complex. The E complex is ATP-independent and commits the pre-mRNA to the splicing pathway.
Next, U2 snRNP joins in an ATP-dependent manner, base-pairing its snRNA with the branch point sequence. This displaces SF1 and forms the A complex (also called the prespliceosome). The base-pairing between U2 snRNA and the branch point is extensive—about 6–8 base pairs—and it bulges out the branch point adenosine, positioning it for the first catalytic step.
The U4/U6.U5 tri-snRNP then joins to form the B complex. The tri-snRNP is a pre-assembled particle in which U4 and U6 snRNAs are extensively base-paired with each other, and U5 snRNP is associated with the U4/U6 complex. At this stage, the spliceosome is catalytically inactive; the U6 snRNA is held in an inactive conformation by its pairing with U4 snRNA.
Activation and Catalytic Center Formation
The transition from the B complex to the activated B* complex is a dramatic remodeling event. The RNA helicase Prp28 (in humans, also called DDX23) destabilizes the U1 snRNA–5′ splice site interaction, allowing U6 snRNA to base-pair with the 5′ splice site. Simultaneously, the helicase Brr2 unwinds the U4/U6 duplex, releasing U4 snRNP from the complex. U6 snRNA then refolds into its catalytically active conformation, forming an intramolecular stem-loop and base-pairing with U2 snRNA to create the U2/U6 helix I. This U2/U6 structure forms the core of the catalytic center, coordinating the magnesium ions required for both transesterification steps.
The B* complex undergoes further rearrangements to form the B* catalytically active complex (often called B*act), which performs the first transesterification. After step 1, the complex is called C. The helicase Prp16 then remodels the complex for step 2, and the resulting complex (C*) catalyzes exon ligation. The post-catalytic complex, designated P, contains the ligated exons and the lariat intron.
Disassembly and Recycling
After exon ligation, the spliceosome must be disassembled to release the mRNA and to recycle the snRNPs for subsequent rounds of splicing. The helicase Prp22 is required to release the mRNA from the spliceosome, and the helicases Prp43 (with its cofactors Ntr1 and Ntr2) then disassemble the remaining complex, releasing the lariat intron and the U2, U5, and U6 snRNPs. The lariat is debranched by DBR1 and degraded. The U4/U6.U5 tri-snRNP must be reassembled for the next round, which requires the annealing of U4 and U6 snRNAs by the protein SART3 (also called Tip110) and the re-association of U5 snRNP.
The entire assembly–catalysis–disassembly cycle occurs in minutes in vivo. In a typical in vitro splicing reaction using HeLa cell nuclear extract, splicing is assayed at 30 °C for 60–90 minutes, and the products are resolved by denaturing polyacrylamide gel electrophoresis.
Splice Site Recognition and Consensus Sequences
The accuracy of splicing depends on the recognition of conserved sequence elements at the intron boundaries. These elements are short and degenerate, and their recognition involves both RNA–RNA base-pairing with snRNAs and protein–RNA interactions. The consensus sequences for human splice sites are shown in Table 1.
| Element | Location | Consensus Sequence (5′ to 3′) | Recognized By | |
|---|---|---|---|---|
| 5′ splice site | Intron start | AG | GURAGU (R = purine) | U1 snRNA, U6 snRNA |
| Branch point | 18–40 nt upstream of 3′ SS | YNYURAY (Y = pyrimidine, N = any, R = purine) | U2 snRNA, SF1 | |
| Polypyrimidine tract | Between branch point and 3′ SS | U-rich (10–20 nt) | U2AF65 | |
| 3′ splice site | Intron end | YAG | G | U2AF35, U5 snRNA |
Table 1. Consensus sequences at human splice sites. The vertical bar indicates the exon–intron boundary. The branch point adenosine is underlined in the branch point consensus.
5′ Splice Site and U1 snRNA
The 5′ splice site consensus is AG|GURAGU, where the GU dinucleotide at the intron start is nearly invariant. The U1 snRNA contains a sequence at its 5′ end that is complementary to the 5′ splice site: U1 snRNA nucleotides 3–10 are 3′-ACUUACCUG-5′, which base-pair with the 5′ splice site. This base-pairing is the first recognition event in spliceosome assembly and is required for the definition of the intron.
However, U1 snRNA base-pairing alone is not sufficient for accurate 5′ splice site selection. The U1-70K and U1-C proteins stabilize the interaction, and the RNA helicase Prp28 subsequently disrupts the U1–5′ splice site duplex to allow U6 snRNA to take over. U6 snRNA base-pairs with the 5′ splice site through its conserved ACAGAGA box, and this interaction is essential for catalysis. The transition from U1 to U6 binding is a critical proofreading step; if the 5′ splice site is not correctly recognized, the helicase Prp28 fails to act, and the complex is discarded.
Branch Point and U2 snRNA
The branch point sequence in humans is degenerate, but the consensus is YNYURAY, where the final A is the branch point adenosine. U2 snRNA contains a sequence complementary to the branch point (U2 snRNA nucleotides 34–38 are GUAGUA), and base-pairing between U2 snRNA and the branch point bulges out the branch point adenosine. The protein SF3b, a component of the U2 snRNP, binds the branch point and stabilizes the bulged conformation. The branch point adenosine must be bulged, not base-paired, because its 2′-hydroxyl must be free to attack the 5′ splice site.
The distance between the branch point and the 3′ splice site is variable (18–40 nucleotides), and this flexibility is accommodated by the U2 snRNP-associated proteins. The selection of the branch point is also influenced by the polypyrimidine tract: U2AF65 binds the polypyrimidine tract and recruits U2 snRNP to the nearby branch point.
3′ Splice Site and U2AF
The 3′ splice site is defined by three elements: the branch point, the polypyrimidine tract, and the terminal AG dinucleotide. U2AF65 binds the polypyrimidine tract with high affinity, and U2AF35 recognizes the AG dinucleotide at the 3′ splice site. The interaction between U2AF35 and the AG is weak, and it is often stabilized by interactions with other splicing factors, such as the SR proteins bound to nearby exonic splicing enhancers.
The requirement for the AG dinucleotide is not absolute for the first step of splicing; in some introns, the first step can occur without a recognizable AG, and the AG is only required for the second step. This is because the AG is recognized by U2AF35 during initial assembly, but the 3′ splice site is also recognized by U5 snRNA during the second step, when U5 loop 1 base-pairs with the exon sequences flanking the intron.
Alternative Splicing: Expanding the Proteome
Alternative splicing is the process by which different combinations of exons are joined to produce multiple mRNA isoforms from a single gene. It is estimated that over 95% of human multi-exon genes undergo alternative splicing, and this process is a major source of proteomic diversity. The regulation of alternative splicing is achieved by the combinatorial action of cis-acting elements and trans-acting factors.
Types of Alternative Splicing
There are five basic modes of alternative splicing:
- Exon skipping: An entire exon is excluded from the mRNA. This is the most common mode in humans.
- Alternative 5′ splice site selection: Two or more 5′ splice sites compete for pairing with a downstream 3′ splice site.
- Alternative 3′ splice site selection: Two or more 3′ splice sites compete for pairing with an upstream 5′ splice site.
- Intron retention: An intron is retained in the mature mRNA. This is common in plants and lower eukaryotes but rare in mammals.
- Mutually exclusive exons: Two or more adjacent exons are spliced such that only one is included at a time.
These modes can be combined in complex ways; for example, the DSCAM gene in Drosophila can theoretically produce over 38,000 isoforms through the combinatorial use of mutually exclusive exon clusters.
Regulatory Elements and SR Proteins
The decision to include or skip an exon is governed by cis-acting elements in the pre-mRNA: exonic splicing enhancers (ESEs), exonic splicing silencers (ESSs), intronic splicing enhancers (ISEs), and intronic splicing silencers (ISSs). These elements are recognized by trans-acting protein factors that either promote or repress spliceosome assembly.
The major positive regulators are the SR proteins, a family of serine/arginine-rich proteins that includes SF2/ASF (SRSF1), SC35 (SRSF2), and SRp40 (SRSF5). SR proteins bind ESEs through their RNA recognition motifs (RRMs) and recruit the U1 snRNP to the 5′ splice site and U2AF to the 3′ splice site through their RS domains. The RS domains are rich in alternating serine and arginine residues and mediate protein–protein interactions. SR proteins are also involved in constitutive splicing, where they help define exons in genes with weak splice sites.
The activity of SR proteins is regulated by phosphorylation. SR protein kinases, such as SRPK1 and Clk/Sty, phosphorylate the RS domains, and this phosphorylation is required for SR protein nuclear import and for their recruitment to transcription sites. Dephosphorylation by phosphatases such as PP1 is required for catalytic steps of splicing. The phosphorylation state of SR proteins therefore provides a link between cellular signaling pathways and splicing outcomes.
hnRNPs as Silencers
The major negative regulators are the heterogeneous nuclear ribonucleoproteins (hnRNPs), a large family of RNA-binding proteins that includes hnRNP A1, hnRNP C, and polypyrimidine tract binding protein (PTB, also called hnRNP I). hnRNPs bind ESSs and ISSs and repress splicing by blocking the access of SR proteins and snRNPs to the splice sites.
hnRNP A1 is a well-studied silencer that binds high-affinity sequences containing the motif UAGGG(A). When hnRNP A1 binds an ESS, it can antagonize the binding of SR proteins to nearby ESEs, thereby promoting exon skipping. hnRNP A1 can also act at a distance by looping out the intervening RNA, bringing distantly bound hnRNP A1 molecules into proximity and sterically blocking spliceosome assembly.
PTB is a particularly important silencer that regulates the alternative splicing of many genes, including the α-tropomyosin gene and the FGFR2 gene. PTB binds to pyrimidine-rich sequences and represses splicing by competing with U2AF65 for binding to the polypyrimidine tract. The repression of splicing by PTB is often tissue-specific, as the related protein nPTB (neuronal PTB) is expressed in neurons and has different target specificity.
Methods to Study Splicing and the Spliceosome
In Vitro Splicing Assays
The classical method for studying splicing is the in vitro splicing assay, developed in the early 1980s. In this assay, a radiolabeled pre-mRNA substrate is incubated with HeLa cell nuclear extract, which contains all the necessary splicing factors, in the presence of ATP, creatine phosphate, and magnesium chloride (typically 2.5 mM MgCl₂, 0.5 mM ATP, 20 mM creatine phosphate) at 30 °C for 60–90 minutes. The RNA products are then extracted with phenol-chloroform, precipitated with ethanol, and resolved by denaturing polyacrylamide gel electrophoresis (typically 8–15% polyacrylamide, 8 M urea). The products—the lariat intron, the spliced mRNA, and the free 5′ exon—are detected by autoradiography.
This assay has been used to define the order of spliceosome assembly, to identify the transesterification intermediates, and to test the function of individual splicing factors by depletion and add-back experiments. For example, immunodepletion of U2 snRNP from nuclear extract abolishes splicing, and splicing is restored by adding back purified U2 snRNP.
High-Throughput Sequencing of Splicing
The advent of high-throughput RNA sequencing (RNA-seq) has revolutionized the study of splicing. RNA-seq can quantify the inclusion levels of individual exons across the entire transcriptome, allowing the identification of differentially spliced exons between conditions. The standard metric is the percent spliced-in (PSI, ψ) value, which ranges from 0 (exon always skipped) to 1 (exon always included).
RNA-seq has been combined with crosslinking and immunoprecipitation (CLIP) to identify the binding sites of splicing factors genome-wide. In CLIP, cells are irradiated with UV light (254 nm) to crosslink RNA-binding proteins to their target RNAs. The protein of interest is immunoprecipitated, the RNA is partially digested, and the crosslinked RNA fragments are sequenced. This approach has been used to map the binding sites of SR proteins, hnRNPs, and the snRNP components, revealing that splicing factors often bind in clusters near alternative exons.
Cryo-EM Structural Studies
The determination of the spliceosome structure by cryo-electron microscopy (cryo-EM) has been one of the major achievements in molecular biology over the past decade. Because the spliceosome is large (~2–5 MDa) and conformationally dynamic, it was refractory to X-ray crystallography. Cryo-EM, which does not require crystals and can handle conformational heterogeneity, has allowed the determination of structures of the Saccharomyces cerevisiae and human spliceosome at various stages of the assembly and catalytic cycle.
These structures have revealed the molecular architecture of the catalytic center, showing that U6 snRNA and the protein Prp8 form the heart of the active site, with two magnesium ions coordinated by U6 snRNA. The structures have also revealed how the branch point adenosine is positioned for the first step and how U5 snRNA aligns the exons for the second step. For a comprehensive overview of these structures, see Spliceosome Structure.
Spliceosome Dysfunction and Human Disease
Given the central role of splicing in gene expression, it is not surprising that defects in splicing cause or contribute to a wide range of human diseases. These defects can arise from mutations in the splice sites themselves, in the splicing factors, or in the regulatory elements that control alternative splicing.
Spliceosomal Mutations in Cancer
Recurrent somatic mutations in genes encoding core spliceosome components have been identified in several cancers. The most notable are mutations in SF3B1, U2AF1, SRSF2, and ZRSR2, which are found in myelodysplastic syndromes (MDS), chronic lymphocytic leukemia (CLL), and uveal melanoma. These mutations are heterozygous and are thought to act through a dominant-negative or neomorphic mechanism, altering splice site selection rather than abolishing splicing entirely.
Mutations in SF3B1, which encodes a component of the U2 snRNP, are the most common spliceosomal mutations in MDS, occurring in about 20–30% of cases. The most frequent mutation is K700E, which alters the branch point recognition. Cells carrying SF3B1 mutations show aberrant 3′ splice site selection, with a preference for cryptic branch points that are closer to the 3′ splice site. This leads to the mis-splicing of hundreds of genes, including genes involved in iron metabolism and mitochondrial function.
Mutations in U2AF1 (also called U2AF35) occur at two hot spots: S34 and Q157. These mutations alter the sequence specificity of U2AF35 for the 3′ splice site AG dinucleotide, causing a shift in 3′ splice site selection. The consequence is widespread mis-splicing, including the aberrant inclusion of a poison exon in the STRAP gene, which leads to its downregulation.
Inherited Splicing Disorders
Inherited mutations that affect splicing can cause a variety of genetic diseases. One of the best-characterized is spinal muscular atrophy (SMA), a neurodegenerative disease caused by loss-of-function mutations in the SMN1 gene. The SMN1 gene encodes the SMN protein, which is required for the assembly of the Sm core of snRNPs. Humans have a nearly identical paralog, SMN2, which produces only about 10% of functional SMN protein because a C-to-T transition in exon 7 disrupts an exonic splicing enhancer, causing exon 7 skipping. The resulting protein is unstable and rapidly degraded. SMA is therefore a disease of snRNP biogenesis: reduced SMN levels lead to reduced snRNP assembly, which in turn impairs splicing of specific target genes, particularly in motor neurons.
Retinitis pigmentosa (RP) is a group of inherited retinal degenerative diseases caused by mutations in many genes, including several that encode core spliceosome components. Mutations in PRPF31, PRPF8, PRPF3, RP9, and SNRNP200 (which encodes the Brr2 helicase) cause autosomal dominant RP. The mechanism is thought to involve haploinsufficiency: the reduced dosage of these splicing factors makes the retina particularly vulnerable, possibly because photoreceptor cells have very high demands for splicing of genes involved in phototransduction.
Common Pitfalls and Exam Tips for Splicing
Misconception: Splicing Occurs in the Cytoplasm
A common error is to assume that splicing occurs in the cytoplasm. In fact, splicing of nuclear pre-mRNA occurs in the nucleus, co-transcriptionally or shortly after transcription. The spliceosome is a nuclear machine, and the snRNPs are predominantly nuclear. The mature mRNA is exported to the cytoplasm only after splicing is complete. An exception is the splicing of some mRNAs in the cytoplasm of platelets and certain viruses, but this is not the rule.
Remembering the Order of snRNPs
Students often struggle to remember the order of snRNP addition. A useful mnemonic is "U1, U2, U4/U6.U5" — the order of assembly. U1 binds first, then U2, then the U4/U6.U5 tri-snRNP. After activation, U1 and U4 are released, leaving U2, U5, and U6 in the catalytic complex. The key point is that U6 is the catalytic snRNA, and U4 is its inhibitor; U4 must be removed for U6 to become active.
Key Terms to Master
- snRNP: Small nuclear ribonucleoprotein; the RNA–protein complex that recognizes splice sites.
- Lariat: The intron intermediate in which the 5′ end is linked to the branch point via a 2′–5′ bond.
- Transesterification: The chemical reaction that breaks and forms phosphodiester bonds.
- Branch point: The adenosine whose 2′-hydroxyl attacks the 5′ splice site.
- ESE/ESS: Exonic splicing enhancer/silencer; cis-acting elements that regulate alternative splicing.
- SR protein: A family of splicing activators with serine/arginine-rich domains.
- hnRNP: Heterogeneous nuclear ribonucleoprotein; a family of splicing repressors.
Frequently Asked Questions
What is the spliceosome and what does it do?
The spliceosome is a large ribonucleoprotein complex that catalyzes the removal of introns from pre-mRNA and the joining of exons. It is composed of five snRNPs (U1, U2, U4, U5, U6) and numerous accessory proteins. The spliceosome assembles de novo on each intron, performs two transesterification reactions, and is then disassembled. For a concise definition, see Spliceosome Definition.
How does the spliceosome catalyze splicing?
The spliceosome catalyzes splicing via two transesterification reactions. In the first, the 2′-hydroxyl of the branch point adenosine attacks the 5′ splice site, forming a lariat intermediate. In the second, the 3′-hydroxyl of the 5′ exon attacks the 3′ splice site, ligating the exons and releasing the lariat intron. The catalytic center is formed by U6 snRNA, which coordinates two magnesium ions.
What are the main components of the spliceosome?
The main components are the five snRNPs: U1, U2, U4, U5, and U6. Each snRNP contains a snRNA and associated proteins. U1 recognizes the 5′ splice site, U2 recognizes the branch point, U4/U6 form a duplex that is unwound during activation, and U5 aligns the exons for ligation. The spliceosome also contains numerous non-snRNP proteins, including the SF3a/SF3b complexes, Prp19 complex, and RNA helicases. See Spliceosome Proteins and Spliceosome Made for more details.
What is the difference between constitutive and alternative splicing?
Constitutive splicing removes introns and joins exons in the same way for every mRNA molecule; it is the default pathway. Alternative splicing allows different combinations of exons to be joined, producing multiple mRNA isoforms from a single gene. Alternative splicing is regulated by cis-acting elements and trans-acting factors, and it greatly expands the coding capacity of the genome.
Why is alternative splicing important?
Alternative splicing is important because it generates proteomic diversity. Over 95% of human multi-exon genes undergo alternative splicing, allowing a single gene to produce multiple protein isoforms with different functions, localizations, or activities. Alternative splicing is also critical for development and tissue-specific gene expression, and its dysregulation contributes to many diseases.
What are the consensus sequences at splice sites?
The 5′ splice site consensus is AG|GURAGU, the branch point consensus is YNYURAY, the polypyrimidine tract is a U-rich sequence, and the 3′ splice site consensus is YAG|G. These sequences are recognized by U1 snRNA, U2 snRNA, U2AF65, and U2AF35, respectively.
How is splicing studied experimentally?
Splicing is studied using in vitro splicing assays with HeLa nuclear extract, RNA-seq to quantify exon inclusion levels genome-wide, CLIP to map RNA-binding protein sites, and cryo-EM to determine the structures of spliceosomal complexes. Genetic screens in yeast and human cells have also identified many splicing factors.
What diseases are caused by splicing defects?
Splicing defects cause or contribute to many diseases, including spinal muscular atrophy (caused by SMN1 mutations), retinitis pigmentosa (caused by mutations in PRPF31, PRPF8, and other splicing factors), and various cancers (caused by somatic mutations in SF3B1, U2AF1, SRSF2, and ZRSR2).
Key Takeaways
- Splicing removes introns and joins exons via two transesterification reactions, catalyzed by the spliceosome, a ribonucleoprotein machine composed of five snRNPs (U1, U2, U4, U5, U6).
- The chemistry of splicing involves a branch point adenosine attacking the 5′ splice site (step 1) and the 5′ exon attacking the 3′ splice site (step 2), producing a lariat intron intermediate.
- Spliceosome assembly is ordered and ATP-dependent: E → A → B → B* → C → P, with U6 snRNA providing the catalytic core after U4 is released.
- Splice site recognition depends on conserved consensus sequences: the 5′ splice site (AG|GURAGU), branch point (YNYURAY), polypyrimidine tract, and 3′ splice site (YAG|G), recognized by U1, U2, U2AF65, and U2AF35 respectively.
- Alternative splicing, regulated by SR proteins (enhancers) and hnRNPs (silencers), generates over 95% of human proteomic diversity and is controlled by ESEs, ESSs, ISEs, and ISSs.
- Splicing is studied via in vitro assays, RNA-seq (PSI values), CLIP, and cryo-EM, which have revealed the atomic structure of the catalytic center.
- Spliceosome dysfunction causes disease: SF3B1 and U2AF1 mutations drive cancers, SMN1 loss causes spinal muscular atrophy, and PRPF mutations cause retinitis pigmentosa.
Further Reading
- Wilkinson ME, Charenton C, Nagai K. RNA Splicing by the Spliceosome. Annual review of biochemistry. 2020. PubMed 31794245
- Will CL, Lührmann R. Spliceosome structure and function. Cold Spring Harbor perspectives in biology. 2011. PubMed 21441581
- Rogalska ME et al. Transcriptome-wide splicing network reveals specialized regulatory functions of the core spliceosome. Science (New York, N.Y.). 2024. PubMed 39480945
- Gahete MD et al. Dysregulation of splicing variants and spliceosome components in breast cancer. Endocrine-related cancer. 2022. PubMed 35728261
- Jenkins JL, Kielkopf CL. Splicing Factor Mutations in Myelodysplasias: Insights from Spliceosome Structures. Trends in genetics : TIG. 2017. PubMed 28372848
- Wahl MC, Will CL, Lührmann R. The spliceosome: design principles of a dynamic RNP machine. Cell. 2009. PubMed 19239890