Introns and Exons: How Genes Are Split and Spliced

By Dr. Zubair Khalid, DVM, MS, PhD ·

Introns and Exons: How Genes Are Split and Spliced

The central dogma of molecular biology—DNA makes RNA makes protein—is often taught as if a gene were a continuous, unbroken stretch of DNA. In reality, for most genes in complex organisms, this is a simplification. The coding information is fragmented. A gene is interrupted by long stretches of non-coding sequence that must be removed before the final product is made. These interruptions are called introns, and the retained coding segments are called exons. Understanding how these pieces fit together and how they are processed is fundamental to understanding gene expression.

What Are Introns and Exons?

A gene is a segment of DNA that contains the instructions for making a functional product, typically a protein or a functional RNA molecule. In eukaryotic organisms—those with a nucleus, such as plants, animals, and fungi—the coding sequence of most genes is not continuous. Instead, it is split into alternating segments.

Exons are the segments of a gene that are retained in the final mature RNA molecule. For protein-coding genes, exons contain the sequences that specify the amino acid sequence of the protein. They are the "expressed" regions, which is the origin of their name (from "expressed region").

Introns are the segments of a gene that are transcribed into RNA but are removed during RNA processing and are not present in the final mature RNA molecule. The name comes from "intragenic region"—they are regions within a gene that interrupt the coding sequence. They are also sometimes called intervening sequences.

To use an analogy, consider a sentence written with extra words inserted between the letters of meaningful words: "Th*xyz*e cat s*abc*at on the m*def*at." The meaningful letters are the exons; the nonsense insertions are the introns. To read the sentence, you must remove the insertions and join the meaningful letters together.

A critical point is that introns are transcribed. RNA polymerase reads through the entire gene, including both exons and introns, producing a long precursor messenger RNA (pre-mRNA). The introns are then removed from this RNA molecule, not from the DNA. The DNA sequence of the intron remains in the genome, but the RNA copy of it is destroyed during the splicing process.

The Discovery of Split Genes

For decades after the structure of DNA was solved in 1953, the prevailing assumption was that genes were continuous. The genetic code was deciphered in the 1960s as a series of triplets (codons), each specifying an amino acid. It seemed logical that a gene encoding a protein of 300 amino acids would be a continuous stretch of 900 nucleotides.

This assumption was shattered in 1977 by two independent research groups working on adenovirus, a virus that infects human cells. Phillip Sharp at MIT and Richard Roberts at Cold Spring Harbor Laboratory both made the same startling discovery: the genes of adenovirus were not contiguous.

The key experiment involved hybridizing adenovirus messenger RNA (mRNA) to the viral DNA. When mRNA is hybridized to its complementary DNA template, it forms a stable double-stranded region. Under an electron microscope, this appears as a continuous DNA-RNA hybrid. However, Sharp and Roberts observed something unexpected: loops of single-stranded DNA that were not hybridized to the mRNA.

These loops represented regions of the DNA that were transcribed into the pre-mRNA but were absent from the mature mRNA. The mRNA was complementary to only certain segments of the DNA, meaning the gene was split into pieces. The pieces that matched the mRNA were the exons; the loops that did not match were the introns.

This discovery was revolutionary. It meant that the genetic information for a single protein could be scattered across the genome, and that the RNA transcript had to be processed to bring the pieces together. For this work, Sharp and Roberts were awarded the Nobel Prize in Physiology or Medicine in 1993. The discovery of split genes fundamentally changed our understanding of gene structure and laid the foundation for the field of RNA processing.

The Structure of a Gene: Exons and Introns

A typical eukaryotic protein-coding gene has a complex architecture. The gene is not just a string of exons and introns; it also contains regulatory sequences that control when, where, and how much the gene is expressed.

The gene begins with a promoter, a region of DNA upstream of the transcription start site. The promoter contains binding sites for Transcription Factor proteins that recruit RNA polymerase II, the enzyme that transcribes protein-coding genes. The promoter is not transcribed itself; it is a regulatory sequence.

The transcribed region begins at the transcription start site (the +1 position). The first exon often contains a 5' untranslated region (5' UTR)—a sequence that is transcribed and retained in the mRNA but is not translated into protein. This region contains signals for ribosome binding and mRNA stability.

The coding sequence is then interrupted by introns. A typical human gene has an average of about 8-9 introns, but some genes have many more. The human titina gene (TTN), which encodes a muscle protein, has 363 exons and 362 introns, making it one of the largest genes in the human genome.

The final exon contains a 3' untranslated region (3' UTR) followed by a signal for polyadenylation. The polyadenylation signal (typically the sequence AAUAAA) directs the addition of a poly(A) tail to the mRNA, which is important for stability and translation.

The boundaries between exons and introns are marked by conserved sequences. At the 5' end of an intron (the splice donor site), there is almost always the dinucleotide GU. At the 3' end (the splice acceptor site), there is almost always the dinucleotide AG. A branch point sequence, containing an adenine nucleotide, is located 18-40 nucleotides upstream of the 3' splice site. These sequences are essential for the splicing machinery to recognize where to cut and join.

How Introns Are Removed: RNA Splicing

RNA splicing is the process by which introns are removed from the pre-mRNA and exons are joined together. This is a precise, multi-step process that must occur with high fidelity—a single nucleotide error can shift the reading frame and produce a nonfunctional protein.

The chemistry of splicing involves two transesterification reactions. In the first reaction, the 2' hydroxyl group of the adenine at the branch point attacks the phosphate at the 5' splice site, cleaving the RNA and forming a lariat structure (a loop with a tail). In the second reaction, the 3' hydroxyl group of the released 5' exon attacks the phosphate at the 3' splice site, joining the two exons and releasing the intron as a lariat, which is subsequently degraded.

The Spliceosome

The vast majority of introns in eukaryotic genes are removed by a large, dynamic ribonucleoprotein complex called the spliceosome. The spliceosome is composed of five small nuclear ribonucleoproteins (snRNPs), named U1, U2, U4, U5, and U6, along with numerous associated proteins. Each snRNP contains a small nuclear RNA (snRNA) and several proteins.

The assembly of the spliceosome is a highly ordered, stepwise process:

  1. Recognition of the 5' splice site: The U1 snRNP binds to the 5' splice site of the intron via base pairing between its U1 snRNA and the GU sequence at the intron's start.
  2. Recognition of the branch point: The U2 snRNP binds to the branch point sequence, with the U2 snRNA base pairing with the branch point. The conserved adenine is bulged out, making it available for the first transesterification reaction.
  3. Formation of the tri-snRNP complex: The U4/U6 and U5 snRNPs join the complex as a single unit. U4 and U6 are extensively base-paired to each other, and U5 interacts with the 5' exon.
  4. Rearrangement and catalysis: A major conformational rearrangement occurs. U1 is released, U4 is released, and U6 base pairs with U2 and the 5' splice site. This rearrangement positions the reactive groups for the first transesterification reaction, which occurs at this point.
  5. First transesterification: The 2' hydroxyl of the branch point adenine attacks the 5' splice site, cleaving the RNA and forming the lariat.
  6. Second transesterification: The 3' hydroxyl of the 5' exon attacks the 3' splice site, joining the exons and releasing the lariat intron.
  7. Disassembly: The spliceosome disassembles, releasing the mature mRNA and the lariat intron. The snRNPs are recycled for another round of splicing.

This entire process is ATP-dependent. The spliceosome uses energy from ATP hydrolysis to drive the conformational changes required for each step. The spliceosome is one of the most complex molecular machines in the cell, comparable in complexity to the ribosome.

Self-Splicing Introns

Not all introns require the spliceosome. Some introns are self-splicing, meaning they can catalyze their own removal from the RNA molecule without the aid of proteins. These are found in certain genes in bacteria, mitochondria, chloroplasts, and some viruses.

There are two main classes of self-splicing introns:

Group I introns use an external guanosine cofactor. The 3' hydroxyl of the free guanosine attacks the 5' splice site, adding the guanosine to the intron. The 3' hydroxyl of the 5' exon then attacks the 3' splice site, joining the exons and releasing the intron with the guanosine attached to its 5' end.

Group II introns use an internal adenine, similar to the branch point in spliceosomal splicing, to form a lariat structure. The mechanism is essentially identical to that of the spliceosome, which has led to the hypothesis that the spliceosome evolved from group II introns. The snRNAs of the spliceosome may be the descendants of the catalytic RNA of ancient group II introns.

The discovery of self-splicing introns by Thomas Cech and Sidney Altman (who won the Nobel Prize in 1989) provided the first evidence that RNA can act as a catalyst, a finding that has profound implications for the origin of life.

Why Do Introns Exist?

The existence of introns raises an obvious question: why would an organism maintain large stretches of non-coding DNA within its genes, requiring a complex molecular machine to remove them? The answer is that introns provide several significant evolutionary and functional advantages.

Alternative Splicing

The most important function of introns is that they enable alternative splicing. This is a process by which different combinations of exons are joined together to produce different mRNA molecules from the same gene. By including or excluding different exons, a single gene can produce multiple protein isoforms with different functions.

For example, the Drosophila gene Dscam (Down syndrome cell adhesion molecule) has 95 exons arranged in clusters. Through alternative splicing, this single gene can theoretically produce over 38,000 different mRNA isoforms, each encoding a different protein. These proteins are involved in neural wiring, and the diversity generated by alternative splicing allows each neuron to have a unique identity.

In humans, it is estimated that over 95% of multi-exon genes undergo alternative splicing. This greatly expands the coding capacity of the genome. The human genome contains roughly 20,000 protein-coding genes, but it is estimated to produce over 100,000 different proteins, largely due to alternative splicing.

There are several modes of alternative splicing:

  1. Exon skipping: An entire exon is excluded from the mature mRNA. This is the most common mode in humans.
  2. Alternative 5' splice site: A different 5' splice site is used, resulting in a longer or shorter exon.
  3. Alternative 3' splice site: A different 3' splice site is used.
  4. Intron retention: An intron is retained in the mature mRNA. This can lead to a truncated protein if the intron contains a stop codon.
  5. Mutually exclusive exons: Only one of two or more exons is included.

Alternative splicing is regulated by proteins that bind to splicing enhancers and silencers—sequences within exons and introns that promote or inhibit spliceosome assembly at nearby splice sites. This regulation allows different cell types to produce different protein isoforms from the same gene.

Regulatory Roles

Introns are not merely passive spacers between exons. They contain numerous regulatory elements that influence gene expression at multiple levels.

Many introns contain enhancers—DNA sequences that bind activator proteins and increase the rate of transcription from the promoter. These intronic enhancers can be located thousands of base pairs away from the promoter and can act over long distances. For example, the immunoglobulin heavy chain gene contains an enhancer within its second intron that is critical for high-level expression in B cells.

Introns also contain binding sites for regulatory RNAs. MicroRNAs (miRNAs) and other small regulatory RNAs can bind to sequences within introns and regulate mRNA stability or translation. Some introns themselves encode functional RNAs, such as small nucleolar RNAs (snoRNAs) that are involved in ribosome biogenesis.

The process of splicing itself can regulate gene expression. Nonsense-mediated decay (NMD) is a quality control mechanism that degrades mRNAs containing premature stop codons. If a splicing error introduces a stop codon, the mRNA is targeted for degradation. This provides a mechanism for regulating gene expression—by controlling the efficiency of splicing, the cell can control the amount of functional mRNA produced.

Increasing Genetic Diversity

Introns increase genetic diversity in several ways. First, as discussed, alternative splicing generates protein diversity. Second, introns provide a substrate for recombination—the exchange of genetic material between chromosomes. Because introns are under less selective pressure than exons, they accumulate mutations more rapidly. This variation can lead to new splice sites, new exons, and new regulatory elements over evolutionary time.

The process by which new exons arise from intronic sequences is called exonization. Mutations can create new splice sites within an intron, causing a portion of the intron to be included in the mature mRNA. If this new exon encodes a functional protein domain, it may be preserved by natural selection.

Introns also facilitate exon shuffling, a process by which exons from different genes are combined to create new genes. Because exons often correspond to protein domains—functional units of a protein—recombination within introns can shuffle these domains between genes, creating proteins with new combinations of functions. This is thought to have been a major force in the evolution of eukaryotic genomes.

Introns and Exons in Different Organisms

The distribution and characteristics of introns vary dramatically across the tree of life.

Prokaryotes (bacteria and archaea) generally lack introns in their protein-coding genes. Their genes are typically continuous, with no intervening sequences. This is partly because prokaryotes lack a nucleus, so transcription and translation are coupled—ribosomes begin translating the mRNA while it is still being transcribed. There is no time for splicing to occur. Some prokaryotes do have self-splicing introns in tRNA and rRNA genes, but these are rare.

Eukaryotes have introns in most of their protein-coding genes, but the density and size vary greatly between species:

OrganismApproximate Number of Introns per GeneAverage Intron Size
Yeast (Saccharomyces cerevisiae)~0.04~270 bp
Nematode (Caenorhabditis elegans)~4~300 bp
Fruit fly (Drosophila melanogaster)~4~1,500 bp
Human (Homo sapiens)~8-9~3,500 bp
Thale cress (Arabidopsis thaliana)~5~160 bp

Yeast is notable for having very few introns—only about 5% of its genes contain introns, and most of those have only one. This makes yeast an attractive model organism for studying splicing, as the system is simpler.

The size of introns also varies enormously. Some human introns are only a few dozen base pairs, while others are over 100,000 base pairs. The largest known intron is in the human dystrophin gene, which spans over 2 million base pairs and contains 79 exons. The introns in this gene are so large that transcription takes over 16 hours to complete.

Plants have introns that are generally smaller than those in animals, but they have a similar density. The model plant Arabidopsis thaliana has a compact genome with small introns, while some plant genomes, such as those of wheat and maize, have much larger introns.

The evolutionary origin of introns is still debated. The introns-early hypothesis proposes that introns were present in the common ancestor of all life and were lost in prokaryotes. The introns-late hypothesis proposes that introns were inserted into eukaryotic genes after the divergence from prokaryotes. The current evidence suggests a middle ground: some introns are ancient, while others have been inserted more recently. The spliceosomal machinery itself appears to have evolved from group II introns, which were likely present in the ancestor of eukaryotes.

How Scientists Study Introns and Exons

Identifying introns and exons and understanding splicing is a major focus of genomics and molecular biology. Several complementary approaches are used.

cDNA sequencing is a classic method. Messenger RNA is isolated from cells and converted into complementary DNA (cDNA) using the enzyme reverse transcriptase. Because the mRNA has already been spliced, the cDNA represents only the exons. By comparing the cDNA sequence to the genomic DNA sequence, scientists can identify intron-exon boundaries. The cDNA sequence will match the genomic sequence in exons but will skip over the introns.

RNA sequencing (RNA-seq) is a high-throughput version of this approach. Total RNA is isolated, converted to cDNA, and sequenced in bulk. The resulting reads are aligned to the reference genome. Reads that span exon-exon junctions—called junction reads—provide direct evidence of splicing. The number of junction reads at each splice site can be quantified to measure splicing efficiency and to identify alternatively spliced isoforms.

Genome annotation is the process of identifying genes and their structures within a genome sequence. This is done using computational tools that look for signatures of genes, such as open reading frames, splice site consensus sequences, and homology to known genes from other species. However, computational prediction is not perfect, and experimental evidence from cDNA and RNA-seq is essential for accurate annotation.

Splice site prediction is a specific computational challenge. The splice donor (GU) and acceptor (AG) dinucleotides are necessary but not sufficient for splicing—they occur frequently in the genome by chance. Computational tools use more complex models, such as position weight matrices and machine learning, to predict which GU/AG pairs are functional splice sites. These tools are trained on known splice sites and can achieve high accuracy, but they are not perfect.

Minigene reporter assays are used to study splicing in the laboratory. A fragment of a gene containing the exons and introns of interest is cloned into a reporter plasmid, which is then transfected into cells. The cells process the pre-mRNA, and the resulting spliced mRNA can be analyzed by RT-PCR (reverse transcription followed by polymerase chain reaction). This allows researchers to test the effects of mutations on splicing.

In vitro splicing assays use nuclear extracts from cells to study splicing in a test tube. A radiolabeled pre-mRNA substrate is incubated with the extract, and the products are analyzed by gel electrophoresis. This approach has been used to dissect the biochemical requirements of splicing, such as the ATP requirement and the role of specific snRNPs.

Common Misconceptions and Pitfalls

Several misconceptions about introns and exons are common, even among advanced students.

Misconception 1: Introns are "junk" DNA with no function. This is incorrect. While introns do not encode protein, they contain regulatory elements, encode functional RNAs, and enable alternative splicing. The term "junk DNA" is misleading and is falling out of favor. A better term is "non-coding DNA," which is neutral about function.

Misconception 2: Exons always code for proteins. This is also incorrect. Exons are defined as sequences that are retained in the mature RNA, but not all retained sequences are translated. The 5' and 3' untranslated regions are exonic but do not code for protein. Additionally, some genes produce non-coding RNAs, such as tRNAs and rRNAs, which have exons and introns but no protein product.

Misconception 3: Introns are removed during transcription. Introns are transcribed into the pre-mRNA and are removed after transcription, during RNA processing. The entire gene, including introns, is transcribed into RNA. The introns are then excised from the RNA molecule.

Misconception 4: Splicing is a simple cut-and-paste process. Splicing is a highly regulated, multi-step process involving dozens of proteins and RNA molecules. It is not a simple enzymatic reaction but a dynamic assembly and disassembly of a large molecular machine.

Misconception 5: All introns are the same. Introns vary enormously in size, sequence, and function. Some are self-splicing, some require the spliceosome, and some are never spliced at all (in the case of intron retention).

Misconception 6: Prokaryotes have no introns. While most prokaryotic protein-coding genes lack introns, some prokaryotes have self-splicing introns in tRNA and rRNA genes. Additionally, some archaeal genes have introns that are removed by a different mechanism involving a protein enzyme called a bulge-helix-bulge endonuclease.

Pitfall in studying splicing: Assuming that the genomic sequence is sufficient to predict splicing. Splicing is regulated by many factors, including the cellular context, the presence of splicing factors, and the rate of transcription. The same pre-mRNA can be spliced differently in different cell types or under different conditions. This makes it difficult to predict splicing outcomes from sequence alone.

Frequently Asked Questions

What are introns and exons?

Introns are non-coding sequences within a gene that are transcribed into RNA but are removed during RNA processing. Exons are the sequences that are retained in the mature RNA. For protein-coding genes, exons contain the information for the amino acid sequence of the protein.

Can introns be exons?

Yes, through a process called alternative splicing, a sequence that is an intron in one context can be an exon in another. This can happen through intron retention, where an intron is not removed and becomes part of the mature mRNA. It can also happen over evolutionary time, where mutations create new splice sites that cause a portion of an intron to be included as an exon.

What is the function of introns and exons?

Exons contain the coding information for proteins and functional RNAs. Introns have multiple functions: they enable alternative splicing, contain regulatory elements that control gene expression, encode functional RNAs, and provide a substrate for recombination and exon shuffling.

Are introns removed during transcription?

No. Introns are transcribed into the pre-mRNA along with exons. They are removed after transcription is complete, during a process called RNA splicing. The entire gene, including introns, is copied into RNA by RNA polymerase.

Do prokaryotes have introns?

Most prokaryotic protein-coding genes lack introns. However, some prokaryotes have self-splicing introns in tRNA and rRNA genes. These introns are removed by RNA catalysis, not by a spliceosome.

What is the difference between introns and exons?

The key difference is that exons are retained in the mature RNA molecule, while introns are removed. Exons typically contain coding sequence, while introns are non-coding. Introns are also generally longer than exons in higher eukaryotes, and they are removed by the spliceosome.

Why are introns called 'junk DNA'?

The term "junk DNA" was coined in the 1970s to describe non-coding DNA, including introns, that appeared to have no function. This term is now considered misleading. Introns have been shown to have numerous functions, including regulating gene expression and enabling alternative splicing. The term "non-coding DNA" is preferred.

Key Takeaways

  • Introns are non-coding sequences that interrupt genes; exons are the coding sequences retained in mature RNA.
  • The discovery of split genes by Sharp and Roberts in 1977 revolutionized molecular biology and earned them the Nobel Prize.
  • Introns are transcribed into pre-mRNA and removed by the spliceosome, a complex of five snRNPs, through two transesterification reactions.
  • Some introns are self-splicing and can catalyze their own removal without proteins.
  • Introns enable alternative splicing, allowing a single gene to produce multiple protein isoforms.
  • Introns contain regulatory elements and encode functional RNAs, so they are not "junk."
  • Prokaryotes generally lack introns, while eukaryotes have them in most genes, with density and size varying greatly between species.
  • Scientists study introns and exons using cDNA sequencing, RNA-seq, genome annotation, and in vitro splicing assays.

Related Topics

Related Clinical & Scientific Guides