Base Pairing: Rules, Mechanisms, and Biological Significance

By Dr. Zubair Khalid, DVM, MS, PhD ·

Base Pairing: Rules, Mechanisms, and Biological Significance

Introduction to Base Pairing

What Is Base Pairing?

Base pairing is the specific, non-covalent association between nitrogenous bases on complementary nucleic acid strands. In its most fundamental form, base pairing involves the formation of hydrogen bonds between a purine on one strand and a pyrimidine on the opposing strand. This interaction is the central organizing principle of DNA double-helix structure and underpins nearly every process in molecular biology, from faithful genome replication to the decoding of messenger RNA during translation.

The term "base pairing" refers specifically to the hydrogen-bonded interaction between bases, not to the covalent phosphodiester bonds that link nucleotides within a single strand. Each base pair is planar, and the planes of successive base pairs stack perpendicular to the helical axis. The specificity of base pairing—adenine pairs only with thymine (or uracil in RNA), and guanine pairs only with cytosine—arises from the precise arrangement of hydrogen bond donors and acceptors on each base's edge.

The Canonical Base Pairs

The standard, or canonical, base pairs are:

  • Adenine–Thymine (A-T): Held together by two hydrogen bonds.
  • Guanine–Cytosine (G-C): Held together by three hydrogen bonds.
  • Adenine–Uracil (A-U): Found in RNA, also held by two hydrogen bonds.

The G-C pair is thermodynamically more stable than A-T because it forms an additional hydrogen bond. This difference has measurable consequences: DNA regions rich in G-C pairs denature (melt) at higher temperatures than A-T-rich regions, a property exploited in polymerase chain reaction (PCR) primer design and in techniques such as DNA melting curve analysis.

Chemical Basis of Base Pairing

Hydrogen Bond Donors and Acceptors

Hydrogen bonding between bases requires a precise geometric arrangement of donor and acceptor groups. A hydrogen bond donor is a group containing an electronegative atom (typically N or O) covalently bonded to a hydrogen atom, such as an amino group (–NH₂) or a ring nitrogen bearing a hydrogen (–NH–). A hydrogen bond acceptor is an electronegative atom with a lone pair of electrons, such as a carbonyl oxygen (C=O) or a ring nitrogen without a hydrogen (–N=).

For adenine and thymine, the pairing involves:

  • Adenine N1 (acceptor) with thymine N3–H (donor)
  • Adenine C6–NH₂ (donor) with thymine C4=O (acceptor)

For guanine and cytosine:

  • Guanine C6=O (acceptor) with cytosine C4–NH₂ (donor)
  • Guanine N1–H (donor) with cytosine N3 (acceptor)
  • Guanine C2–NH₂ (donor) with cytosine C2=O (acceptor)

The donor–acceptor patterns are complementary: each purine presents a pattern that matches exactly the acceptor–donor pattern of its cognate pyrimidine. This complementarity is the chemical basis of the base pairing rules.

Purine-Pyrimidine Complementarity

The requirement that a purine (a fused two-ring system) pair with a pyrimidine (a single ring) is not arbitrary—it maintains a constant width of the double helix. A purine–purine pair would be too wide, and a pyrimidine–pyrimidine pair too narrow, to maintain the regular helical geometry observed in B-form DNA. The Watson-Crick geometry places the glycosidic bonds (the bonds connecting each base to its sugar) at a fixed distance apart, approximately 10.85 Å, regardless of which base pair is formed.

This geometric constraint, combined with hydrogen bond specificity, ensures that the only stable pairings under physiological conditions are A-T and G-C (or A-U in RNA). Alternative pairings, such as Hoogsteen base pairs, can form under certain conditions (e.g., in triple helices or in some damaged DNA structures), but they are not the standard geometry found in B-DNA.

Tautomerism—the reversible isomerization of a base between keto and enol (or amino and imino) forms—can transiently alter hydrogen bonding patterns. For example, if thymine adopts a rare enol tautomer, it can pair with guanine. Such rare tautomeric forms are estimated to occur at frequencies of roughly 10⁻⁴ to 10⁻⁵ per base per replication, contributing to spontaneous mutation rates. However, DNA repair systems and the kinetic proofreading activity of DNA polymerases reduce the effective error rate to approximately 10⁻⁹ to 10⁻¹⁰ per base pair per replication.

Base Pairing in DNA Double Helix

Antiparallel Strands and Base Stacking

In double-stranded DNA, the two strands run in opposite directions: one in the 5′→3′ direction and the other in the 3′→5′ direction. This antiparallel arrangement is essential for the hydrogen bonding geometry of the canonical base pairs. The base pairs themselves are nearly planar and are stacked on top of one another along the helix axis, separated by approximately 3.4 Å.

Base stacking—the van der Waals and hydrophobic interactions between the planar faces of adjacent base pairs—contributes significantly to the overall stability of the double helix. In fact, stacking interactions, not hydrogen bonds, are the dominant stabilizing force in aqueous solution. The hydrogen bonds between paired bases provide specificity, but the stacking interactions provide the bulk of the free energy of stabilization. This is why DNA duplex stability correlates with G-C content (G-C pairs stack more favorably than A-T pairs) even though the hydrogen bond difference alone (one bond) does not fully account for the observed thermodynamic differences.

The double helix undergoes structural transitions depending on environmental conditions. B-form DNA, the standard right-handed helix with approximately 10.5 base pairs per turn, is the predominant form under physiological conditions (approximately 150 mM monovalent salt, pH 7.4, 37°C). A-form DNA, also right-handed but with a wider, shorter helix (approximately 11 base pairs per turn), forms under conditions of low humidity or in RNA-DNA hybrids. Z-form DNA is a left-handed helix that can form in alternating purine-pyrimidine sequences under high salt or specific protein binding conditions. These structural variants are discussed further in the context of DNA Supercoiling, which describes how torsional stress affects helix geometry.

Grooves and Protein Recognition

The antiparallel arrangement of strands and the specific geometry of base pairs create two grooves on the surface of the double helix: the major groove and the minor groove. These grooves arise because the glycosidic bonds attaching bases to the sugar-phosphate backbone are not symmetrically positioned relative to the base pair.

The major groove is wider (approximately 22 Å in B-DNA) and exposes the edges of the bases in a pattern that is unique for each of the four base pairs. This allows sequence-specific DNA-binding proteins, such as transcription factors, to "read" the DNA sequence without unwinding the helix. For example, the helix-turn-helix motif proteins, including the lac repressor and the λ repressor, make specific contacts with base edges in the major groove.

The minor groove is narrower (approximately 12 Å in B-DNA) and presents a less information-rich pattern. However, some proteins, such as the TATA-box binding protein (TBP), bind primarily through minor groove contacts. TBP binds to the TATA box sequence (consensus TATAAA) in promoters, inducing a sharp bend in the DNA. Small molecules, such as the antibiotic netropsin and the fluorescent dye DAPI, also bind in the minor groove, preferentially at A-T-rich sequences.

The pattern of hydrogen bond donors and acceptors in the major groove is distinct for each base pair. For example, an A-T base pair presents a hydrogen bond acceptor (adenine N7), a hydrogen bond acceptor (adenine N6 is a donor in the free base but becomes an acceptor when paired), and a hydrogen bond acceptor (thymine O4) in the major groove. A G-C pair presents a different pattern: acceptor (guanine N7), acceptor (guanine O6), donor (cytosine N4). These patterns enable proteins to distinguish between A-T and G-C pairs with high fidelity.

Base Pairing in RNA Structures

RNA vs DNA Base Pairing

RNA differs from DNA in two fundamental chemical respects: the sugar is ribose rather than 2′-deoxyribose, and thymine is replaced by uracil. Uracil is chemically similar to thymine (it lacks the 5-methyl group) and pairs with adenine through the same two hydrogen bonds. Therefore, the canonical RNA base pairs are A-U and G-C.

The presence of the 2′-hydroxyl group on ribose has profound structural consequences. It makes RNA less stable than DNA under alkaline conditions (RNA is hydrolyzed rapidly at pH > 10), and it restricts the conformational flexibility of the sugar. As a result, RNA duplexes typically adopt the A-form helix rather than the B-form. A-form RNA helices have a narrower major groove and a wider minor groove compared to B-DNA, making the major groove largely inaccessible to protein side chains. Consequently, RNA-binding proteins often recognize RNA through minor groove and backbone contacts.

RNA molecules are typically single-stranded, but they fold into complex three-dimensional structures stabilized by intramolecular base pairing. The most common secondary structure elements are:

  • Hairpins: A stem-loop structure formed when a single strand folds back on itself, pairing complementary regions separated by a loop of unpaired nucleotides.
  • Bulges: Regions where one strand has extra nucleotides that cannot pair, causing a distortion in the helix.
  • Internal loops: Regions where both strands have unpaired nucleotides opposite each other.

These structures are critical for RNA function. Transfer RNA (tRNA) adopts a cloverleaf secondary structure with four stems and three loops, which folds into an L-shaped tertiary structure. The anticodon loop contains the three-nucleotide sequence that base pairs with the mRNA codon during translation. Ribosomal RNA (rRNA) is extensively base-paired and forms the structural core of the ribosome. Regulatory RNAs, such as riboswitches, use base pairing to sense metabolite concentrations and control gene expression.

Wobble Base Pairing

The wobble base pair is a non-canonical pairing that occurs primarily between the third nucleotide of a codon (the 3′ position) and the first nucleotide of the anticodon (the 5′ position, position 34 of tRNA). The most common wobble pair is G-U, which is held together by two hydrogen bonds. This pairing is geometrically distinct from Watson-Crick pairs but fits within the same overall helical framework.

The Wobble Base Hypothesis, proposed by Francis Crick in 1966, explains how a limited number of tRNAs can decode all 61 sense codons. The rules are:

  1. The first two codon positions pair strictly according to Watson-Crick rules.
  2. The third codon position can tolerate non-canonical pairing.

Specifically, the following wobble pairings are allowed at the third position:

Anticodon base (position 34)Codon base (position 3)
GU or C
UA or G
I (inosine)A, U, or C
AU only
CG only

Inosine, a deaminated adenine derivative, is found at the wobble position of many tRNAs and can pair with A, U, or C. This degeneracy means that a single tRNA can recognize up to three different codons. For example, tRNA^Ile with the anticodon IAU can decode the codons AUU, AUC, and AUA.

Wobble pairing is not limited to codon-anticodon interactions. G-U wobble pairs also occur in RNA secondary structures, where they contribute to structural stability. The G-U wobble pair is particularly common in rRNA and in the acceptor stems of tRNAs. It is important to note that wobble pairing is a property of RNA; it does not occur in standard DNA duplexes under physiological conditions.

Base Pairing Rules and Chargaff's Findings

Chargaff's Rules

In the late 1940s, Erwin Chargaff and his colleagues at Columbia University analyzed the base composition of DNA from various organisms using paper chromatography. Their findings, published between 1949 and 1952, established two empirical rules:

  1. Chargaff's first rule: In double-stranded DNA, the amount of adenine equals the amount of thymine (A = T), and the amount of guanine equals the amount of cytosine (G = C). Consequently, the total purine content equals the total pyrimidine content (A + G = T + C).
  1. Chargaff's second rule: The base composition of DNA varies between species. For example, human DNA is approximately 30% A, 30% T, 20% G, and 20% C, while the bacterium Streptomyces coelicolor has a G-C content of approximately 72%.

The first rule was immediately suggestive of a pairing relationship between A and T, and between G and C. However, Chargaff himself did not propose a structural model. The significance of his data was recognized by Watson and Crick, who used Chargaff's ratios as a key constraint in building their model of DNA structure.

Watson-Crick Model

In 1953, James Watson and Francis Crick proposed the double-helical model of DNA based on three lines of evidence:

  1. Chargaff's rules (A = T, G = C)
  2. X-ray diffraction data from Rosalind Franklin and Maurice Wilkins, which indicated a helical structure with a repeat of 34 Å and a diameter of approximately 20 Å
  3. The chemical knowledge that purines and pyrimidines can form hydrogen bonds

The Watson-Crick model specified:

  • Two antiparallel polynucleotide chains
  • A right-handed helix with approximately 10 base pairs per turn
  • The sugar-phosphate backbones on the outside, bases on the inside
  • Specific hydrogen bonding between A and T (two bonds) and between G and C (three bonds)
  • A constant helix diameter maintained by purine-pyrimidine pairing

The model explained Chargaff's rules directly: the equality of A and T, and of G and C, is a necessary consequence of specific base pairing. The model also suggested a mechanism for replication: each strand can serve as a template for the synthesis of its complement, a proposal that was confirmed by the Meselson-Stahl experiment in 1958.

Methods to Study Base Pairing

X-ray Crystallography and NMR

X-ray crystallography has been the primary method for determining the three-dimensional structure of DNA and RNA at atomic resolution. The first DNA fiber diffraction patterns, obtained by Franklin and Wilkins, revealed the helical parameters but could not resolve individual base pairs. High-resolution crystal structures of DNA oligonucleotides, first reported by Richard Dickerson and colleagues in 1981 for the dodecamer d(CGCGAATTCGCG), provided the first atomic-level view of B-DNA, including the precise geometry of A-T and G-C base pairs.

For RNA, crystallographic structures of tRNAs, ribozymes, and ribosomes have revealed the full repertoire of base pairing interactions, including non-canonical pairs. The ribosome structures, solved at resolutions of 2.4–3.5 Å, have shown how base pairing between mRNA codons and tRNA anticodons is monitored by the ribosome during translation.

Nuclear magnetic resonance (NMR) spectroscopy is complementary to crystallography. NMR can determine structures of nucleic acids in solution, which may differ from crystal structures. It is particularly useful for studying:

  • The dynamics of base pair opening and closing
  • The stability of individual base pairs (measured by imino proton exchange rates)
  • The structure of RNA hairpins and other secondary structure elements

NMR studies have shown that base pairs in DNA open transiently at rates of approximately 10–100 s⁻¹ at 37°C, with A-T pairs opening more frequently than G-C pairs.

UV Melting Curves and Thermodynamics

The thermal stability of nucleic acid duplexes is routinely measured by UV absorbance spectroscopy. When DNA is heated, the double helix denatures into single strands, a process accompanied by an increase in absorbance at 260 nm (the hyperchromic effect). The midpoint of the melting transition is called the melting temperature (Tm).

The Tm depends on:

  • G-C content: Each G-C pair contributes approximately 2–3°C more stability than an A-T pair.
  • Salt concentration: Higher ionic strength stabilizes the duplex by screening the repulsion between negatively charged phosphate groups. Typical conditions use 10–100 mM NaCl or 1–10 mM MgCl₂.
  • Duplex length: Longer duplexes have higher Tm values.
  • Sequence context: Nearest-neighbor interactions affect stability.

For a short oligonucleotide duplex, the Tm can be estimated by the Wallace rule: Tm = 2(A+T) + 4(G+C) for oligos shorter than 14 nucleotides. More accurate predictions use nearest-neighbor thermodynamic parameters, which account for the fact that the stability of a base pair depends on its neighboring base pairs. These parameters are tabulated for all 10 unique nearest-neighbor combinations and are used in primer design software.

Thermodynamic analysis of melting curves provides the enthalpy (ΔH) and entropy (ΔS) of duplex formation. For a typical 12-base-pair duplex, ΔH is approximately −80 to −100 kcal/mol and ΔS is approximately −220 to −280 cal/(mol·K). The free energy of formation at 37°C is typically −10 to −15 kcal/mol, indicating that duplex formation is favorable but readily reversible.

Computational methods, including molecular dynamics simulations and free energy calculations, complement experimental approaches. These methods can predict the structure and stability of nucleic acid duplexes from sequence alone, and they are widely used in the design of PCR primers, antisense oligonucleotides, and CRISPR guide RNAs.

Biological Significance of Base Pairing

Replication and Transcription

Base pairing is the mechanistic basis of all template-directed nucleic acid synthesis. During DNA replication, the enzyme DNA polymerase reads the template strand and incorporates complementary nucleotides into the nascent strand. The polymerase selects nucleotides based on the hydrogen bonding pattern of the template base: an A in the template directs incorporation of T, a G directs incorporation of C, and so on.

The fidelity of replication depends on two factors:

  1. Base selection: DNA polymerases discriminate between correct and incorrect nucleotides by a factor of approximately 10³–10⁴, based primarily on the geometric fit of the incoming nucleotide in the active site.
  2. Proofreading: Many DNA polymerases possess 3′→5′ exonuclease activity that removes mismatched nucleotides immediately after incorporation. This increases fidelity by an additional factor of approximately 10²–10³.

The combined fidelity of replication in E. coli is approximately 10⁻⁹ to 10⁻¹⁰ errors per base pair per generation. In humans, the error rate is similar, although the larger genome size means that each cell division introduces roughly 10–100 new mutations.

During transcription, RNA polymerase reads the template strand of DNA and synthesizes a complementary RNA strand. The rules are the same as in replication, except that uracil is incorporated opposite adenine. Transcription is less accurate than replication, with an error rate of approximately 10⁻⁴ to 10⁻⁵, because RNA polymerases lack proofreading activity (though some have a low-fidelity cleavage activity).

Mutations and Repair

Errors in base pairing during replication can lead to mutations. A transition is a substitution of one purine for another (A↔G) or one pyrimidine for another (C↔T). A transversion is a substitution of a purine for a pyrimidine or vice versa. Transitions are more common than transversions, partly because they can arise from tautomeric shifts that produce mispairings with similar geometry.

Mispairings that escape proofreading can become fixed as mutations in the next round of replication. For example, if a G is incorporated opposite a template T (a G-T mispair), the next replication will produce a G-C pair in one daughter molecule and an A-T pair in the other, resulting in a T→C transition in one lineage.

Cells have multiple DNA repair pathways that recognize and correct mispaired bases:

  • Mismatch repair (MMR): Recognizes and corrects replication errors, including base-base mismatches and small insertion-deletion loops. In E. coli, the MutS protein recognizes the mismatch, MutL recruits MutH, and the daughter strand (distinguished by its transient lack of methylation at GATC sites) is excised and resynthesized.
  • Base Excision Repair: Removes damaged bases, such as uracil (from cytosine deamination) or 8-oxoguanine (from oxidative damage), and replaces them with the correct nucleotides.
  • Nucleotide excision repair (NER): Removes bulky adducts, such as pyrimidine dimers caused by UV light, by excising a short oligonucleotide containing the damage.

The importance of these repair systems is underscored by the consequences of their dysfunction. Mutations in human mismatch repair genes (e.g., MLH1, MSH2) cause Lynch syndrome, a hereditary predisposition to colorectal and other cancers. Defects in NER cause xeroderma pigmentosum, a condition characterized by extreme sensitivity to UV light and a >1000-fold increased risk of skin cancer.

Base pairing also underlies newer genome engineering technologies. Base Editing uses a catalytically impaired Cas9 fused to a deaminase enzyme to convert one base pair to another without creating a double-strand break. For example, the adenine base editor converts A-T pairs to G-C pairs by deaminating adenine to inosine, which then pairs with cytosine during replication. Base Pair Substitution is the general term for these and other changes that replace one base pair with another.

Common Pitfalls and Misconceptions

Base Pairing vs Base Stacking

A frequent source of confusion is the distinction between base pairing (hydrogen bonding between bases on opposite strands) and base stacking (van der Waals and hydrophobic interactions between adjacent base pairs on the same strand). These are distinct phenomena with different energetic contributions.

Base pairing provides specificity: it ensures that A pairs with T and G pairs with C. Base stacking provides stability: it is the major contributor to the free energy of duplex formation. A common exam question asks which interaction is primarily responsible for holding the two strands together. The answer is that both contribute, but stacking is the dominant stabilizing force in aqueous solution. The hydrogen bonds in base pairs are largely offset by the cost of desolvating the polar groups, whereas stacking interactions are entropically favorable due to the release of ordered water molecules from the hydrophobic faces of the bases.

RNA vs DNA Pairing Mistakes

Students often forget that RNA uses uracil instead of thymine. In RNA, the canonical pairs are A-U and G-C. Thymine is not used in RNA (except in some tRNAs, where it occurs as ribothymidine in the TΨC loop). Conversely, uracil is not used in DNA (except as a transient intermediate during base excision repair or as a result of cytosine deamination).

Another common error is assuming that the base pairing rules apply identically to RNA and DNA. While the Watson-Crick pairs are the same in both, RNA has additional pairing possibilities, including G-U wobble pairs and a wider repertoire of non-canonical pairs in structured RNAs. These non-canonical pairs are essential for RNA folding and function but are not part of the standard DNA base pairing rules.

Misapplying Chargaff's Rules

Chargaff's rules apply to double-stranded DNA. They do not apply to:

  • Single-stranded DNA (e.g., viral genomes, PCR products)
  • RNA (which is typically single-stranded, though double-stranded RNA viruses exist)
  • DNA-RNA hybrids (where A pairs with U, not T)

A common misconception is that Chargaff's rules imply that the base composition of all organisms is the same. In fact, Chargaff's second rule states the opposite: base composition varies widely between species. The G-C content of bacterial genomes ranges from approximately 25% (Mycoplasma genitalium) to approximately 75% (Streptomyces coelicolor). Humans have a G-C content of approximately 41%.

Confusing the Number of Hydrogen Bonds

A-T pairs have two hydrogen bonds; G-C pairs have three. This is often tested in exams, and students sometimes reverse the numbers. A useful mnemonic: G and C are both "curvy" letters and form the stronger pair with three bonds. The additional hydrogen bond in G-C pairs explains why G-C-rich DNA has a higher melting temperature.

Assuming Base Pairing Is Covalent

Base pairing is non-covalent. The hydrogen bonds that hold base pairs together are individually weak (approximately 1–3 kcal/mol each) compared to covalent bonds (approximately 80–100 kcal/mol). This weakness is essential for biology: it allows the two strands of DNA to be separated during replication and transcription without breaking covalent bonds. The enzyme helicase uses ATP hydrolysis to actively separate the strands, but the energy required is modest because the hydrogen bonds are weak and the stacking interactions are disrupted progressively.

Summary and Study Tips

Key Takeaways

  • Base pairing is the specific hydrogen bonding between purines and pyrimidines: A-T (two bonds), G-C (three bonds), and A-U (two bonds) in RNA.
  • The specificity of base pairing arises from the complementary arrangement of hydrogen bond donors and acceptors on each base.
  • Base stacking, not hydrogen bonding, is the dominant stabilizing force in the DNA double helix.
  • The antiparallel arrangement of DNA strands and the constant width of the helix are consequences of purine-pyrimidine complementarity.
  • Chargaff's rules (A = T, G = C) are a direct consequence of base pairing and were critical for the Watson-Crick model.
  • RNA uses uracil instead of thymine and can form wobble base pairs (especially G-U) that expand the coding capacity of the genetic code.
  • Base pairing is the mechanistic basis of DNA replication, transcription, and translation, and errors in base pairing are corrected by multiple DNA repair pathways.

Exam Preparation Tips

  1. Draw the base pairs: Practice drawing A-T and G-C pairs, labeling the hydrogen bond donors and acceptors. This will help you remember the number of hydrogen bonds and the chemical basis of specificity.
  1. Understand the energetics: Know that stacking interactions dominate duplex stability, and that G-C content correlates with melting temperature.
  1. Compare DNA and RNA: Make a table comparing DNA and RNA base pairing, including the sugars, the bases, and the types of helices formed.
  1. Connect to processes: For each process (replication, transcription, translation, repair), ask: "How does base pairing make this process possible?"
  1. Work through problems: Calculate the base composition of a DNA molecule given partial information (e.g., if A = 30%, then T = 30%, and G = C = 20%). Determine the Tm of an oligonucleotide using the Wallace rule.
  1. Know the exceptions: Be aware that wobble pairing occurs in RNA, that Hoogsteen pairs exist in special contexts, and that Chargaff's rules apply only to double-stranded DNA.

Frequently Asked Questions

What is base pairing?

Base pairing is the specific hydrogen bonding between nitrogenous bases on complementary nucleic acid strands. In DNA, adenine pairs with thymine (A-T) and guanine pairs with cytosine (G-C). In RNA, adenine pairs with uracil (A-U). Base pairing provides the specificity that allows genetic information to be stored, copied, and expressed.

What are the base pairing rules?

The base pairing rules state that adenine pairs with thymine (or uracil in RNA) through two hydrogen bonds, and guanine pairs with cytosine through three hydrogen bonds. These rules apply to double-stranded nucleic acids and are the basis of the complementarity between the two strands of DNA.

What are the types of base pairing?

The main types are:

  • Watson-Crick base pairing: The standard A-T, G-C, and A-U pairs found in double-stranded DNA and RNA.
  • Wobble base pairing: Non-canonical pairing, most commonly G-U, that occurs at the third position of codons during translation.
  • Hoogsteen base pairing: An alternative geometry in which the purine base is rotated 180° relative to the Watson-Crick orientation. Hoogsteen pairs occur in triple helices and some damaged DNA structures.

How does base pairing work?

Base pairing works through the formation of hydrogen bonds between complementary bases. Each base has a specific pattern of hydrogen bond donors (groups that donate a hydrogen atom) and acceptors (groups that accept a hydrogen atom). The patterns of adenine and thymine are complementary, as are those of guanine and cytosine. The geometric fit of the bases, combined with the hydrogen bonding pattern, ensures that only the correct pairs form stably.

What does base pairing mean in DNA replication?

During DNA replication, the two strands of the double helix are separated, and each strand serves as a template for the synthesis of a new complementary strand. DNA polymerase reads the template strand and incorporates nucleotides that base pair with the template: A pairs with T, and G pairs with C. This process is called semiconservative replication because each daughter molecule contains one original strand and one newly synthesized strand.

What is an example of base pairing?

A simple example: if one DNA strand has the sequence 5′-ATGC-3′, the complementary strand has the sequence 3′-TACG-5′. The A in the first strand pairs with T in the second, the T pairs with A, the G pairs with C, and the C pairs with G. In RNA, the same template would produce the sequence 3′-UACG-5′.

Why is base pairing important?

Base pairing is essential for all of molecular biology. It provides the structural basis of the DNA double helix, enables the faithful replication of genetic information, allows the transcription of DNA into RNA, and permits the decoding of mRNA during translation. Errors in base pairing can lead to mutations, which are the raw material of evolution but also the cause of many diseases, including cancer. Understanding base pairing is fundamental to molecular biology, genetics, and biotechnology.

Further Reading

  • Crooke ST et al. RNA-Targeted Therapeutics. Cell metabolism. 2018. PubMed 29617640
  • Dowdy SF et al. Delivery of RNA Therapeutics: The Great Endosomal Escape!. Nucleic acid therapeutics. 2022. PubMed 35612432
  • Hinnebusch AG. The scanning mechanism of eukaryotic translation initiation. Annual review of biochemistry. 2014. PubMed 24499181
  • Raina M et al. Dual-Function RNAs. Microbiology spectrum. 2018. PubMed 30191807
  • Delaunay S, Helm M, Frye M. RNA modifications in physiology and disease: towards clinical applications. Nature reviews. Genetics. 2024. PubMed 37714958
  • Cech TR, Steitz JA. The noncoding RNA revolution-trashing old rules to forge new ones. Cell. 2014. PubMed 24679528

Related Clinical & Scientific Guides