Nucleosome Protein: Structure, Function, and Assembly

By Dr. Zubair Khalid, DVM, MS, PhD ·

Nucleosome Protein: Structure, Function, and Assembly

Introduction to Nucleosome Proteins

What Are Nucleosome Proteins?

Nucleosome proteins are the histone proteins that package eukaryotic DNA into nucleosomes, the fundamental repeating unit of chromatin. The term "nucleosome protein" is functionally synonymous with "histone protein," though it emphasizes the role these proteins play specifically within the nucleosome complex rather than their broader biochemical identity. Histones are small, highly basic proteins rich in lysine and arginine residues, which give them a strong positive charge at physiological pH. This positive charge enables electrostatic interactions with the negatively charged phosphate backbone of DNA.

The nucleosome itself consists of approximately 147 base pairs of DNA wrapped around a protein core, forming a structure often described as "beads on a string" when viewed under low-resolution microscopy. This packaging is not merely a storage solution; it is a dynamic regulatory platform that governs DNA accessibility for replication, transcription, repair, and recombination. Without nucleosome proteins, the roughly two meters of human genomic DNA could not fit into a nucleus that measures only about 6 micrometers in diameter.

The Nucleosome Core Particle

The nucleosome core particle (NCP) is the minimal repeating unit of chromatin. It comprises an octamer of core histone proteins—two copies each of H2A, H2B, H3, and H4—around which 147 base pairs of DNA are wrapped in 1.65 left-handed superhelical turns. The octamer forms a disk-like structure approximately 11 nanometers in diameter and 5.7 nanometers in height. The DNA enters and exits the nucleosome at defined positions, and the overall architecture is remarkably conserved across all eukaryotes, from yeast to humans.

The linker histone H1 binds to the DNA entry/exit site of the nucleosome, stabilizing the wrapping and facilitating higher-order chromatin folding. The complete unit—the core particle plus linker DNA and H1—is sometimes called a chromatosome, though in common usage "nucleosome" often refers to the core particle alone. For a more detailed architectural description, refer to the Nucleosome Structure resource.

Histone Protein Family

Core Histones (H2A, H2B, H3, H4)

The four core histones share a common structural motif called the histone fold, a conserved three-helix domain (α1-α2-α3) connected by two loop regions (L1 and L2). This domain mediates histone-histone interactions through a "handshake" arrangement, where the α2 helix of one histone pairs with the α1 and α3 helices of another. The histone fold is essential for the dimerization of H3 with H4 and H2A with H2B, and these dimers further assemble into the octamer.

H3 and H4 are among the most evolutionarily conserved proteins known. Human H3 and H4 share over 90% sequence identity with their yeast orthologs. H3 has a molecular weight of approximately 15.4 kDa and is 135 amino acids long. H4 is smaller at 11.3 kDa and 102 amino acids. Both possess long, unstructured N-terminal tails (30–40 amino acids) that protrude from the nucleosome surface and are subject to extensive post-translational modification.

H2A and H2B are slightly less conserved. H2A is approximately 14 kDa (129 amino acids), and H2B is approximately 13.8 kDa (125 amino acids). H2A is distinguished by a C-terminal docking domain that interacts with H3 within the octamer, and it also has a unique "acidic patch" on its surface—a cluster of negatively charged residues that serves as a binding site for various chromatin-associated proteins and histone chaperones.

The stoichiometry of the octamer is strictly maintained: two H3-H4 dimers form a tetramer through a four-helix bundle, and two H2A-H2B dimers flank this tetramer. The assembly order is hierarchical, as described in the Histone Nucleosome overview.

Linker Histone H1

Linker histone H1 is structurally distinct from the core histones. It has a tripartite structure: a short N-terminal domain, a central globular winged-helix domain, and a long, intrinsically disordered C-terminal tail rich in lysine. H1 binds to the nucleosome at the dyad axis, where the DNA enters and exits, and interacts with the linker DNA between nucleosomes. This binding stabilizes the nucleosome and promotes chromatin compaction into higher-order structures such as the 30-nanometer fiber.

H1 is present at approximately one molecule per nucleosome in most somatic cells, though its stoichiometry varies with cell type and metabolic state. Unlike core histones, H1 is not required for the basic wrapping of DNA around the octamer; rather, it modulates the degree of compaction and the dynamics of nucleosome sliding. There are multiple H1 variants (H1.0 through H1.5 in humans), each with slightly different affinities for chromatin and distinct expression patterns during development.

Histone Variants

Beyond the canonical histones, which are expressed primarily during S phase and incorporated into chromatin during DNA replication, cells express histone variants that are incorporated at specific genomic locations or under specific conditions. These variants alter nucleosome stability, DNA wrapping, and the surface available for protein interactions.

  • H3.3 differs from canonical H3 by only four amino acids but is incorporated into chromatin in a replication-independent manner, often at transcriptionally active loci and regulatory elements.
  • CENP-A is an H3 variant specific to centromeres, where it replaces H3 in the nucleosomes that form the kinetochore attachment site.
  • H2A.Z is an H2A variant that localizes to promoters and enhancers, where it destabilizes nucleosomes and facilitates transcription factor binding.
  • H2A.X is phosphorylated at serine 139 (γ-H2A.X) at sites of DNA double-strand breaks, serving as a recruitment signal for DNA repair factors.
  • macroH2A contains a large C-terminal macrodomain and is enriched on the inactive X chromosome in female mammals, where it contributes to transcriptional silencing.

The incorporation of variants is mediated by specific chaperone complexes and often occurs at the boundaries of nucleosome-free regions, influencing the local Nucleosome Model of chromatin organization.

Nucleosome Structure and DNA Wrapping

Histone Fold and Octamer Assembly

The assembly of the histone octamer proceeds through a defined pathway. First, H3 and H4 form a heterodimer via their histone folds. Two H3-H4 dimers then associate through a four-helix bundle formed by the α3 helices of the two H3 molecules, creating an H3₂-H4₂ tetramer. This tetramer has a central dyad axis of symmetry and provides the scaffold onto which DNA is initially wrapped.

Next, two H2A-H2B dimers dock onto the tetramer. Each H2A-H2B dimer interacts with the tetramer through the H2A docking domain, which contacts the H3 α2 helix, and through the H2B L1 loop, which contacts the H4 α2 helix. The resulting octamer has a molecular weight of approximately 108 kDa and a net charge of +164 at neutral pH, reflecting its high arginine and lysine content.

The octamer is not a rigid sphere; it has a distinctive "spool" shape with a central channel that accommodates the DNA superhelix. The surface of the octamer is decorated with positively charged residues that form discrete DNA-binding sites, each contacting the phosphate backbone at regular intervals along the wrapped DNA.

DNA-Histone Interactions

The 147 base pairs of DNA in the nucleosome are wrapped in a left-handed superhelix, making 1.65 turns around the octamer. The DNA is bent sharply at multiple points, with a mean bend radius of approximately 4.2 nanometers—far tighter than the persistence length of free DNA (about 50 nanometers). This extreme bending is stabilized by two classes of interactions:

  1. Electrostatic interactions between the DNA phosphate backbone and basic residues (lysine and arginine) on the histone surface. These occur every 10–11 base pairs, where the minor groove faces inward toward the octamer. Approximately 120 direct hydrogen bonds and over 350 water-mediated contacts stabilize the complex.
  1. Hydrophobic interactions and van der Waals contacts between the DNA sugar moieties and nonpolar histone residues, particularly at positions where the minor groove is compressed.

The histone tails are not resolved in most crystal structures because they are highly flexible, but they are known to emerge from the nucleosome at specific points and contact linker DNA, neighboring nucleosomes, or regulatory proteins. The N-terminal tails of H3 and H4, in particular, exit near the DNA entry/exit sites and can be crosslinked to the DNA, contributing to nucleosome stability.

The precise positioning of DNA on the octamer is determined by the sequence-dependent deformability of the DNA double helix. A-T-rich regions tend to be positioned where the minor groove faces inward, while G-C-rich regions face outward, because the former are more easily compressed. This sequence preference contributes to the phenomenon of nucleosome positioning, where certain genomic sequences favor nucleosome formation while others (such as poly(dA:dT) tracts) resist it.

Nucleosome Assembly and Disassembly

Histone Chaperones

Nucleosome assembly is a carefully orchestrated process that prevents the formation of nonspecific histone-DNA aggregates. Histone chaperones are proteins that bind histones, shield their positive charge, and deliver them to DNA in a controlled manner. They are classified by their histone specificity:

  • H3-H4 chaperones: The best characterized is CAF-1 (chromatin assembly factor 1), which deposits H3-H4 tetramers onto newly replicated DNA in a PCNA-dependent manner. HIRA (histone regulator A) deposits H3.3-H4 at transcriptionally active loci in a replication-independent manner. ASF1 (anti-silencing function 1) binds H3-H4 dimers and hands them off to CAF-1 or HIRA.
  • H2A-H2B chaperones: NAP1 (nucleosome assembly protein 1) and FACT (facilitates chromatin transcription) bind H2A-H2B dimers and deposit them onto pre-formed H3-H4 tetramers. FACT also plays a role in disassembling nucleosomes ahead of the transcription machinery and reassembling them behind it.

The assembly process proceeds in two steps. First, an H3-H4 tetramer is deposited onto DNA, forming a tetrasome (a nucleosome lacking H2A-H2B). Second, two H2A-H2B dimers are added to complete the octamer. This two-step pathway is conserved from yeast to humans and ensures that nucleosome assembly is coupled to DNA replication and repair.

Chromatin Remodeling Complexes

Nucleosomes are not static structures; they can be moved, ejected, or restructured by ATP-dependent chromatin remodeling complexes. These are large, multi-subunit enzymes that use the energy of ATP hydrolysis to alter histone-DNA interactions. All remodelers share a conserved ATPase domain of the SWI2/SNF2 family, but they are divided into four major families based on additional domains and biological functions:

  1. SWI/SNF family (e.g., yeast SWI/SNF, human BAF): Slides and ejects nucleosomes, generally promoting chromatin accessibility.
  2. ISWI family (e.g., yeast ISW1/ISW2, human ACF/RSF): Slides nucleosomes along DNA to generate regularly spaced arrays.
  3. CHD family (e.g., yeast Chd1, human Mi-2/NuRD): Slides nucleosomes and, in some cases, promotes histone variant exchange.
  4. INO80/SWR1 family: Remodels nucleosomes at specific loci, including the exchange of H2A for H2A.Z.

The mechanism of nucleosome sliding involves the ATPase domain translocating along the DNA, creating a DNA bulge that propagates around the octamer, effectively moving the histone core relative to the DNA sequence. This process, termed Nucleosome Sliding, occurs at rates of approximately 1–5 base pairs per second under saturating ATP conditions (1–2 mM ATP, 5 mM MgCl₂, 37°C).

Nucleosome disassembly occurs when the octamer is removed from DNA entirely. This requires the coordinated action of remodelers and chaperones: the remodeler destabilizes the histone-DNA contacts, and chaperones such as FACT or NAP1 capture the released H2A-H2B dimers and H3-H4 tetramers. Disassembly is essential for processes that require complete access to the underlying DNA sequence, such as transcription initiation at strong promoters and DNA repair.

Post-Translational Modifications of Nucleosome Proteins

Histone Acetylation and Deacetylation

Histone acetylation is the addition of an acetyl group to the ε-amino group of lysine residues, catalyzed by histone acetyltransferases (HATs). This modification neutralizes the positive charge of lysine, weakening the electrostatic interaction between the histone tail and DNA. The result is a more open chromatin structure that is permissive for transcription.

The major HAT families include:

  • GNAT family (e.g., Gcn5, PCAF): Acetylates H3 lysines 9 and 14.
  • MYST family (e.g., Tip60, MOF): Acetylates H4 lysine 16, among others.
  • p300/CBP: A large, multifunctional HAT that acetylates multiple sites on H3 and H4.

Acetylation is reversed by histone deacetylases (HDACs), which remove acetyl groups and restore the positive charge on lysine. HDACs are divided into four classes: class I (HDAC1–3, 8), class II (HDAC4–7, 9–10), class III (sirtuins SIRT1–7, which require NAD⁺), and class IV (HDAC11). The balance between HAT and HDAC activity is tightly regulated, and its disruption is implicated in many cancers. Inhibitors of HDACs, such as vorinostat and romidepsin, are approved anticancer drugs.

Histone Methylation and Demethylation

Histone methylation occurs on lysine and arginine residues and is catalyzed by histone methyltransferases (HMTs). Unlike acetylation, methylation does not change the charge of the residue; instead, it alters the hydrophobicity and hydrogen-bonding capacity of the side chain, creating binding sites for reader proteins.

Lysine residues can be mono-, di-, or trimethylated. The functional consequence depends on the specific residue and the degree of methylation:

  • H3K4me3 (trimethylation of H3 lysine 4) is associated with active gene promoters.
  • H3K36me3 marks the body of actively transcribed genes.
  • H3K9me3 and H3K27me3 are associated with transcriptional repression and heterochromatin.

Methylation is reversed by histone demethylases, which fall into two families: the flavin-dependent LSD1 (lysine-specific demethylase 1), which removes mono- and dimethyl groups, and the Fe(II)/α-ketoglutarate-dependent JmjC domain-containing demethylases (e.g., JHDM, UTX), which can remove trimethyl groups.

Histone Code Hypothesis

The histone code hypothesis proposes that combinations of post-translational modifications on histone tails act as a "code" that is read by effector proteins to direct specific downstream functions. For example, the combination of H3K4me3 and H3K27ac at a promoter recruits the transcription machinery, while H3K9me3 recruits heterochromatin protein 1 (HP1) to establish silent chromatin.

This hypothesis has been refined over time. It is now clear that individual modifications do not act independently; rather, they are interpreted in context, and the same modification can have different effects depending on the genomic location and the presence of other marks. The concept remains useful, however, for understanding how chromatin states are established and maintained.

The enzymes that add ("writers"), remove ("erasers"), and recognize ("readers") histone modifications are themselves targets of regulation. Readers typically contain conserved domains such as bromodomains (which bind acetyl-lysine) or chromodomains, PHD fingers, and Tudor domains (which bind methyl-lysine). The interplay between writers, erasers, and readers creates a dynamic regulatory network that responds to cellular signals.

Methods to Study Nucleosome Proteins

Structural Biology Techniques

X-ray crystallography has been the cornerstone of nucleosome structural biology. The first high-resolution crystal structure of the nucleosome core particle was solved at 2.8 Å resolution in 1997 by Karolin Luger and colleagues, using nucleosomes reconstituted from recombinant Xenopus histones and a 146-base-pair DNA fragment derived from human α-satellite DNA. Subsequent structures have been solved at resolutions approaching 1.8 Å, revealing the precise atomic contacts between histones and DNA.

Cryo-electron microscopy (cryo-EM) has become increasingly important for studying nucleosomes in complex with binding partners, such as chromatin remodelers, histone chaperones, and transcription factors. Single-particle cryo-EM can now achieve resolutions of 3–4 Å for nucleosome complexes, and it has the advantage of not requiring crystallization, which is often difficult for large, flexible complexes.

Hydrogen-deuterium exchange mass spectrometry (HDX-MS) and crosslinking mass spectrometry provide complementary information about nucleosome dynamics and protein-protein interactions in solution, where the structure may differ from the crystal lattice.

Genome-Wide Mapping of Nucleosomes

MNase-seq (micrococcal nuclease digestion followed by sequencing) is the standard method for mapping nucleosome positions genome-wide. The protocol involves:

  1. Crosslinking cells with 1% formaldehyde for 10 minutes at room temperature to preserve chromatin architecture.
  2. Digesting chromatin with micrococcal nuclease (MNase), which cleaves linker DNA preferentially over nucleosomal DNA. Typical digestion uses 0.1–1 U of MNase per microgram of chromatin for 5–15 minutes at 37°C.
  3. Purifying the protected DNA fragments (~147 base pairs) and sequencing them.
  4. Aligning the reads to the reference genome and identifying nucleosome positions based on the midpoint of the protected fragments.

ChIP-seq (chromatin immunoprecipitation followed by sequencing) is used to determine the genomic localization of specific histone modifications or histone variants. The procedure involves crosslinking, sonication to fragment chromatin to 200–600 base pairs, immunoprecipitation with an antibody against the modification of interest, reversal of crosslinks, and sequencing of the enriched DNA. ChIP-seq can identify the positions of H3K4me3 at promoters, H3K27me3 at silenced loci, or the distribution of H2A.Z across the genome.

ATAC-seq (assay for transposase-accessible chromatin) uses the Tn5 transposase to tag accessible chromatin regions, providing an indirect measure of nucleosome occupancy. Open chromatin (nucleosome-free regions) is preferentially tagged, and the resulting data reveal regulatory elements such as promoters and enhancers.

Common Misconceptions and Pitfalls

Nucleosome vs. Chromatin vs. Chromosome

A common source of confusion is the relationship between nucleosomes, chromatin, and chromosomes. A nucleosome is a discrete complex of eight histone proteins and ~147 base pairs of DNA. Chromatin is the higher-order structure formed by the linear array of nucleosomes and associated proteins; it is the physiological form of DNA in the nucleus. A chromosome is a single, continuous DNA molecule (with its associated chromatin proteins) that is organized and condensed during cell division. The hierarchy is: DNA → nucleosome → chromatin fiber → looped domains → chromosome.

Histones vs. Protamines

During spermatogenesis, most histones are replaced by protamines, small arginine-rich proteins that package DNA much more tightly than histones. Protamines are not nucleosome proteins; they do not form octamers, and the resulting DNA-protamine complex is not organized into nucleosomes. This distinction is important because protamine-bound DNA is transcriptionally inert and represents a fundamentally different packaging strategy. A small fraction of histones (about 5–15% in humans) is retained in sperm at specific loci, often at developmentally important genes, and this retention has functional significance for embryonic development.

Nucleosome vs. Protein

A nucleosome is not a protein; it is a nucleoprotein complex. The protein components are the histones, but the nucleosome also contains DNA. Referring to a nucleosome as a "protein" is imprecise and can lead to confusion about its composition and function. Similarly, the term "nucleosome protein" specifically refers to the histone components, not the entire complex.

Common Experimental Pitfalls

When studying nucleosomes, several technical pitfalls can lead to erroneous conclusions:

  • MNase overdigestion: Excessive MNase digestion can cause nucleosome sliding or the loss of unstable nucleosomes, leading to an overestimation of nucleosome-free regions. Titrating the enzyme and digestion time is essential.
  • Antibody cross-reactivity in ChIP: Many histone modification antibodies cross-react with related modifications or with unmodified histones. Validation using peptide arrays or knockout cells is critical.
  • Salt-dependent artifacts: Nucleosomes are unstable at high ionic strength. Purification or biochemical assays performed at NaCl concentrations above 300 mM can cause histone dissociation or octamer rearrangement.
  • Assuming all nucleosomes are identical: Nucleosomes containing histone variants or modified histones have different stabilities and properties. Treating all nucleosomes as equivalent can obscure biologically meaningful differences.

Summary and Key Takeaways

Nucleosome proteins—the core histones H2A, H2B, H3, and H4, plus the linker histone H1—are the fundamental protein components of eukaryotic chromatin. They package DNA into nucleosomes, regulate DNA accessibility, and serve as platforms for post-translational modifications that control gene expression. The assembly and disassembly of nucleosomes are mediated by histone chaperones and ATP-dependent remodelers, and the study of nucleosome structure and function relies on a combination of structural biology and genome-wide mapping techniques.

For further reading on the architectural details, consult the Nucleosome Definition and Nucleosome Chromatin entries. The relationship between nucleosome proteins and other DNA-binding factors, such as Z DNA Binding Protein 1, highlights the diversity of protein-DNA interactions in the nucleus.

Frequently Asked Questions

Is a nucleosome a protein?

No. A nucleosome is a nucleoprotein complex consisting of DNA and histone proteins. The protein components are the histones, but the nucleosome as a whole is not a protein. It is the fundamental repeating unit of chromatin.

What are nucleosome proteins?

Nucleosome proteins are the histone proteins that make up the protein core of the nucleosome. These include the core histones H2A, H2B, H3, and H4, which form the octamer around which DNA is wrapped, and the linker histone H1, which binds to the DNA entry/exit site.

Are histones the same as nucleosome proteins?

Yes, in practical terms, "histones" and "nucleosome proteins" refer to the same family of proteins. The term "nucleosome protein" emphasizes their functional role within the nucleosome, while "histone" is the broader biochemical name for this family of basic, DNA-packaging proteins.

How many proteins are in a nucleosome?

A canonical nucleosome core particle contains eight histone proteins: two copies each of H2A, H2B, H3, and H4. If the linker histone H1 is included, the total is nine proteins, though H1 is not part of the core particle.

What is the function of nucleosome proteins?

Nucleosome proteins package DNA into a compact form that fits within the nucleus, protect DNA from damage, and regulate access to the genetic information. They also serve as platforms for post-translational modifications that control gene expression, DNA replication, and DNA repair.

Are nucleosome proteins found in prokaryotes?

No. Nucleosome proteins are found in eukaryotes. Prokaryotes do not have histones or nucleosomes. Instead, their DNA is organized by other proteins, such as HU, H-NS, and SMC proteins, which do not form nucleosome structures. Some archaea do have histone-like proteins that form nucleosome-like structures, but these are distinct from eukaryotic nucleosomes.

What is the difference between a nucleosome and a protein?

A protein is a polymer of amino acids. A nucleosome is a complex of DNA and histone proteins. The nucleosome contains proteins, but it is not itself a protein. The distinction is important for understanding the composition and function of chromatin.

Key Takeaways

  • Nucleosome proteins are the histones H2A, H2B, H3, and H4, which form an octamer around which 147 base pairs of DNA are wrapped.
  • The linker histone H1 binds to the nucleosome at the DNA entry/exit site and stabilizes higher-order chromatin folding.
  • The histone fold domain mediates histone-histone interactions and is essential for octamer assembly.
  • Nucleosome assembly is a two-step process involving H3-H4 tetramer deposition followed by H2A-H2B dimer addition, facilitated by histone chaperones.
  • ATP-dependent chromatin remodelers slide, eject, or restructure nucleosomes to regulate DNA accessibility.
  • Post-translational modifications of histone tails—including acetylation, methylation, and phosphorylation—constitute a regulatory code that controls chromatin state and gene expression.
  • Nucleosome proteins are eukaryotic; prokaryotes use different DNA-packaging proteins, and protamines replace histones during sperm maturation.

Further Reading

  • Lobbia VR, Trueba Sanchez MC, van Ingen H. Beyond the Nucleosome: Nucleosome-Protein Interactions and Higher Order Chromatin Structure. Journal of molecular biology. 2021. PubMed 33460684
  • Campbell AM, Cotter RI. The molecular weight of nucleosome protein by laser light scattering. FEBS letters. 1976. PubMed 99206280759-4)
  • Malinina DK et al. Hmo1 Protein Affects the Nucleosome Structure and Supports the Nucleosome Reorganization Activity of Yeast FACT. Cells. 2022. PubMed 36230893
  • Mondal A, Kolomeisky AB. Role of Nucleosome Sliding in the Protein Target Search for Covered DNA Sites. The journal of physical chemistry letters. 2023. PubMed 37527481
  • Crippa MP, Alfonso PJ, Bustin M. Nucleosome core binding region of chromosomal protein HMG-17 acts as an independent functional domain. Journal of molecular biology. 1992. PubMed 145345590833-6)
  • Azzaz AM et al. Human heterochromatin protein 1α promotes nucleosome associations that drive chromatin condensation. The Journal of biological chemistry. 2014. PubMed 24415761

Related Clinical & Scientific Guides