Chromosome Structure: From DNA to Condensed Chromosomes
By Dr. Zubair Khalid, DVM, MS, PhD ·

The genetic information of every living organism is stored in DNA, but the naked double helix is not the form in which this information is packaged, protected, and partitioned. In cells, DNA exists as part of a dynamic nucleoprotein complex called a chromosome. The structure of chromosomes is not a static, uniform entity; rather, it represents a hierarchy of folding states that changes with the cell cycle and with the transcriptional activity of the underlying DNA. Understanding chromosome structure requires tracing the path from the 2-nanometer-wide double helix to the roughly 1,400-nanometer-wide metaphase chromosome, a compaction ratio of nearly 10,000-fold. This article describes the molecular players and the architectural principles that achieve this remarkable feat of packaging.
What Is a Chromosome?
A chromosome is a discrete, physical unit of genetic material composed of a single, continuous DNA molecule associated with proteins that package and organize it. The word "chromosome" comes from the Greek chroma (color) and soma (body), reflecting the fact that these structures stain intensely with basic dyes and were first visualized as colored bodies in dividing cells in the late 19th century.
The fundamental distinction in chromosome organization lies between prokaryotes and eukaryotes. Prokaryotic organisms—bacteria and archaea—typically possess a single, circular chromosome that resides in the cytoplasm in a region called the nucleoid. This chromosome is not enclosed by a membrane, and its DNA is associated with a different set of proteins than eukaryotic chromosomes, including histone-like proteins such as HU and H-NS in Escherichia coli. The prokaryotic chromosome is compacted through supercoiling and the formation of large loop domains, but it lacks the regular, repeating nucleosome structure found in eukaryotes.
Eukaryotic chromosomes, by contrast, are linear DNA molecules housed within the membrane-bound nucleus. Each eukaryotic species has a characteristic chromosome number: humans have 46 (23 pairs), the fruit fly Drosophila melanogaster has 8, and the model plant Arabidopsis thaliana has 10. Eukaryotic chromosomes are associated with histones, the family of proteins that form the fundamental packaging unit, and they undergo dramatic condensation during cell division. The distinction between prokaryotic and eukaryotic chromosome organization is not merely a matter of scale; it reflects fundamentally different strategies for managing gene expression, DNA replication, and segregation. For a comparative overview, see DNA Chromosome Structure.
DNA Packaging: The Need for Condensation
The Double Helix
The starting point for understanding chromosome structure is the DNA double helix itself. Each strand of DNA is a polymer of nucleotides linked by phosphodiester bonds between the 5' phosphate of one nucleotide and the 3' hydroxyl of the next. The two strands run antiparallel and are held together by hydrogen bonds between complementary base pairs: adenine pairs with thymine (two hydrogen bonds), and guanine pairs with cytosine (three hydrogen bonds). The double helix has a diameter of approximately 2 nanometers and a helical repeat of about 10.5 base pairs per turn.
The genetic information is encoded in the linear sequence of bases along the strand. In humans, the haploid genome contains approximately 3.1 billion base pairs distributed across 23 chromosomes. If stretched end-to-end, the DNA in a single human cell would measure roughly 2 meters. This length must fit inside a nucleus that is typically 10 to 20 micrometers in diameter—a volume roughly one million times smaller than the volume the DNA would occupy if fully extended.
The Challenge of Fitting DNA into a Cell
The problem of fitting 2 meters of DNA into a 10-micrometer nucleus is not merely a geometric puzzle; it is a functional necessity. DNA must be compacted enough to fit within the cell, yet remain accessible to the molecular machines that read, copy, and repair it. RNA polymerases must access promoter sequences, replication origins must be recognized by initiator proteins, and DNA repair enzymes must reach damaged bases. The solution is a hierarchical packaging system that achieves compaction while preserving access.
The packaging problem is solved through multiple levels of organization, each building on the previous one. The first level wraps DNA around histone proteins to form nucleosomes. The second level folds nucleosome arrays into higher-order fibers. The third level organizes these fibers into loop domains attached to a protein scaffold. The final level, achieved only during cell division, condenses the entire structure into the classic X-shaped metaphase chromosome. Each level of packaging represents a trade-off between compaction and accessibility, and the cell dynamically regulates this balance. The principles of this hierarchy are detailed in Chromosome Structure Organization.
Nucleosomes: The First Level of Packaging
Histone Proteins
The first and most fundamental level of eukaryotic DNA packaging is the nucleosome, the repeating unit of chromatin. The nucleosome core particle consists of 147 base pairs of DNA wrapped in 1.65 left-handed superhelical turns around an octamer of histone proteins. This octamer contains two copies each of the four core histones: H2A, H2B, H3, and H4.
Histones are small, highly basic proteins rich in lysine and arginine residues. These positively charged amino acids interact electrostatically with the negatively charged phosphate backbone of DNA. Each core histone has a characteristic fold: a three-helix domain called the histone fold, which mediates histone-histone interactions within the octamer, and an unstructured N-terminal tail that extends outward from the nucleosome core. These tails are the targets of extensive post-translational modifications—acetylation, methylation, phosphorylation, and ubiquitination—that regulate chromatin structure and gene expression.
The assembly of the histone octamer is a stepwise process. An H3-H4 tetramer forms first, followed by the addition of two H2A-H2B dimers. The DNA wraps around this octamer in a specific register, with the minor groove facing inward at regular intervals. The wrapping is not uniform; certain positions along the DNA make more extensive contacts with the histone surface than others, creating a defined rotational positioning of the DNA on the nucleosome.
The Beads-on-a-String Model
When chromatin is gently lysed and examined under the electron microscope at low ionic strength, it appears as a series of beads connected by thin threads of DNA. This "beads-on-a-string" conformation is the 10-nanometer fiber, the first level of DNA compaction. Each bead is a nucleosome core particle, and the connecting thread is linker DNA—the segment between nucleosomes that is not wrapped around the histone octamer.
The length of linker DNA varies between species and between cell types. In most somatic cells, the average nucleosome repeat length is approximately 200 base pairs (147 base pairs of core DNA plus roughly 50 base pairs of linker DNA), but this can range from 165 base pairs in some fungi to 240 base pairs in sea urchin sperm. The positioning of nucleosomes along the DNA is not random; it is influenced by the underlying DNA sequence, by ATP-dependent chromatin remodelers that slide or eject nucleosomes, and by the presence of other DNA-binding proteins.
The 10-nanometer fiber achieves a compaction ratio of approximately 6-fold relative to naked DNA. This is far from sufficient to fit the genome into the nucleus, but it represents the fundamental repeating unit upon which all higher-order structures are built. The detailed architecture of this level is explored in Chromosome Histone Structure.
Chromatin Fibers and Higher-Order Folding
The 30-nm Fiber
The next level of packaging involves folding the beads-on-a-string array into a more compact fiber. Under conditions of moderate ionic strength, nucleosome arrays condense into a fiber with a diameter of approximately 30 nanometers. The precise structure of this 30-nm fiber has been debated for decades, with two main models proposed: the solenoid model and the zigzag model.
In the solenoid model, nucleosomes are arranged in a one-start helix with approximately six nucleosomes per turn, and the linker DNA bends smoothly between adjacent nucleosomes. In the zigzag model, the linker DNA crosses back and forth, and nucleosomes that are not adjacent in the linear sequence come into contact. Evidence from electron microscopy, small-angle X-ray scattering, and chromatin conformation capture experiments has largely favored the zigzag model, at least for nucleosome arrays with longer linker DNA.
The formation of the 30-nm fiber requires the presence of linker histone H1. H1 binds to the entry and exit points of DNA on the nucleosome core, stabilizing the wrapping and promoting compaction. Unlike the core histones, H1 is not part of the nucleosome core particle; it is a "linker" histone that sits outside the core and facilitates higher-order folding.
It is important to note that the 30-nm fiber is not a universal or constitutive structure. In many organisms, including the budding yeast Saccharomyces cerevisiae, a canonical 30-nm fiber is not observed, and in mammalian cells, the fiber appears to be highly dynamic and locally regulated. The 30-nm fiber may represent one possible conformation of nucleosome arrays rather than a fixed, obligatory intermediate in chromosome packaging.
Loop Domains and Scaffold Attachment
The 30-nm fiber is further organized into loop domains of approximately 50 to 200 kilobases. These loops are anchored at their bases to a protein scaffold, historically described as the nuclear matrix or chromosome scaffold. The proteins that mediate this attachment include topoisomerase II and members of the structural maintenance of chromosomes (SMC) family, particularly cohesin and condensin.
The loop domain model proposes that the chromatin fiber is organized into a series of radial loops emanating from a central protein axis. This organization is thought to be functionally important for several reasons. First, it brings distantly separated regulatory elements, such as enhancers and promoters, into close spatial proximity. Second, it organizes the genome into topologically associating domains (TADs), which are regions of the genome that preferentially interact with themselves rather than with neighboring regions. Third, it provides a framework for the further condensation that occurs during mitosis.
The molecular machinery that establishes and maintains loop domains is the SMC family of protein complexes. Cohesin, a ring-shaped complex, is loaded onto chromatin during the G1 phase of the cell cycle and is thought to extrude DNA loops in an ATP-dependent manner. The CTCF protein (CCCTC-binding factor) acts as a boundary element that limits the extent of loop extrusion, defining the borders of TADs. This loop extrusion model has become the dominant paradigm for understanding higher-order chromatin organization. The functional implications of these structures are discussed in Chromosome Structure Function.
Chromosome Territories and Interphase Organization
Nuclear Organization
During interphase, the period between cell divisions, chromosomes are not randomly distributed throughout the nucleus. Instead, each chromosome occupies a distinct, non-overlapping region called a chromosome territory. This was first demonstrated by chromosome painting, a technique in which fluorescently labeled DNA probes complementary to specific chromosomes are hybridized to fixed cells, revealing the spatial extent of each chromosome.
The positioning of chromosome territories within the nucleus is non-random. In human cells, gene-dense chromosomes tend to be located toward the interior of the nucleus, while gene-poor chromosomes are often found at the nuclear periphery. This radial positioning correlates with gene expression: genes at the nuclear periphery are generally less active than those in the nuclear interior. The nuclear periphery is enriched in heterochromatin and is associated with the nuclear lamina, a meshwork of intermediate filament proteins (lamins) that lines the inner nuclear membrane.
Within each chromosome territory, the chromatin is further organized into TADs, which are megabase-scale regions of preferential self-interaction. TADs are separated by boundary regions enriched in CTCF and housekeeping genes. The average size of a TAD in mammalian cells is approximately 1 megabase, and the boundaries are largely conserved across cell types, suggesting that they represent a fundamental organizing principle of the genome.
Euchromatin vs. Heterochromatin
Interphase chromatin is classically divided into two cytological states: euchromatin and heterochromatin. Euchromatin is less condensed, stains lightly, and is generally transcriptionally active. It is enriched in genes and is accessible to RNA polymerase and regulatory factors. Heterochromatin is more condensed, stains darkly, and is largely transcriptionally silent. It is enriched in repetitive sequences, transposons, and genes that are developmentally silenced.
Heterochromatin is further subdivided into constitutive and facultative forms. Constitutive heterochromatin is permanently condensed and includes centromeric and telomeric regions, as well as other highly repetitive sequences. It is marked by specific histone modifications, particularly trimethylation of lysine 9 on histone H3 (H3K9me3), and is bound by heterochromatin protein 1 (HP1). Facultative heterochromatin is conditionally condensed; it can be converted to euchromatin in response to developmental or environmental signals. The classic example is the inactive X chromosome in female mammals, which is largely heterochromatic but contains genes that escape silencing.
The distinction between euchromatin and heterochromatin is not binary but rather a spectrum. Chromatin exists in multiple states with varying degrees of compaction and accessibility, and these states are dynamically regulated by histone modifications, DNA methylation, and chromatin remodeling complexes. The organization of interphase chromatin is summarized in Chromatin Chromosome Structure.
Mitotic Chromosomes: The Most Condensed Form
Sister Chromatids
When a cell enters mitosis, the interphase chromatin undergoes a dramatic condensation to form the classic metaphase chromosome. At this stage, each chromosome consists of two identical sister chromatids, the products of DNA replication during S phase. The sister chromatids are held together along their entire length by cohesin complexes, but they are most tightly associated at the centromere, where they remain joined until anaphase.
The condensation of interphase chromatin into mitotic chromosomes is driven by the condensin complex, another member of the SMC family. Condensin is loaded onto chromatin at the onset of mitosis and uses ATP hydrolysis to introduce positive supercoils and compact the chromatin fiber. The activity of condensin is regulated by phosphorylation, particularly by the cyclin-dependent kinase CDK1, which is the master regulator of mitosis.
The final mitotic chromosome is roughly 700 nanometers in diameter and, in human cells, contains approximately 100 million base pairs of DNA per chromatid. The compaction ratio from naked DNA to mitotic chromosome is approximately 10,000-fold. This extreme condensation is essential for the faithful segregation of chromosomes during cell division; a less condensed chromosome would be prone to breakage and mis-segregation.
Centromeres and Kinetochores
The centromere is the region of the chromosome where the sister chromatids are most tightly associated and where the kinetochore, the protein complex that attaches to spindle microtubules, is assembled. In most eukaryotes, the centromere is defined epigenetically rather than by DNA sequence. The centromere-specific histone H3 variant CENP-A (centromere protein A) replaces canonical H3 in nucleosomes at the centromere, marking the site for kinetochore assembly.
The kinetochore is a large protein complex, with a mass of approximately 5 to 10 megadaltons in vertebrates, that forms on the outer surface of the centromere. It contains more than 100 different proteins, including the KNL1-Mis12-Ndc80 (KMN) network, which directly binds to spindle microtubules. The kinetochore performs several critical functions: it attaches chromosomes to spindle microtubules, it generates force to move chromosomes during anaphase, and it activates the spindle assembly checkpoint, which delays anaphase until all chromosomes are correctly attached.
The position of the centromere along the chromosome determines the chromosome's morphology. Metacentric chromosomes have a centrally located centromere, giving them a symmetrical X shape. Submetacentric chromosomes have a centromere slightly off-center, producing arms of unequal length. Acrocentric chromosomes have a centromere near one end, with a very short arm and a long arm. Telocentric chromosomes have the centromere at the very end, although these are not found in humans. These morphological distinctions are described in Chromosome Structure Type.
Telomeres
The telomere is the specialized structure at each end of a linear chromosome. Telomeres consist of tandem repeats of the hexanucleotide sequence TTAGGG in vertebrates, extending for 5 to 15 kilobases in human cells. The G-rich strand extends beyond the C-rich strand, forming a single-stranded 3' overhang of 50 to 300 nucleotides.
Telomeres serve two essential functions. First, they protect chromosome ends from being recognized as double-strand breaks by the DNA damage response machinery. This protection is achieved through the formation of a specialized structure called the T-loop, in which the single-stranded overhang invades the double-stranded region of the telomere, forming a lariat-like structure. The T-loop is stabilized by the shelterin protein complex, which includes TRF1, TRF2, POT1, TIN2, TPP1, and RAP1.
Second, telomeres solve the end-replication problem. DNA polymerase cannot replicate the very end of a linear DNA molecule because it requires an RNA primer to initiate synthesis, and the primer at the 5' end of the lagging strand cannot be replaced. As a result, telomeres shorten with each cell division. In most somatic cells, this shortening eventually triggers cellular senescence. In germ cells, stem cells, and cancer cells, the enzyme telomerase, a reverse transcriptase that carries its own RNA template, adds telomeric repeats to chromosome ends, maintaining telomere length and allowing continued proliferation.
The overall structure of the mitotic chromosome, including the centromere, telomeres, and sister chromatids, is illustrated in Chromosome Structure Diagram.
Methods Used to Study Chromosome Structure
Microscopy
The study of chromosome structure has been driven by advances in microscopy. Light microscopy, using DNA-binding dyes such as DAPI (4',6-diamidino-2-phenylindole) or Giemsa stain, reveals the overall morphology of chromosomes, including the number, size, and centromere position. Chromosome banding techniques, such as G-banding, produce a characteristic pattern of light and dark bands along each chromosome that is used for karyotyping and for identifying chromosomal abnormalities.
Electron microscopy provides higher resolution and has been used to visualize nucleosomes, the 30-nm fiber, and the radial loop organization of mitotic chromosomes. Scanning electron microscopy of metaphase chromosomes reveals a highly folded surface, while transmission electron microscopy of spread chromatin shows the beads-on-a-string conformation.
Fluorescence in situ hybridization (FISH) allows specific DNA sequences to be visualized within chromosomes. In this technique, a fluorescently labeled DNA probe is hybridized to denatured chromosomal DNA, and the position of the probe is detected by fluorescence microscopy. FISH is used to map genes to specific chromosomal locations, to identify chromosomal rearrangements, and to visualize chromosome territories in interphase nuclei.
Chromatin Conformation Capture
Chromatin conformation capture (3C) and its derivatives have revolutionized the study of chromosome structure by providing a molecular readout of physical proximity between genomic loci. The basic principle of 3C is as follows:
- Cells are treated with formaldehyde to cross-link proteins to DNA and proteins to proteins, freezing the three-dimensional contacts that exist at that moment.
- The cross-linked chromatin is digested with a restriction enzyme, such as HindIII or EcoRI, which cuts at specific recognition sequences.
- The digested fragments are ligated under dilute conditions that favor intramolecular ligation. Fragments that are physically close in three-dimensional space, even if far apart in linear sequence, become joined.
- The cross-links are reversed, and the ligated junctions are detected by PCR or sequencing.
The frequency of ligation between two loci reflects their spatial proximity in the nucleus. This approach has revealed that the genome is organized into TADs, that enhancers physically contact their target promoters, and that the spatial organization of the genome is cell-type specific.
Hi-C and 3D Genome Mapping
Hi-C is a genome-wide variant of 3C in which all ligation junctions are sequenced in a high-throughput manner. In a typical Hi-C experiment, the ligated DNA is sheared, biotinylated, and enriched for fragments containing ligation junctions before sequencing. The resulting data are used to construct a contact matrix, in which each entry represents the frequency of interaction between two genomic regions.
Hi-C data have revealed several key features of genome organization. At the megabase scale, the genome is partitioned into A and B compartments, which correspond roughly to euchromatin and heterochromatin. The A compartment is gene-rich, transcriptionally active, and located toward the nuclear interior; the B compartment is gene-poor, transcriptionally silent, and located at the nuclear periphery. At the sub-megabase scale, the genome is organized into TADs, which appear as squares along the diagonal of the contact matrix. At the finest scale, Hi-C reveals loops between CTCF-bound sites, which are thought to be formed by cohesin-mediated loop extrusion.
The resolution of Hi-C is limited by the restriction enzyme used and by sequencing depth. Standard Hi-C experiments achieve resolutions of 10 to 40 kilobases, while higher-resolution variants, such as Micro-C, which uses a non-specific nuclease to digest chromatin, can achieve near-base-pair resolution. These techniques have provided unprecedented insight into the three-dimensional organization of the genome.
Common Misconceptions and Study Tips
Chromatin vs. Chromosome
A common point of confusion is the distinction between chromatin and chromosome. Chromatin is the complex of DNA and proteins (primarily histones) that makes up the genetic material of eukaryotic cells. It is the substance of which chromosomes are made. A chromosome is a single, discrete unit of chromatin that contains one continuous DNA molecule. In interphase, the chromatin is distributed throughout the nucleus, and individual chromosomes are not visible as distinct structures. During mitosis, the chromatin condenses, and each chromosome becomes a visible, discrete body. Thus, the relationship is hierarchical: chromatin is the material, and chromosomes are the structural units.
DNA Is Not Always Condensed
Another misconception is that DNA is always tightly packed into the condensed chromosome form. In fact, the degree of condensation varies dramatically across the cell cycle and across different regions of the genome. During interphase, most of the genome is in a relatively decondensed state that allows access to the transcriptional machinery. Only a fraction of the genome, primarily heterochromatin, remains highly condensed. The extreme condensation seen in metaphase chromosomes is a transient state that exists only during cell division. Even within mitotic chromosomes, the condensation is not uniform; regions that were transcriptionally active in interphase remain less condensed than inactive regions.
Histones Are Not Just Structural
Histones are often described as "spools" around which DNA is wrapped, implying a purely structural role. This view is incomplete. Histones are dynamic regulators of genome function. The N-terminal tails of histones extend beyond the nucleosome core and are subject to a wide array of post-translational modifications that affect chromatin structure and gene expression. Acetylation of lysine residues neutralizes the positive charge of the tail, weakening histone-DNA interactions and promoting a more open chromatin state. Methylation of lysine or arginine residues can either activate or repress transcription, depending on the specific residue and the degree of methylation. Phosphorylation of serine and threonine residues is involved in chromatin condensation during mitosis and in the DNA damage response.
Histone variants also play important roles. In addition to the canonical histones, which are expressed primarily during S phase, cells express variant histones that can be incorporated into chromatin at any time. H3.3 is deposited at transcriptionally active genes and regulatory elements. H2A.Z is enriched at promoters and enhancers and is thought to poise chromatin for transcriptional activation. CENP-A, as discussed above, marks the centromere. These variants confer specialized properties on the nucleosomes that contain them.
Frequently Asked Questions
What is the structure of a chromosome?
A chromosome is a single, continuous DNA molecule complexed with proteins. In its most condensed form, the metaphase chromosome, it consists of two sister chromatids joined at a centromere. Each chromatid contains one DNA molecule packaged with histones into nucleosomes, which are folded into higher-order fibers and loop domains. The ends of the chromosome are capped by telomeres.
Are chromosomes structures found in all cells?
All cellular organisms have chromosomes, but their structure differs between prokaryotes and eukaryotes. Prokaryotes typically have a single, circular chromosome in the cytoplasm. Eukaryotes have multiple, linear chromosomes within a membrane-bound nucleus. Viruses are an exception; they may have DNA or RNA genomes that are not organized into chromosomes in the same sense.
What are the different types of chromosome structure?
Chromosomes can be classified by centromere position: metacentric (centromere in the middle), submetacentric (centromere slightly off-center), acrocentric (centromere near one end), and telocentric (centromere at the end). They can also be classified by their state: interphase chromosomes are decondensed and occupy territories, while mitotic chromosomes are highly condensed and visible by light microscopy.
How does DNA packaging into chromosomes work?
DNA packaging proceeds through several levels. First, DNA wraps around histone octamers to form nucleosomes (10-nm fiber). Second, nucleosome arrays fold into a 30-nm fiber with the help of linker histone H1. Third, the fiber forms loop domains anchored to a protein scaffold. Finally, during mitosis, condensin compacts the loops into the metaphase chromosome.
What is the difference between chromatin and a chromosome?
Chromatin is the complex of DNA and proteins that constitutes the genetic material. A chromosome is a single, discrete unit of chromatin containing one DNA molecule. Chromatin is the material; chromosomes are the structural units made of that material.
Why do chromosomes have a specific structure?
Chromosome structure solves the problem of fitting a very long DNA molecule into a small cell while maintaining access to the genetic information. The structure also serves functional roles: it brings regulatory elements into proximity, protects chromosome ends, and ensures faithful segregation during cell division.
What is the simplest way to describe chromosome structure?
The simplest description is that a chromosome is a long DNA molecule wrapped around histone proteins, folded into loops, and condensed into a compact structure. The DNA contains genes, the histones package the DNA, and the overall structure is dynamic, changing with the cell cycle.
What are the parts of a chromosome?
The main parts of a metaphase chromosome are the centromere (the constricted region where sister chromatids are joined and the kinetochore forms), the arms (the regions on either side of the centromere), the telomeres (the specialized ends), and the sister chromatids (the two identical copies of the chromosome produced by DNA replication).
Key Takeaways
- Chromosomes are DNA-protein complexes that carry genetic information; prokaryotes have a single circular chromosome, while eukaryotes have multiple linear chromosomes in the nucleus.
- DNA packaging is hierarchical: the double helix wraps around histones to form nucleosomes, which fold into fibers, which organize into loop domains, which condense into metaphase chromosomes.
- Nucleosomes are the fundamental repeating unit of chromatin, consisting of 147 base pairs of DNA wrapped around an octamer of H2A, H2B, H3, and H4 histones.
- Interphase chromosomes occupy distinct territories in the nucleus and are organized into topologically associating domains (TADs) by cohesin-mediated loop extrusion.
- Mitotic chromosomes are the most condensed form, with sister chromatids joined at the centromere, protected by telomeres, and segregated by the kinetochore-spindle apparatus.
- Chromatin exists in a spectrum of states from euchromatin (open, active) to heterochromatin (condensed, silent), regulated by histone modifications and chromatin remodeling.
- Modern techniques such as Hi-C and high-resolution microscopy have revealed that chromosome structure is dynamic and functionally important for gene regulation, DNA replication, and genome stability.