# X-Ray Crystallography: Principles, Workflow, and Applications

## Introduction to X-Ray Crystallography

### What is X-Ray Crystallography?

X-ray crystallography is a biophysical technique used to determine the three-dimensional arrangement of atoms within a crystalline material. When applied to biological macromolecules such as proteins, nucleic acids, and their complexes, it provides atomic-level structural information that underpins our understanding of molecular function. The technique exploits the fact that X-rays have wavelengths on the order of 1 Å (0.1 nm), comparable to the distances between bonded atoms, allowing them to be diffracted by the periodic arrangement of molecules in a crystal.

The output of an X-ray crystallography experiment is an electron density map—a three-dimensional representation of the average electron distribution within the crystal—from which an atomic model is built. This model reveals the positions of individual atoms, the conformation of side chains, the location of bound ligands, and the overall fold of the macromolecule. For a protein of 300 amino acids, the resulting model typically contains several thousand non-hydrogen atoms, each positioned with an accuracy of 0.1–0.3 Å when diffraction data extend to high resolution.

### Historical Context and Importance

The foundations of X-ray crystallography were laid in 1912 when Max von Laue demonstrated that X-rays are diffracted by crystals, confirming both the wave nature of X-rays and the periodic arrangement of atoms in crystals. Within a year, William Henry Bragg and William Lawrence Bragg derived the equation that bears their name and solved the first crystal structures—simple inorganic salts such as sodium chloride. The technique remained confined to small molecules for decades because of the phase problem (discussed in Section 5), which was only overcome for macromolecules in the late 1950s.

In 1958, John Kendrew determined the first [protein structure](/knowledge/bioinformatics/protein-structure-biophysical-levels-folding)—sperm whale myoglobin—at 6 Å resolution, followed by a 2 Å structure in 1959. Max Perutz solved the structure of hemoglobin in 1960, and both shared the 1962 Nobel Prize in Chemistry. Since then, X-ray crystallography has produced the vast majority of macromolecular structures in the [Protein Data Bank](/knowledge/bioinformatics/protein-data-bank-formats-archival-validation) (PDB), which currently holds over 200,000 entries. The technique has been instrumental in revealing the mechanisms of enzymes such as lysozyme, the structure of DNA polymerases, the architecture of ribosomes, and the molecular basis of diseases including HIV/AIDS and cancer. Despite the rise of cryo-electron microscopy (cryo-EM), X-ray crystallography remains the method of choice for high-resolution structures of proteins smaller than ~100 kDa and for structure-guided drug design.

## Basic Principles of X-Ray Diffraction

### Bragg's Law

A crystal is a periodic array of molecules repeated by translation in three dimensions. This periodicity defines a set of planes—called lattice planes—that pass through the repeating units. When a beam of monochromatic X-rays strikes a crystal, the electric field of the X-rays causes electrons in the atoms to oscillate and re-emit X-rays of the same wavelength in all directions. This process is called elastic scattering (or Thomson scattering). For most directions, the scattered waves from different atoms are out of phase and cancel by destructive interference. However, for specific angles, the waves scattered from successive lattice planes are in phase and reinforce each other, producing a diffracted beam.

The condition for constructive interference is given by Bragg's law:

**nλ = 2d sinθ**

where **n** is an integer (the order of diffraction), **λ** is the X-ray wavelength (typically 0.5–1.5 Å for macromolecular crystallography), **d** is the interplanar spacing, and **θ** is the angle between the incident beam and the lattice plane. Bragg's law states that diffraction occurs only when the path difference between waves scattered from adjacent planes (2d sinθ) equals an integer multiple of the wavelength. Because sinθ cannot exceed 1, the minimum spacing that can be resolved is λ/2. For a wavelength of 1.0 Å, this corresponds to a resolution of 0.5 Å—sufficient to resolve individual atoms.

The set of all possible lattice planes in a crystal is described by Miller indices (h, k, l), which are integers that specify the orientation and spacing of each plane family. Each set of planes gives rise to one reflection (also called a Bragg peak) at a specific angle. A typical protein crystal diffracts to 2–3 Å resolution, producing tens of thousands to millions of measurable reflections.

### Diffraction Pattern and Reciprocal Space

The diffraction pattern of a crystal is not a direct image of the molecule; rather, it is a map of the crystal's Fourier transform sampled at discrete points. Each reflection corresponds to a wave with an amplitude and a phase. The amplitude is related to the intensity of the diffracted beam, which is measured directly. The phase, however, is lost during measurement—this is the phase problem (Section 5).

The mathematical framework for interpreting diffraction patterns is most naturally expressed in reciprocal space, a construct in which distances are inversely proportional to real-space distances. A reflection with indices (h, k, l) corresponds to a point in reciprocal space at position (h/a, k/b, l/c), where a, b, and c are the real-space unit cell dimensions. The intensity of each reflection, I(hkl), is proportional to the square of the amplitude of the structure factor, F(hkl), which is the sum of scattering contributions from all atoms in the unit cell:

**F(hkl) = Σ fⱼ exp[2πi(hxⱼ + kyⱼ + lzⱼ)]**

where fⱼ is the atomic scattering factor of atom j (proportional to its electron count), and (xⱼ, yⱼ, zⱼ) are its fractional coordinates. The electron density at any point in the unit cell, ρ(x, y, z), is the inverse Fourier transform of the structure factors:

**ρ(x, y, z) = (1/V) Σ F(hkl) exp[−2πi(hx + ky + lz)]**

This equation is central to crystallography: if we know the amplitudes (from measured intensities) and the phases (from experimental or computational methods), we can calculate the electron density map and build the atomic model.

## The Crystallization Process

### Protein Purification and Concentration

Crystallization requires a highly pure, homogeneous protein sample. Impurities, including degradation products, post-translationally modified variants, and aggregation states, can poison crystal nucleation or produce crystals with poor order. The protein is typically purified to >95% purity as assessed by SDS-PAGE. Affinity tags, such as a hexahistidine tag, are commonly used for initial purification (see [His Tag Protein Purification](/knowledge/molecular-biology/his-tag-protein-purification)), followed by size-exclusion chromatography (see [Gel Permeation Chromatography](/knowledge/molecular-biology/gel-permeation-chromatography)) to remove aggregates and exchange the buffer.

The purified protein must be concentrated to 5–20 mg/mL, though the optimal concentration varies with the protein's molecular weight and surface properties. Concentration is achieved using centrifugal concentrators with molecular weight cutoffs (typically 10–30 kDa) and is monitored by absorbance at 280 nm (see [Automated Protein Quantification](/knowledge/molecular-biology/automated-protein-quantification)). The buffer should be simple—typically 10–50 mM HEPES or Tris at pH 7.0–8.0, with 100–300 mM NaCl—and free of reducing agents that might interfere with crystallization. The protein must be monodisperse, meaning it exists as a single oligomeric state in solution; dynamic light scattering (DLS) is used to verify this.

### Crystallization Techniques

[Protein crystallization](/knowledge/molecular-biology/protein-crystallization) is a phase transition in which the protein leaves solution and forms an ordered solid. The goal is to bring the solution to supersaturation—a state where the protein concentration exceeds its solubility—without inducing amorphous precipitation. This is achieved by adding a precipitant (e.g., ammonium sulfate, polyethylene glycol, or 2-methyl-2,4-pentanediol) that reduces the protein's solubility.

The most common method is **vapor diffusion**, performed in either hanging-drop or sitting-drop format. In a typical hanging-drop experiment, 1–2 µL of protein solution is mixed with an equal volume of reservoir solution containing the precipitant. The drop is sealed over a reservoir of 500–1000 µL of the same precipitant solution. Because the precipitant concentration in the reservoir is higher than in the drop, water vapor diffuses from the drop to the reservoir, gradually increasing the precipitant concentration in the drop and driving the protein toward supersaturation. Over days to weeks, crystals may nucleate and grow.

Other methods include **microbatch** (where protein and precipitant are mixed under oil to prevent evaporation) and **dialysis** (where the precipitant diffuses through a semipermeable membrane). For membrane proteins, **lipid cubic phase** crystallization is used, in which the protein is reconstituted into a lipid bilayer that mimics its native environment (see [Protein Crystallization](/knowledge/molecular-biology/protein-crystallization) for a detailed protocol).

Crystallization conditions are screened using commercially available sparse-matrix screens that sample a wide range of pH, precipitant type and concentration, and additives. Each condition is tested in a single drop, and hits are then optimized by varying the precipitant concentration, pH, temperature, and protein concentration in small increments.

### Crystal Quality Assessment

A good crystal is not simply one that diffracts; it must be single (not twinned or clustered), have well-defined faces, and be of adequate size (typically 50–300 µm in each dimension). Crystals are first examined under a light microscope with polarized light; protein crystals are birefringent and appear brightly colored against a dark background. The most definitive test is to observe diffraction: a crystal is mounted on a loop, flash-cooled to 100 K in liquid nitrogen (to reduce radiation damage), and exposed to an X-ray beam. The diffraction pattern reveals the resolution limit, the symmetry (space group), and the mosaicity (the spread of crystal plane orientations). A well-ordered crystal produces sharp, intense spots to high resolution; a poorly ordered crystal produces diffuse, weak spots that fade quickly with increasing scattering angle. For practical advice on improving crystal quality, see [Protein Crystallography Improving Crystals](/knowledge/molecular-biology/protein-crystallography-improving-crystals).

## Data Collection and Processing

### X-Ray Sources and Detectors

X-rays for macromolecular crystallography are produced either by rotating-anode generators in home laboratories or, far more commonly, by synchrotron radiation sources. Synchrotrons accelerate electrons to near-light speeds in a storage ring; bending magnets and insertion devices (wigglers and undulators) cause the electrons to emit intense, tunable X-ray beams. Modern synchrotron beamlines deliver 10¹²–10¹³ photons per second focused to a spot of 10–50 µm, allowing data collection from microcrystals in seconds to minutes.

The crystal is mounted on a goniometer that rotates it during data collection, typically in 0.1–1.0° increments over a total rotation range of 180–360°. As the crystal rotates, different lattice planes come into diffracting orientation, and the detector records the positions and intensities of the resulting reflections. Detectors are either charge-coupled devices (CCDs), pixel-array detectors (e.g., Pilatus, Eiger), or hybrid photon-counting detectors. These detectors have short readout times (milliseconds) and high dynamic range, enabling rapid collection of complete datasets.

Data collection is performed at cryogenic temperatures (100 K) to mitigate radiation damage, which is caused by the absorption of X-rays and the generation of free radicals that break disulfide bonds, decarboxylate acidic residues, and disorder the crystal lattice. Cryoprotectants such as glycerol (20–30% v/v), ethylene glycol, or sucrose are added to the mother liquor to prevent ice formation during flash-cooling.

### Data Reduction and Scaling

The raw diffraction images contain thousands of spots, each characterized by its position (x, y) on the detector, its intensity, and the crystal orientation at the time of exposure. Data processing involves three steps:

1. **Indexing**: The positions of reflections are used to determine the unit cell dimensions and the space group—the symmetry of the crystal lattice. This step assigns Miller indices (h, k, l) to each reflection.

2. **Integration**: The intensity of each reflection is estimated by summing the pixel counts within the spot and subtracting the background. This step also corrects for the Lorentz factor (a geometric correction for the time each reflection spends in diffracting orientation) and polarization effects.

3. **Scaling and merging**: Reflections measured multiple times (due to symmetry or crystal rotation) are averaged, and the dataset is scaled to account for variations in beam intensity, crystal decay, and absorption. The output is a list of unique reflections with their mean intensities, I(hkl), and associated uncertainties, σ(I).

The quality of the dataset is assessed by several statistics. The **R-merge** (or R-sym) value reports the agreement between multiply measured reflections; values below 10% are generally acceptable. The **completeness** (the fraction of theoretically observable reflections that were measured) should exceed 90%, ideally 99%. The **I/σ(I)** ratio—the signal-to-noise of the intensities—should be high, particularly in the highest-resolution shell, where a value of 2 or greater is typically required to justify including those reflections in refinement.

## Solving the Phase Problem

The measured intensities provide the amplitudes of the structure factors, |F(hkl)|, but not their phases, φ(hkl). Without phases, the electron density map cannot be calculated directly. This is the **phase problem**, the central challenge of crystallography. Several methods exist to determine phases, each exploiting different physical or computational principles.

### Isomorphous Replacement

Isomorphous replacement involves soaking or co-crystallizing the protein with heavy atoms (e.g., mercury, platinum, gold, or uranium) that bind specifically to cysteine thiols, histidine imidazoles, or methionine sulfurs. The heavy atoms must bind at a small number of sites without altering the crystal lattice—the crystals must be isomorphous (same unit cell and space group). The presence of a heavy atom with many electrons (e.g., mercury, Z = 80) changes the scattered intensities measurably, and the differences between native and derivative datasets can be used to locate the heavy atoms and estimate their contribution to the structure factors.

In **multiple isomorphous replacement (MIR)**, data from two or more different heavy-atom derivatives are combined. The heavy-atom positions are determined from difference Patterson maps, and the phases are estimated using the Harker construction, which combines the known heavy-atom contributions with the measured amplitudes to constrain the protein phase. MIR was the method used by Perutz and Kendrew to solve the first protein structures and remains useful for proteins with no homologous model.

### Anomalous Scattering

Anomalous scattering exploits the fact that atoms absorb X-rays at specific wavelengths, causing the atomic scattering factor to become complex—it acquires an imaginary component that introduces a wavelength-dependent phase shift. This effect is negligible for light atoms (C, N, O, S) at typical X-ray wavelengths but is significant for heavier elements such as selenium, bromine, and metals.

**Multi-wavelength anomalous dispersion (MAD)** is the most powerful variant. The protein is expressed in a medium containing selenomethionine, in which the sulfur atoms of methionine are replaced by selenium. Data are collected at three wavelengths: the absorption peak (where the imaginary component is maximal), the inflection point (where the real component changes most rapidly), and a remote wavelength. The wavelength-dependent intensity differences allow the selenium positions to be located and the phases to be determined ab initio. **Single-wavelength anomalous dispersion (SAD)** uses data at one wavelength and is now widely used, particularly at synchrotrons with tunable beamlines.

### Molecular Replacement

Molecular replacement (MR) is the most common method for proteins with a known homologous structure. If a protein shares >30% sequence identity with a protein of known structure, the known model can be used as a search probe. The probe is placed in the unknown unit cell by a six-dimensional search (three rotations and three translations) to find the orientation and position that best explains the observed diffraction intensities. The rotation function searches for the orientation that maximizes the overlap between the probe's Patterson map (the autocorrelation of the electron density) and the observed Patterson map; the translation function then positions the oriented molecule within the unit cell.

MR is fast and requires no experimental phase information, but it is biased by the search model: if the probe is too different from the target, the solution may be wrong, and the resulting electron density map will reflect the model's errors. Programs such as Phaser and MolRep implement sophisticated likelihood-based MR algorithms that score solutions based on the probability of observing the data given the model.

## Model Building and Refinement

### Electron Density Maps

Once phases are obtained, the electron density map is calculated using the Fourier synthesis equation. The initial map is often noisy and may contain errors, particularly if the phases are poor. To improve the map, **density modification** techniques are applied: solvent flattening (assuming that the solvent region has uniform, low electron density), histogram matching (adjusting the map's density distribution to match a theoretical distribution), and non-crystallographic symmetry averaging (if multiple copies of the molecule are present in the asymmetric unit).

The quality of the map is judged by its **figure of merit** (a measure of phase accuracy) and by visual inspection. At 2.5–3.0 Å resolution, the polypeptide backbone is visible as a continuous tube, and large side chains (tryptophan, tyrosine, arginine) can be identified. At 2.0 Å or better, individual atoms are resolved, and water molecules and small ligands are visible. At 1.5 Å or better, the map is of sufficient quality to distinguish between alternative conformations of side chains.

### Model Building

Model building is the process of interpreting the electron density map by placing atoms at positions consistent with the density. This is done interactively using graphics programs such as Coot. The sequence of the protein is threaded through the density, starting with recognizable features such as the carbonyl oxygen of peptide bonds (which protrudes from the backbone) and bulky aromatic side chains. The model is built residue by residue, with the backbone traced first and side chains fitted into the corresponding density.

For high-resolution maps, automated building programs such as ARP/wARP or Buccaneer can build a large fraction of the model automatically. For lower-resolution maps, manual building is required, and the process is iterative: build a partial model, refine it, calculate new phases, and rebuild.

### Refinement and Validation

Refinement is the optimization of the atomic model against the observed diffraction data. The target function is the sum of the crystallographic residual (the difference between observed and calculated structure factor amplitudes) and geometric restraints (bond lengths, bond angles, planarity, and chirality that must conform to known stereochemistry). The most common refinement algorithms are **maximum likelihood** (implemented in REFMAC5 and phenix.refine) and **simulated annealing** (in CNS), the latter being particularly effective at escaping local minima.

The progress of refinement is monitored by two statistics:

- **R-factor**: R = Σ||F(obs)| − |F(calc)|| / Σ|F(obs)|. A well-refined structure at 2.0 Å resolution typically has an R-factor of 15–20%. An R-factor above 30% indicates serious errors in the model or data.

- **R-free**: Calculated identically to the R-factor but using a subset of reflections (typically 5%) that are excluded from refinement. The R-free is an unbiased measure of model quality; it should be within 3–5% of the R-factor. A large gap between R and R-free indicates overfitting.

Validation also includes checking the geometry of the model: Ramachandran plots (which show the allowed φ/ψ backbone torsion angles) should have >95% of residues in favored regions and no residues in disallowed regions. Side-chain rotamers, bond lengths, and bond angles are compared to standard values, and the fit of the model to the electron density is assessed using real-space correlation coefficients.

## Applications in Protein Biochemistry

### Structure-Function Relationships

The primary application of X-ray crystallography is to explain how [protein structure](/blog/guides/protein-structure-levels-how-primary-secondary-tertiary-and-quaternary-structure) determines function. For example, the structure of **lysozyme** (the first enzyme structure solved, by David Phillips in 1965) revealed a deep active-site cleft that binds the peptidoglycan substrate, with two catalytic glutamates (Glu35 and Asp52) positioned to donate and accept protons during glycosidic bond cleavage. This structure provided the first direct evidence for the mechanism of enzyme catalysis at atomic resolution.

Crystallography has also revealed the conformational changes that accompany ligand binding and catalysis. The structures of **hexokinase** in its open (unliganded) and closed (glucose-bound) states showed that substrate binding induces a large domain movement that brings the two lobes of the enzyme together, excluding water from the active site and positioning the ATP for phosphoryl transfer. Similarly, structures of **HIV-1 protease** in complex with peptide substrates and inhibitors have illuminated the catalytic mechanism of this aspartic protease and guided the design of drugs that block its activity.

The technique is also essential for understanding macromolecular assemblies. The ribosome—a complex of RNA and more than 50 proteins—was solved by crystallography at atomic resolution in 2000 (by Venki Ramakrishnan, Thomas Steitz, and Ada Yonath, who shared the 2009 Nobel Prize in Chemistry). These structures revealed the details of mRNA decoding, peptide bond formation, and the action of antibiotics that target the ribosome.

### Drug Discovery and Design

[Structure-based drug design](/knowledge/bioinformatics/structure-based-drug-design-bioinformatics) is one of the most impactful applications of X-ray crystallography. The workflow is as follows:

1. Determine the high-resolution structure of the target protein, ideally in complex with a known ligand or substrate.
2. Identify the binding site and the key interactions that stabilize ligand binding.
3. Use computational docking and medicinal chemistry to design candidate compounds.
4. Co-crystallize the protein with each candidate and determine the structure of the complex.
5. Iterate: use the structures to guide chemical modifications that improve affinity, selectivity, and pharmacokinetic properties.

This approach has been used to develop drugs against HIV protease (e.g., saquinavir, ritonavir), influenza neuraminidase (oseltamivir, zanamivir), and kinases involved in cancer (imatinib, erlotinib). In each case, crystallography revealed the precise geometry of the binding site and the conformational changes induced by inhibitor binding, enabling the design of compounds with nanomolar affinities. The technique is also used in fragment-based drug discovery, where low-molecular-weight fragments are soaked into crystals and their binding modes are determined by crystallography, providing starting points for fragment linking and growing.

## Limitations and Complementary Methods

### Challenges and Artifacts

The most fundamental limitation of X-ray crystallography is the requirement for diffraction-quality crystals. Many proteins—particularly membrane proteins, [intrinsically disordered proteins](/knowledge/bioinformatics/intrinsically-disordered-proteins-and-computational-structural-classification), and large multi-subunit complexes—are difficult or impossible to crystallize. Membrane proteins require detergents or lipidic environments to maintain solubility, and their hydrophobic surfaces hinder the formation of ordered lattices. Despite decades of effort, only a small fraction of the human membrane proteome has been structurally characterized.

Crystallization also imposes artificial conditions. The high protein concentrations (5–20 mg/mL), the presence of precipitants, and the low temperature (100 K during data collection) may trap the protein in a conformation that differs from its solution state. Crystal contacts—the interactions between molecules in the lattice—can distort the structure, particularly at the protein surface. The lattice may also select for a particular conformation from an ensemble of states, providing a static snapshot that may not represent the dynamic behavior of the protein in solution.

Radiation damage is another concern. Even at cryogenic temperatures, X-ray exposure causes specific damage: disulfide bonds are broken, acidic residues are decarboxylated, and metal centers are reduced. These effects can be minimized by collecting data at lower doses, using multiple crystals, or adding radical scavengers such as ascorbate, but they cannot be entirely eliminated.

### Complementary Structural Biology Techniques

Several techniques complement X-ray crystallography by providing structural information under different conditions or for different types of samples:

- **Nuclear magnetic resonance (NMR) spectroscopy** determines structures of proteins in solution, typically up to ~40 kDa. NMR can also probe dynamics and interactions at atomic resolution, providing information that crystallography cannot. However, NMR requires high protein concentrations (0.1–1 mM) and is limited by spectral overlap for larger proteins.

- **Cryo-electron microscopy (cryo-EM)** has emerged as a powerful alternative for large complexes and membrane proteins. In cryo-EM, the protein is embedded in vitreous ice and imaged directly in a transmission electron microscope. Single-particle analysis reconstructs the 3D structure from thousands of 2D images. Recent advances in direct electron detectors and image-processing algorithms have pushed cryo-EM to resolutions below 2 Å, rivaling crystallography. Cryo-EM does not require crystals and can capture multiple conformational states from a single sample.

- **Small-angle X-ray scattering (SAXS)** provides low-resolution (10–50 Å) information about the shape and oligomeric state of proteins in solution. SAXS is useful for studying flexible and dynamic systems that resist crystallization.

- **Hydrogen-deuterium exchange mass spectrometry (HDX-MS)** and **cross-linking mass spectrometry (XL-MS)** provide low-resolution constraints on protein dynamics and interactions that can be combined with crystallographic structures to build models of complexes.

For a broader view of how these methods are used to study protein interactions, see [Yeast Two Hybrid System](/knowledge/molecular-biology/yeast-two-hybrid-system) and [Yeast Two Hybrid Assay](/knowledge/molecular-biology/yeast-two-hybrid-assay), which describe complementary approaches for detecting and characterizing protein-protein interactions in living cells.

## Common Pitfalls and Practical Tips

### Misinterpreting Resolution

Students frequently confuse the resolution of a crystal with the resolution of the structure. The **resolution** of a crystal structure (e.g., 2.5 Å) is the minimum interplanar spacing for which diffraction was observed and is determined by the quality of the crystal and the completeness of the data. It is not a measure of the accuracy of the model, nor does it indicate the size of the molecule. At 3 Å resolution, the backbone is visible but side chains are poorly defined; at 1.5 Å, individual atoms are resolved. When reading a structure paper, always check the resolution and consider what level of detail is justified.

### Overlooking Crystal Quality

A common error is to assume that any crystal that diffracts is suitable for structure determination. Crystals can be twinned (two or more lattices intergrown in specific orientations), have high mosaicity (broad rocking curves), or be composed of multiple lattices. These defects produce overlapping or split reflections that degrade the data quality. Before investing time in data collection, assess the diffraction pattern: sharp, well-separated spots to high resolution indicate a good crystal; streaky, broad, or overlapping spots indicate problems. The **mosaicity** (typically 0.1–1.0°) and the **B-factor** (a measure of atomic displacement) from data processing provide quantitative measures of crystal order.

### Understanding R-factors and Electron Density

Students often misinterpret R-factors. An R-factor of 20% does not mean that 20% of the structure is wrong; it means that the calculated amplitudes differ from the observed amplitudes by 20% on average. The R-free is a more reliable indicator of model quality because it is computed from reflections not used in refinement. A structure with an R-factor of 18% and an R-free of 28% is overfit—the model has been adjusted to fit noise. Always compare R and R-free.

Similarly, the electron density map is not a direct image of the molecule. It is a Fourier synthesis that depends on the phases, which are imperfect. Features in the map may be artifacts of phase errors, and the absence of density for a loop or side chain does not necessarily mean the atoms are absent—they may be disordered (mobile) and thus have smeared, low-density features. When examining a structure, look at the **real-space correlation coefficient** (RSCC) for each residue, which quantifies how well the model fits the density.

### Practical Tips for Studying Crystallography

- **Learn the mathematics conceptually**: You do not need to derive the Fourier transform, but you must understand that the diffraction pattern is the Fourier transform of the electron density, and that phases are required to invert the transform.
- **Use the PDB**: Download structures of well-known proteins (lysozyme, hemoglobin, GFP) and view them in PyMOL or ChimeraX. Examine the electron density maps (available from the Electron Density Server) to see what maps at different resolutions look like.
- **Understand the statistics**: Know what resolution, completeness, I/σ(I), R-factor, and R-free mean and how they are calculated. Be able to interpret a crystallography table in a paper.
- **Connect structure to function**: When studying an enzyme, ask how the active-site geometry explains the catalytic mechanism. The structure is a hypothesis about function; test it against biochemical data.

## Frequently Asked Questions

### What is X-ray crystallography?

X-ray crystallography is a technique for determining the three-dimensional arrangement of atoms in a crystalline material by analyzing the diffraction pattern produced when X-rays pass through the crystal. For biological macromolecules, it provides atomic-resolution structures of proteins, nucleic acids, and their complexes, revealing the molecular basis of function.

### What is the principle of X-ray crystallography?

The principle is that a crystal is a periodic array of molecules that acts as a three-dimensional diffraction grating for X-rays. When monochromatic X-rays strike the crystal, they are scattered by the electrons of the atoms. The scattered waves interfere constructively only at specific angles that satisfy Bragg's law (nλ = 2d sinθ). The resulting diffraction pattern contains information about the amplitudes of the structure factors, and the electron density map is calculated by Fourier synthesis using these amplitudes and the phases.

### How does X-ray crystallography work?

The workflow involves: (1) purifying the protein to homogeneity and concentrating it to 5–20 mg/mL; (2) crystallizing the protein by vapor diffusion or other methods; (3) collecting diffraction data at a synchrotron or home source; (4) processing the data to obtain intensities; (5) solving the phase problem by molecular replacement, isomorphous replacement, or anomalous scattering; (6) calculating the electron density map; (7) building the atomic model; and (8) refining and validating the model against the data.

### What is the mechanism of X-ray crystallography?

The mechanism is based on the wave nature of X-rays and the periodic structure of crystals. X-rays are elastically scattered by electrons in the crystal. For a given set of lattice planes, the scattered waves from successive planes are in phase only when the path difference equals an integer multiple of the wavelength (Bragg's law). The intensity of each reflection is proportional to the square of the amplitude of the corresponding structure factor, which is the sum of scattering contributions from all atoms. The electron density is the inverse Fourier transform of the structure factors, but because phases are lost during measurement, they must be determined by experimental or computational methods.

### What is a X-ray crystallography diagram?

A typical diagram shows a crystal mounted on a loop in the X-ray beam, with the diffracted beams recorded on a detector. The diagram also illustrates the geometry of Bragg's law: parallel lattice planes with spacing d, the incident beam at angle θ, and the diffracted beam at the same angle. In a more detailed diagram, the diffraction pattern is shown as a grid of spots, each corresponding to a reflection with indices (h, k, l), and the electron density map is depicted as a mesh or surface surrounding the atomic model.

### Why is X-ray crystallography important?

X-ray crystallography is the primary method for determining atomic-resolution structures of biological macromolecules. It has revealed the mechanisms of enzymes, the architecture of molecular machines such as the ribosome, the details of protein-DNA and protein-protein interactions, and the molecular basis of disease. It is essential for structure-based drug design, enabling the development of drugs against HIV, influenza, cancer, and other diseases. Over 85% of the structures in the [Protein Data Bank](/blog/guides/protein-data-bank) were determined by X-ray crystallography.

### What are the limitations of X-ray crystallography?

The main limitations are: (1) the requirement for diffraction-quality crystals, which is difficult or impossible for many proteins, especially membrane proteins and [intrinsically disordered proteins](/knowledge/bioinformatics/intrinsically-disordered-proteins-computational-challenges); (2) the need for large amounts of pure, homogeneous protein; (3) the possibility that the crystal lattice traps a non-native conformation or introduces artifacts from crystal contacts; (4) radiation damage during data collection; and (5) the phase problem, which requires additional experimental or computational effort. Complementary methods such as cryo-EM and NMR can overcome some of these limitations.

## Key Takeaways

- X-ray crystallography determines atomic-resolution structures by analyzing the diffraction of X-rays by crystals, with the resolution limit set by Bragg's law (nλ = 2d sinθ).
- The diffraction pattern is the Fourier transform of the electron density; the measured intensities provide amplitudes, but phases must be obtained by molecular replacement, isomorphous replacement, or anomalous scattering.
- Crystallization requires highly pure, monodisperse protein at 5–20 mg/mL and is achieved by vapor diffusion, microbatch, or lipid cubic phase methods; crystal quality is assessed by the diffraction pattern.
- Data collection at synchrotrons with modern detectors yields high-completeness datasets; data processing involves indexing, integration, and scaling, with quality assessed by R-merge, completeness, and I/σ(I).
- Model building and refinement are iterative processes that optimize the fit of the atomic model to the electron density, monitored by R-factor and R-free; validation checks geometry and density fit.
- X-ray crystallography is essential for understanding structure-function relationships and for structure-based drug design, but it is limited by the need for crystals and may not capture dynamic or heterogeneous states.
- Complementary methods—cryo-EM, NMR, SAXS, and cross-linking mass spectrometry—provide structural information for systems that resist crystallization and for studying dynamics in solution.

## Further Reading

- Ahmed MH, Ghatge MS, Safo MK. *Hemoglobin: Structure, Function and Allostery*. Sub-cellular biochemistry. 2020. [PubMed 32189307](https://doi.org/10.1007/978-3-030-41769-7_14)
- Catterall WA. *Voltage gated sodium and calcium channels: Discovery, structure, function, and Pharmacology*. Channels (Austin, Tex.). 2023. [PubMed 37983307](https://doi.org/10.1080/19336950.2023.2281714)
- Gushchin I, Gordeliy V. *Microbial Rhodopsins*. Sub-cellular biochemistry. 2018. [PubMed 29464556](https://doi.org/10.1007/978-981-10-7757-9_2)
- Kato S, Matsui T, Tanaka Y. *Molluscan Hemocyanins*. Sub-cellular biochemistry. 2020. [PubMed 32189300](https://doi.org/10.1007/978-3-030-41769-7_7)
- Powell HR. *X-ray data processing*. Bioscience reports. 2017. [PubMed 28899925](https://doi.org/10.1042/BSR20170227)
- Lanyi JK. *Bacteriorhodopsin*. Annual review of physiology. 2004. [PubMed 14977418](https://doi.org/10.1146/annurev.physiol.66.032102.150049)



<div data-calculator="toxicity"></div>

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)