# How Mass Spec Identifies Proteins: A Practical Guide

## Introduction to Mass Spectrometry for Protein Identification

Mass spectrometry (MS) is an analytical technique that measures the mass-to-charge ratio (\(m/z\)) of gas-phase ions. In protein biochemistry, MS has become the cornerstone of proteomics—the large-scale study of proteins—because it can identify thousands of proteins from a complex mixture in a single experiment, determine post-translational modifications, and quantify relative or absolute protein abundance. The fundamental principle is deceptively simple: ionize a molecule, measure its \(m/z\), and use that measurement to infer its molecular mass and, with fragmentation, its sequence.

For proteins, the challenge is that intact proteins are large, heterogeneous, and difficult to ionize efficiently. The field has therefore converged on a "bottom-up" strategy: digest proteins into peptides, measure the peptides, and use the peptide sequences as unique identifiers for their parent proteins. This approach, sometimes called shotgun proteomics, is the dominant workflow in modern protein identification.

### Why Mass Spectrometry for Proteins?

Proteins are the functional units of the cell, but they cannot be amplified like DNA. There is no protein equivalent of PCR. To identify a protein, you must directly analyze its physical properties. Mass spectrometry offers three decisive advantages over older methods like Edman degradation or immunoblotting:

1. **Sensitivity**: Modern instruments can detect femtomole to attomole quantities of peptide, corresponding to protein amounts far below what a Coomassie-stained gel can visualize.
2. **Throughput**: A single liquid chromatography–tandem mass spectrometry (LC-MS/MS) run can identify thousands of proteins from a cell lysate.
3. **Generality**: Unlike antibodies, which require prior knowledge of the target, MS is unbiased. It identifies whatever is present, including unexpected proteins, isoforms, and modifications.

The core measurement—\(m/z\)—is made by a [Mass Spectrometer](/knowledge/molecular-biology/mass-spectrometer), which consists of three essential components: an ion source, a mass analyzer, and a detector. The ion source converts neutral peptides into charged ions; the mass analyzer separates those ions by their \(m/z\); and the detector records the abundance of each ion. The resulting mass spectrum is a plot of ion intensity versus \(m/z\).

### Overview of the Workflow

A typical [bottom-up proteomics](/knowledge/bioinformatics/bottom-up-proteomics-principles-workflow-and-applications) experiment proceeds through five stages:

1. **Sample preparation**: Proteins are extracted from cells or tissues, purified, and digested into peptides.
2. **Peptide separation**: The peptide mixture is fractionated, usually by reversed-phase liquid chromatography (RPLC), to reduce complexity before MS analysis.
3. **Ionization**: Peptides are ionized by electrospray ionization (ESI) or matrix-assisted laser desorption/ionization (MALDI).
4. **Mass analysis and fragmentation**: The mass analyzer measures intact peptide \(m/z\) (MS1), then selects individual peptides for fragmentation to generate sequence-informative spectra (MS/MS).
5. **Data analysis**: MS/MS spectra are matched against a [protein database](/blog/guides/protein-database) using search algorithms to assign peptide sequences and infer protein identities.

Each step is a potential source of error, and understanding the mechanism at each stage is essential for interpreting results correctly. The remainder of this guide walks through each stage in detail.

## Sample Preparation: From Cells to Peptides

The quality of protein identification is determined by the quality of the sample entering the [mass spectrometer](/knowledge/molecular-biology/mass-spectrometer). Poor sample preparation cannot be rescued by even the most expensive instrument. The goal is to produce a clean, soluble, and completely digested peptide mixture that is free of detergents, salts, and other contaminants that suppress ionization.

### Protein Extraction and Purification

Cells or tissues are first lysed to release their protein content. The lysis buffer typically contains:

- **A chaotropic agent** such as 8 M urea or 6 M guanidine hydrochloride to denature proteins and disrupt hydrogen bonding.
- **A reducing agent** such as 5–10 mM dithiothreitol (DTT) or tris(2-carboxyethyl)phosphine (TCEP) to break disulfide bonds.
- **An alkylating agent** such as 15–25 mM iodoacetamide (IAA) to irreversibly modify free cysteine thiols and prevent disulfide reformation.
- **A protease inhibitor cocktail** (e.g., 1 mM phenylmethylsulfonyl fluoride, PMSF) to prevent endogenous proteases from degrading the sample.
- **A buffer** such as 50 mM Tris-HCl or 100 mM ammonium bicarbonate to maintain pH 7.5–8.5, which is optimal for trypsin activity later.

After lysis, insoluble material (membranes, DNA, cell debris) is removed by centrifugation at 20,000 × g for 15–30 minutes at 4 °C. If the sample is a complex mixture such as a whole-cell lysate, additional fractionation may be performed—for example, SDS-PAGE with in-gel digestion, or strong cation exchange (SCX) chromatography—to reduce complexity. However, for many experiments, a "filter-aided sample preparation" (FASP) approach is used, where proteins are buffer-exchanged into digestion buffer using a 30 kDa molecular weight cutoff filter. This simultaneously removes detergents and small-molecule contaminants.

### Enzymatic Digestion (Trypsin)

Proteins must be digested into peptides because the mass analyzer operates optimally on molecules below ~4 kDa. Trypsin is the enzyme of choice for most proteomics experiments. It cleaves peptide bonds on the C-terminal side of lysine (K) and arginine (R) residues, unless the next residue is proline (P). This specificity is advantageous because:

- Lysine and arginine are abundant (about 1 in 10 residues), producing peptides of 6–20 amino acids—ideal for MS.
- Each peptide (except the C-terminal peptide of the protein) carries at least one basic residue, which protonates readily during ESI.
- The predictable cleavage pattern allows search algorithms to compute theoretical peptide masses with high confidence.

Digestion is performed with sequencing-grade modified trypsin at an enzyme-to-substrate ratio of 1:50 to 1:100 (w/w), typically at 37 °C for 12–18 hours. The reaction is carried out in 50–100 mM ammonium bicarbonate (pH 8.0) or 50 mM Tris-HCl (pH 8.0). Urea must be diluted to below 1 M before adding trypsin, because urea above 2 M denatures trypsin itself. After digestion, the reaction is quenched by acidification with formic acid to a final concentration of 1% (v/v), which also prepares the sample for LC-MS.

### Peptide Cleanup and Fractionation

The digested peptide mixture contains salts, residual reagents, and trypsin autolysis products that interfere with ionization. Solid-phase extraction (SPE) using reversed-phase C18 cartridges is the standard cleanup method. Peptides bind to the hydrophobic C18 resin in aqueous solution (e.g., 0.1% trifluoroacetic acid, TFA), salts wash through, and peptides are eluted with 60–80% acetonitrile (ACN) containing 0.1% formic acid. The eluted peptides are then dried in a vacuum concentrator and resuspended in 0.1% formic acid for LC-MS analysis.

For highly complex samples (e.g., whole proteomes), a single LC-MS run may not provide sufficient depth. Offline fractionation—such as SCX chromatography, high-pH reversed-phase chromatography, or isoelectric focusing—can be performed before the final low-pH RPLC step. Each fraction is analyzed separately, multiplying the number of peptides identified. A typical deep proteome experiment might fractionate into 10–24 fractions, each analyzed by a 2-hour LC-MS/MS gradient.

## Ionization Techniques: ESI and MALDI

Peptides are non-volatile and thermally labile; they cannot be vaporized by heating without degradation. Ionization must therefore occur directly from solution (ESI) or from a solid matrix (MALDI). Both techniques are "soft" ionization methods, meaning they produce intact molecular ions without extensive fragmentation.

### Electrospray Ionization (ESI)

ESI is the most widely used ionization method for LC-MS-based proteomics. The peptide solution flows through a narrow capillary (typically 10–100 μm inner diameter) held at a high voltage (2–5 kV relative to the inlet of the [mass spectrometer](/knowledge/molecular-biology/mass-spectrometer)). The strong electric field disperses the liquid into a fine spray of charged droplets. As solvent evaporates, droplet size decreases and charge density increases until Coulombic repulsion exceeds surface tension, causing droplet fission. Eventually, gas-phase peptide ions are released, carrying multiple protons. A peptide of mass 1,000 Da might carry 1–3 protons, giving \(m/z\) values of 1,001, 501, or 334 for the +1, +2, or +3 charge states, respectively.

The key advantage of ESI is that it couples directly to liquid chromatography, enabling online separation of complex peptide mixtures. The multiple charging also means that large peptides (up to ~4 kDa) fall within the \(m/z\) range of most analyzers (typically 300–2,000). ESI is a continuous ionization method, producing a steady stream of ions that the mass analyzer can sample repeatedly.

### MALDI

Matrix-assisted laser desorption/ionization (MALDI) is a pulsed ionization method. The peptide sample is co-crystallized with a large molar excess of an organic matrix—commonly α-cyano-4-hydroxycinnamic acid (CHCA) for peptides—on a metal plate. A pulsed UV laser (typically 337 nm nitrogen or 355 nm Nd:YAG) is fired at the crystal. The matrix absorbs the laser energy, desorbs into the gas phase, and transfers protons to the analyte peptides. MALDI predominantly produces singly charged ions, which simplifies spectral interpretation but limits the mass range for fragmentation.

MALDI is inherently a batch technique: samples are spotted on a plate and analyzed one at a time. It is therefore less suited to high-throughput LC-MS workflows but remains valuable for imaging mass spectrometry, microbial identification (e.g., MALDI-TOF biotyping), and rapid quality control of purified proteins.

### Choosing Between ESI and MALDI

| Feature | ESI | MALDI |
|---|---|---|
| Ionization mechanism | From solution, via charged droplets | From solid matrix, via laser desorption |
| Charge states | Multiple (2+ to 5+ for peptides) | Predominantly singly charged |
| Coupling to LC | Direct (online) | Offline (spot collection) |
| Tolerance to salts | Low (requires cleanup) | Moderate |
| Typical applications | Shotgun proteomics, LC-MS/MS | Imaging, microbial ID, QC |
| Throughput | Continuous, high with LC | Batch, moderate |

For bottom-up protein identification, ESI is the default choice because of its seamless integration with RPLC. MALDI is preferred when sample complexity is low and speed of analysis per sample is critical.

## Mass Analyzers: Measuring Mass-to-Charge Ratio

The mass analyzer separates ions by their \(m/z\) and measures their abundance. Four types dominate modern proteomics: quadrupole, time-of-flight (TOF), Orbitrap, and ion trap. Each has distinct strengths in resolution, mass accuracy, and scan speed.

### Quadrupole Mass Filters

A quadrupole consists of four parallel rods with alternating radiofrequency (RF) and direct current (DC) voltages. Ions travel along the axis between the rods, and only ions with a specific \(m/z\) have a stable trajectory for a given set of voltages. By scanning the voltages, the quadrupole acts as a mass filter, transmitting ions of one \(m/z\) at a time. Quadrupoles have unit resolution (they can distinguish \(m/z\) values differing by 1) and moderate mass accuracy (~0.1 Da). They are inexpensive, robust, and fast, making them ideal for the first stage of a triple quadrupole instrument used in selected reaction monitoring (SRM) or for precursor ion selection in hybrid instruments.

### Time-of-Flight (TOF) Analyzers

A TOF analyzer measures the time it takes for an ion to travel a fixed distance (typically 1–2 meters) in a field-free region. Ions are accelerated by a pulsed electric field to a kinetic energy of \(zeV\), where \(z\) is the charge, \(e\) is the elementary charge, and \(V\) is the acceleration voltage. Since kinetic energy is equal for all ions, velocity is inversely proportional to the square root of \(m/z\). Lighter ions arrive at the detector sooner than heavier ions. The flight time \(t\) is related to \(m/z\) by:

\[
\frac{m}{z} = \frac{2eV}{L^2} t^2
\]

where \(L\) is the flight path length. Modern TOF instruments use a reflectron—an electrostatic mirror that doubles the flight path and corrects for initial kinetic energy spread—achieving resolving powers of 20,000–50,000 and mass accuracy below 5 ppm. TOF analyzers are paired with MALDI sources (MALDI-TOF) and with quadrupoles in Q-TOF hybrid instruments.

### Orbitrap and Ion Trap Analyzers

The Orbitrap is an electrostatic ion trap that measures \(m/z\) by detecting the axial oscillation frequency of ions orbiting a central spindle electrode. Ions are injected into the trap and oscillate along the axis with a frequency \(\omega\) given by:

\[
\omega = \sqrt{\frac{k}{m/z}}
\]

where \(k\) is a constant determined by the instrument geometry and field strength. The detector records the image current produced by the oscillating ions, and a Fourier transform converts the time-domain signal into a frequency spectrum, which is then converted to \(m/z\). The Orbitrap routinely achieves resolving powers of 60,000–240,000 at \(m/z\) 400 and mass accuracy below 3 ppm with [internal calibration](/knowledge/diagnostics/molecular/internal-calibration). This high resolution is critical for distinguishing peptides of similar mass and for assigning charge states unambiguously.

Ion trap analyzers (linear or 3D) capture ions in a RF field and sequentially eject them to a detector by ramping the RF amplitude. They are fast and sensitive but have lower resolution (~10,000) and mass accuracy (~0.1 Da). Ion traps are often used as the second analyzer in LTQ-Orbitrap hybrids, where the Orbitrap performs high-resolution MS1 scans and the ion trap performs rapid MS/MS fragmentation scans.

## Tandem Mass Spectrometry (MS/MS) for Peptide Sequencing

A single MS scan measures the mass of intact peptides, but that information alone cannot identify a protein. Many peptides share the same nominal mass, and the mass of a peptide does not reveal its sequence. Tandem mass spectrometry (MS/MS) solves this problem by fragmenting a selected peptide and measuring the masses of the resulting fragment ions. The fragment masses encode the [amino acid sequence](/blog/guides/amino-acid-sequence).

### Peptide Fragmentation (CID, HCD, ETD)

In an MS/MS experiment, the instrument first performs an MS1 scan to detect all peptide ions. The data system selects a precursor ion of interest (typically the most intense ion above a threshold) and isolates it in a mass filter or trap. The isolated ion is then fragmented, and the fragment ions are mass-analyzed in a second scan (MS2).

The most common fragmentation method is collision-induced dissociation (CID). The precursor ion is accelerated and collided with an inert gas (nitrogen, helium, or argon). The collisions convert kinetic energy into internal vibrational energy, which distributes throughout the peptide. When the internal energy exceeds the activation energy of the peptide backbone, the amide bond breaks. This cleavage produces two fragment series:

- **b-ions**: fragments containing the N-terminus, with the charge retained on the N-terminal side.
- **y-ions**: fragments containing the C-terminus, with the charge retained on the C-terminal side.

The mass difference between consecutive b-ions (or consecutive y-ions) corresponds to the mass of one amino acid residue. By reading the series of mass differences, the peptide sequence can be deduced.

Higher-energy collisional dissociation (HCD) is a variant of CID performed in a dedicated collision cell (rather than an ion trap), which allows detection of low-mass fragment ions (below \(m/z\) 200) that are lost in conventional ion trap CID. HCD is the standard fragmentation method on Orbitrap instruments.

Electron transfer dissociation (ETD) is an alternative that fragments peptides by transferring an electron from a radical anion (e.g., fluoranthene) to the peptide cation. ETD cleaves the N–Cα bond, producing c-ions and z-ions. ETD is particularly useful for highly charged peptides (e.g., from histidine-rich or lysine-rich regions) and for preserving labile post-translational modifications such as phosphorylation, which are often lost during CID.

### Interpreting MS/MS Spectra

A peptide MS/MS spectrum is a plot of fragment ion intensity versus \(m/z\). The spectrum is interpreted by matching observed fragment ion masses to the theoretical masses of b- and y-ions for a candidate peptide sequence. For a peptide of length \(n\), there are \(n-1\) possible b-ions and \(n-1\) possible y-ions (excluding the immonium ions and internal fragments). A high-quality spectrum will show a nearly complete series of y-ions, which is the most informative series because y-ions are usually more abundant than b-ions.

Consider a peptide with sequence AGPDR. The y-ion series would include:

- y1: R (mass 175.12)
- y2: DR (mass 289.16)
- y3: PDR (mass 386.21)
- y4: GPDR (mass 443.23)

The mass difference between y2 and y3 is 97.05 Da, which corresponds to proline (P). The difference between y3 and y4 is 57.02 Da, corresponding to glycine (G). By walking the y-ion series, the sequence is read from the C-terminus toward the N-terminus.

Manual interpretation is feasible for simple spectra, but modern proteomics relies on automated database searching to assign sequences to spectra.

## Database Searching and Protein Identification

The goal of database searching is to determine which peptide sequence best explains each MS/MS spectrum. This is a computational problem: the search algorithm generates theoretical spectra for all peptides in a protein database that match the precursor mass, then scores each candidate against the observed spectrum.

### Sequence Databases and Search Algorithms

The search database is a FASTA-formatted file containing protein sequences from the organism of interest (e.g., *Homo sapiens* from UniProt). The algorithm performs an *in silico* digestion of every protein in the database using the same enzyme specificity as the experimental digestion (e.g., trypsin, allowing for up to two missed cleavages). It then calculates the theoretical \(m/z\) of each peptide and its fragment ions.

Common search engines include:

- **Mascot**: Uses a probability-based Mowse scoring algorithm. The score reflects the probability that the observed match is a random event; higher scores indicate greater confidence.
- **Sequest**: Uses cross-correlation scoring, comparing the observed spectrum to the theoretical spectrum and to a background noise model.
- **MaxQuant (Andromeda)**: An open-source engine integrated with the MaxQuant software suite, widely used for label-free quantification and SILAC.

All search engines require several parameters to be set correctly:

- **Enzyme specificity**: Trypsin (cleave after K/R, not before P).
- **Missed cleavages**: Typically 2 allowed.
- **Precursor mass tolerance**: ±10 ppm for high-resolution instruments; ±0.5 Da for ion traps.
- **Fragment mass tolerance**: ±0.02 Da for Orbitrap HCD; ±0.5 Da for ion trap CID.
- **Fixed modifications**: Carbamidomethylation of cysteine (+57.021 Da) from iodoacetamide treatment.
- **Variable modifications**: Oxidation of methionine (+15.995 Da), acetylation of protein N-termini (+42.011 Da), and phosphorylation of serine/threonine/tyrosine (+79.966 Da) if relevant.

### Scoring and False Discovery Rates

Each peptide-spectrum match (PSM) receives a score. The search engine also calculates a false discovery rate (FDR) by searching the data against a decoy database—a reversed or shuffled version of the target database. Matches to decoy sequences represent false positives. The FDR is calculated as:

\[
\text{FDR} = \frac{\text{Number of decoy hits}}{\text{Number of target hits}}
\]

A common threshold is 1% FDR at the peptide level and 1% at the protein level. This means that among the identified peptides, 1% are expected to be false positives. The target-decoy approach is the standard for controlling error in proteomics, and it is essential for reporting results in publications.

After peptides are identified, they are assembled into proteins. A protein is considered identified if it has at least one unique peptide—a peptide that maps to only one protein in the database. Proteins that share peptides (e.g., isoforms or homologs) are grouped, and the principle of parsimony is applied: the minimal set of proteins that explains all observed peptides is reported.

## Quantitative Mass Spectrometry Approaches

Identifying which proteins are present is often only half the question. Researchers frequently need to know how protein abundance changes between conditions (e.g., treated vs. untreated, healthy vs. diseased). Quantitative proteomics methods fall into two broad categories: label-free and stable isotope labeling.

### Label-Free Quantification

Label-free quantification compares peptide ion intensities (MS1 peak areas) or spectral counts (number of MS/MS spectra assigned to a protein) across different LC-MS runs. The logic is straightforward: more abundant proteins produce more intense peptide peaks and more MS/MS spectra. Label-free methods are simple, inexpensive, and applicable to any sample type, but they require highly reproducible LC-MS conditions and careful normalization between runs. The most common software for label-free analysis is MaxQuant (with the LFQ algorithm) or Skyline for targeted approaches.

### Stable Isotope Labeling (SILAC)

Stable isotope labeling by amino acids in cell culture (SILAC) is a metabolic labeling method. Cells are grown in media containing either "light" (natural) or "heavy" (e.g., \(^{13}\)C\(_6\)-lysine and \(^{13}\)C\(_6\)-arginine) amino acids. The heavy amino acids are incorporated into newly synthesized proteins. After mixing equal amounts of light and heavy cell lysates, the proteins are digested and analyzed. Each peptide appears as a pair of peaks separated by a known mass difference (e.g., 6 Da for lysine-containing peptides). The ratio of light to heavy peak intensities reflects the relative protein abundance between the two conditions. SILAC is highly accurate because the samples are mixed before digestion, eliminating run-to-run variability. Its limitation is that it requires metabolically active cells; it cannot be applied to tissues or clinical samples directly.

### Isobaric Tags (TMT/iTRAQ)

Isobaric tags for relative and absolute quantification (iTRAQ) and tandem mass tags (TMT) are chemical labeling methods that can multiplex up to 16 samples (TMTpro). Each sample is labeled with a tag of identical total mass but containing a reporter ion of unique mass (e.g., 126–134 Da for TMT). The tags are attached to primary amines (lysine side chains and N-termini) after digestion. Labeled peptides from all samples are combined and analyzed in a single LC-MS run. In MS1, the peptides from all samples appear as a single peak because the tags are isobaric. During MS/MS fragmentation, the tag cleaves to release the reporter ion, and the relative intensities of the reporter ions in the low-mass region of the MS/MS spectrum reflect the relative abundance of the peptide in each sample. TMT enables high-throughput multiplexed quantification and is the method of choice for large clinical cohorts. For a deeper treatment of these methods, see [Protein Quantification Mass Spectrometry](/knowledge/molecular-biology/protein-quantification-mass-spectrometry) and the broader topic of [Quantitative Determination of Proteins](/knowledge/molecular-biology/quantitative-determination-of-proteins).

## Common Pitfalls and Troubleshooting in Protein Identification

Even experienced researchers encounter failures in proteomics experiments. Understanding the most common failure modes is essential for troubleshooting and for interpreting results critically.

### Sample Contamination

Keratin from skin and hair is the most notorious contaminant in proteomics. It appears in nearly every sample, regardless of the source, and produces intense peptide peaks that consume MS/MS acquisition time and inflate the protein list. Prevention is the best cure: wear gloves and a lab coat, use dedicated reagents, and avoid opening tubes unnecessarily. If keratin is detected in the results, it is usually present in the blank and can be subtracted computationally, but the best practice is to minimize its introduction. Other common contaminants include trypsin autolysis peptides (from the digestion enzyme itself) and bovine serum albumin (BSA) if it was used as a blocking agent or carrier.

### Incomplete Digestion

Missed cleavages occur when trypsin fails to cleave at every K/R site. This is more common when digestion time is too short, the enzyme-to-substrate ratio is too low, or the protein is resistant to denaturation. Missed cleavages produce longer peptides that may be outside the optimal mass range for MS and that complicate database searching. The search engine can accommodate up to two missed cleavages, but excessive missed cleavages indicate a digestion problem. Solutions include increasing digestion time to 18 hours, adding more trypsin, or including a chaotrope (e.g., 0.1% RapiGest or 20% acetonitrile) to improve enzyme accessibility.

### Database Search Pitfalls

The most common search-related errors are:

- **Wrong database**: Searching a human sample against a mouse database will produce few matches. Always verify the organism and the database version.
- **Incorrect enzyme specificity**: If the sample was digested with Lys-C (cleaves after K only) but the search specifies trypsin, many peptides will not be found.
- **Incorrect modification settings**: Forgetting to set carbamidomethylation of cysteine as a fixed modification will cause a systematic mass error of 57 Da on every cysteine-containing peptide, preventing identification.
- **Too narrow or too wide mass tolerance**: A precursor tolerance of ±10 ppm is appropriate for Orbitrap data; using ±0.5 Da when the instrument is capable of 3 ppm will produce false negatives. Conversely, using ±20 ppm when the instrument is poorly calibrated will produce false positives.
- **Ignoring the FDR**: Reporting all matches without applying a 1% FDR threshold will include many false identifications.

A related issue is the misinterpretation of protein groups. When a peptide maps to multiple proteins, the search engine groups them. Reporting all members of a group as "identified" overstates the confidence. The convention is to report only the leading protein (the one with the most unique peptides) and to note the others as "indistinguishable."

Finally, beware of the "one-hit wonder"—a protein identified by a single peptide. While this can be a true identification, it is less reliable than a protein identified by multiple unique peptides. For critical conclusions, require at least two unique peptides per protein.

## Frequently Asked Questions

### How does mass spec identify proteins?

Mass spectrometry identifies proteins by digesting them into peptides, measuring the mass-to-charge ratio of those peptides, fragmenting them to generate sequence-informative MS/MS spectra, and matching those spectra against a protein database. The peptide sequences are then assembled into protein identifications. This is the bottom-up approach, and it is the standard method in proteomics.

### What is the basic principle of mass spectrometry?

The basic principle is the measurement of the mass-to-charge ratio (\(m/z\)) of gas-phase ions. A sample is ionized, the ions are separated by their \(m/z\) in a mass analyzer, and their abundance is recorded. The resulting mass spectrum provides the molecular mass of the analyte, and with fragmentation, structural information.

### Why do we digest proteins into peptides before mass spec?

Intact proteins are too large and heterogeneous for efficient ionization and fragmentation. Peptides (6–20 amino acids) are within the optimal mass range for mass analyzers, ionize efficiently, and fragment predictably. Digestion with trypsin produces peptides with a basic residue at the C-terminus, which enhances ionization and simplifies database searching.

### What is the difference between ESI and MALDI?

ESI ionizes peptides from solution and produces multiply charged ions; it couples directly to liquid chromatography. MALDI ionizes peptides from a solid matrix using a laser and produces predominantly singly charged ions; it is a batch technique. ESI is preferred for LC-MS/MS shotgun proteomics; MALDI is used for imaging, microbial identification, and rapid analysis.

### What is MS/MS and why is it important?

MS/MS (tandem mass spectrometry) involves isolating a peptide ion, fragmenting it, and measuring the fragment ion masses. The fragment masses encode the [amino acid sequence](/blog/guides/amino-acid-sequence), allowing the peptide to be identified. Without MS/MS, a peptide mass alone cannot uniquely identify a protein.

### How does database searching work in proteomics?

Database searching generates theoretical peptide masses and fragment ion masses from protein sequences in a FASTA database, then scores each observed MS/MS spectrum against these theoretical spectra. The best match is reported as a peptide-spectrum match (PSM), and the FDR is controlled using a decoy database.

### What are common mistakes in mass spec protein identification?

Common mistakes include keratin contamination, incomplete trypsin digestion, incorrect database or search parameters, ignoring the FDR, and overinterpreting single-peptide protein identifications. Proper sample handling, careful parameter setting, and adherence to FDR thresholds prevent most errors.

## Key Takeaways

- Mass spectrometry identifies proteins by measuring the \(m/z\) of peptide ions and fragmenting them to derive sequence information.
- The bottom-up workflow—protein extraction, trypsin digestion, LC-MS/MS, and database searching—is the standard approach.
- ESI is the ionization method of choice for LC-MS; MALDI is used for batch and imaging applications.
- Mass analyzers (quadrupole, TOF, Orbitrap, ion trap) differ in resolution, accuracy, and speed; Orbitrap offers the highest performance for proteomics.
- MS/MS fragmentation produces b- and y-ions whose mass differences reveal the peptide sequence.
- Database searching with target-decoy FDR control is essential for confident protein identification.
- Quantitative methods (label-free, SILAC, TMT) extend MS from identification to measuring protein abundance changes.
- Sample contamination, incomplete digestion, and incorrect search parameters are the most common sources of error; rigorous controls and parameter validation are critical.

## Further Reading

- Marshall AG, Hendrickson CL. *High-resolution mass spectrometers*. Annual review of analytical chemistry (Palo Alto, Calif.). 2008. [PubMed 20636090](https://doi.org/10.1146/annurev.anchem.1.031207.112945)
- Woods AG et al. *Mass Spectrometry for Proteomics-Based Investigation*. Advances in experimental medicine and biology. 2019. [PubMed 31347039](https://doi.org/10.1007/978-3-030-15950-4_1)
- Woods AG et al. *Mass spectrometry for proteomics-based investigation*. Advances in experimental medicine and biology. 2014. [PubMed 24952176](https://doi.org/10.1007/978-3-319-06068-2_1)
- Andersen JS, Mann M. *[Functional genomics](/blog/guides/functional-genomics) by mass spectrometry*. FEBS letters. 2000. [PubMed 10967324](https://doi.org/10.1016/s0014-5793(00)01773-7)
- Yates JR 3rd. *Mass spectrometry and the age of the proteome*. Journal of mass spectrometry : JMS. 1998. [PubMed 9449829](https://doi.org/10.1002/(SICI)1096-9888(199801)33:1<1::AID-JMS624>3.0.CO;2-9)
- Zhu P et al. *Mass spectrometry of peptides and proteins from human blood*. Mass spectrometry reviews. 2011. [PubMed 24737629](https://doi.org/10.1002/mas.20291)

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)