# Mass Spectrometry to Identify Proteins: A Practical Guide

## Introduction to Mass Spectrometry for Protein Identification

### What is Mass Spectrometry?

Mass spectrometry is an analytical technique that measures the mass-to-charge ratio (\( m/z \)) of gas-phase ions. In protein identification, the fundamental goal is to determine the [amino acid sequence](/blog/guides/amino-acid-sequence) of a protein by measuring the masses of its constituent peptides and their fragments. The instrument accomplishes this by converting neutral molecules into charged ions, separating those ions based on their \( m/z \) values in a mass analyzer, and detecting them to produce a mass spectrum—a plot of ion abundance versus \( m/z \).

The power of mass spectrometry lies in its precision. Modern instruments can measure peptide masses with an accuracy of 1–5 parts per million (ppm), meaning a peptide of 1000 daltons (Da) can be measured to within 0.001–0.005 Da. This precision allows researchers to infer elemental composition and, critically, to match experimental data against predicted masses derived from [protein sequence](/blog/guides/protein-sequence) databases.

### Why Use Mass Spectrometry for Proteins?

Proteins are not directly amplifiable like nucleic acids. There is no protein equivalent of PCR. Mass spectrometry is the only method that provides direct, sequence-specific information about proteins at scale. Edman degradation, the classical sequencing method, is limited to single purified proteins and fails for N-terminally modified proteins. Mass spectrometry overcomes these limitations: it can analyze complex mixtures, identify post-translational modifications (PTMs), and determine the relative abundance of thousands of proteins in a single experiment. For these reasons, mass spectrometry has become the cornerstone of modern proteomics—the large-scale study of proteins expressed by a genome, cell, or tissue.

## The Mass Spectrometry Workflow for Proteins

The typical [bottom-up proteomics](/knowledge/bioinformatics/bottom-up-proteomics-principles-workflow-and-applications) workflow—so named because proteins are digested into peptides before analysis—follows a series of defined steps. Each step is critical, and failures at any stage compromise the final identification.

### Protein Digestion and Peptide Generation

Proteins are too large and heterogeneous for direct mass spectrometric analysis in most workflows. A 50 kDa protein, when ionized, produces a complex charge-state envelope that is difficult to fragment predictably. The solution is to digest proteins into peptides of 7–35 amino acids, a size range optimal for both chromatography and mass spectrometry.

The enzyme of choice is trypsin, a serine protease that cleaves peptide bonds C-terminal to arginine (R) and lysine (K) residues, unless followed by proline (P). Trypsin digestion generates peptides with a basic residue at the C-terminus, which ionizes efficiently in positive-ion mode—a property that makes tryptic peptides ideal for mass spectrometry. The digestion is typically performed at 37°C for 12–18 hours in 50 mM ammonium bicarbonate buffer (pH 8.0) at an enzyme-to-protein ratio of approximately 1:50 (w/w). Reduction and alkylation of cysteine residues precede digestion: dithiothreitol (DTT) at 5–10 mM reduces disulfide bonds, and iodoacetamide at 15–25 mM alkylates the free thiols to prevent disulfide reformation. This step is essential because disulfide-linked peptides complicate both chromatography and database searching.

### Liquid Chromatography Separation

A typical proteomics sample contains thousands of distinct peptides. The [mass spectrometer](/knowledge/molecular-biology/mass-spectrometer) cannot analyze all of them simultaneously with adequate sensitivity. Therefore, peptides are separated by reverse-phase liquid chromatography (RP-LC) before entering the [mass spectrometer](/knowledge/molecular-biology/mass-spectrometer). The separation uses a hydrophobic stationary phase, typically C18 silica particles packed into a fused-silica capillary column (75 µm inner diameter, 15–25 cm length). Peptides are loaded in an aqueous mobile phase containing 0.1% formic acid (v/v) and eluted with a gradient of increasing acetonitrile (typically 5% to 35% over 60–120 minutes). The 0.1% formic acid serves a dual purpose: it protonates peptides for positive-ion mode detection and provides the volatile counter-ion necessary for stable electrospray.

The elution order is governed by hydrophobicity: hydrophilic peptides elute early, hydrophobic peptides late. This online coupling of liquid chromatography to mass spectrometry (LC-MS) reduces sample complexity by delivering peptides to the mass spectrometer in order of increasing hydrophobicity, allowing the instrument to acquire spectra on relatively few peptides at any given moment.

### Ionization Techniques: ESI and MALDI

Two ionization techniques dominate protein mass spectrometry: electrospray ionization (ESI) and matrix-assisted laser desorption/ionization (MALDI). Both are "soft" ionization methods that produce intact gas-phase ions without extensive fragmentation.

**Electrospray ionization (ESI)** operates at atmospheric pressure. The liquid eluent from the LC column passes through a narrow capillary held at a high voltage (2–5 kV relative to the inlet). The applied electric field disperses the liquid into a fine mist of charged droplets. As solvent evaporates, droplet size decreases and charge density increases until Coulombic repulsion exceeds surface tension, causing droplet fission and ultimately releasing desolvated peptide ions. ESI produces multiply charged ions: a 1500 Da peptide typically carries 2–3 protons, giving \( m/z \) values of 500–750, well within the detection range of most analyzers. ESI is the method of choice for LC-MS because it is compatible with continuous liquid flow.

**Matrix-assisted laser desorption/ionization (MALDI)** uses a pulsed UV laser (typically 337 nm nitrogen or 355 nm Nd:YAG) to desorb and ionize peptides co-crystallized with an organic matrix, most commonly α-cyano-4-hydroxycinnamic acid (CHCA) for peptides. The matrix absorbs the laser energy, causing rapid heating and sublimation that carries peptides into the gas phase. MALDI predominantly produces singly charged ions, which simplifies spectral interpretation but limits fragmentation efficiency in MS/MS. MALDI is inherently a batch technique—samples are spotted onto a target plate—making it less amenable to online LC coupling, though off-line fraction collection is possible.

The choice between ESI and MALDI depends on the application. ESI is preferred for high-throughput shotgun proteomics because of its seamless LC integration. MALDI is valuable for rapid sample profiling, imaging mass spectrometry, and analysis of samples that tolerate batch processing. For a detailed comparison of these ionization methods, see [Mass Spectrometry Work for Proteins](/knowledge/molecular-biology/mass-spectrometry-work-for-proteins).

## Key Mass Analyzers and Their Principles

The mass analyzer separates ions by their \( m/z \) ratio. Four analyzer types dominate protein proteomics, each with distinct physical principles, strengths, and limitations.

### Quadrupole Mass Analyzers

A quadrupole consists of four parallel metal rods arranged in a square array. Opposite rods are connected electrically, and a combination of direct current (DC) and radiofrequency (RF) voltages is applied. The RF voltage creates an oscillating electric field that causes ions to follow complex helical trajectories. For a given set of voltages, only ions of a specific \( m/z \) have stable trajectories and reach the detector; all others collide with the rods and are lost. By scanning the voltages, the quadrupole sequentially transmits ions of increasing \( m/z \), producing a mass spectrum.

Quadrupoles are robust, inexpensive, and fast, but they offer limited resolution (typically 0.5–1 Da) and mass accuracy (100–500 ppm). They are rarely used alone for protein identification. Instead, they serve as mass filters in hybrid instruments: a triple quadrupole (Q1-Q2-Q3) is the workhorse for targeted quantification, and a quadrupole coupled to a time-of-flight analyzer (Q-TOF) provides high-resolution MS/MS data.

### Time-of-Flight (TOF) Analyzers

A time-of-flight analyzer measures the time ions take to travel a fixed distance (typically 1–2 meters) under the influence of a constant electric field. All ions are accelerated to the same kinetic energy, so their velocity depends on their mass: lighter ions travel faster and arrive at the detector sooner. The flight time is converted to \( m/z \) using the equation \( m/z = 2eVt^2/d^2 \), where \( e \) is the elementary charge, \( V \) is the acceleration voltage, \( t \) is flight time, and \( d \) is flight path length.

Modern TOF analyzers incorporate a reflectron—an electrostatic mirror that reverses the ion beam and compensates for small differences in initial kinetic energy. This design doubles the flight path and improves resolution to 20,000–50,000. TOF analyzers offer high mass accuracy (2–10 ppm) and a virtually unlimited \( m/z \) range, making them ideal for MALDI-MS and for intact protein analysis. Their main limitation is a lower dynamic range compared to ion traps.

### Ion Trap and Orbitrap Analyzers

**Ion traps** store ions in a three-dimensional RF field and then eject them sequentially by mass. The two main types are the 3D quadrupole ion trap (QIT) and the linear ion trap (LTQ). Ion traps excel at MS/MS because they can isolate a precursor ion, fragment it, and analyze the product ions in the same device. They offer high sensitivity and fast scan rates but have limited resolution (~10,000) and mass accuracy (~100 ppm), and they suffer from space-charge effects—ion-ion repulsion that degrades performance when too many ions are stored.

**Orbitrap** analyzers represent a breakthrough in high-resolution mass spectrometry. The Orbitrap consists of a central spindle electrode surrounded by an outer barrel electrode. Ions injected into the trap orbit the spindle while oscillating axially. The frequency of this axial oscillation is independent of initial ion velocity and inversely proportional to the square root of \( m/z \). A detector measures the image current produced by the oscillating ions, and a Fourier transform converts the time-domain signal into a mass spectrum.

The Orbitrap achieves resolving powers of 100,000–1,000,000 and mass accuracy below 1 ppm with [internal calibration](/knowledge/diagnostics/molecular/internal-calibration). This performance enables confident peptide identification and accurate determination of charge states. Orbitraps are almost always coupled to a linear ion trap or quadrupole for precursor selection and fragmentation, forming hybrid instruments such as the Q-Exactive and Orbitrap Fusion series. These instruments dominate contemporary proteomics because they combine high resolution, high sensitivity, and fast scan speeds.

| Analyzer | Resolution | Mass Accuracy | MS/MS Capability | Typical Use |
|----------|------------|---------------|------------------|-------------|
| Quadrupole | 1,000–2,000 | 100–500 ppm | Yes (triple quad) | Targeted quantification (SRM/MRM) |
| TOF | 20,000–50,000 | 2–10 ppm | Yes (Q-TOF) | MALDI-MS, intact proteins, discovery |
| Ion Trap | 5,000–10,000 | 50–100 ppm | Yes (excellent) | MS/MS-intensive discovery |
| Orbitrap | 100,000–1,000,000 | <1–5 ppm | Yes (hybrid) | High-resolution discovery, PTM analysis |

## Tandem Mass Spectrometry (MS/MS) for Peptide Sequencing

A mass spectrum of intact peptides provides only their masses—insufficient information for unambiguous protein identification. Tandem mass spectrometry (MS/MS) adds a second dimension: fragmentation. In an MS/MS experiment, a peptide of interest (the precursor ion) is isolated in the mass analyzer, fragmented by collision with inert gas molecules, and the resulting fragment ions are mass-analyzed. The fragment masses encode the amino acid sequence.

### Peptide Fragmentation Patterns

The most common fragmentation method in proteomics is collision-induced dissociation (CID), also called collisionally activated dissociation (CAD). In CID, the isolated precursor ion is accelerated and collided with nitrogen or helium gas. The collisions convert kinetic energy into internal vibrational energy, which distributes throughout the peptide and causes cleavage of the amide (peptide) bonds.

Fragmentation of a peptide at the amide bond generates two complementary ion series. If the charge is retained on the N-terminal fragment, the ion is designated a **b-ion**; if retained on the C-terminal fragment, it is a **y-ion**. The nomenclature follows the system proposed by Roepstorff and Fohlman: a, b, and c ions retain the N-terminus, while x, y, and z ions retain the C-terminus. In practice, CID predominantly produces b- and y-ions.

The mass difference between consecutive y-ions (or b-ions) corresponds to the mass of one amino acid residue. For example, if a peptide yields y-ions at \( m/z \) 175.12 (y₁, glycine), 262.15 (y₂, glycine + alanine), and 389.21 (y₃, glycine + alanine + glutamine), the C-terminal sequence is Gly-Ala-Gln. By reading the mass differences across the entire fragment ion series, the complete peptide sequence can be deduced. In practice, the spectrum is not perfectly complete—some fragment ions are weak or absent—so the sequence is inferred by matching the entire fragmentation pattern against a database rather than by manual de novo sequencing.

### Database Searching with MS/MS Data

The experimental MS/MS spectrum is a pattern of fragment ion masses. The challenge is to determine which peptide sequence produced this pattern. The standard approach is database searching: a search engine predicts the MS/MS spectra for all peptides in a protein database that match the precursor mass, then scores how well each predicted spectrum matches the experimental one.

The search process begins with the precursor ion mass, which constrains the candidate peptide mass to within a tolerance window (typically ±10 ppm for high-resolution instruments). The enzyme specificity (e.g., trypsin) further restricts candidates to peptides with appropriate C-terminal arginine or lysine. For each candidate peptide, the search engine generates theoretical b- and y-ion masses and compares them to the experimental fragment ions. A peptide is identified when its predicted fragment ions match a significant number of experimental peaks.

The output of a database search is a list of peptide-spectrum matches (PSMs), each with a score reflecting the quality of the match. The most widely used search engines are MASCOT, Sequest, and Andromeda (the engine integrated into MaxQuant). Each uses a different scoring algorithm: MASCOT uses a probability-based score, Sequest uses cross-correlation, and Andromeda uses a score based on the number and intensity of matched fragment ions. For a deeper explanation of how these tools fit into the broader workflow, see __MASK_2__.

## Protein Identification via Database Searching

### Sequence Databases and Search Engines

The choice of protein sequence database is a critical decision. The most commonly used databases are UniProt (which includes Swiss-Prot, a manually curated, reviewed subset, and TrEMBL, an automatically annotated, unreviewed set) and NCBI's RefSeq. For human samples, the standard is the UniProt human proteome, containing approximately 20,000 canonical protein sequences. For model organisms, the corresponding proteome databases are used.

The database must contain the protein sequences present in the sample. If the organism is not well represented in the database, or if the sample contains contaminants from another species (e.g., bovine serum albumin from cell culture media), identifications will fail or be misattributed. A common practice is to append common contaminants (trypsin, keratins, serum proteins) to the database so that these frequent contaminants are identified rather than causing false matches.

Search engines also require parameters that define the search space. Key parameters include:

1. **Enzyme specificity**: typically trypsin with up to 2 missed cleavages (sites where trypsin failed to cleave).
2. **Precursor mass tolerance**: ±10 ppm for Orbitrap data, ±20 ppm for Q-TOF data.
3. **Fragment mass tolerance**: ±0.02 Da for high-resolution MS/MS, ±0.5 Da for ion trap data.
4. **Fixed modifications**: e.g., carbamidomethylation of cysteine (+57.021 Da) from iodoacetamide treatment.
5. **Variable modifications**: e.g., oxidation of methionine (+15.995 Da), acetylation of protein N-termini (+42.011 Da).

### Scoring and Statistical Validation

A raw search score is insufficient to establish confidence. The number of candidate peptides in a database is enormous—a typical search evaluates millions of peptide candidates—so random matches are inevitable. Statistical validation is therefore essential.

The most widely used approach is the **target-decoy search strategy**. The search is performed against a concatenated database containing both the target sequences (the real proteome) and decoy sequences (reversed or shuffled versions of the target sequences). Decoy peptides cannot be truly present in the sample, so any PSM matching a decoy sequence represents a false positive. The false discovery rate (FDR) is calculated as:

\[ \text{FDR} = \frac{\text{Number of decoy hits}}{\text{Number of target hits}} \]

For example, if a search yields 1,000 target PSMs and 10 decoy PSMs at a given score threshold, the FDR is 10/1000 = 1%. The standard threshold for reporting protein identifications is 1% FDR at the peptide level and 1% at the protein level. This statistical framework ensures that reported identifications are reproducible and not artifacts of random matching.

## Quantitative Proteomics and Beyond

Identifying which proteins are present is only half the story. Biological questions often require knowing how protein abundance changes between conditions—between healthy and diseased tissue, for example. Mass spectrometry provides several strategies for quantitative proteomics, broadly divided into label-free and stable isotope labeling approaches. For a comprehensive overview of quantification strategies, see __MASK_3__ and __MASK_4__.

### Label-Free Quantification

Label-free quantification compares peptide ion intensities or spectral counts across different LC-MS runs. In the intensity-based approach, the area under the chromatographic peak for each peptide is integrated and used as a proxy for abundance. In the spectral count approach, the number of MS/MS spectra assigned to a protein is counted; more abundant proteins generate more spectra. Label-free methods are simple and cost-effective, requiring no additional reagents, but they suffer from run-to-run variability in chromatography and ionization efficiency. Normalization strategies, such as adjusting total ion current across runs, partially compensate for this variability.

### Stable Isotope Labeling (SILAC, TMT)

Stable isotope labeling introduces mass differences into peptides that can be distinguished in the mass spectrometer. **SILAC** (Stable Isotope Labeling by Amino acids in Cell culture) involves growing cells in media containing either light (¹²C₆-lysine and ¹²C₆-arginine) or heavy (¹³C₆-lysine and ¹³C₆-arginine) amino acids. After several cell doublings, all proteins incorporate the labeled amino acids. The light and heavy cell populations are mixed, digested, and analyzed together. Each peptide appears as a pair of peaks separated by a known mass difference (6 Da for lysine, 10 Da for arginine), and the intensity ratio of the pair reflects the relative abundance of the protein in the two conditions.

**TMT** (Tandem Mass Tag) labeling uses isobaric chemical tags that react with primary amines (lysine side chains and peptide N-termini). Each tag has the same total mass but contains a reporter group of different mass (126–131 Da) and a balancer group of complementary mass. When labeled peptides from different samples are mixed, they co-elute and appear as a single precursor peak. Upon fragmentation in MS/MS, the reporter ions are released, and their relative intensities reflect the relative abundance of the peptide in each sample. TMT enables multiplexed quantification of up to 16 samples in a single experiment, making it a powerful tool for large-scale comparative studies. For more on the principles of isotope-based quantification, see __MASK_5__ and __MASK_6__.

Beyond quantification, mass spectrometry is the premier method for characterizing post-translational modifications. Phosphorylation, acetylation, ubiquitination, and glycosylation can all be identified by the characteristic mass shifts they impart on peptides. For example, phosphorylation adds 79.966 Da (HPO₃) to a peptide, and this modification is confirmed by the presence of a neutral loss of 98 Da (H₃PO₄) in the MS/MS spectrum. Enrichment strategies, such as immobilized metal affinity chromatography (IMAC) for phosphopeptides, are typically required to detect these low-abundance modifications.

## Common Pitfalls and Troubleshooting in Protein Identification

Even with a well-designed experiment, protein identification can fail. The following are the most frequent sources of error and practical strategies to avoid them.

### Sample Contamination and Cleanup

Keratin contamination from skin, hair, and dust is the most common source of false protein identifications. Human keratins (KRT1, KRT2, KRT9, KRT10) appear in nearly every sample if proper precautions are not taken. To minimize contamination:

- Wear gloves and a lab coat at all times.
- Use HPLC-grade solvents and fresh reagents.
- Avoid opening tubes unnecessarily.
- Include a "blank" sample (buffer only) processed in parallel to identify contaminant proteins.
- Add common contaminants to the search database so they are identified as such rather than causing false matches.

Sample cleanup is equally important. Salts, detergents, and other non-volatile contaminants suppress ionization and degrade chromatography. C18 reverse-phase cleanup using pipette tips (e.g., ZipTips) or spin columns removes salts and hydrophilic contaminants. For samples containing high concentrations of SDS or other detergents, a precipitation step (acetone or chloroform-methanol) or a detergent removal column is necessary before digestion.

### Optimizing Digestion and Separation

Incomplete digestion is a frequent problem. Missed cleavages increase peptide mass and reduce the efficiency of both chromatography and MS/MS. Common causes include insufficient enzyme, short digestion time, or the presence of chaotropes (e.g., urea) that denature trypsin. The standard protocol—1:50 trypsin-to-protein ratio, 37°C, 12–18 hours—should be followed. If digestion is consistently incomplete, increase the enzyme ratio to 1:20 or add a second aliquot of trypsin after 4 hours.

Over-digestion is less common but can occur with prolonged incubation, producing peptides too short for reliable identification. Peptides shorter than 7 amino acids are generally not identified because they provide insufficient fragment ions.

Chromatography problems manifest as poor peak shape, low signal, or carryover. Ensure the LC column is properly equilibrated, use fresh mobile phases, and include a blank run between samples to monitor carryover. If the column pressure is elevated, the column may be clogged; replace the frit or column.

### Interpreting Search Results Correctly

The most common interpretation error is over-reporting identifications. A single peptide-spectrum match (PSM) does not constitute a confident protein identification, especially if the peptide sequence is shared among multiple proteins. The standard is to require at least two unique peptides per protein, each with a PSM FDR below 1%. Proteins identified by a single peptide should be flagged as "one-hit wonders" and treated with caution.

Another error is ignoring the distinction between protein groups. When a peptide sequence is shared by multiple proteins (e.g., isoforms or homologs), search engines group these proteins together. Reporting a specific isoform as identified when the evidence only supports the shared peptide is an over-interpretation. The correct practice is to report the protein group and note that the specific isoform cannot be distinguished.

Finally, be aware of the "protein inference problem": a peptide may match multiple proteins in the database, and the search engine must assign the peptide to the most parsimonious set of proteins. Tools like MaxQuant and Proteome Discoverer handle this automatically, but the user must understand the output to avoid misannotation.

## Practical Summary: From Sample to Identified Protein

### Checklist for a Successful Experiment

A successful bottom-up proteomics experiment follows a logical sequence. Use this checklist to ensure no step is overlooked:

1. **Sample preparation**: Lyse cells or tissue in a buffer containing 8 M urea or 2% SDS, with protease and phosphatase inhibitors. Quantify total protein using a BCA or Bradford assay.
2. **Reduction and alkylation**: Add DTT to 5–10 mM, incubate at 56°C for 30 minutes. Add iodoacetamide to 15–25 mM, incubate in the dark at room temperature for 30 minutes.
3. **Digestion**: Dilute the urea to below 1 M (or remove SDS), add trypsin at 1:50 (w/w), incubate at 37°C for 12–18 hours. Quench with formic acid to 1% (v/v).
4. **Cleanup**: Desalt using C18 tips or columns. Dry the eluted peptides in a vacuum centrifuge.
5. **LC-MS/MS**: Resuspend peptides in 0.1% formic acid. Load 0.5–2 µg onto the column. Run a 60–120 minute gradient from 5% to 35% acetonitrile.
6. **Database searching**: Search the raw data against the appropriate proteome database with trypsin specificity, 2 missed cleavages, carbamidomethylation (C) as fixed, oxidation (M) as variable. Set precursor tolerance to ±10 ppm and fragment tolerance to ±0.02 Da.
7. **Validation**: Filter to 1% FDR at the peptide and protein levels. Require at least two unique peptides per protein.
8. **Interpretation**: Review the identified proteins in the biological context. Check for expected contaminants and verify that the identified proteins are consistent with the sample type.

### Resources for Further Learning

The field of proteomics is vast, and the resources below provide entry points for deeper study:

- **MaxQuant** (maxquant.org): Free, open-source software for quantitative proteomics, integrates the Andromeda search engine.
- **Proteome Discoverer** (Thermo Fisher): Commercial software with a graphical interface for data analysis.
- **UniProt** (uniprot.org): The primary protein sequence database.
- **PRIDE Archive** (ebi.ac.uk/pride): A public repository for mass spectrometry data, useful for exploring published datasets.
- **The Global Proteome Machine** (thegpm.org): A free search engine and data repository.

For a more detailed treatment of the instrumentation and its capabilities, see __MASK_7__ and __MASK_8__.

## Frequently Asked Questions

### How does mass spectrometry identify proteins?

Mass spectrometry identifies proteins by digesting them into peptides, measuring the mass-to-charge ratios of those peptides, fragmenting the peptides in MS/MS to generate sequence-informative fragment ions, and matching the resulting spectra against a protein sequence database. The match is scored statistically, and proteins are reported when the cumulative evidence from multiple peptides exceeds a defined confidence threshold (typically 1% FDR).

### What is the basic principle of mass spectrometry?

Mass spectrometry measures the mass-to-charge ratio (\( m/z \)) of gas-phase ions. The instrument has three essential components: an ion source that converts neutral molecules into ions, a mass analyzer that separates ions by \( m/z \) using electric or magnetic fields, and a detector that records ion abundance. The resulting mass spectrum plots ion intensity against \( m/z \), providing both qualitative (what is present) and quantitative (how much is present) information.

### What are the main steps in a typical proteomics experiment?

The main steps are: (1) protein extraction and quantification, (2) reduction and alkylation of cysteines, (3) enzymatic digestion (typically with trypsin), (4) peptide cleanup and desalting, (5) liquid chromatography separation, (6) electrospray ionization and mass analysis, (7) tandem mass spectrometry (MS/MS) for peptide fragmentation, (8) database searching, and (9) statistical validation and biological interpretation.

### What is the difference between ESI and MALDI?

Electrospray ionization (ESI) ionizes peptides from a liquid solution at atmospheric pressure, producing multiply charged ions, and is directly compatible with liquid chromatography. Matrix-assisted laser desorption/ionization (MALDI) ionizes peptides from a solid matrix using a pulsed laser, predominantly producing singly charged ions, and is a batch technique requiring sample spotting on a target plate. ESI is preferred for high-throughput LC-MS/MS; MALDI is used for rapid profiling and imaging applications.

### What is tandem mass spectrometry (MS/MS)?

Tandem mass spectrometry is a two-stage mass analysis. In the first stage, a peptide of interest (the precursor ion) is isolated by its \( m/z \). In the second stage, the isolated peptide is fragmented—typically by collision with inert gas—and the resulting fragment ions are mass-analyzed. The fragment ion masses, primarily b- and y-ions, provide sequence information that allows the peptide to be identified.

### How do search engines like MASCOT or Sequest work?

Search engines compare experimental MS/MS spectra against theoretical spectra predicted from a protein sequence database. They first filter candidate peptides by precursor mass and enzyme specificity, then generate predicted fragment ion masses for each candidate. The experimental and predicted spectra are scored for similarity—MASCOT uses a probability-based score, Sequest uses cross-correlation, and Andromeda uses a weighted match score. The highest-scoring peptide-spectrum match is reported, and its statistical significance is assessed using the target-decoy approach.

### What is a false discovery rate (FDR) in proteomics?

The false discovery rate is the expected proportion of false positive identifications among all identifications reported above a given score threshold. It is estimated by searching the data against a decoy database containing reversed or shuffled protein sequences. Decoy hits are known to be false, so the FDR is calculated as the ratio of decoy hits to target hits. An FDR of 1% means that approximately 1 in 100 reported identifications is expected to be incorrect.

### Why is protein digestion necessary before mass spectrometry?

Protein digestion is necessary for three reasons. First, intact proteins are too large for effective chromatographic separation and produce complex charge-state envelopes that complicate mass analysis. Second, peptides of 7–35 amino acids ionize more efficiently and fragment more predictably than intact proteins. Third, database searching relies on matching peptide masses and fragment ions to predicted sequences, which is computationally tractable for peptides but not for intact proteins. Digestion converts the protein identification problem into a peptide identification problem, which is far more amenable to high-throughput analysis.

## Key Takeaways

- Mass spectrometry identifies proteins by measuring the mass-to-charge ratios of peptide ions and their fragments, then matching the data against protein sequence databases.
- The bottom-up workflow—digestion, chromatography, ionization, mass analysis, and database searching—is the standard approach for protein identification.
- Trypsin is the enzyme of choice because it generates peptides with C-terminal arginine or lysine, which ionize efficiently in positive-ion mode.
- ESI is preferred for LC-MS/MS because of its compatibility with liquid flow; MALDI is used for batch analysis and imaging.
- The Orbitrap provides the highest resolution and mass accuracy, enabling confident identification and PTM analysis.
- MS/MS fragmentation generates b- and y-ions whose mass differences encode the peptide sequence.
- Database searching with target-decoy validation ensures that reported identifications are statistically significant, with FDR typically controlled at 1%.
- Quantitative proteomics extends mass spectrometry beyond identification, using label-free or stable isotope labeling (SILAC, TMT) to measure changes in protein abundance.
- Contamination, incomplete digestion, and over-interpretation of single-peptide identifications are the most common pitfalls; rigorous sample handling and statistical validation mitigate these risks.

## Further Reading

- Woods AG et al. *Mass Spectrometry for Proteomics-Based Investigation*. Advances in experimental medicine and biology. 2019. __MASK_9__
- Woods AG et al. *Mass spectrometry for proteomics-based investigation*. Advances in experimental medicine and biology. 2014. __MASK_10__
- Andersen JS, Mann M. *Functional genomics by mass spectrometry*. FEBS letters. 2000. __MASK_11__01773-7)
- Yates JR 3rd. *Mass spectrometry and the age of the proteome*. Journal of mass spectrometry : JMS. 1998. __MASK_12__1096-9888(199801)33:1<1::AID-JMS624>3.0.CO;2-9)
- Zhu P et al. *Mass spectrometry of peptides and proteins from human blood*. Mass spectrometry reviews. 2011. [PubMed 24737629](https://doi.org/10.1002/mas.20291)
- Shevchenko A et al. *In-gel digestion for mass spectrometric characterization of proteins and proteomes*. Nature protocols. 2006. [PubMed 17406544](https://doi.org/10.1038/nprot.2006.468)

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)