# Neoantigen Discovery Using Proteogenomics: Integrating Mass Spectrometry and Genomic Data to Identify Immunogenic Peptides

## At a Glance

Proteogenomics combines genomic and transcriptomic sequencing with mass spectrometry-based proteomics to identify and validate neoantigen peptides from tumor samples. This approach addresses a distinct limitation of genomics-only pipelines: standard exome-based variant calling can predict mutation-derived neoantigens, but it cannot detect peptides arising from post-transcriptional regulation, noncanonical translation, intron retention, or aberrantly expressed noncoding regions. Mass spectrometry provides direct peptide evidence that genomic prediction alone cannot supply.

The table below summarizes the core workflow components, their primary functions, and the practical decisions researchers must make at each stage.

| Workflow Component | Primary Function | Key Decision Point |
| --- | --- | --- |
| Sample collection and nucleic acid extraction | Preserve DNA, RNA, and protein from the same tumor specimen | Coordinate tissue processing so that all three molecular layers come from the same region of the tumor |
| Genomic and transcriptomic sequencing | Identify somatic mutations, expression levels, and splice variants | Choose sequencing depth and read length appropriate for variant calling and transcript reconstruction |
| Custom protein sequence database construction | Translate genomic variants and noncanonical transcripts into peptide sequences for database searching | Decide which variant types and genomic regions to include in the search space |
| Mass spectrometry acquisition | Generate peptide fragmentation spectra from tumor lysates or immunopeptidomes | Select between whole-proteome analysis and MHC-associated peptide analysis |
| Database searching and peptide validation | Match spectra to peptide sequences and control false discovery rate | Set false discovery rate thresholds and decide whether to use target-decoy strategies |
| Neoantigen prioritization and reporting | Rank candidate peptides by immunogenicity and presentation likelihood | Combine binding prediction, expression evidence, and experimental validation data |

Researchers who adopt this workflow gain the ability to detect tumor-specific antigens that originate from noncoding regions, which standard exome-based approaches miss. In two murine cancer cell lines and seven human primary tumors, a proteogenomic approach identified 40 tumor-specific antigens, with about 90% derived from allegedly noncoding regions. Most of these antigens came from nonmutated yet aberrantly expressed transcripts such as endogenous retroelements, and many could be shared across multiple tumor types. The same pattern appears in acute myeloid leukemia, where analysis of the MHC class I-associated immunopeptidome from 19 primary samples identified 58 tumor-specific antigens bearing no mutations, with 86% derived from supposedly noncoding genomic regions and 48% resulting from intron retention and translation.

## The Problem with Genomics-Only Neoantigen Prediction

### Why Mutation-Derived Prediction Falls Short

Standard neoantigen discovery begins with whole-exome or whole-genome sequencing to identify nonsynonymous somatic mutations. The mutant peptide sequences are then tested in silico for predicted binding to the patient's HLA alleles. This approach has produced clinically meaningful results, but it operates under a narrow definition of what constitutes a tumor-specific antigen.

The genomics-only framework assumes that neoantigens are mutant peptides. It does not account for peptides that are tumor-specific because they arise from transcripts that are normally silenced or expressed at very low levels in healthy tissue. It also misses peptides generated through alternative splicing, intron retention, or translation of noncoding regions. These sources of antigenic peptides are invisible to variant calling because they do not involve a DNA sequence change.

The consequence is a systematic blind spot. A genomics-only pipeline can prioritize genomically inferred targets such as copy-number drivers and mutation-derived neoantigens, but it cannot detect the substantial fraction of tumor-specific antigens that originate from noncanonical sources. The colon cancer proteogenomic study demonstrated that integrating proteomic and phosphoproteomic data with genomic analysis produces a catalog of cancer-associated proteins and phosphosites that includes known and putative new biomarkers, drug targets, and cancer/testis antigens. This integration also revealed associations that genomics alone could not establish, such as the link between decreased CD8 T cell infiltration and increased glycolysis in microsatellite instability-high tumors.

### The False-Positive and False-Negative Problem

Current neoantigen prediction tools suffer from low accuracy and high false-positive rates. The challenge is not simply that predictions are imperfect. It is that the entire validation pathway depends on the initial candidate list. If the candidate list is built only from mutated peptides, then every true neoantigen arising from noncanonical translation is absent from the list before any filtering begins.

Mass spectrometry adds a direct experimental layer. When peptide spectra are matched to a custom database that includes variant peptides and noncanonical translations, the resulting identifications carry physical evidence that the peptide exists in the tumor sample. This evidence is qualitatively different from a binding prediction. A predicted binder may or may not be processed and presented. A peptide identified by mass spectrometry from the MHC-associated repertoire has already survived the antigen processing and presentation pathway.

The proteogenomic concept emerged in 2004 and has evolved with increased sequencing depth from next-generation technologies and maturation of mass spectrometry-based proteomics. The approach now allows discovery of high-confidence candidate neoantigens that would be missed by prediction-only workflows.

## Core Principles of Proteogenomic Neoantigen Discovery

### Coordinate Collection of DNA, RNA, and Protein

The foundation of any proteogenomic study is the coordinated collection of nucleic acids and protein from the same tumor specimen. DNA provides the variant information. RNA provides expression evidence and splice isoform structure. Protein provides the direct evidence that a peptide is translated and presented.

Tissue heterogeneity complicates this coordination. A tumor contains multiple cell populations, and the molecular profile of one region may differ from another. Researchers should document the exact tissue region used for each extraction and note any visible heterogeneity in the specimen. This documentation becomes critical when interpreting discordant results between genomic and proteomic data.

### Custom Database Construction

The search space for mass spectrometry data must include the peptide sequences that could arise from the tumor's specific genomic alterations. A standard reference proteome database will not contain variant peptides or peptides from noncanonical transcripts. The custom database must be built from the patient's own sequencing data.

The database construction step requires decisions about which variant types to include. Nonsynonymous single nucleotide variants are the most straightforward. Insertions and deletions create frameshift peptides that can be highly immunogenic because they are entirely novel to the immune system. Gene fusions can generate chimeric peptides spanning the breakpoint. Alternative splicing events and intron retention produce peptides from sequences that are not part of the canonical protein.

The size of the search space directly affects the false discovery rate. A larger database increases the chance of random spectrum-to-peptide matches. Researchers must balance sensitivity against specificity by controlling which genomic regions and variant types enter the database.

### Mass Spectrometry Acquisition Strategy

Two acquisition strategies serve different purposes in neoantigen discovery. Whole-proteome analysis identifies which proteins are expressed in the tumor. Immunopeptidomics analysis isolates and sequences the peptides that are actually bound to MHC molecules on the cell surface.

Immunopeptidomics provides the most direct evidence for neoantigen presentation. The peptides identified from MHC pull-down experiments have already passed through antigen processing. They are the peptides that T cells could potentially recognize. The tradeoff is that immunopeptidomics requires specialized protocols for MHC-peptide complex isolation and produces data that are more difficult to analyze than standard proteomics data.

Whole-proteome analysis is more accessible and can identify tumor-specific proteins that may be processed into presented peptides. The evidence is indirect, but the workflow is more established and the data analysis tools are more mature.

### Integration and Prioritization

The final step integrates all evidence layers to prioritize candidate neoantigens. A peptide that is identified by mass spectrometry, predicted to bind the patient's HLA alleles, and derived from a transcript with tumor-specific expression receives higher priority than a peptide that satisfies only one of these criteria.

The integration step can also identify defects in the antigen presentation machinery. Multiomics integration in the NeoDisc pipeline demonstrated that combining immunopeptidomics, genomics, and transcriptomics can reveal when a tumor has lost the ability to present antigens effectively. This information influences the interpretation of the entire neoantigen landscape and has direct implications for whether a patient is likely to benefit from T cell-based therapies.

## Practical Workflow for Proteogenomic Neoantigen Discovery

### Step 1: Sample Collection and Quality Assessment

Collect tumor tissue and matched normal tissue from the same patient. The normal tissue provides the germline reference needed to distinguish somatic mutations from inherited variants. Document the tissue source, collection time, and storage conditions for each sample.

Assess the quality of DNA, RNA, and protein before proceeding. Degraded RNA will compromise transcript reconstruction. Cross-linked or degraded protein will reduce the number of identifiable peptides. Quality thresholds should be established before the experiment begins, and samples that fail these thresholds should be flagged for potential exclusion.

### Step 2: Sequencing and Variant Calling

Sequence the tumor and normal DNA to identify somatic variants. The sequencing strategy depends on the research question. Whole-exome sequencing is more economical and focuses on protein-coding regions. Whole-genome sequencing captures variants in noncoding regions that may create novel open reading frames.

RNA sequencing provides expression data and enables detection of splice variants and fusion transcripts. The RNA data also provide the evidence needed to determine whether a variant is actually transcribed in the tumor.

Variant calling requires a bioinformatics pipeline that aligns reads to the reference genome, identifies differences between tumor and normal samples, and annotates the functional impact of each variant. The [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible tutorials for variant calling and RNA sequencing analysis that can help researchers build reproducible pipelines. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline configuration and usage that support reproducible analysis.

### Step 3: Custom Database Construction

Build the protein sequence database that will be searched against the mass spectrometry data. This database must include:

- The reference proteome
- Peptides containing nonsynonymous variants
- Frameshift peptides from insertions and deletions
- Chimeric peptides from gene fusions
- Peptides from retained introns
- Peptides from noncoding regions that show tumor-specific expression

The [Bioconductor project](https://bioconductor.org/) provides packages for genomic annotation and sequence manipulation that support custom database construction. The [EMBL-EBI training](https://www.ebi.ac.uk/training) resources offer practical education on sequence databases and analysis approaches.

### Step 4: Mass Spectrometry Data Acquisition

Prepare the tumor protein lysate or isolate the MHC-peptide complex. The choice depends on whether the goal is whole-proteome analysis or immunopeptidomics.

For whole-proteome analysis, digest the protein lysate into peptides and analyze by liquid chromatography-tandem mass spectrometry. For immunopeptidomics, immunoprecipitate the MHC molecules, elute the bound peptides, and analyze those peptides by mass spectrometry.

The mass spectrometry instrument settings affect the quality of the data. Higher resolution instruments provide more accurate precursor masses. Fragmentation settings determine the quality of the tandem mass spectra. These settings should be optimized for the specific instrument and sample type.

### Step 5: Database Searching and Peptide Validation

Search the mass spectrometry data against the custom database. The search algorithm matches each experimental spectrum to a theoretical spectrum generated from the database peptides. The output is a list of peptide-spectrum matches with scores reflecting the quality of each match.

Control the false discovery rate using a target-decoy strategy. A decoy database is created by reversing or shuffling the target sequences. Matches to the decoy database estimate the number of false positives in the target database. The false discovery rate threshold should be set before the search and applied consistently.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference sequences and search systems that support the validation of identified peptides. The [The Carpentries lessons](https://carpentries.org/lessons) offer foundational training in the computing skills needed to manage and analyze large datasets reproducibly.

### Step 6: Neoantigen Prioritization

Combine the mass spectrometry identifications with HLA binding predictions and expression data. Rank the candidate peptides by:

- Confidence in the mass spectrometry identification
- Predicted binding affinity to the patient's HLA alleles
- Tumor-specific expression of the source transcript
- Evidence of immunogenicity from T cell assays if available

The prioritization step should be transparent about the criteria used and the limitations of each evidence layer. A peptide with strong mass spectrometry evidence but weak predicted binding may still be immunogenic. A peptide with strong predicted binding but no mass spectrometry evidence may not be processed and presented.

## Options and Tradeoffs in Workflow Design

### Whole-Exome versus Whole-Genome Sequencing

Whole-exome sequencing is more economical and produces smaller data files. It captures the majority of known disease-relevant variants. However, it cannot detect variants in noncoding regions that may create novel open reading frames. The finding that most tumor-specific antigens derive from allegedly noncoding regions argues for whole-genome sequencing when the research question is about the full antigenic landscape.

Whole-genome sequencing produces larger data files and requires more computational resources. The [nf-core documentation](https://nf-co.re/docs) describes community pipelines that can handle whole-genome data at scale. The tradeoff is between cost and completeness.

### Whole-Proteome versus Immunopeptidomics

Whole-proteome analysis is more accessible and uses established protocols. It identifies which proteins are expressed but does not directly identify which peptides are presented on MHC molecules.

Immunopeptidomics directly identifies the presented peptides. This is the most relevant evidence for neoantigen discovery because it captures the actual antigenic landscape. The tradeoff is that immunopeptidomics requires specialized protocols and produces data that are more difficult to analyze. The NeoDisc pipeline was developed specifically to integrate immunopeptidomics data with genomic and transcriptomic data, demonstrating that this integration is feasible and produces clinically relevant results.

### Custom Database Size and False Discovery Rate

A larger custom database increases sensitivity by including more potential peptide sequences. It also increases the false discovery rate because more random matches are possible. The database must be large enough to include the relevant variant and noncanonical peptides but small enough to maintain statistical power.

One approach is to filter the database using RNA expression evidence. Peptides from transcripts that are not expressed in the tumor are unlikely to be present and can be excluded. This reduces the search space without sacrificing relevant candidates.

### Rule-Based versus Machine-Learning Prioritization

The NeoDisc pipeline supports both rule-based and machine-learning approaches for antigen prioritization. Rule-based approaches apply explicit criteria such as binding affinity thresholds and expression cutoffs. Machine-learning approaches learn patterns from training data and can capture complex interactions between features.

The choice depends on the available data and the research question. Rule-based approaches are more transparent and easier to interpret. Machine-learning approaches may achieve better performance but require larger training datasets and careful validation.

## Observations and Measurements

### What Mass Spectrometry Actually Detects

Mass spectrometry detects peptides that are present in the sample. The detection limit depends on the instrument, the sample preparation, and the abundance of the peptide. Low-abundance peptides may be missed even if they are immunologically relevant.

The immunopeptidome is the set of peptides that are bound to MHC molecules and presented on the cell surface. This is the biologically relevant set for T cell recognition. Immunopeptidomics analysis of primary AML samples identified 58 tumor-specific antigens, demonstrating that the approach can detect clinically relevant peptides directly from patient samples.

### Expression Evidence and Its Limitations

RNA expression data provide evidence that a transcript is present in the tumor. The correlation between RNA expression and protein abundance is imperfect. Post-transcriptional regulation can prevent translation of abundant transcripts. Conversely, some proteins are translated from transcripts with low steady-state RNA levels.

The colon cancer proteogenomic study demonstrated the value of integrating protein-level data with genomic data. The proteomic analysis identified associations that could not be detected from genomics alone, including the relationship between Rb phosphorylation and proliferation and the link between glycolysis and CD8 T cell infiltration in MSI-H tumors.

### T Cell Evidence

The ultimate test of neoantigen immunogenicity is whether T cells recognize the peptide. This evidence comes from T cell assays that measure activation, proliferation, or cytokine release in response to peptide stimulation.

The AML study found that the predicted number of tumor-specific antigens correlated with spontaneous expansion of cognate T cell receptor clonotypes, accumulation of activated cytotoxic T cells, immunoediting, and improved survival. This correlation provides evidence that the identified antigens are immunologically relevant in patients.

## Records and Documentation

### What to Record at Each Step

The reproducibility of a proteogenomic study depends on detailed documentation. Record the following information at each step:

- Sample collection: tissue source, collection time, storage conditions, visible heterogeneity
- Sequencing: library preparation protocol, sequencing platform, read length, coverage
- Variant calling: alignment tool, variant caller, filtering thresholds, annotation database
- Database construction: reference proteome version, variant inclusion criteria, noncanonical transcript sources
- Mass spectrometry: instrument settings, digestion protocol, fractionation method, acquisition mode
- Database searching: search algorithm, precursor and fragment mass tolerances, false discovery rate threshold
- Prioritization: binding prediction tool, expression cutoff, immunogenicity criteria

The [Galaxy Training Network](https://training.galaxyproject.org/) emphasizes reproducibility through documented workflows. The [nf-core documentation](https://nf-co.re/docs) describes standards for pipeline configuration that support reproducible analysis.

### Version Control and Provenance

Every tool and database used in the analysis should be versioned. Tool updates can change results. Database updates can change annotations. The analysis should be reproducible from the raw data using the documented versions.

The [The Carpentries lessons](https://carpentries.org/lessons) provide training in version control with Git and reproducible computing practices. These skills are essential for managing the complexity of proteogenomic analyses.

## Quality Controls and Validation

### False Discovery Rate Control

The false discovery rate is the expected proportion of false positives among the identified peptides. A target-decoy strategy estimates this rate by searching against a decoy database. The false discovery rate threshold should be set before the search and applied consistently across all samples.

A common threshold is 1% at the peptide level. Stricter thresholds reduce false positives but may also reduce sensitivity. The appropriate threshold depends on the downstream use of the data. Clinical applications may require stricter thresholds than exploratory research.

### Validation of Custom Database Peptides

Peptides identified from a custom database require additional validation because the database itself is constructed from predictions. A variant peptide identified by mass spectrometry should be confirmed by examining the spectra for the specific amino acid change. The fragmentation pattern should be consistent with the variant sequence.

For noncanonical peptides, validation may include checking the genomic location of the source sequence and confirming that the transcript is expressed in the tumor. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to genomic and transcript sequences that support this validation.

### Replicate Analysis

Biological replicates from the same tumor and technical replicates from the same sample provide confidence in the reproducibility of the identifications. Peptides identified in multiple replicates are more reliable than peptides identified in a single run.

The cost of replicates must be balanced against the value of the additional confidence. For clinical applications, the cost is justified. For exploratory research, a single run may be sufficient for hypothesis generation.

## Common Failure Patterns

### Database Construction Errors

The most common failure in proteogenomic neoantigen discovery is an incomplete or incorrect custom database. If the database does not contain the relevant variant peptides, the mass spectrometry search cannot identify them. This is a false negative that is invisible in the results.

Common errors include:

- Using an outdated reference proteome
- Missing variants due to low sequencing coverage
- Excluding noncanonical transcripts from the database
- Including too many irrelevant sequences, which inflates the false discovery rate

### Mass Spectrometry Data Quality Issues

Poor quality mass spectrometry data produce few peptide identifications regardless of the database quality. Common issues include:

- Low protein yield from the tumor sample
- Inefficient digestion
- Contamination from buffers or plastics
- Instrument calibration drift

These issues should be identified early through quality control metrics such as the number of identified peptides, the distribution of peptide scores, and the chromatographic peak shape.

### Overinterpretation of Binding Predictions

Binding predictions are computational estimates, not experimental measurements. A peptide predicted to bind a given HLA allele may not bind in practice. Conversely, a peptide with a weak predicted binding score may still be presented and immunogenic.

The AML study demonstrated that tumor-specific antigens can be identified without relying on mutation prediction. The antigens were validated through their correlation with T cell responses and improved survival. This validation is stronger than binding prediction alone.

### Ignoring the Antigen Presentation Machinery

A tumor may have defects in the antigen presentation pathway that prevent presentation of otherwise immunogenic peptides. The NeoDisc pipeline demonstrated that multiomics integration can identify these defects. Ignoring the presentation machinery leads to overestimation of the number of actionable neoantigens.

## Limitations of Proteogenomic Approaches

### Technical Limitations

Mass spectrometry cannot detect every peptide in a sample. The dynamic range of protein abundance in a tumor lysate is enormous, and low-abundance peptides may fall below the detection limit. Immunopeptidomics improves the detection of presented peptides but still misses low-abundance MHC-peptide complexes.

The custom database approach is limited by the completeness of the genomic and transcriptomic data. Variants in regions with low sequencing coverage may be missed. Transcripts that are expressed at very low levels may not be detected by RNA sequencing.

### Biological Limitations

The presence of a peptide in the immunopeptidome does not guarantee that it will elicit a T cell response. T cell recognition depends on the T cell repertoire of the individual patient. A peptide that is presented but not recognized by any T cell is not an effective neoantigen.

The AML study found that the strength of antitumor responses after TSA vaccination was influenced by TSA expression and the frequency of TSA-responsive T cells in the preimmune repertoire. These parameters can be estimated in humans and could serve for TSA prioritization in clinical studies.

### Interpretation Limitations

The correlation between RNA expression and protein abundance is imperfect. A transcript may be expressed at high levels but not translated. Conversely, a protein may be translated from a transcript with low steady-state RNA levels.

The interpretation of proteogenomic data requires careful consideration of these limitations. A peptide identified by mass spectrometry is strong evidence that the peptide exists in the sample. The absence of a peptide from the mass spectrometry data is weaker evidence that the peptide does not exist.

## Safety and Regulatory Context

### Data Handling and Privacy

Tumor sequencing data contain sensitive patient information. The data must be handled according to applicable privacy regulations and institutional policies. Raw sequencing data and mass spectrometry data should be stored securely and access should be restricted to authorized personnel.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide secure repositories for sequence data with controlled access options. Researchers should follow the data deposition and access policies of their institution and funding agency.

### Clinical Translation Considerations

The use of proteogenomic neoantigen discovery for clinical decision-making requires validation beyond the research setting. The [EMBL-EBI training](https://www.ebi.ac.uk/training) resources provide education on the standards and practices for clinical bioinformatics.

The NeoDisc pipeline was developed as an end-to-end clinical proteogenomic pipeline, demonstrating that the approach can be implemented in a clinical context. However, the pipeline's outputs require interpretation by qualified professionals who understand the limitations of each evidence layer.

### Professional Escalation Criteria

Researchers should escalate to a supervisor or collaborator when:

- The mass spectrometry data quality is consistently poor across multiple samples
- The custom database construction produces unexpected results that cannot be explained
- The false discovery rate cannot be controlled at the desired threshold
- The integration of genomic and proteomic data produces discordant results that suggest a technical problem
- The results have clinical implications that require interpretation by a qualified clinician

## Decision Framework for Selecting a Proteogenomic Discovery Strategy

### Matching Workflow Design to Research Objectives

The choice between genomics-only prediction, whole-proteome proteogenomics, and immunopeptidomics-based proteogenomics depends on the specific research question, available resources, and downstream application. A structured decision framework helps researchers avoid committing to an expensive or technically demanding workflow that does not match their actual objectives.

The first decision point is whether the research question requires direct peptide evidence. If the goal is to generate a broad list of candidate neoantigens for initial screening, a genomics-only approach with rigorous binding prediction may be sufficient. If the goal is to identify peptides that are actually processed and presented, immunopeptidomics provides the most direct evidence. If the goal is to discover tumor-specific antigens from noncanonical sources, proteogenomics with a custom database is required because standard exome-based approaches will miss these peptides entirely.

The second decision point is the tumor type and sample availability. Some tumor types have well-characterized mutational landscapes where mutation-derived neoantigens are the dominant immunogenic species. Other tumor types, such as acute myeloid leukemia, have relatively low mutation burdens but produce tumor-specific antigens from intron retention and epigenetic changes. The AML study demonstrated that 48% of tumor-specific antigens resulted from intron retention and translation, with RNA expression correlating with mutations of epigenetic modifiers such as DNMT3A. Researchers working with low-mutation-burden tumors should prioritize proteogenomic approaches that capture noncanonical antigen sources.

The third decision point is the available infrastructure. Immunopeptidomics requires specialized protocols for MHC-peptide complex isolation, access to appropriate antibodies for immunoprecipitation, and expertise in handling the resulting data. Whole-proteome analysis uses more established protocols and a wider range of available tools. The [EMBL-EBI training](https://www.ebi.ac.uk/training) resources provide practical education on both approaches and can help researchers assess the infrastructure requirements before committing to a workflow.

### Resource Allocation and Cost-Benefit Analysis

The cost of a proteogenomic study scales with the complexity of the workflow. Whole-exome sequencing is more economical than whole-genome sequencing. Whole-proteome analysis is less expensive than immunopeptidomics because it uses standard digestion protocols and does not require the additional immunoprecipitation step. The custom database construction and bioinformatics analysis add computational costs that are often underestimated.

A practical approach is to stage the investment. Begin with whole-exome sequencing and whole-proteome analysis to establish the feasibility of the workflow and generate preliminary data. If the preliminary data show that noncanonical antigens are likely to be relevant for the tumor type under study, invest in whole-genome sequencing and immunopeptidomics to capture the full antigenic landscape. This staged approach reduces the risk of committing substantial resources to a workflow that does not match the research question.

The colon cancer proteogenomic study demonstrated the value of integrating multiple data layers in a single cohort. The study produced a catalog of cancer-associated proteins and phosphosites, including known and putative new biomarkers, drug targets, and cancer/testis antigens. The integration prioritized genomically inferred targets such as copy-number drivers and mutation-derived neoantigens while also yielding novel findings from the proteomic and phosphoproteomic data. This comprehensive approach required substantial resources but produced results that no single data layer could provide.

### Criteria for Choosing Between Whole-Proteome and Immunopeptidomics

The choice between whole-proteome analysis and immunopeptidomics should be guided by the specific evidence needed for the research question. Whole-proteome analysis answers the question of which proteins are expressed in the tumor. Immunopeptidomics answers the question of which peptides are presented on MHC molecules.

For neoantigen discovery, immunopeptidomics provides the most relevant evidence because the identified peptides have already survived antigen processing and presentation. The NeoDisc pipeline was developed specifically to integrate immunopeptidomics data with genomic and transcriptomic data, demonstrating that this integration is feasible and produces clinically relevant results. The pipeline combines state-of-the-art publicly available and in-house software for immunopeptidomics, genomics, and transcriptomics with in silico tools for identification, prediction, and prioritization of tumor-specific and immunogenic antigens.

However, immunopeptidomics has limitations that researchers must consider. The yield of MHC-peptide complexes from tumor samples can be low, particularly for small biopsies. The dynamic range of peptide abundances in the immunopeptidome is wide, and low-abundance peptides may fall below the detection limit. The data analysis is more complex than whole-proteome analysis because the peptide length distribution is constrained by the MHC binding groove and the search space must account for this constraint.

Whole-proteome analysis is more accessible and can identify tumor-specific proteins that may be processed into presented peptides. The evidence is indirect, but the workflow is more established and the data analysis tools are more mature. For exploratory research or hypothesis generation, whole-proteome analysis may be sufficient. For clinical applications where the goal is to identify actionable neoantigens for vaccine design or T cell therapy, immunopeptidomics provides the strongest evidence.

### Decision Matrix for Workflow Selection

The following decision matrix summarizes the key considerations for selecting a proteogenomic discovery strategy. Researchers should evaluate their research question, tumor type, sample availability, and infrastructure before committing to a workflow.

| Research Objective | Recommended Approach | Key Evidence Generated | Primary Limitation |
| --- | --- | --- | --- |
| Initial screening of mutation-derived neoantigens | Genomics-only with binding prediction | Predicted mutant peptides | No evidence of processing or presentation |
| Identification of expressed tumor-specific proteins | Whole-proteome proteogenomics | Protein expression evidence | Indirect evidence of presentation |
| Direct identification of presented neoantigens | Immunopeptidomics proteogenomics | MHC-bound peptide evidence | Requires specialized protocols and expertise |
| Discovery of noncanonical tumor-specific antigens | Proteogenomics with custom database | Peptides from noncoding regions and intron retention | Requires whole-genome or deep RNA sequencing |
| Clinical vaccine design | Immunopeptidomics with multiomics integration | Prioritized immunogenic antigens | Requires comprehensive infrastructure and validation |

### Implementation Steps for the Decision Framework

The first implementation step is to document the research objective in explicit terms. Write down the specific question the study is designed to answer. This documentation should include the expected downstream use of the data, whether for exploratory research, vaccine design, or clinical decision-making.

The second step is to assess the tumor type and sample availability. Review the literature for the tumor type to understand the expected mutational burden and the likelihood of noncanonical antigen sources. The finding that most tumor-specific antigens derive from allegedly noncoding regions in both murine cell lines and human primary tumors suggests that noncanonical sources are relevant across multiple tumor types. Assess the amount of tumor tissue available and whether it is sufficient for the required extractions.

The third step is to evaluate the available infrastructure. Review the mass spectrometry instruments, antibodies for immunoprecipitation, and bioinformatics expertise available in the laboratory or through collaboration. The [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible tutorials for variant calling and RNA sequencing analysis that can help researchers build reproducible pipelines. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline configuration and usage that support reproducible analysis.

The fourth step is to estimate the cost and timeline for each workflow option. Include the cost of sequencing, mass spectrometry, antibodies, and bioinformatics analysis. Include the time required for protocol optimization, data acquisition, and analysis. Compare the cost and timeline against the expected value of the data for the research question.

The fifth step is to select the workflow and document the rationale. The documentation should include the research objective, the tumor type considerations, the infrastructure assessment, and the cost-benefit analysis. This documentation becomes part of the study records and supports the interpretation of the results.

### Common Failure Patterns in Workflow Selection

The most common failure in workflow selection is choosing a genomics-only approach when the research question requires direct peptide evidence. This failure produces a candidate list that cannot distinguish between predicted and confirmed neoantigens. The false-positive rate of binding prediction tools is well documented, and the absence of mass spectrometry evidence means that the candidate list contains many peptides that are not processed or presented.

The second common failure is choosing immunopeptidomics without the necessary infrastructure or expertise. The specialized protocols for MHC-peptide complex isolation require optimization for each tumor type and sample type. The data analysis requires expertise in immunopeptidomics-specific search strategies and quality control. Without this expertise, the experiment may produce poor quality data that cannot be interpreted.

The third common failure is underestimating the importance of the custom database. The database construction step determines which peptides can be identified. An incomplete database produces false negatives that are invisible in the results. A database that is too large inflates the false discovery rate and reduces the confidence in the identified peptides. The database must be built from the patient's own sequencing data and filtered using RNA expression evidence to balance sensitivity and specificity.

The fourth common failure is ignoring the antigen presentation machinery. A tumor may have defects in the antigen presentation pathway that prevent presentation of otherwise immunogenic peptides. The NeoDisc pipeline demonstrated that multiomics integration can identify these defects. Ignoring the presentation machinery leads to overestimation of the number of actionable neoantigens and may lead to incorrect conclusions about the immunogenicity of the tumor.

### Escalation Criteria for Workflow Decisions

Researchers should escalate to a supervisor or collaborator when the workflow selection involves substantial resource commitments that exceed the approved budget. The decision to move from whole-exome to whole-genome sequencing or from whole-proteome to immunopeptidomics should be reviewed by the research team before implementation.

Escalation is also appropriate when the tumor type or sample availability introduces uncertainty that cannot be resolved from the literature. A pathologist or clinician with expertise in the specific tumor type can provide guidance on the expected molecular characteristics and the feasibility of the proposed workflow.

Escalation is required when the results of the study will be used for clinical decision-making. The interpretation of proteogenomic data for clinical applications requires expertise that goes beyond the technical analysis. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference sequences and search systems that support the validation of identified peptides, but the clinical interpretation requires qualified professionals who understand the limitations of each evidence layer.

### Records and Documentation for Workflow Decisions

The workflow selection process should be documented in the study records. The documentation should include the research objective, the tumor type assessment, the infrastructure evaluation, the cost-benefit analysis, and the rationale for the selected workflow. This documentation supports the interpretation of the results and provides context for future studies.

The documentation should also include the criteria used to evaluate the success of the workflow. These criteria should be established before the experiment begins and should include specific metrics such as the number of identified peptides, the false discovery rate, and the number of prioritized neoantigens. The [The Carpentries lessons](https://carpentries.org/lessons) provide training in reproducible computing practices that support the documentation and version control of the analysis.

The [Bioconductor project](https://bioconductor.org/) provides packages for genomic annotation and sequence manipulation that support the custom database construction and the integration of genomic and proteomic data. The [EMBL-EBI training](https://www.ebi.ac.uk/training) resources offer practical education on sequence databases and analysis approaches that support the workflow implementation.

## Frequently Asked Questions

### What is the main advantage of proteogenomics over genomics-only neoantigen prediction?

Proteogenomics provides direct experimental evidence that a peptide is present in the tumor sample. Genomics-only prediction identifies candidate peptides from DNA variants but cannot confirm that the peptide is transcribed, translated, processed, and presented. Mass spectrometry data from the immunopeptidome provide evidence that the peptide has survived the antigen processing and presentation pathway. Studies in colon cancer, AML, and other tumor types have identified tumor-specific antigens that would be missed by exome-based approaches, with most of these antigens deriving from noncoding regions.

### What types of peptides can proteogenomics detect that genomics-only approaches miss?

Proteogenomics can detect peptides arising from intron retention, translation of noncoding regions, aberrantly expressed endogenous retroelements, alternative splicing, and other noncanonical sources. In the AML study, 86% of identified tumor-specific antigens derived from supposedly noncoding genomic regions and 48% resulted from intron retention and translation. These peptides are invisible to standard variant calling because they do not involve DNA sequence changes.

### How do I build a custom protein sequence database for mass spectrometry searching?

The custom database must include the reference proteome plus peptides containing nonsynonymous variants, frameshift peptides from insertions and deletions, chimeric peptides from gene fusions, and peptides from retained introns or noncoding regions with tumor-specific expression. The database is built from the patient's own sequencing data. RNA expression evidence can be used to filter the database to reduce the search space and control the false discovery rate. The [Bioconductor project](https://bioconductor.org/) provides packages for genomic annotation and sequence manipulation that support this process.

### What is the difference between whole-proteome analysis and immunopeptidomics?

Whole-proteome analysis identifies which proteins are expressed in the tumor by digesting the total protein lysate and analyzing the resulting peptides. Immunopeptidomics isolates the peptides that are bound to MHC molecules on the cell surface and analyzes those peptides specifically. Immunopeptidomics provides more direct evidence for neoantigen presentation because the identified peptides have already passed through antigen processing. The tradeoff is that immunopeptidomics requires specialized protocols and produces data that are more difficult to analyze.

### How do I control the false discovery rate in proteogenomic neoantigen discovery?

The standard approach is a target-decoy strategy. A decoy database is created by reversing or shuffling the target sequences. Matches to the decoy database estimate the number of false positives in the target database. The false discovery rate threshold should be set before the search and applied consistently. A common threshold is 1% at the peptide level, but the appropriate threshold depends on the downstream use of the data.

### What evidence is needed to validate a neoantigen identified by proteogenomics?

Validation requires multiple evidence layers. The mass spectrometry identification must meet the false discovery rate threshold. The source transcript should show tumor-specific expression. The peptide should be predicted to bind the patient's HLA alleles. The strongest validation is T cell evidence showing that patient T cells recognize the peptide. The AML study found that the predicted number of tumor-specific antigens correlated with spontaneous expansion of cognate T cell receptor clonotypes and improved survival.

### What are the most common reasons a proteogenomic neoantigen discovery experiment fails?

The most common failures are incomplete custom databases, poor mass spectrometry data quality, and overinterpretation of binding predictions. An incomplete database produces false negatives that are invisible in the results. Poor quality mass spectrometry data produce few peptide identifications regardless of the database. Binding predictions are computational estimates that may not reflect actual peptide presentation or T cell recognition.

### When should I escalate to a supervisor or collaborator?

Escalate when the mass spectrometry data quality is consistently poor across multiple samples, when the custom database construction produces unexpected results, when the false discovery rate cannot be controlled at the desired threshold, when genomic and proteomic data produce discordant results suggesting a technical problem, or when the results have clinical implications requiring interpretation by a qualified clinician.

## Related Bioinformatics Guides

- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Genomic Epidemiology: Integrating Pathogen Genomics into Outbreak Investigations](/knowledge/bioinformatics/genomic-epidemiology-integrating-pathogen-genomics-into-outbreak-investigations)
- [Proteomics Mass Spectrometry: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/proteomics-mass-spectrometry-from-sample-preparation-to-data-analysis)
- [Mass Spectrometry-Based Proteomics: Data Analysis Pipelines and Tools](/knowledge/bioinformatics/mass-spectrometry-based-proteomics-data-analysis-pipelines-and-tools)
- [Spatial Proteomics Mass Spectrometry: Techniques and Applications](/knowledge/bioinformatics/spatial-proteomics-mass-spectrometry-techniques-and-applications)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Proteogenomic Analysis of Human Colon Cancer Reveals New Therapeutic Opportunities.](https://pubmed.ncbi.nlm.nih.gov/31031003). Cell, 2019.
- [A comprehensive proteogenomic pipeline for neoantigen discovery to advance personalized cancer immunotherapy.](https://pubmed.ncbi.nlm.nih.gov/39394480). Nature biotechnology, 2025.
- [Noncoding regions are the main source of targetable tumor-specific antigens.](https://pubmed.ncbi.nlm.nih.gov/30518613). Science translational medicine, 2018.
- [Atypical acute myeloid leukemia-specific transcripts generate shared and immunogenic MHC class-I-associated epitopes.](https://pubmed.ncbi.nlm.nih.gov/33740418). Immunity, 2021.
- [Proteogenomics offers a novel avenue in neoantigen identification for cancer immunotherapy.](https://pubmed.ncbi.nlm.nih.gov/39270345). International immunopharmacology, 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.