Mass Spectrometry-Based Proteomics: Data Analysis Pipelines and Tools
Mass spectrometry-based proteomics generates large volumes of complex data that require specialized bioinformatics processing to convert raw instrument output into biologically meaningful protein identifications and quantifications. This article provides researchers, students, and analysts with a practical framework for navigating proteomics data analysis, covering database searching, quantification strategies, statistical validation, and software selection. The guidance focuses on concrete decisions at each pipeline stage, from raw data quality assessment to final biological interpretation, with attention to reproducibility and common failure points.
At a Glance: Proteomics Data Analysis Pipeline Overview
The table below summarizes the major stages of a typical bottom-up proteomics data analysis pipeline, the primary decisions at each stage, and the tools commonly used. Bottom-up proteomics, which analyzes peptides derived from protein digestion, remains the most widely used workflow and consists of sample preparation, LC-MS/MS analysis, and data analysis steps [8].
| Pipeline Stage | Primary Decisions | Representative Tools | Key Output |
|---|---|---|---|
| Raw data processing | Peak picking, centroiding, quality filtering | Vendor software, OpenMS, ProteoWizard | Processed spectra ready for searching |
| Peptide identification | Search engine selection, database choice, FDR control | MSFragger, MaxQuant, Comet, OpenSWATH | Peptide-spectrum matches with confidence scores |
| Protein inference and quantification | Label-free versus labeled approaches, spectral counting versus intensity | MaxQuant, DIA-NN, Spectronaut, Philosopher | Protein group quantities across samples |
| Statistical analysis | Normalization, differential expression testing, batch correction | MSstats, POMAShiny, SWATH2stats | Statistically validated protein abundance changes |
| Biological interpretation | Pathway analysis, network reconstruction, functional annotation | Various R packages, web platforms | Mechanistic insights and candidate biomarkers |
Understanding Mass Spectrometry Data Types and Acquisition Strategies
Mass spectrometry is unmatched in its versatility for studying practically any aspect of the proteome, yet the foundations of the technique span multiple scientific fields, creating a perceived high barrier to entry [5]. A relatively simple quantitative proteomic experiment moves from sample preparation through data acquisition to analysis of protein group quantities, and understanding how data are acquired, processed, and analyzed is essential for making sound bioinformatics decisions [5].
Data-Dependent Acquisition
Data-dependent acquisition (DDA) selects precursor ions for fragmentation based on their intensity in a survey scan. This approach generates tandem mass spectra that are typically searched against protein sequence databases. DDA remains a standard method for discovery proteomics, and most search engines were originally designed for this data type. The stochastic nature of precursor selection means that DDA can suffer from missing values across samples, particularly for low-abundance peptides.
Data-Independent Acquisition
Data-independent acquisition (DIA) has emerged as a powerful technology for high-throughput, accurate, and reproducible quantitative proteomics [11]. DIA systematically fragments all precursor ions within defined isolation windows instead of selecting individual precursors. DIA acquisition schemes are categorized based on the design of precursor isolation windows, including wide-window, overlapping-window, narrow-window, scanning quadrupole-based, and parallel accumulation-serial fragmentation-enhanced methods [11]. For DIA data analysis, major strategies include spectrum reconstruction, sequence-based search, library-based search, de novo sequencing, and sequencing-independent approaches [11].
The choice between DDA and DIA affects downstream analysis tools substantially. DIA data require specialized software that can deconvolve the complex multiplexed spectra, whereas DDA data can be processed by a broader range of traditional search engines.
Top-Down Proteomics
Top-down mass spectrometry analyzes intact proteins instead of digested peptides, offering advantages for studying complex proteoforms [9]. TopPIC Suite is a widely used software package for top-down MS-based proteoform identification and quantification [9]. Researchers working with top-down data must use tools designed for this data type, as bottom-up search engines are not appropriate for intact protein spectra.
Core Principles of Proteomics Data Analysis
The Role of Bioinformatics in Proteomics
Bioinformatics is fundamental for elaborating mass spectrometry data because of the sheer volume of data the technique produces [12]. New software packages and algorithms are continuously developed to improve protein identification and characterization in terms of high-throughput and statistical accuracy, yet limitations exist concerning bioinformatics spectral data elaboration [12]. Researchers should expect to invest time in understanding the strengths and weaknesses of their chosen analysis tools.
The Central Dogma of Proteomics Informatics
The analysis pipeline follows a logical progression from raw spectra to biological insight. First, raw data files are converted to a searchable format. Second, peptide identification matches experimental spectra to theoretical spectra derived from a protein database. Third, protein inference assembles peptides into protein groups. Fourth, quantification measures abundance differences across conditions. Fifth, statistical analysis determines which changes are significant. Finally, biological interpretation places results in a mechanistic context.
The Importance of Experimental Design
Quantification of differences between two or more physiological states of a biological system is among the most important but also most challenging technical tasks in proteomics [7]. Experimental design decisions made before data acquisition, including the choice of labeling strategy, number of replicates, and randomization of sample processing order, directly impact the statistical power of downstream analysis. Mass-spectrometry-based quantification methods employ differential stable isotope labeling to create specific mass tags recognized by the mass spectrometer, with tags introduced metabolically, chemically, enzymatically, or through spiked synthetic peptide standards [7]. Label-free quantification approaches instead correlate the mass spectrometric signal of intact proteolytic peptides or the number of peptide sequencing events with relative or absolute protein quantity [7].
Building a Data Analysis Pipeline: Step-by-Step Workflow
Step 1: Raw Data Quality Assessment
Before any database searching, assess the quality of raw data files. Check total ion chromatograms for consistent peak shapes and intensities across runs. Examine base peak chromatograms for signs of column degradation or sample contamination. Review precursor ion intensities and charge state distributions. Poor quality raw data cannot be rescued by sophisticated downstream analysis, so early detection of problems saves substantial time.
Step 2: Data Conversion and Processing
Most search engines require data in specific formats. Vendor raw files often need conversion to open formats such as mzML. Tools like ProteoWizard facilitate this conversion. During conversion, consider whether to apply centroiding, which collapses profile data into discrete peaks. Some search engines work with either profile or centroid data, but consistency across the dataset is essential.
Step 3: Database Searching for Peptide Identification
Database searching compares experimental tandem mass spectra against theoretical spectra generated from a protein sequence database. The choice of database is critical. For standard proteomics experiments, a species-specific database from a resource such as the NCBI Data Resources or UniProt is appropriate [2]. For metaproteomics or samples containing multiple species, a combined database may be necessary. The P4PP pipeline demonstrates a specialized approach for virus identification, enabling the identification of 1896 virus species from 32 virus families based on multiple identified discriminatory peptides [13].
Search engine parameters must be set carefully. Key parameters include precursor mass tolerance, fragment mass tolerance, enzyme specificity, missed cleavage sites, fixed modifications, and variable modifications. Setting tolerances too loosely increases false discoveries, while setting them too tightly may miss true identifications.
Step 4: False Discovery Rate Control
False discovery rate (FDR) control is essential for reliable peptide and protein identification. The target-decoy approach, where searches are performed against a database containing reversed or shuffled sequences, allows estimation of the number of false positives. Most modern search engines and post-processing tools implement FDR estimation. A typical threshold is 1% FDR at the peptide level and 1% at the protein level, though researchers should report their specific thresholds.
Step 5: Protein Inference and Quantification
Protein inference assigns identified peptides to proteins. Peptides shared between multiple proteins complicate this process, and software packages handle shared peptides differently. Protein grouping algorithms aim to report the minimal set of proteins that explains all observed peptides.
Quantification strategies fall into two broad categories. Label-based approaches use stable isotope tags introduced during sample preparation, allowing multiplexed analysis of multiple conditions in a single run. Label-free approaches compare peptide intensities or spectral counts across separate runs. Each approach has distinct advantages and challenges, and the choice depends on experimental goals, sample availability, and budget.
Step 6: Statistical Analysis and Differential Expression
Statistical analysis transforms quantitative data into biological conclusions. Normalization corrects for systematic variation between runs. Common methods include median normalization, quantile normalization, and variance stabilization. After normalization, differential expression testing identifies proteins with statistically significant abundance changes between conditions.
POMAShiny provides a user-friendly web-based workflow for the visualization, exploration, and statistical analysis of metabolomics and proteomics data, integrating several statistical methods and based on the POMA R/Bioconductor package [20]. This tool increases reproducibility and flexibility of analyses outside the web environment [20].
Step 7: Data Interpretation and Biological Context
The final pipeline stage places quantitative results in biological context. This may involve pathway enrichment analysis, protein-protein interaction network reconstruction, or integration with other omics data types. Quantitative protein data can be used to reconstruct protein interactions and signaling networks [6]. Machine learning methods have been applied to proteomics data for classification and biomarker discovery, as demonstrated by a LightGBM-based pipeline that discriminated malignant from benign pulmonary nodules using plasma protein panels [14].
Software Tools Comparison and Selection Criteria
Philosopher Toolkit
Philosopher is a free, open-source, versatile, and robust data analysis toolkit designed to bring easy access to a powerful and comprehensive set of computational tools for shotgun proteomics data analysis [16]. It integrates high-performance algorithms and existing tools and is dependency-free, fast, and comprehensive, able to rapidly process complex proteomics datasets with efficient resource management [16]. Philosopher addresses the challenge that existing tools such as the Trans-Proteomic Pipeline, MaxQuant, and PeptideShaker require installation and depend on specific operating systems, libraries, and other software [16].
DIA Analysis Tools
A comparative analysis of five DIA data analysis tools, including OpenSWATH, EncyclopeDIA, Skyline, DIA-NN, and Spectronaut, assessed their performance using six DIA datasets obtained from TripleTOF, Orbitrap, and TimsTOF Pro instruments [17]. The study found that library-free approaches outperformed library-based methods when the spectral library had limited comprehensiveness, but constructing a comprehensive library still offers benefits for most DIA analyses [17]. This comparison provides guidance for both experienced and novice users of DIA-mass spectrometry technology [17].
MSFragger-DIA and FragPipe
MSFragger-DIA is a fast and sensitive approach for direct peptide identification from DIA data, leveraging the speed of the fragment ion indexing-based search engine MSFragger [18]. Unlike most existing methods, MSFragger-DIA conducts a database search of DIA tandem mass spectra prior to spectral feature detection and peak tracing across the liquid chromatography dimension [18]. The tool is integrated into the FragPipe computational platform for seamless support of peptide identification and spectral library building from DIA, DDA, or both data types combined [18]. MSFragger-DIA demonstrates fast, sensitive, and accurate performance across a variety of sample types and data acquisition schemes, including single-cell proteomics, phosphoproteomics, and large-scale tumor proteome profiling studies [18].
SWATH2stats
SWATH-MS is an acquisition and analysis technique of targeted proteomics that enables measuring several thousand proteins with high reproducibility and accuracy across many samples [19]. OpenSWATH is popular open-source software for peptide identification and quantification from SWATH-MS data, but the transfer of data from OpenSWATH to downstream statistical tools is technically challenging [19]. SWATH2stats is an R/Bioconductor package that allows convenient processing of the data into a format directly readable by downstream analysis tools, including MSstats, mapDIA, and aLFQ [19]. It also allows annotation, analysis of variation and reproducibility, FDR estimation, and advanced filtering before submitting processed data to downstream tools [19].
Tool Selection Criteria
When selecting analysis software, consider the following factors. First, match the tool to your data type. DDA data can be processed by many search engines, while DIA data requires specialized tools. Second, evaluate the learning curve and documentation quality. Third, consider whether the tool supports your computational infrastructure, including operating system and available memory. Fourth, assess the tool's active development and community support. Fifth, verify that the tool produces outputs compatible with your downstream statistical analysis software.
Label-Free Versus Label-Based Quantification
Label-Based Quantification
Stable isotope labeling methods create specific mass tags that can be recognized by a mass spectrometer and provide the basis for quantification [7]. These tags can be introduced into proteins or peptides metabolically, by chemical means, enzymatically, or through spiked synthetic peptide standards [7]. Label-based approaches allow multiple samples to be combined and analyzed in a single run, reducing run-to-run variability. However, labeling reagents add cost and sample preparation complexity.
Label-Free Quantification
Label-free quantification approaches aim to correlate the mass spectrometric signal of intact proteolytic peptides or the number of peptide sequencing events with the relative or absolute protein quantity directly [7]. Label-free methods are simpler and less expensive than labeling approaches, and they can accommodate any number of samples. However, they are more susceptible to run-to-run variability and require careful normalization and quality control.
Choosing Between Approaches
The choice between label-based and label-free quantification depends on experimental design. For experiments with a limited number of conditions and sufficient sample amount, label-based approaches may provide more precise quantification. For large cohort studies or experiments with many samples, label-free approaches are often more practical. Researchers should consider the trade-offs between precision, cost, throughput, and sample availability.
Spectral Library Construction and Use
The Role of Spectral Libraries in DIA Analysis
Spectral libraries are critical resources for DIA analysis [11]. A spectral library contains the expected retention times and fragment ion spectra for peptides, enabling targeted extraction of quantitative information from DIA data. Libraries can be generated from DDA data, from DIA data directly, or from a combination of both data types.
Library Generation Strategies
The generation and optimization of spectral libraries are discussed extensively in the DIA literature [11]. Library-free approaches have gained popularity because they do not require a separate DDA run for library construction. However, the comparative analysis of DIA tools found that constructing a comprehensive library still offers benefits for most DIA analyses [17]. Researchers should evaluate whether their sample type and experimental design warrant the additional effort of library construction.
Library Quality Considerations
A spectral library is only as good as the data used to build it. Libraries generated from limited samples may miss peptides present in the experimental dataset. Libraries generated from different instrument types may not transfer optimally to other instruments. Researchers should validate library coverage against their experimental data and consider supplementing libraries with additional DDA runs when coverage is insufficient.
Statistical Analysis and Quality Control
Normalization Strategies
Normalization is a critical step in quantitative proteomics data analysis. The goal is to remove systematic technical variation while preserving biological differences. Common approaches include total intensity normalization, median normalization, and more sophisticated methods that account for differences in sample composition. The choice of normalization method can substantially affect downstream results, and researchers should evaluate multiple approaches when possible.
Handling Missing Values
Missing values are a common challenge in proteomics data, particularly in label-free workflows. Missing values can arise from peptides falling below the detection limit, from stochastic precursor selection in DDA, or from technical issues. The choice of missing value handling strategy, including imputation methods, can substantially affect statistical results. Researchers should understand the pattern of missingness in their data and choose appropriate strategies.
Batch Effects and Experimental Design
Batch effects, systematic technical variation introduced by processing samples in different batches, can confound biological signals. Careful experimental design, including randomization of sample processing order and inclusion of technical replicates, helps mitigate batch effects. Computational methods for batch correction are available, but they should be applied cautiously and their assumptions validated.
Quality Control Metrics
Several metrics help assess data quality throughout the analysis pipeline. These include the number of identified peptides and proteins, the distribution of peptide intensities, the coefficient of variation for technical replicates, and the correlation between replicate samples. SWATH2stats provides functionality for analyzing variation and reproducibility of measurements, FDR estimation, and advanced filtering before submitting processed data to downstream tools [19]. Monitoring these metrics across batches helps identify technical problems early.
Specialized Analysis Scenarios
Single-Cell Proteomics Analysis
Single-cell proteomics by mass spectrometry can now quantify hundreds to thousands of proteins per cell, but the field lacks standardized analytical pipelines that accommodate the diversity of instruments, sample preparation workflows, and biological contexts encountered in practice [15]. Existing workflows largely adapted from single-cell transcriptomics do not account for the informative missingness, pervasive ambient protein contamination, and limited feature space that distinguish proteomic from transcriptomic data [15]. Cell type annotation remains a manual bottleneck that is subjective, difficult to reproduce, and hard to scale [15]. An end-to-end pipeline integrating adaptive quality control, entropy-guided iterative batch correction, multi-modal marker discovery, and context-aware annotation by large language models has been developed to address these challenges [15].
Limited Proteolysis-Coupled Mass Spectrometry
Limited proteolysis combined with mass spectrometry (LiP-MS) facilitates probing structural changes on a proteome-wide scale by leveraging differences in proteinase K accessibility of native protein structures [10]. Distinguishing different contributions to the LiP-MS signal, such as changes in protein abundance or chemical modifications, from structural protein alterations remains challenging [10]. A comprehensive computational pipeline using a two-step approach first removes unwanted variations from the LiP signal that are not caused by protein structural effects, then infers the effects of variables of interest on the remaining signal [10]. This framework provides a powerful approach for deconvolving LiP-MS signals and separating protein structural changes from changes in protein abundance, posttranslational modifications, and alternative splicing [10].
Virus Detection and Pandemic Preparedness
The P4PP pipeline enables the identification of 1896 virus species from 32 virus families based on multiple identified discriminatory peptides, in which at least one human infectious virus is described [13]. The pipeline was evaluated using different datasets of cell-cultivated viruses generated at different institutes, measured with different instruments, and prepared with different sample preparation methods [13]. Shotgun proteomics in combination with the developed data analysis approach can identify all types of virus species after cultivation in a cell line [13].
Reproducibility and Data Sharing
Reproducible Analysis Workflows
Reproducibility is a fundamental requirement for scientific research. Proteomics data analysis should be documented thoroughly, including software versions, parameter settings, and database versions. Containerization tools and workflow managers help ensure that analyses can be reproduced exactly. Philosopher was initially built and deployed with Docker containers with different applications for proteomics, which in part inspired the creation of the BioContainers resource for different bioinformatics fields [16].
Data Sharing Policies
Researchers should be aware of data sharing requirements from funding agencies and journals. The NIH Genomic Data Sharing Policy outlines expectations for sharing genomic data generated with NIH funding [3]. While this policy specifically addresses genomic data, similar principles apply to proteomics data. Researchers should plan for data deposition in appropriate repositories and ensure that analysis scripts and parameter files are shared alongside raw data.
FAIR Guiding Principles
The FAIR Guiding Principles provide a framework for making data findable, accessible, interoperable, and reusable [4]. Applying these principles to proteomics data involves using standardized file formats, depositing data in public repositories, providing rich metadata, and using community-standard ontologies and identifiers. Adherence to FAIR principles increases the value of generated data for the broader scientific community.
Common Failure Patterns and Troubleshooting
Low Peptide Identification Rates
When database searching yields fewer identifications than expected, several factors may be responsible. Check precursor and fragment mass tolerances against instrument calibration. Verify that the correct protease specificity and missed cleavage settings are used. Confirm that the protein database is appropriate for the sample species. Examine whether post-translational modifications are present that were not included in the search parameters.
Poor Quantitative Reproducibility
High variability between technical replicates indicates problems in sample preparation, chromatography, or instrument performance. Review sample preparation protocols for consistency. Check liquid chromatography performance, including column condition and gradient reproducibility. Monitor instrument sensitivity and calibration across the experiment.
Discrepancies Between Search Engines
Different search engines may produce different identification results for the same dataset. This is expected because each engine uses different scoring algorithms and parameter handling. When discrepancies arise, examine the confidence scores and spectra for borderline identifications. Consider using multiple search engines and combining results with tools designed for this purpose.
Batch Effects in Large Studies
When samples are processed in multiple batches, batch effects can obscure biological signals. Examine principal component analysis or clustering plots for batch-related patterns. If batch effects are present, consider computational correction methods, but first evaluate whether experimental design changes could prevent batch effects in future experiments.
Limitations and Interpretation Caveats
Technical Limitations of Mass Spectrometry
Mass spectrometry cannot detect all proteins in a complex sample. Dynamic range limitations mean that low-abundance proteins may escape detection. Sequence coverage is typically incomplete, and some proteins are difficult to detect due to their physicochemical properties. Researchers should interpret the absence of a protein as lack of detection instead of proof of absence.
Statistical Limitations
Proteomics data have unique characteristics that challenge standard statistical methods. The high dimensionality, correlated features, and missing value patterns require specialized approaches. Multiple testing correction is essential when testing thousands of proteins simultaneously. Machine learning methods applied to proteomics data require careful validation to avoid overfitting, as demonstrated by the two-stage feature selection strategy used in the LightGBM-based pipeline for pulmonary nodule classification [14].
Biological Interpretation Limitations
Quantitative proteomics data provide information about protein abundance but not directly about protein activity, localization, or interactions. Post-translational modifications, protein-protein interactions, and subcellular localization require specialized analysis approaches. Integrating proteomics data with other omics data types can provide additional biological context but introduces additional complexity.
Professional Escalation Criteria
Researchers should seek expert assistance when encountering specific situations. If search engine results are consistently poor across multiple parameter settings, consult with a bioinformatics specialist or the software developers. If statistical analysis reveals unexpected patterns that cannot be explained by experimental design, review the analysis with a biostatistician. If results are intended to guide clinical decisions or regulatory submissions, involve experts in the relevant regulatory framework. If data sharing requirements are unclear, consult institutional research offices or funding agency guidance.
Records and Documentation
Maintain detailed records of all analysis steps. Document software versions, including operating system and dependency versions. Record all parameter settings for each search engine and analysis tool. Save the exact database version and download date. Document normalization and statistical analysis methods. Keep versions of analysis scripts and note any changes made during the analysis. This documentation supports reproducibility and facilitates troubleshooting when problems arise.
Frequently Asked Questions
What is the difference between DDA and DIA data analysis?
DDA data analysis searches tandem mass spectra acquired from individually selected precursor ions against a protein database. DIA data analysis must first deconvolve multiplexed spectra containing fragments from multiple precursors, then identify and quantify peptides. DIA analysis tools use strategies including spectrum reconstruction, library-based search, and library-free direct search. The choice of analysis approach depends on the DIA acquisition scheme and available spectral libraries [11].
How do I choose between library-based and library-free DIA analysis?
Library-free approaches outperform library-based methods when the spectral library has limited comprehensiveness, but constructing a comprehensive library still offers benefits for most DIA analyses [17]. If a high-quality library is available for your sample type and instrument, library-based analysis may provide better sensitivity. If no suitable library exists, library-free approaches are a practical starting point.
What is the target-decoy approach for FDR estimation?
The target-decoy approach searches spectra against a database containing both forward protein sequences and reversed or shuffled decoy sequences. False discoveries are estimated from the number of matches to decoy sequences. This approach allows calculation of q-values and FDR control at both peptide and protein levels. Most modern search engines implement this approach automatically.
How should I handle missing values in quantitative proteomics data?
Missing value handling depends on the pattern and mechanism of missingness. Values missing because peptides fell below the detection limit may be imputed with small values representing the detection threshold. Values missing due to stochastic precursor selection in DDA require different treatment. Evaluate the missingness pattern in your data and choose an approach appropriate for your experimental design.
What normalization method should I use for label-free quantification?
The choice of normalization method depends on the data characteristics and experimental design. Median normalization is robust to outliers and works well when most proteins do not change between conditions. More sophisticated methods may be needed when sample composition differs substantially between conditions. Evaluate multiple normalization approaches and assess their effect on downstream results.
Can I use the same analysis pipeline for DDA and DIA data?
Some tools support both DDA and DIA data, but the analysis strategies differ substantially. MSFragger-DIA supports peptide identification and spectral library building from DIA, DDA, or both data types combined [18]. However, researchers should understand the specific requirements of each data type and configure their pipeline accordingly.
What information should I include when reporting proteomics data analysis methods?
Report all software versions, database versions and download dates, search parameters including tolerances and modifications, FDR thresholds, normalization methods, statistical tests, and multiple testing correction approaches. Include this information in the methods section of publications and in data repository submissions to support reproducibility.
How do I validate results from machine learning analysis of proteomics data?
Machine learning analysis requires rigorous validation to avoid overfitting. Use independent validation cohorts whenever possible, as demonstrated by the LightGBM-based classifier that maintained strong performance in an independent validation cohort [14]. Apply cross-validation during model development and report performance metrics with appropriate confidence intervals.
Related Bioinformatics Guides
- Structure-Based Drug Design in Bioinformatics: Computational Pipelines, Active Site Grid Mapping, and Virtual Screening Workflows
- The KEGG Database and Pathway Analysis
- Alternative Splicing Analysis from RNA-Seq Data
- GWAS QC Steps: Structural Analysis and Computational Methodologies in Bioinformatics
- Qiime2 Taxonomy: Structural Analysis and Computational Methodologies in Bioinformatics
References and Further Reading
- EMBL-EBI Training. European Bioinformatics Institute.
- NCBI Data Resources. National Center for Biotechnology Information.
- Genomic Data Sharing Policy. National Institutes of Health.
- The FAIR Guiding Principles. Scientific Data.
- An Introduction to Mass Spectrometry-Based Proteomics.. Journal of proteome research, 2023.
- Bioinformatics Methods for Mass Spectrometry-Based Proteomics Data Analysis.. International journal of molecular sciences, 2020.
- Quantitative mass spectrometry in proteomics: a critical review.. Analytical and bioanalytical chemistry, 2007.
- Bottom-Up Proteomics: Advancements in Sample Preparation.. International journal of molecular sciences, 2023.
- Top-Down Mass Spectrometry Data Analysis Using TopPIC Suite.. Methods in molecular biology (Clifton, N.J.), 2022.
- Analysis of Limited Proteolysis-Coupled Mass Spectrometry Data.. Molecular & cellular proteomics : MCP, 2025.
- Acquisition and Analysis of DIA-Based Proteomic Data: A Comprehensive Survey in 2023.. Molecular & cellular proteomics : MCP, 2024.
- Bioinformatics in mass spectrometry data analysis for proteomics studies.. Expert review of proteomics, 2004.
- P4PP: A Universal Shotgun Proteomics Data Analysis Pipeline for Virus Identification.. 2025.
- Artificial intelligence-assisted proteomic signatures for discriminating malignant from benign pulmonary nodules.. 2026.
- A Context-Aware Single-Cell Proteomics Analysis pipeline. 2026.
- Philosopher: a versatile toolkit for shotgun proteomics data analysis. Nature Methods, 2020.
- A Comparative Analysis of Data Analysis Tools for Data-Independent Acquisition Mass Spectrometry. Molecular & Cellular Proteomics, 2023.
- Analysis of DIA proteomics data using MSFragger-DIA and FragPipe computational platform. Nature Communications, 2023.
- SWATH2stats: An R/Bioconductor Package to Process and Convert Quantitative SWATH-MS Proteomics Data for Downstream Analysis Tools. PLoS ONE, 2016.
- POMAShiny: A user-friendly web-based workflow for metabolomics and proteomics data analysis. PLoS Comput. Biol., 2021.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.