# De Novo Sequencing vs. Database Searching: When to Use Each Approach in Mass Spectrometry-Based Proteomics

Database searching and de novo sequencing are the two principal computational strategies for interpreting tandem mass spectrometry (MS/MS) data in proteomics. Database searching matches experimental spectra against predicted peptide sequences derived from a protein database, while de novo sequencing derives peptide sequences directly from the spectral data without requiring a reference database. The choice between these approaches depends on database availability, sample complexity, research objectives, and the tolerance for computational cost. This article provides a decision framework for researchers who need to select the appropriate analysis strategy for their MS/MS datasets, with particular attention to organisms lacking complete databases and to novel protein discovery.

## At a Glance

The table below summarizes the primary decision criteria for choosing between database searching and de novo sequencing.

| Decision Factor | Database Searching | De Novo Sequencing |
| --- | --- | --- |
| Database availability | Requires a complete or near-complete protein database for the organism under study | Does not require a reference database, works with any organism |
| Novel sequence discovery | Cannot identify peptides absent from the database | Can identify novel peptides, including those from unannotated open reading frames |
| Post-translational modification handling | Requires predefined modification search parameters, combinatorial expansion increases search time | Emerging tools can detect modifications without exhaustive predefined search space |
| Computational cost | Generally fast with indexed databases, filtration algorithms reduce search time | Historically slower, newer non-autoregressive models reduce runtime substantially |
| Best use case | Well-characterized organisms, quantitative proteomics, large-scale screening | Non-model organisms, novel proteins, antibody sequencing, modified peptide discovery |

## Understanding the Two Approaches

### Database Searching: Matching Spectra to Predicted Peptides

Database searching remains the standard method for peptide identification in most proteomics experiments. The approach works by generating theoretical spectra from every peptide in a protein database and then scoring how well each experimental spectrum matches these predictions. Search engines such as SEQUEST and Mascot perform this conceptual task, but they differ from sequence alignment tools like BLAST in a critical way. The key algorithmic idea of filtration, which makes BLAST fast by rapidly eliminating candidate sequences while retaining the true one, was never implemented in these early search tools. As a result, MS/MS protein identification tools became time-consuming for many applications, including searches for post-translationally modified peptides. Matching millions of spectra against all known proteins creates a bottleneck similar to the one that made genome versus genome comparisons too slow for BLAST. Filtration techniques for MS/MS database searches dramatically reduce running time and effectively remove bottlenecks in searching the large space of protein modifications. A tag generation algorithm based on a probability model for determining sequence tag accuracy achieves superior results compared to GutenTag, a popular tag generation algorithm at the time. This tag generating algorithm, along with the de novo sequencing algorithm PepNovo, was made accessible through the Peptide.ucsd.edu web resource [7].

The practical implication is that database searching works well when the database is complete and when the search space is manageable. For well-annotated organisms such as human, mouse, or yeast, database searching provides fast and reliable peptide identification. The approach also integrates naturally with quantitative proteomics workflows because the identified peptides map directly to known proteins, enabling label-free or isobaric-label quantification.

### De Novo Sequencing: Deriving Sequences from Spectra Alone

De novo sequencing determines peptide sequences directly from MS/MS spectra without consulting a protein database. This approach offers a database-free alternative to traditional search methods, but it struggles with accurately modeling complex MS/MS spectra. Most current tools use autoregressive decoding, which predicts sequence tokens one at a time in a left-to-right fashion. This decoding strategy is prone to error propagation, where an incorrect early prediction influences all subsequent predictions, and it is computationally slow. PowerNovo2, a non-autoregressive model based on generative normalizing flows, addresses these limitations by using variational inference to capture intricate token dependencies and peptide-level uncertainties. This model outperforms existing de novo tools in accuracy and speed, matching state-of-the-art autoregressive models like Casanovo while being 4.3 times faster. It also demonstrates competitive performance against other non-autoregressive methods such as π-PrimeNovo, particularly on long peptides and low-resolution spectra. As the first flow-based de novo sequencer, PowerNovo2 provides a scalable and accurate solution for large-scale proteomic applications [8].

The significance of this development is that de novo sequencing is no longer confined to small-scale experiments. The computational speed improvements make it feasible to apply de novo sequencing to large datasets, opening the door for its use in discovery-oriented research where database completeness cannot be assumed.

## Core Principles for Method Selection

### Database Completeness as the Primary Determinant

The most important factor in choosing between database searching and de novo sequencing is the completeness of the protein database for the organism under investigation. For model organisms with well-annotated genomes, database searching is the appropriate first-line approach because the vast majority of peptides will be present in the database. For non-model organisms, organisms with incomplete annotations, or samples containing novel proteins, database searching will leave many spectra unmatched. These unmatched spectra represent potential novel peptides that de novo sequencing can recover.

The National Center for Biotechnology Information (NCBI) maintains a comprehensive set of sequence databases and search systems that researchers can use to assess database completeness for their organism of interest. NCBI provides access to nucleotide and protein sequences, along with search tools that allow researchers to determine whether their organism has been sequenced and annotated [1]. Before deciding on an analysis strategy, researchers should query these resources to evaluate the depth of sequence coverage for their organism.

### Research Objectives Shape the Analytical Approach

The research question determines which approach is more appropriate. If the goal is to quantify known proteins across conditions, database searching is the standard choice because it provides direct mapping between peptides and proteins. If the goal is to discover novel proteins, identify unannotated open reading frames, or sequence antibodies with unknown sequences, de novo sequencing is necessary.

A proteogenomic approach illustrates this principle in practice. Researchers studying hepatocellular carcinoma used high-quality Ribo-seq translatomic datasets to generate an extensive database of human liver long noncoding RNA-derived open reading frames (lncORFs). They applied this database to proteomics data from tumor-adjacent normal tissue pairs. Using the new database, they discovered 104 novel lncRNA-derived microproteins, including 46 that were differentially expressed between tumor and nontumor tissues and 13 with significant correlation with prognosis. Combining the expression of these microproteins with canonical proteins in a LASSO regression model improved predictive performance for recurrence, increasing the area under the curve by 0.005 to 0.085 across three recurrence time points [10].

This example demonstrates that database searching can be extended to novel protein discovery when the database is constructed from appropriate evidence. However, the database construction requires substantial bioinformatics work, including the analysis of Ribo-seq data to define translated open reading frames. For researchers who lack the resources to build custom databases, de novo sequencing offers a more direct path to novel peptide discovery.

### Post-Translational Modification Complexity

Post-translational modifications (PTMs) play a central role in cellular regulation and are implicated in numerous diseases. Database searching remains the standard for identifying modified peptides from tandem mass spectra, but it is hindered by the combinatorial expansion of modification types and sites. Each additional modification considered in the search increases the search space, and the computational cost grows accordingly. De novo peptide sequencing offers an attractive alternative, yet existing methods have been limited to unmodified peptides or a narrow set of PTMs.

Recent work has expanded the scope of de novo sequencing to include modified peptides. A large dataset of spectra from endogenous and synthetic peptides spanning 19 biologically relevant amino acid-PTM combinations, covering phosphorylation, acetylation, and ubiquitination, was used to develop Modanovo, an extension of the Casanovo transformer architecture for de novo peptide sequencing. Modanovo achieved robust performance across these amino acid-PTM combinations with a median area under the precision-coverage curve of 0.92, while maintaining performance on unmodified peptides at 0.93, nearly identical to Casanovo at 0.94. The model outperformed π-PrimeNovo-PTM and InstaNovo-P and showed increased precision and complementarity to the database search tool MSFragger. Applied to a phosphoproteomics dataset from monkeypox virus-infected cells, Modanovo recovered numerous confident peptides not reported by database search, including new viral phosphosites supported by spectral evidence [9].

This development is important for researchers studying phosphorylation, acetylation, or ubiquitination in systems where the modification sites are not well characterized. Database searching requires the researcher to specify which modifications to search for and where they might occur. De novo sequencing with modification-aware tools can detect modifications without this prior specification, making it valuable for discovery-oriented PTM analysis.

## Practical Workflow for Method Selection

### Step 1: Assess Database Availability and Completeness

Begin by evaluating whether a suitable protein database exists for your organism. Query NCBI databases to determine the depth of sequence coverage [1]. Check whether the organism has a reference genome, whether gene models are annotated, and whether the predicted proteome has been deposited in a public database. For well-studied organisms, the database will contain the vast majority of expressed proteins. For poorly characterized organisms, the database may be sparse or contain only homologs from related species.

The European Bioinformatics Institute (EMBL-EBI) provides training resources that cover database searching and sequence analysis, which can help researchers understand the strengths and limitations of available databases [2]. These resources are particularly useful for researchers who are new to proteomics bioinformatics and need to understand how database choice affects downstream analysis.

### Step 2: Evaluate Sample Complexity and Research Goals

Consider the complexity of your sample and the specific questions you need to answer. For quantitative proteomics experiments comparing protein abundance across conditions in a well-annotated organism, database searching is the appropriate choice. For experiments aimed at discovering novel proteins, characterizing unannotated open reading frames, or sequencing antibodies, de novo sequencing is required.

The antibody sequencing case is particularly instructive. Elucidating antibody sequences by mass spectrometry-based de novo sequencing is essential but technically challenging. XA-Novo, an accurate and high-throughput de novo sequencing solution, integrates a single-pot multi-enzymatic gradient digestion method with a beam search-based assembler called Fusion to reconstruct full-length antibody sequences directly from bottom-up mass spectrometry data. Benchmarking across well-characterized antibodies from multiple species demonstrates that XA-Novo outperforms commercial solutions in identification sensitivity, sequence completeness, and reconstruction accuracy. The tool successfully reconstructs six immunotherapeutic antibodies with unknown sequences, and in vitro and in vivo assays validate that these generated antibodies exhibit functionality equivalent to their commercial counterparts. XA-Novo achieves over 99.54 percent accurate sequence coverage in distinguishing mixed COVID-19 neutralizing antibodies, exceeding the performance of current assemblers reported for single-antibody sequencing [11].

This example shows that de novo sequencing is also a fallback for missing databases. It is the method of choice for applications where the sequence itself is the research output, such as therapeutic antibody development.

### Step 3: Consider Computational Resources and Throughput

Database searching is generally faster than de novo sequencing, particularly when the database is indexed and filtration algorithms are used. However, the computational landscape is changing. Non-autoregressive de novo sequencing models such as PowerNovo2 achieve speeds that are competitive with database searching while providing the advantage of not requiring a reference database [8].

For large-scale proteomic applications, the throughput of the analysis pipeline matters. Researchers should evaluate whether their computational infrastructure can support the chosen approach. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover the practical aspects of running proteomics analyses, including resource requirements and reproducibility considerations [4]. The nf-core documentation describes community pipeline standards and usage, which can help researchers implement reproducible workflows for either database searching or de novo sequencing [5].

### Step 4: Validate Results with Complementary Approaches

Regardless of the primary analysis strategy, validation is essential. For database searching, validation typically involves controlling the false discovery rate using target-decoy approaches. For de novo sequencing, validation may involve comparing the derived sequences with homologous proteins from related organisms or confirming novel peptides with synthetic standards.

The complementarity of database searching and de novo sequencing is well documented. In the Modanovo study, the de novo sequencing approach recovered numerous confident peptides not reported by database search, including new viral phosphosites supported by spectral evidence. This demonstrates that the two approaches identify different subsets of peptides, and combining them provides more complete coverage than either approach alone [9].

## Options and Tradeoffs

### Database Searching with Custom Databases

One option for organisms lacking complete databases is to construct a custom database from transcriptomics or translatomics data. This approach was used successfully in the hepatocellular carcinoma study, where Ribo-seq data were used to generate a database of lncRNA-derived open reading frames [10]. The advantage of this approach is that it retains the speed and statistical rigor of database searching while extending coverage to novel sequences. The disadvantage is that it requires additional experimental data and substantial bioinformatics effort to construct and validate the database.

### De Novo Sequencing with Assembly

For applications such as antibody sequencing, de novo sequencing followed by assembly is the standard approach. The XA-Novo workflow integrates multi-enzymatic digestion with a beam search-based assembler to reconstruct full-length sequences from bottom-up data [11]. This approach is more complex than simple peptide identification because it requires assembling overlapping peptide sequences into full-length proteins. The computational cost is higher, but the output is a complete sequence instead of a list of peptide identifications.

### Hybrid Approaches

In practice, many researchers use both approaches. Database searching identifies the peptides that match known proteins, and de novo sequencing recovers peptides from the unmatched spectra. This hybrid approach maximizes coverage and is particularly valuable for samples containing a mixture of known and novel proteins. The complementarity demonstrated in the Modanovo study supports this strategy [9].

## Observations and Measurements

### Quality Metrics for Database Searching

For database searching, the primary quality metrics are the number of peptide-spectrum matches, the false discovery rate, and the number of identified proteins. These metrics are well established and supported by most search engines. Researchers should report these metrics in publications to enable comparison across studies.

### Quality Metrics for De Novo Sequencing

For de novo sequencing, quality metrics include the precision-coverage curve, which measures the tradeoff between the accuracy of predicted sequences and the fraction of the sequence covered. The Modanovo study used the area under the precision-coverage curve as the primary metric, reporting a median of 0.92 for modified peptides and 0.93 for unmodified peptides [9]. Researchers should be familiar with this metric and report it when publishing de novo sequencing results.

### Runtime and Throughput Measurements

Computational runtime is an important practical consideration. PowerNovo2 matches the accuracy of state-of-the-art autoregressive models while being 4.3 times faster [8]. For large-scale applications, this speed difference can be the deciding factor in whether de novo sequencing is feasible. Researchers should benchmark the runtime of their chosen tools on their own data and hardware before committing to a large-scale analysis.

## Records and Documentation

### Maintaining Analysis Records

Reproducibility requires careful documentation of the analysis workflow. The Carpentries lessons provide foundational training in computing, data management, shell, Git, and programming that supports reproducible analysis practices [6]. Researchers should record the software versions, parameter settings, and database versions used in each analysis. This information should be included in publications to enable others to reproduce the results.

### Workflow Management

Workflow management systems help ensure reproducibility by capturing the entire analysis pipeline in a structured format. The nf-core documentation describes community pipeline standards that support reproducible workflow implementation [5]. The Galaxy Training Network provides accessible workflow training that covers the practical aspects of running analyses reproducibly [4]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation for researchers who prefer to work in the R environment [3].

### Data Deposition

Depositing raw mass spectrometry data and analysis results in public repositories is standard practice in proteomics. NCBI provides data resources that support data deposition and retrieval [1]. Researchers should deposit their data to enable validation and reanalysis by others.

## Common Failure Patterns

### Failure Pattern 1: Using an Incomplete Database for a Non-Model Organism

The most common failure is applying database searching to an organism with an incomplete database. This results in a large fraction of unmatched spectra and a biased view of the proteome. The solution is to either construct a custom database from transcriptomics data or use de novo sequencing for the unmatched spectra.

### Failure Pattern 2: Ignoring Post-Translational Modifications

Database searching without specifying relevant modifications will miss modified peptides. However, specifying too many modifications creates a combinatorial explosion in the search space and increases computational cost. De novo sequencing with modification-aware tools such as Modanovo can detect modifications without this exhaustive search specification [9].

### Failure Pattern 3: Overlooking Novel Peptides in Well-Annotated Organisms

Even in well-annotated organisms, the database may be incomplete. The hepatocellular carcinoma study demonstrated that lncRNA-derived microproteins were largely unidentified, resulting in unmatchable MS/MS spectra [10]. Researchers working with well-annotated organisms should not assume that all spectra will match the database. Unmatched spectra should be investigated with de novo sequencing or custom database construction.

### Failure Pattern 4: Using Autoregressive De Novo Tools for Large Datasets

Autoregressive de novo sequencing tools are computationally slow and prone to error propagation. For large datasets, non-autoregressive tools such as PowerNovo2 provide substantial speed improvements without sacrificing accuracy [8]. Researchers should evaluate the computational requirements of their chosen tools before analyzing large datasets.

## Limitations and Interpretation Constraints

### De Novo Sequencing Accuracy Limits

De novo sequencing is not perfectly accurate. The precision-coverage curves reported in the literature show that accuracy decreases as the required sequence coverage increases. Researchers should understand that de novo sequences are predictions that require validation, particularly when the sequences are used for downstream applications such as antibody development.

### Database Searching Cannot Discover Novel Sequences

Database searching is fundamentally limited by the contents of the database. Peptides that are not present in the database cannot be identified, regardless of the quality of the spectra. This limitation is inherent to the approach and cannot be overcome by improving the search algorithm.

### Modification Detection Remains Challenging

While modification-aware de novo sequencing tools have improved, they are not yet comprehensive. The Modanovo study covered 19 amino acid-PTM combinations, but many other modifications exist [9]. Researchers studying less common modifications may still need to rely on database searching with carefully specified modification parameters.

### Computational Resource Requirements

De novo sequencing, particularly with deep learning models, requires substantial computational resources. The training and inference of transformer-based models require GPUs or specialized hardware. Researchers without access to such resources may need to use cloud computing or institutional high-performance computing facilities.

## Safety and Regulatory Context

### Data Integrity and Reproducibility

In proteomics research, data integrity is essential. The field has established standards for data deposition, analysis documentation, and reporting. Researchers should follow these standards to ensure that their results can be validated and reproduced. The training resources provided by EMBL-EBI cover data management and reproducibility best practices [2].

### Ethical Use of Sequence Data

De novo sequencing of antibodies and other proteins raises intellectual property considerations. Researchers should be aware of the patent landscape for antibody sequences and consult with their institutional technology transfer offices before pursuing commercial applications.

### Clinical Translation Considerations

For research with clinical implications, such as the hepatocellular carcinoma biomarker study, additional regulatory considerations apply. The discovery of novel microproteins with prognostic correlation requires validation in independent cohorts before clinical use can be considered [10]. Researchers should be cautious about overinterpreting exploratory findings.

## Professional Escalation Criteria

Researchers should escalate to more specialized expertise or alternative approaches in the following situations:

1. When more than 30 percent of high-quality spectra remain unmatched after database searching, consider de novo sequencing or custom database construction. If the unmatched fraction remains high after both approaches, consult with a proteomics bioinformatics specialist.

2. When de novo sequencing results are inconsistent across replicate runs or across different tools, the spectra may be of poor quality or the sample may contain unexpected modifications. Consult with a mass spectrometry specialist to evaluate spectral quality.

3. When antibody sequencing results are intended for therapeutic development, the sequences must be validated with functional assays. The XA-Novo study demonstrated that in vitro and in vivo assays are necessary to confirm that reconstructed antibodies exhibit functionality equivalent to their commercial counterparts [11].

4. When computational resources are insufficient for the chosen analysis approach, consult with institutional bioinformatics support or consider cloud-based analysis platforms.

5. When the research involves clinical samples and the results may influence patient care decisions, consult with clinical proteomics specialists and ensure compliance with applicable regulations.

## Building a Practical Decision Framework for Method Selection

Selecting between de novo sequencing and database searching requires a structured evaluation that goes beyond general principles. A practical decision framework helps researchers make consistent, defensible choices based on observable data about their samples, databases, and research objectives. This section provides a scoring system, a tiered workflow, and a record-keeping structure that researchers can adapt to their specific proteomics projects.

### The Database Completeness Score

The first step in any method selection is quantifying database completeness for the organism under study. instead of relying on subjective impressions, researchers should calculate a Database Completeness Score based on three measurable components.

The first component is the annotation status of the reference genome. Check whether the organism has a reference genome in NCBI, whether gene models are annotated, and whether a predicted proteome has been deposited [1]. Score this component from 0 to 3, where 0 means no genome sequence available, 1 means a genome assembly exists without annotation, 2 means annotated gene models exist, and 3 means a curated reference proteome is available.

The second component is the phylogenetic distance to well-annotated relatives. If the organism lacks its own complete database, determine whether closely related species have high-quality proteomes. Score this from 0 to 2, where 0 means no close relatives with annotated proteomes, 1 means relatives exist but with substantial evolutionary distance, and 2 means a closely related species has a well-annotated proteome.

The third component is the expected novelty of the sample. Consider whether the sample contains proteins from the organism itself, from pathogens, from environmental sources, or from engineered constructs. Score this from 0 to 2, where 0 means the sample is expected to contain only known proteins from the target organism, 1 means some novel sequences are expected, and 2 means the sample is expected to contain substantial novel sequence content.

The total Database Completeness Score ranges from 0 to 7. A score of 6 or higher indicates that database searching should be the primary approach. A score of 3 to 5 indicates that a hybrid approach is appropriate, where database searching handles the known fraction and de novo sequencing recovers novel peptides. A score below 3 indicates that de novo sequencing should be the primary approach, with database searching used only for confirmation of known sequences.

### The Sample Complexity Assessment

Sample complexity affects both the feasibility and the interpretation of each approach. Researchers should evaluate three dimensions of complexity before committing to a method.

The first dimension is the dynamic range of protein abundances. Samples with a wide dynamic range, such as plasma or serum, produce spectra dominated by high-abundance proteins, leaving low-abundance proteins underrepresented. Database searching handles this situation well because the search space is fixed and the scoring statistics account for the uneven distribution of peptide matches. De novo sequencing may struggle with low-quality spectra from low-abundance peptides, producing lower confidence sequence predictions.

The second dimension is the number of distinct proteins expected in the sample. A purified protein complex or a single antibody sample has low complexity and is well suited to de novo sequencing with assembly. A whole-cell lysate from a mammalian tissue has high complexity and is better handled by database searching for the known fraction, with de novo sequencing reserved for unmatched spectra.

The third dimension is the presence of post-translational modifications. If the sample is expected to contain phosphorylated, acetylated, or ubiquitinated peptides, the researcher must decide whether to specify these modifications in a database search or use a modification-aware de novo tool. The combinatorial expansion of modification types and sites makes database searching computationally expensive when many modifications are considered simultaneously [9]. Modification-aware de novo tools such as Modanovo can detect modifications without exhaustive search space specification, making them attractive for discovery-oriented PTM analysis [9].

### The Research Objective Matrix

Research objectives determine which analytical output is most valuable. The table below maps common research objectives to the appropriate primary approach and the complementary approach.

| Research Objective | Primary Approach | Complementary Approach |
| --- | --- | --- |
| Quantify known proteins across conditions | Database searching | De novo sequencing for unmatched spectra |
| Identify novel proteins from unannotated regions | De novo sequencing | Database searching with custom database |
| Sequence antibodies with unknown sequences | De novo sequencing with assembly | Database searching for framework regions |
| Characterize post-translational modifications | Modification-aware de novo sequencing | Database searching with specified modifications |
| Build a proteome catalog for a non-model organism | De novo sequencing | Database searching against related species |
| Validate predicted open reading frames | Database searching with custom database | De novo sequencing for confirmation |

The research objective matrix helps researchers avoid the common mistake of applying a single approach to all samples in a study. A study that combines quantitative comparisons of known proteins with discovery of novel proteins requires both approaches, applied to different fractions of the data.

### The Tiered Analysis Workflow

A tiered workflow provides a practical structure for applying the decision framework. This workflow proceeds through three tiers, with each tier adding analytical depth only when the previous tier leaves questions unanswered.

Tier 1 is the database search. Run a standard database search against the best available protein database for the organism. Use a target-decoy approach to control the false discovery rate. Record the number of peptide-spectrum matches, the false discovery rate, and the fraction of high-quality spectra that remain unmatched. The European Bioinformatics Institute provides training resources that cover the practical aspects of running database searches and interpreting the results [2].

Tier 2 is the unmatched spectra analysis. Collect the high-quality spectra that did not match the database. Apply de novo sequencing to these spectra to recover novel peptides. The complementarity of database searching and de novo sequencing is well documented, with de novo tools recovering numerous confident peptides not reported by database search [9]. For samples from organisms with incomplete databases, this tier often reveals substantial novel sequence content.

Tier 3 is the custom database construction. If the de novo sequences from Tier 2 reveal a substantial number of novel peptides, consider constructing a custom database from transcriptomics or translatomics data. The hepatocellular carcinoma study used high-quality Ribo-seq translatomic datasets to generate a database of lncRNA-derived open reading frames, which was then applied to proteomics data [10]. This approach retains the speed and statistical rigor of database searching while extending coverage to novel sequences.

The tiered workflow ensures that researchers do not jump directly to de novo sequencing when a database search would suffice, and it ensures that researchers do not stop at database searching when substantial novel sequence content remains unidentified.

### The Method Selection Scorecard

A scorecard provides a structured way to document the method selection decision. Researchers should complete the scorecard before beginning the analysis and update it as results accumulate.

The scorecard should record the Database Completeness Score, the sample complexity assessment, the research objective, and the selected primary approach. It should also record the expected runtime for each approach, the computational resources available, and the validation strategy.

For database searching, record the database version, the search engine, the precursor and fragment mass tolerances, the enzyme specificity, the number of missed cleavages allowed, and the modifications specified in the search. For de novo sequencing, record the tool version, the model type, the precision-coverage threshold, and the validation approach.

The scorecard serves as a documentation tool that supports reproducibility. The Carpentries lessons provide foundational training in data management and reproducible analysis practices that support this documentation effort [6]. The nf-core documentation describes community pipeline standards that support reproducible workflow implementation [5].

### Records and Measurements for Method Comparison

Comparing the performance of database searching and de novo sequencing on the same dataset requires consistent measurements. Researchers should record the following metrics for each approach.

For database searching, record the number of peptide-spectrum matches at a specified false discovery rate, the number of identified proteins, the fraction of high-quality spectra matched, and the runtime. For de novo sequencing, record the number of peptides sequenced at a specified precision threshold, the median precision-coverage area, the fraction of spectra yielding confident sequences, and the runtime.

The precision-coverage curve is the standard metric for evaluating de novo sequencing tools. The Modanovo study reported a median area under the precision-coverage curve of 0.92 for modified peptides and 0.93 for unmodified peptides [9]. Researchers should report this metric when publishing de novo sequencing results.

Runtime measurements are particularly important for large-scale applications. PowerNovo2 matches the accuracy of state-of-the-art autoregressive models while being 4.3 times faster [8]. For datasets with hundreds of thousands of spectra, this speed difference can be the deciding factor in whether de novo sequencing is feasible.

### Common Failure Patterns in Method Selection

Several recurring failure patterns emerge when researchers apply the decision framework incorrectly.

The first failure pattern is over-reliance on database searching for non-model organisms. Researchers working with organisms that lack complete databases often run database searches against related species and accept a high fraction of unmatched spectra. This approach produces a biased view of the proteome, missing the most interesting novel sequences. The solution is to apply the tiered workflow, using de novo sequencing for unmatched spectra.

The second failure pattern is premature de novo sequencing for well-annotated organisms. Researchers sometimes apply de novo sequencing to all spectra without first running a database search, producing lower-confidence identifications for peptides that could be identified with high confidence by database searching. The solution is to use database searching as the first tier and reserve de novo sequencing for unmatched spectra.

The third failure pattern is ignoring post-translational modifications in the method selection. Researchers who study phosphorylation but use a database search without specifying phosphorylation as a variable modification will miss modified peptides. Researchers who specify too many modifications create a combinatorial explosion in the search space. The solution is to use modification-aware de novo tools such as Modanovo for PTM discovery [9].

The fourth failure pattern is neglecting validation of de novo sequences. De novo sequences are predictions that require validation, particularly when used for downstream applications such as antibody development. The XA-Novo study demonstrated that in vitro and in vivo assays are necessary to confirm that reconstructed antibodies exhibit functionality equivalent to their commercial counterparts [11].

### Practical Implementation Steps

Implementing the decision framework requires several practical steps that researchers can complete in a single session.

First, query NCBI to assess database availability for the organism under study [1]. Record the annotation status and the availability of a reference proteome. This step takes approximately 15 minutes and provides the foundation for the Database Completeness Score.

Second, complete the sample complexity assessment by reviewing the experimental design. Determine the expected dynamic range, the number of distinct proteins, and the expected post-translational modifications. This step requires knowledge of the sample preparation and the biological system under study.

Third, complete the research objective matrix by writing a clear statement of the research question. Determine whether the primary output is quantitative comparison of known proteins, discovery of novel sequences, or sequence determination for unknown proteins.

Fourth, select the primary approach based on the Database Completeness Score, the sample complexity assessment, and the research objective matrix. Document the selection in the method selection scorecard.

Fifth, run the tiered workflow, starting with database searching and proceeding to de novo sequencing for unmatched spectra. Record all metrics in the scorecard.

Sixth, validate the results using complementary approaches. For database searching, validate with false discovery rate control. For de novo sequencing, validate with synthetic peptide standards or functional assays where applicable.

### Professional Escalation Criteria

Researchers should escalate to specialized expertise when the decision framework produces ambiguous results or when the analysis reveals unexpected complexity.

Escalate when the Database Completeness Score is between 3 and 5 and the tiered workflow produces conflicting results between database searching and de novo sequencing. A proteomics bioinformatics specialist can help interpret the discrepancies and determine whether the database is incomplete or the de novo sequences are unreliable.

Escalate when more than 30 percent of high-quality spectra remain unmatched after both database searching and de novo sequencing. This situation may indicate unexpected modifications, unusual digestion patterns, or sample contamination. A mass spectrometry specialist should evaluate spectral quality and sample preparation.

Escalate when de novo sequencing results are inconsistent across replicate runs or across different tools. This inconsistency may indicate poor spectral quality or the presence of unexpected modifications. Consult with a mass spectrometry specialist to evaluate the spectra.

Escalate when antibody sequencing results are intended for therapeutic development. The sequences must be validated with functional assays before any clinical or commercial application [11]. Consult with institutional technology transfer offices and regulatory specialists.

Escalate when computational resources are insufficient for the chosen approach. Non-autoregressive de novo tools such as PowerNovo2 provide substantial speed improvements, but they still require appropriate hardware [8]. Consult with institutional bioinformatics support or consider cloud-based analysis platforms.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover the practical aspects of running proteomics analyses, including resource requirements and reproducibility considerations [4]. Bioconductor provides official package, workflow, installation, and reproducible genomic-analysis documentation for researchers who prefer to work in the R environment [3]. These resources can help researchers implement the decision framework and troubleshoot common problems.

## Frequently Asked Questions

### What is the main difference between de novo sequencing and database searching?

Database searching matches experimental MS/MS spectra against predicted peptides from a protein database, while de novo sequencing derives peptide sequences directly from the spectra without a reference database. Database searching requires a complete database for the organism under study, whereas de novo sequencing works for any organism. The choice depends on database availability, sample complexity, and research goals.

### When should I use de novo sequencing instead of database searching?

Use de novo sequencing when working with organisms lacking complete databases, when searching for novel proteins or unannotated open reading frames, when sequencing antibodies with unknown sequences, or when studying post-translational modifications that are not well characterized. Database searching is preferred for well-annotated organisms and for quantitative proteomics where direct peptide-to-protein mapping is needed.

### Can de novo sequencing detect post-translational modifications?

Yes, recent tools such as Modanovo can detect post-translational modifications including phosphorylation, acetylation, and ubiquitination. Modanovo was trained on spectra spanning 19 biologically relevant amino acid-PTM combinations and achieved a median area under the precision-coverage curve of 0.92 for modified peptides [9]. This capability makes de novo sequencing valuable for PTM discovery in systems where modification sites are not well characterized.

### How does computational speed compare between the two approaches?

Database searching is generally faster, particularly with indexed databases and filtration algorithms. However, non-autoregressive de novo sequencing models such as PowerNovo2 match the accuracy of state-of-the-art autoregressive models while being 4.3 times faster [8]. For large-scale applications, the speed difference between database searching and modern de novo tools is narrowing.

### What is a precision-coverage curve in de novo sequencing?

A precision-coverage curve measures the tradeoff between the accuracy of predicted peptide sequences and the fraction of the sequence covered. The area under this curve is a standard metric for evaluating de novo sequencing tools. Higher values indicate better performance, with values around 0.92 to 0.94 reported for state-of-the-art tools [9].

### Can I use both database searching and de novo sequencing on the same dataset?

Yes, combining both approaches is a common and effective strategy. Database searching identifies peptides that match known proteins, while de novo sequencing recovers peptides from unmatched spectra. The Modanovo study demonstrated that de novo sequencing recovered numerous confident peptides not reported by database search, showing that the two approaches are complementary [9].

### How do I construct a custom database for an organism without a complete reference proteome?

Custom databases can be constructed from transcriptomics or translatomics data. The hepatocellular carcinoma study used high-quality Ribo-seq translatomic datasets to generate a database of lncRNA-derived open reading frames, which was then applied to proteomics data [10]. This approach requires additional experimental data and bioinformatics expertise but retains the speed and statistical rigor of database searching.

### What validation is needed for de novo sequencing results?

De novo sequences should be validated using multiple approaches. For antibody sequencing, functional assays are necessary to confirm that reconstructed antibodies exhibit functionality equivalent to their commercial counterparts [11]. For novel peptides, validation may involve comparing sequences with homologous proteins from related organisms or confirming with synthetic peptide standards.

## Related Bioinformatics Guides

- [Mass Spectrometry-Based Proteomics: Data Analysis Pipelines and Tools](/knowledge/bioinformatics/mass-spectrometry-based-proteomics-data-analysis-pipelines-and-tools)
- [Proteomics Mass Spectrometry: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/proteomics-mass-spectrometry-from-sample-preparation-to-data-analysis)
- [Single-Cell Sequencing Databases: Resources for Data Sharing and Exploration](/knowledge/bioinformatics/single-cell-sequencing-databases-resources-for-data-sharing-and-exploration)
- [Spatial Proteomics Mass Spectrometry: Techniques and Applications](/knowledge/bioinformatics/spatial-proteomics-mass-spectrometry-techniques-and-applications)
- [RNA-Seq Databases: Accessing and Using Public RNA-Seq Data](/knowledge/bioinformatics/rna-seq-databases-accessing-and-using-public-rna-seq-data)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Peptide sequence tags for fast database search in mass-spectrometry.](https://pubmed.ncbi.nlm.nih.gov/16083278). Journal of proteome research, 2005.
- [PowerNovo2: A generative flow-based approach to non-autoregressive de novo peptide sequencing.](https://doi.org/10.1371/journal.pcbi.1014298). 2026.
- [Modanovo: A Unified Model for Post-translational Modification-Aware De Novo Sequencing Using Experimental Spectra From In Vivo and Synthetic Peptides.](https://doi.org/10.1016/j.mcpro.2025.101501). 2026.
- [A Proteogenomic Approach to Discover Novel lncRNA-Derived Microproteins and Their Potential Clinical Utility in Hepatocellular Carcinoma.](https://doi.org/10.1016/j.mcpro.2026.101584). 2026.
- [XA-Novo: high-throughput mass spectrometry-based de novo sequencing technology for monoclonal antibodies and antibody mixtures.](https://doi.org/10.1038/s41467-026-70496-y). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.