Antibody Sequence Database
An antibody sequence database is a curated digital repository that stores the nucleotide or amino acid sequences of immunoglobulin variable domains, constant regions, full length antibodies, or single chain fragments. If you are a researcher, bioinformatician, or therapeutic antibody developer who needs to search for known antibody sequences, compare new sequencing data against reference sets, or annotate clonal families, this guide gives you a practical, source bounded framework for using these databases correctly and critically.
The first step in using any antibody sequence database is to understand that the sequence of an antibody is not static. B cells undergo V(D)J recombination, somatic hypermutation, and class switching, meaning an expressed antibody sequence differs from the germline. Reliable databases distinguish between germline alleles and rearranged immune repertoires, and you must choose the appropriate type for your question. As the NCBI Bookshelf explains, reference sequence collections are essential for standardizing comparisons across samples and experiments NCBI Bookshelf. Similarly, the EMBL EBI training portal provides dedicated resources for navigating immunological data, including step by step guides for searching antibody specific databases EMBL EBI Training. In the first two paragraphs, you have already seen two foundational sources. All subsequent sections will continue to link directly to authoritative materials so that you can verify the guidance yourself.
At a Glance
| Aspect | Key Details |
|---|---|
| Purpose | Store, annotate, and retrieve immunoglobulin variable region sequences for research and engineering |
| Primary databases | IMGT (international ImMunoGeneTics), SAbDab (structural antibody database), AbDb (antibody database), NCBI IgBLAST databases |
| Data types | Germline V(D)J alleles, rearranged receptor sequences, crystal structures with sequences |
| Typical use cases | Clonotyping, humanization, CDR identification, repertoire analysis, lineage tracing |
| Access methods | Web interface, BLAST search, API, local download for high throughput use |
| Limitations | Incomplete species coverage, allelic variation not fully represented, occasional annotation errors |
Decision Criteria for Choosing an Antibody Sequence Database
Not every antibody sequence database serves the same purpose. Your choice depends on three main factors: the question you are asking, the species you are studying, and the resolution you need.
Germline versus rearranged databases. If you need to identify V, D, or J gene usage from a new sequencing read, you must align against a germline database. The international ImMunoGeneTics database (IMGT) is the authoritative reference for germline sequences across many vertebrate species. If you need to compare a newly discovered antibody to known therapeutic antibodies or published crystal structures, you should use a database of rearranged sequences, such as the Structural Antibody Database (SAbDab) which links sequence and 3D structure. The Galaxy Training Network offers a workflow oriented overview of how to choose alignment references for antibody repertoire analysis Galaxy Training Network.
Species coverage. Most public databases focus heavily on human and mouse sequences. If you work with camelid, shark, or avian antibodies, confirm that the database includes those lineages. IMGT provides data for over 40 species, but the completeness varies. For less common species, you may need to search the NCBI Sequence Read Archive (SRA) for raw immune repertoire sequencing and construct your own reference set NCBI Sequence Read Archive.
Level of annotation. Some databases provide only raw sequence strings. Others include curated CDR boundaries, framework regions, and links to functional data or clinical studies. The Bioconductor project contains packages, such as immcc and dplyr based tools, that can integrate multiple data sources and attach annotations programmatically Bioconductor. Choose a database that matches the depth of metadata you need for your downstream analysis.
Practical Workflow for Using Antibody Sequence Databases
The following step by step workflow will help you retrieve and validate antibody sequences for your project. Each step includes a quality control check.
Step 1: Define Your Query
Write down the exact biological question. Are you identifying the germline family of a single rearranged sequence? Are you searching for all sequences that share a particular CDR3 motif? Or are you downloading a reference set for bulk repertoire annotation? Write a clear query before you touch a search box.
Step 2: Select the Appropriate Database
Use the decision criteria above. For germline assignments, go to IMGT/V QUEST or IgBLAST. For known therapeutic antibodies, visit SAbDab or The Antibody Society therapeutic database. For structure informed design, use SAbDab which links PDB structures to antibody sequences. The EMBL EBI training resources include a module titled "Searching for antibody sequences" that walks through these choices EMBL EBI Training.
Step 3: Perform the Search and Retrieve Sequences
Most databases accept FASTA formatted input. For a single sequence, a web BLAST against the chosen database is sufficient. For hundreds or thousands of sequences, use a programmatic interface. For example, IgBLAST can be run locally through the NCBI command line tools. The Galaxy Platform provides ready made workflows for batch alignment of immune repertoires, as documented in the Galaxy Training Network materials Galaxy Training Network.
Quality check: After retrieval, inspect the top hits. Do they match your expected isotype or species? CDR3 length should be consistent with known antibody repertoire features. A mismatch in any of these indicates a database selection error or a contaminated query sequence.
Step 4: Annotate the Sequences
Once you have the raw sequence and the assigned V, D, J genes, you need to demarcate framework regions and CDRs. Use the IMGT numbering scheme or Kabat numbering depending on your downstream use. Many online tools provide CDR annotation automatically. After annotation, verify that the CDR3 start and end conform to the conserved C terminal cysteine and the FGG or FGX motif. The Bioconductor package nestedRates or alakazam can automate this validation for large datasets Bioconductor.
Step 5: Document Your Source
Record the exact database version, date of access, and parameters used. Antibody germline databases are updated periodically as new alleles are discovered. Without version tracking, you cannot reproduce your analysis later.
Step 6: Validate with Literature or Experimental Data
Compare your retrieved sequence against published sequences from the same study or against PDB structures. For example, a recent study on artificial intelligence advancements in monoclonal antibody development used curated sequence databases to train models predicting antigen binding PubMed 42427491. Such validation ensures your database derived sequence is realistic.
Common Mistakes When Using Antibody Sequence Databases
Mistake 1: Using a rearranged database for germline assignments. Aligning a somatic mutated sequence against other rearranged antibodies will obscure the original germline origin. Always use a dedicated germline reference set.
Mistake 2: Ignoring allele variants. Many published sequences are assigned to a gene family but not to a specific allele. When engineering a chimeric antibody, using the wrong allele can affect expression or immunogenicity. Check the IMGT allele number.
Mistake 3: Forgetting that databases have errors. Manual curation reduces errors but does not eliminate them. A sequence labeled as human may contain a few ambiguous residues or an incorrect CDR boundary. Cross reference with at least one other database.
Mistake 4: Using only one database for rare species. If you work with rabbit, rat, or llama antibodies, do not rely solely on IMGT. Supplement with sequences from SRA and with published data. The NCBI SRA can provide raw reads that you can assemble into allele candidates NCBI Sequence Read Archive.
Mistake 5: Over interpreting absence of a sequence. A negative search result does not prove that the antibody does not exist. It may mean that the sequence has not been submitted to the database yet. Always report negative results with caution.
Limits and Uncertainty in Antibody Sequence Databases
No publicly available antibody sequence database covers every possible antibody. Germline databases are incomplete for many species, and rearranged databases are biased toward heavily studied therapeutic antibodies. The quality of sequence records varies. Some entries come from patent filings and may contain intentional obfuscation. Others come from early low throughput sequencing with higher error rates.
The development of new alleles and the discovery of rare V genes means that databases are always a snapshot. When you perform a BLAST search, the top hit may not be the closest biological match, especially if your query comes from an underrepresented population. The IMGT database uses a strict naming convention, but other databases may use different nomenclature, leading to confusion.
Antibody sequence databases also do not capture the functional context. A sequence alone does not tell you affinity, specificity, or stability. You must integrate experimental data to fully characterize an antibody. Recent studies, such as one linking S100A8/S100A9 to cardiac progenitor cell dysfunction, used transcriptomic and single cell analysis to move from sequence to function PubMed 42437006. Similarly, the biology of GPR17+ oligodendrocyte lineage cells involved novel chondroitin sulfate structures that are not encoded solely by antibody sequences PubMed 42419731. These examples remind us that a sequence database entry is a starting point, not an endpoint.
Frequently Asked Questions
1. What is the difference between IMGT and SAbDab? IMGT is the primary reference database for germline V, D, and J gene sequences, alleles, and rearranged sequences, with sequence annotation according to the IMGT numbering system. SAbDab is a structural antibody database that stores sequences and crystal structures of antibodies, with links to PDB files. Use IMGT for gene assignment and SAbDab for structure informed design.
2. Can I BLAST a patient derived antibody sequence against the germline database to find its V gene? Yes. Use IgBLAST or IMGT/V QUEST with the germline reference set. Be aware that somatic hypermutation will cause many mismatches, but the alignment should still identify the closest V allele.
3. How do I assign CDRs from a sequence without a structure? Use a numbering scheme such as IMGT or Kabat. Tools like IMGT/DomainGapAlign or the Bioconductor package CDR3 can delineate CDRs based on conserved flanking residues. Validate by checking the conserved cysteine at position 104 and the tryptophan after CDR3.
4. Are antibody sequence databases useful for studying bispecific antibodies? Yes, but you must treat each variable domain as a separate sequence record. Some databases, like SAbDab, include bispecific constructs. Check the metadata carefully to ensure you have both the targeting sequences and the correct assembly.
References and Further Reading
- NCBI Bookshelf General bioinformatics reference for sequence databases and analysis.
- EMBL EBI Training Courses on antibody sequence searching and immunoinformatics.
- Galaxy Training Network Workflow tutorials for bulk immune repertoire analysis.
- Bioconductor Software packages for antibody repertoire annotation and statistical analysis.
- NCBI Sequence Read Archive Repository for raw immune repertoire sequencing data for custom database construction.
- PubMed 42427491 Artificial intelligence advancements in monoclonal antibody development technology.
- PubMed 42437006 Integrated transcriptomic and single cell study linking S100A8/S100A9 to cardiac progenitor dysfunction (example of moving from sequences to function).
- PubMed 42419731 GPR17+ oligodendrocyte lineage cells and novel chondroitin sulfate structures (example of sequence functional interplay).
- PubMed 42430569 Clinical study of rituximab based regimen for primary CNS lymphoma (example of therapeutic antibody use).
- PubMed 42417171 MET amplified gastric cancers with co amplifications of BRAF (example of cancer genomics where antibody sequence databases inform targeting).