Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

NCBI Genome Submission

If you have sequenced and assembled a genome and want to share it with the scientific community, submitting it to NCBI is the standard route. This guide walks you through the core concepts, decision points, and practical steps needed to submit a genome to NCBI databases, from preparing your data to completing the release. It is written for researchers, bioinformaticians, and lab managers who have generated a genome assembly and require a clear, source bounded framework for submission. NCBI’s own technical resources NCBI Bookshelf provide the authoritative procedures, while training materials from EMBL EBI Training offer complementary best practices for data formatting and metadata.

At a Glance

Aspect Details
Goal Make genome assembly publicly accessible and citable
Primary databases GenBank (submitted genomes), SRA (raw reads), BioProject/BioSample (project and sample metadata)
Required inputs Assembly file (FASTA), annotation file (if applicable), metadata (organism, sequencing platform, etc.)
Submission tools NCBI Genome Submission Portal (web), tbl2asn (command line), Submission Portal API
Typical timeline 1-2 weeks for review, can be faster if automated checks pass
Key validation steps Contig naming, vector contamination screen, assembly statistics check
Release options Immediate release or hold until a linked publication

Decision Criteria

Before beginning, you need to answer several questions that determine the exact submission path.

What level of assembly do you have?
NCBI accepts contig, scaffold, chromosome, and complete genome levels. The choice affects annotation requirements and the final accession format. For a complete genome, you must provide evidence of physical closure.

Which NCBI database should you use?
Raw sequencing data go to the Sequence Read Archive (SRA) as described on the NCBI SRA resource. The genome assembly itself goes to GenBank. You must also register a BioProject and BioSample before submitting the assembly.

Will you provide annotation?
GenBank submissions can include annotation files in Sequin or tbl2asn format. If you do not provide annotation, NCBI may perform an automated structural annotation after submission.

Are you submitting a eukaryotic or prokaryotic genome?
The workflows differ slightly. Prokaryotic genomes often use the Prokaryotic Genome Annotation Pipeline (PGAP) integration, while eukaryotic genomes may undergo manual curation. Refer to the NCBI Bookshelf documentation for the specific pipeline requirements.

Practical Workflow or Implementation Sequence

Below is a step by step sequence for a typical genome submission. The process is iterative, and you may need to revisit earlier steps after validation.

Step 1: Register a BioProject and BioSample

These are umbrella records that link your genome to the project and sample metadata. Go to the NCBI Submission Portal and create a BioProject with a brief title, funding source, and relevant publications. Then create a BioSample with attributes such as organism name, isolation source, and geographic location. Each genome submission requires a unique BioSample.

Step 2: Prepare the Assembly File

Your assembly must be in FASTA format with contig names that follow NCBI naming conventions (no spaces, no special characters like pipe or comma). Each sequence header should have a unique identifier. For example: >scaffold_1. NCBI provides a validation tool on the submission portal that checks for illegal characters and sequence ambiguity. Use the Galaxy Training Network workflow for assembly quality assessment before submission to catch issues early.

Step 3: Prepare Annotation (Optional but Recommended)

If you have gene predictions, create an annotation file in tbl2asn format. The official Bioconductor package GenomicRanges can help you manipulate feature positions and export GFF3, which can then be converted to tbl2asn using NCBI utilities. Without annotation, your submission will undergo automated NCBI annotation after acceptance.

Step 4: Submit Through the Genome Submission Portal

Log in to the NCBI Genome Submission Portal (https://submit.ncbi.nlm.nih.gov). Choose the submission type “Genome Assembly.” You will be prompted to:

  • Link your BioProject and BioSample
  • Upload the FASTA file (and annotation file if any)
  • Enter metadata: sequencing platform, assembly method, coverage depth, and relevant publications

The portal will run automated validation checks. Correct any errors before proceeding. This process is documented in detail on the NCBI Bookshelf.

Step 5: Include Raw Reads in SRA

If your raw reads are not already in SRA, submit them using the SRA Submission Portal. The NCBI SRA resource requires FASTQ or BAM files along with metadata matching the BioSample. This step is critical because NCBI reviewers often cross reference raw data quality with assembly quality.

Step 6: Review and Release

After submission, NCBI curators manually review the assembly for contamination, correct organism identification, and metadata completeness. The review typically takes a few business days. You can choose to release the genome immediately upon acceptance or hold it until a related manuscript is published (or until a specified date).

Quality Checks

Do not rely solely on NCBI’s automated checks. Perform your own verification before submission.

  • Contig naming: Ensure no spaces, pipes, commas, or colons in headers. Use only alphanumeric characters and underscores.
  • Vector contamination: Screen your assembly against the UniVec database using NCBI’s VecScreen tool. Remove any contaminant sequences.
  • Assembly statistics: Compute N50, total length, and number of contigs. Compare with expected values for your organism. The EMBL EBI Training modules on genome assembly provide benchmarks.
  • Annotation consistency: If providing annotation, verify that all CDS features have start and stop codons and that gene IDs match accepted formats.
  • Organism name: Use the exact scientific name from NCBI Taxonomy. A misnamed organism can delay review.

Common Mistakes

  • Incorrect FASTA format: Using line breaks longer than 80 characters or including non ASCII characters. NCBI accepts only plain text ASCII FASTA.
  • Missing metadata: Submitting without a BioProject or BioSample is the most frequent reason for rejection. Always register these first.
  • Not screening for contamination: A common oversight especially for metagenomes or low coverage assemblies. The NCBI SRA records often reveal source contamination.
  • Ignoring automated validation warnings: Many users skip the pre submission validation and then face manual rejection. Respect every warning.
  • Using generic contig names: Names like “Contig1” are acceptable but vague. Using unique identifiers like “scaffold_00001” helps downstream processing.
  • Submitting duplicate assemblies: If you have already submitted a version, use the update function instead of creating a new record.

Limits and Uncertainty

NCBI genome submission is robust but has boundaries. Not all genome types are accepted equally. Metagenome assembled genomes (MAGs) require additional scrutiny and must meet minimum quality standards (e.g., completeness >90%, contamination <5%). For eukaryotic genomes, large numbers of unplaced scaffolds may be rejected if the assembly is too fragmented.

The timeline for review is approximate. During peak submission periods or for complex eukaryotic genomes, curation can take several weeks. Additionally, automated annotation pipelines from NCBI may not perfectly predict gene models for non model organisms. You should view the final annotated record as a best effort, not a definitive interpretation.

For examples of real world submissions, studies such as the diversity of endophytic bacteria from plantain Diversity, molecular identification and plant growth promoting potential of endophytic bacteria from plantain (Plantago lanceolata L.) and the Earth BioGenome Project Phase II The Earth BioGenome Project Phase II: illuminating the eukaryotic tree of life illustrate the range of organisms and submission scales. For genetic variant studies that rely on submitted genomes, see Genetic variants associated with Sjögren's disease subtypes stratified by clinical feature or the use of targeted NGS in lymphoma Beyond Binary FISH Results: Targeted NGS Provides Deeper Characterization of MYC and Broader Genomic Alterations in DLBCL. These papers highlight how submitted genomes underpin further research.

Frequently Asked Questions

Q: Do I need to submit raw sequencing data before my genome assembly?
A: No, but they are strongly encouraged. NCBI prefers that raw reads be in SRA to verify your assembly. Submissions without raw data may still be accepted if you provide an explanation and the assembly is of high quality.

Q: Can I update my genome submission after it has been released?
A: Yes. Use the “Update” function in the Submission Portal to submit a new version of the assembly. The old version will remain accessible with a note about the update.

Q: What file formats does NCBI accept for annotation?
A: NCBI accepts Sequin (.sqn), tbl2asn (.tbl), and GFF3 (converted via tbl2asn). Do not submit raw GFF3 files directly, convert them first using NCBI utilities.

Q: How long does it take for a genome to appear in GenBank after I hit submit?
A: After manual review (typically 1 2 weeks) and acceptance, the genome appears in GenBank within 1 2 business days. If you choose immediate release, it becomes public then. If you choose a hold, it is stored privately until the release date.

References and Further Reading

Related Articles