# How to Submit Your Genome Assembly to NCBI: A Step-by-Step Guide to GenBank and RefSeq Compliance


## Key Takeaways

- Genome assembly submission to NCBI necessitates three interconnected registrations: BioProject (research initiative), BioSample (biological material), and the assembly itself, all requiring consistent metadata and adherence to NCBI's Submission Portal workflow.
- Sequence files must be in FASTA format with unique, valid headers; AGP format is mandatory for assemblies with scaffold structure and gaps, detailing contig arrangement and gap characteristics.
- Assembly quality, assessed by metrics like contig N50 and total gap length, is critical, with RefSeq inclusion demanding higher contiguity and completeness than GenBank, often requiring additional sequencing or assembly improvements.
- Contamination screening against common laboratory and adapter sequences is paramount, as any detected contamination can lead to rejection or delayed processing, particularly for RefSeq submissions.
- Metadata consistency across BioProject, BioSample, and assembly submissions, especially organism scientific names and strain identifiers verified against NCBI Taxonomy, is a primary determinant of successful validation and avoids common delays.
- GenBank accepts all valid assemblies for public access, while RefSeq inclusion requires rigorous quality checks, including contamination screening and annotation completeness, often leveraging pipelines like PGAP for prokaryotes.

---

Submitting a genome assembly to NCBI requires preparation of sequence files, metadata, and adherence to database standards. This guide walks through the complete workflow using the NCBI Submission Portal, covering BioProject registration, BioSample metadata, and assembly submission, with attention to common pitfalls that cause delays or rejection. The target reader is a researcher or laboratory professional who has generated an assembly and needs to deposit it in GenBank or pursue RefSeq inclusion.

## Scope of NCBI Assembly Submission

The National Center for Biotechnology Information operates the Assembly database, which provides stable accessioning and data tracking for genome assembly data. The database model accommodates a range of assembly structures, including sets of unordered contig or scaffold sequences, bacterial genomes consisting of a single complete chromosome, or complex structures such as a human genome with modeled allelic variation. The Assembly database provides an assembly accession and version to unambiguously identify the set of sequences that make up a particular version of an assembly, and tracks changes to updated genome assemblies. It reports metadata such as assembly names, simple statistical reports of the assembly including number of contigs and scaffolds, contiguity metrics such as contig N50, total sequence length, and total gap length, as well as the assembly update history. The database also tracks the relationship between an assembly submitted to the International Nucleotide Sequence Database Consortium (INSDC) and the assembly represented in the NCBI RefSeq project. Users can find assemblies of interest by querying the Assembly Resource directly or by browsing available assemblies for a particular organism. Links in the Assembly Resource allow users to easily download sequence and annotations for current versions of genome assemblies from the NCBI genomes FTP site.

The submission process involves three linked registrations: a BioProject that describes the research initiative, a BioSample that describes the biological material sequenced, and the assembly submission itself. These three components must be consistent with each other and with the sequence data being deposited. The NCBI Submission Portal at https://submit.ncbi.nlm.nih.gov/ serves as the entry point for all three processes.

## At a Glance

| Submission Component | Purpose | Key Requirement | Common Delay Cause |
| --- | --- | --- | --- |
| BioProject | Links the assembly to a research initiative | Unique project title and description | Duplicate or vague project descriptions |
| BioSample | Describes the biological source material | Accurate organism name and isolation metadata | Mismatched organism names between sample and assembly |
| Assembly Submission | Deposits the sequence files | FASTA or AGP format with valid sequence headers | Contig names with illegal characters or duplicate identifiers |
| GenBank vs RefSeq | Determines processing and annotation path | GenBank accepts all valid assemblies, RefSeq requires additional quality checks | Contamination or incomplete annotation for RefSeq requests |

## Preparing Your Assembly Files

### Sequence File Formats

The primary file format for assembly submission is FASTA, containing the assembled contigs or scaffolds as nucleotide sequences. Each sequence record must have a header line beginning with a greater-than symbol followed by a unique identifier. The identifier must be unique within the file, must not contain spaces or special characters beyond standard alphanumeric characters and underscores, and should be short enough to be displayed clearly in database records. NCBI provides validation tools that check for common formatting errors before submission.

For assemblies that include scaffold structure with gaps, the AGP format is required. AGP files describe the arrangement of contigs within scaffolds, including gap sizes and gap types. The AGP specification is maintained by NCBI and must be followed precisely. Errors in AGP files, such as incorrect gap lengths or contig ordering, will cause validation failures.

### Sequence Quality Considerations

The quality of the assembly directly affects the likelihood of acceptance and the usefulness of the deposited data. The Assembly database reports contiguity metrics such as contig N50, total sequence length, and total gap length. These metrics are displayed in the Assembly Resource and are used by the scientific community to evaluate assembly quality. Before submission, calculate these metrics for your assembly and compare them with existing assemblies for the same or related organisms. If your assembly has substantially lower contiguity than expected for your organism and sequencing technology, consider whether additional sequencing or assembly improvement is warranted before submission.

The red flour beetle Tribolium castaneum provides an example of assembly improvement through additional sequencing. The first version of the genome assembly was generated by Sanger sequencing with a small set of RNA sequence data limiting annotation quality. An improved assembly was generated by adding large-distance jumping library DNA sequencing to join scaffolds and fill small gaps, reducing gaps and increasing the N50 to 4753 kbp. This example illustrates that assembly quality can be improved through additional library types and that the improved assembly was accepted by GenBank and subsequently as a RefSeq genome. For your own assembly, consider whether additional long-range or long-read sequencing data could improve contiguity before submission.

### Contamination Screening

Contamination is a frequent cause of assembly rejection or delayed processing. NCBI evaluates assemblies for contamination, and the RefSeq project has become more stringent in evaluating assemblies for contamination and completeness of annotation prior to acceptance. Screen your assembly against databases of common contaminants, including adapter sequences, vector sequences, and sequences from organisms that are common laboratory contaminants. Remove any contigs that match contaminants before submission. Document the screening process and be prepared to provide information about the screening method if requested by NCBI staff.

## Registering a BioProject

The BioProject database organizes submissions around research initiatives. A BioProject can encompass multiple samples and multiple assemblies, such as a project that sequences several strains of a bacterial species or multiple tissues from a single organism. The BioProject registration requires a title, a description of the project's scope, and information about the submitting organization.

When registering a BioProject, provide a specific and descriptive title that distinguishes your project from others. Avoid generic titles that do not convey the organism or the purpose of the sequencing. The description should include the biological question being addressed and the types of data being generated. If your project is part of a larger initiative, such as a genome sequencing consortium, indicate that relationship in the registration.

The BioProject accession, which begins with PRJNA for NCBI-registered projects, will be required for the BioSample registration and the assembly submission. Register the BioProject first, before preparing the BioSample or assembly files, so that the accession is available when needed.

## Registering a BioSample

The BioSample database stores descriptive information about the biological source material used in the sequencing project. This includes the organism name, strain or isolate identifier, collection date, geographic location, and host information for pathogens or symbionts. The level of detail required depends on the organism type and the intended use of the data.

For the organism name, use the complete scientific name as recognized by NCBI Taxonomy. Verify the taxonomy name using the NCBI Taxonomy database before submission. Mismatched organism names between the BioSample and the assembly submission are a common cause of validation errors. If your organism does not have a formal taxonomy entry, you may need to request a new taxonomy ID before proceeding with the submission.

The BioSample registration form varies by organism type. Bacterial and archaeal samples require isolation source, collection date, and geographic location. Eukaryotic samples may require tissue type, developmental stage, or other biological context. Provide as much detail as available, but do not fabricate metadata. If a field is unknown, indicate that it is unknown instead of providing a guess.

The BioSample accession, which begins with SAMN for NCBI-registered samples, will be required for the assembly submission. If you are submitting multiple assemblies from the same biological sample, such as different assembly versions or haplotypes, they can share the same BioSample accession.

## Submitting the Assembly

### Using the Submission Portal

The NCBI Submission Portal at https://submit.ncbi.nlm.nih.gov/ provides a guided interface for assembly submission. The portal will ask for the BioProject accession, the BioSample accession, the assembly type, and the sequence files. The assembly type options include contig, scaffold, and complete genome. Choose the option that matches the highest level of assembly structure you have achieved.

The portal will also ask for an assembly name. The assembly name should be brief and descriptive, typically including the organism abbreviation and a version indicator. For example, the Tribolium castaneum assembly was named Tcas5.2 to indicate the species and the version. The assembly name will appear in the Assembly database and in GenBank records, so choose a name that is informative and unique within your organism.

### Genome Assembly Submission Options

The submission portal offers different paths depending on the type of assembly. For standard genome assemblies, the genome submission wizard guides you through the process. For viral assemblies, including SARS-CoV-2, NCBI has developed an automated pipeline that ensures the collection of contextual information about the virus source, assesses sequence quality, and annotates descriptive biological features such as protein-coding regions and mature peptides. This pipeline promotes standardized nomenclature and creates and publishes fully processed GenBank files within minutes of deposition. If you are submitting a viral assembly, check whether the automated pipeline applies to your virus and use that path if available.

For prokaryotic assemblies, the Prokaryotic Genome Annotation Pipeline (PGAP) is used to annotate nearly all RefSeq assemblies. PGAP is available as a stand-alone tool able to produce GenBank-ready files. If you plan to request RefSeq inclusion, you may choose to run PGAP locally to review the annotation before submission, or you can allow NCBI to run PGAP after submission. The choice depends on whether you want to review and potentially modify the annotation before it becomes public.

### File Upload and Validation

The submission portal accepts FASTA files for contig or scaffold assemblies and AGP files for scaffold assemblies that include gap information. The files must be compressed in a format accepted by the portal, typically gzip. The portal will run validation checks on the files and report any errors or warnings. Common errors include duplicate sequence identifiers, illegal characters in sequence headers, and inconsistent sequence lengths between FASTA and AGP files.

Validation warnings do not necessarily block submission, but they should be reviewed carefully. A warning about low sequence complexity or unusual GC content may indicate a problem with the assembly or contamination. Address warnings before proceeding if possible, or document why the warning is acceptable for your data.

## GenBank vs RefSeq Submission Paths

### GenBank Submission

GenBank accepts all valid genome assemblies that meet the basic requirements for sequence data and metadata. The GenBank path is appropriate for assemblies that are being deposited for public access and citation, regardless of whether they meet the higher quality standards required for RefSeq. GenBank records are accessible through the NCBI nucleotide database and are linked to the Assembly database.

The GenBank submission process includes an annotation step if you are submitting annotated records. For genome assemblies, annotation can be provided by the submitter or generated by NCBI pipelines. If you provide your own annotation, the annotation files must be in a format accepted by NCBI, typically a table file or a GenBank flat file. The annotation will be validated for consistency with the sequence and for correct feature placement.

### RefSeq Inclusion

RefSeq is the Reference Sequence collection at NCBI, which contains over 315,000 bacterial and archaeal genomes and 236 million proteins with up-to-date and consistent annotation. RefSeq inclusion is not automatic upon GenBank submission. The RefSeq project evaluates assemblies for quality, including contamination screening and completeness of annotation, prior to acceptance. Assemblies are now more stringently evaluated for contamination and for completeness of annotation prior to acceptance into RefSeq.

To request RefSeq inclusion, indicate this preference during the assembly submission process. The NCBI staff will evaluate the assembly against RefSeq quality standards. If the assembly meets the standards, it will be processed for RefSeq inclusion, which involves annotation by PGAP for prokaryotes or by the appropriate eukaryotic annotation pipeline. If the assembly does not meet the standards, it will remain in GenBank without RefSeq status.

The Tribolium castaneum example demonstrates the RefSeq path. The improved genome assembly and enhanced genome annotation were submitted to GenBank and accepted as a RefSeq genome by NCBI. The acceptance followed significant improvements in assembly quality and annotation completeness. For your assembly, assess whether it meets the quality standards expected for RefSeq before requesting inclusion. If your assembly has low contiguity, incomplete annotation, or evidence of contamination, focus on improving those aspects before requesting RefSeq processing.

## Required Metadata and File Preparation

### Organism and Taxonomy Metadata

The organism name in the assembly submission must match the organism name in the BioSample and must be a valid NCBI Taxonomy name. If the taxonomy name has changed or if your organism is newly described, verify the current taxonomy status before submission. The NCBI Taxonomy database is accessible through the NCBI website and provides the authoritative list of accepted names.

For strains or isolates, provide the strain or isolate identifier in the BioSample and in the assembly submission. The strain identifier should be consistent across all records. If the strain has a common name or a collection identifier, include that information in the BioSample metadata.

### Assembly Structure Metadata

The assembly submission requires information about the assembly structure, including whether the assembly is a set of contigs, a set of scaffolds, or a complete genome. For complete genomes, indicate whether the assembly includes one chromosome or multiple chromosomes and whether plasmids are included. For scaffold assemblies, provide the number of scaffolds and the total gap length if known.

The assembly level affects the processing path and the validation checks applied. Complete genomes receive additional validation to confirm that the sequence is circular for bacterial chromosomes and that the replication origin is correctly identified. Scaffold assemblies are checked for AGP consistency and gap size accuracy.

### Sequencing Technology and Coverage

Provide information about the sequencing technology used to generate the data. This includes the platform, the read length, the insert size for paired-end libraries, and the sequencing coverage. This information is used by NCBI and by users of the data to assess assembly quality and to plan downstream analyses. The Leishmania infantum study provides an example of the level of detail expected. The study used the Illumina HiSeq 2500 platform with the TruSeq Nano DNA Low Throughput Library Prep Kit and generated single-fragment reads of 2 x 150 bp paired-end with two fragment end-to-end assemblies. The resulting whole-genome sequence of 32,009,137 base pairs from 36 chromosomes was submitted to GenBank and registered under the name Leishmania infantum_TR01. Providing this level of detail in your submission helps reviewers and users understand the data.

## Common Failure Patterns and How to Avoid Them

### Metadata Inconsistencies

The most common cause of submission delays is inconsistent metadata across the BioProject, BioSample, and assembly submission. The organism name, strain identifier, and project title must match across all three records. Before starting the assembly submission, review the BioProject and BioSample records and confirm that the information is accurate and complete. If corrections are needed, make them before submitting the assembly.

### File Format Errors

FASTA files with duplicate sequence identifiers, illegal characters in headers, or inconsistent line wrapping will fail validation. AGP files with incorrect gap sizes or contig ordering will also fail. Run the NCBI validation tools on your files before starting the submission to identify and correct these errors. The validation tools are available through the submission portal and through the NCBI FTP site.

### Contamination

Contamination is a serious problem that can cause rejection or removal of an assembly. Screen your assembly thoroughly before submission. Common contamination sources include adapter sequences, vector sequences, and sequences from organisms that are common in the laboratory environment. The RefSeq project has increased its scrutiny of contamination, so assemblies with any evidence of contamination are unlikely to be accepted for RefSeq inclusion.

### Incomplete Annotation

For assemblies that include annotation, incomplete or inconsistent annotation will cause validation failures. The annotation must include all expected features, and the features must be consistent with the sequence. If you are using PGAP for prokaryotic annotation, review the PGAP output for completeness before submission. If you are providing your own annotation, use the NCBI annotation tools to validate the annotation files before submission.

## Records and Measurements to Maintain

### Assembly Quality Metrics

Maintain a record of assembly quality metrics for your assembly, including the number of contigs or scaffolds, the N50 and L50 values, the total sequence length, and the total gap length. These metrics are reported in the Assembly database and are used to evaluate the assembly. Keep the calculation methods and the versions of the assembly tools used to generate the metrics.

### Sequencing and Assembly Logs

Keep detailed logs of the sequencing and assembly process, including the versions of all software used, the parameters for each step, and the dates of each analysis. This information is important for reproducibility and for responding to questions from NCBI staff or from users of the data. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility, and the nf-core documentation describes community pipeline standards for reproducible workflow context. Following these standards for your own analysis will make the submission process smoother.

### Submission Correspondence

Maintain records of all correspondence with NCBI staff during the submission process. This includes the submission ticket number, the dates of communication, and the responses to any questions or requests for additional information. If the submission is delayed or rejected, the correspondence will help you understand the reason and take corrective action.

## Quality Controls and Validation

### NCBI Validation Tools

NCBI provides validation tools that check assembly files for common errors before submission. These tools check sequence identifiers, sequence length, base composition, and AGP consistency. Run these tools on your files before starting the submission and correct any errors they identify.

### Independent Quality Assessment

In addition to the NCBI validation tools, perform an independent quality assessment of your assembly. This includes checking for contamination, assessing completeness using expected gene content, and evaluating the assembly against the sequencing coverage and read quality. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can support these assessments. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help you develop the skills needed for thorough quality assessment.

### Completeness Assessment

For eukaryotic assemblies, assess completeness using a tool that checks for the presence of conserved single-copy genes expected in the organism's lineage. For prokaryotic assemblies, check for the presence of expected housekeeping genes and for the absence of genes from other organisms. The completeness assessment provides evidence that the assembly represents the full genome and not a partial or contaminated version.

## Professional Escalation Criteria

### When to Contact NCBI Support

Contact NCBI support if the submission portal reports errors that you cannot resolve, if the validation tools identify problems that you do not understand, or if the submission has been pending for an extended period without processing. The NCBI support team can provide guidance on the submission process and can help resolve technical issues.

### When to Seek Additional Sequencing

If your assembly fails quality checks due to low contiguity, excessive gaps, or evidence of missing sequence, consider whether additional sequencing is needed before resubmission. The Tribolium castaneum example shows that adding large-distance jumping library DNA sequencing can join scaffolds and fill small gaps, substantially improving assembly quality. If your assembly has similar limitations, additional sequencing may be more effective than repeated attempts to submit the current assembly.

### When to Revise the Assembly

If the assembly fails validation due to contamination or assembly errors, revise the assembly before resubmission. This may involve removing contaminated contigs, correcting assembly errors, or reassembling with different parameters. Do not submit the same assembly repeatedly without addressing the identified problems.

## Safety and Regulatory Context

### Data Release and Public Access

GenBank submissions become public upon processing, and the data are accessible to anyone through the NCBI website. Consider the implications of public data release before submission. If the data are subject to restrictions, such as those from material transfer agreements or institutional policies, resolve those issues before submission. The NCBI data resources are described at https://www.ncbi.nlm.nih.gov/, and the data release policies are documented on the NCBI website.

### Ethical and Legal Considerations

For human data or data from endangered species, additional ethical and legal considerations apply. Human data may require controlled access or specific consent documentation. Data from endangered species may be subject to international regulations. Consult with your institution's ethics board or legal counsel before submitting such data.

### Pathogen Data

For pathogen data, including SARS-CoV-2, rapid data release is important for public health. The NCBI automated pipeline for SARS-CoV-2 sequences ensures the collection of contextual information about the virus source, assesses sequence quality, and annotates descriptive biological features. The pipeline has processed and published hundreds of thousands of annotated SARS-CoV-2 sequences. If you are submitting pathogen data, use the appropriate submission path and provide complete contextual information to support public health activities.

## Decision Framework for Assembly Submission Path Selection

Choosing the correct submission path requires a structured evaluation of assembly quality, intended data use, and annotation readiness. Many researchers default to requesting RefSeq inclusion without first assessing whether their assembly meets the stricter quality thresholds, which leads to avoidable delays and repeated processing cycles. A practical decision framework helps you match your assembly characteristics to the appropriate submission route and prepares you for the questions NCBI staff will ask during review.

### Step 1: Assess Assembly Completeness and Contiguity

Before selecting a submission path, calculate and record your assembly statistics using the same metrics the Assembly database reports. These include the number of contigs and scaffolds, contig N50, total sequence length, and total gap length. The Assembly database tracks these statistics for every submitted assembly, and NCBI staff use them as an initial screen for RefSeq eligibility.

Compare your metrics against existing assemblies for the same or closely related organisms. The Assembly Resource at https://www.ncbi.nlm.nih.gov/assembly/ allows you to browse assemblies for a particular organism and view their reported statistics. If your assembly falls substantially below the median contiguity for your organism, the GenBank path is more appropriate as an initial submission. You can pursue RefSeq inclusion after improving the assembly with additional sequencing data.

The Tribolium castaneum example illustrates this principle. The original assembly generated by Sanger sequencing had limited contiguity and annotation quality. The research team added large-distance jumping library DNA sequencing to join scaffolds and fill small gaps, which reduced gaps and increased the N50 to 4753 kbp. Only after this improvement was the assembly submitted to GenBank and accepted as a RefSeq genome. For your own assembly, compare your N50 and gap statistics to published assemblies for your organism before requesting RefSeq processing.

### Step 2: Evaluate Contamination Status

The RefSeq project has become more stringent in evaluating assemblies for contamination and completeness of annotation prior to acceptance. This means contamination screening is not optional for RefSeq requests. Run contamination screening against databases of common laboratory contaminants, adapter sequences, and vector sequences. Document the screening method, the database version, and the parameters used.

If your screening identifies any contigs with significant matches to contaminants, remove them before submission. For the GenBank path, minor contamination may be acceptable if you document the screening process and the rationale for retaining specific sequences. For the RefSeq path, any evidence of contamination will likely result in rejection or a request for resubmission after cleanup. The decision framework should therefore include a contamination screening step before you commit to a submission path.

### Step 3: Determine Annotation Readiness

RefSeq inclusion requires complete and consistent annotation. For prokaryotic assemblies, the Prokaryotic Genome Annotation Pipeline (PGAP) is used to annotate nearly all RefSeq assemblies. PGAP is available as a stand-alone tool able to produce GenBank-ready files at https://github.com/ncbi/pgap. You have two options: run PGAP locally to review the annotation before submission, or allow NCBI to run PGAP after submission.

Running PGAP locally gives you the opportunity to review gene calls, check for missing features, and identify potential annotation errors before the assembly enters the RefSeq pipeline. This is particularly valuable if your organism has unusual genomic features or if you have specific knowledge about gene content that the automated pipeline might miss. The Tribolium castaneum project demonstrated the value of enhanced annotation, with the new official gene set including alternative splicing, well defined UTRs, and microRNA target predictions. For eukaryotic assemblies, the annotation requirements are more complex, and you should assess whether you have the transcriptomic evidence needed to support gene model prediction.

If your assembly lacks annotation or if your annotation is incomplete, the GenBank path allows submission without annotation. NCBI will run the appropriate annotation pipeline after submission, and you can request RefSeq inclusion at a later time once the annotation is complete and validated.

### Step 4: Consider the Intended Use of the Data

The intended use of your assembly should influence the submission path. If your goal is to make the sequence publicly available for the scientific community, to support a publication, or to comply with funding agency requirements, the GenBank path is sufficient. GenBank records are accessible through the NCBI nucleotide database and are linked to the Assembly database, providing stable accessioning and data tracking.

If your goal is to establish a reference genome for your organism, to support comparative genomics studies, or to provide a community resource, RefSeq inclusion adds value through curated annotation and consistent naming. The RefSeq collection at NCBI contains over 315,000 bacterial and archaeal genomes and 236 million proteins with up-to-date and consistent annotation. RefSeq assemblies receive ongoing updates and are integrated into NCBI resources such as the genome browser and comparative genomics tools.

The decision framework should also account for the likelihood of future updates. If you anticipate submitting improved versions of the assembly as new sequencing data become available, the GenBank path allows you to deposit each version with an updated assembly accession. The Assembly database tracks changes to updated genome assemblies and provides an assembly accession and version to unambiguously identify the set of sequences that make up a particular version. You can pursue RefSeq inclusion once the assembly reaches a stable version that you do not expect to change substantially.

### Step 5: Apply the Decision Matrix

Use the following decision matrix to select your submission path based on your assembly characteristics.

| Assembly Characteristic | GenBank Path | RefSeq Path |
| --- | --- | --- |
| Contiguity below organism median | Appropriate | Not recommended until improved |
| Contiguity at or above organism median | Appropriate | Appropriate |
| Contamination detected | Appropriate with documentation | Not recommended until cleaned |
| No contamination detected | Appropriate | Appropriate |
| Annotation incomplete or absent | Appropriate | Not recommended until annotated |
| Annotation complete and validated | Appropriate | Appropriate |
| Expected to update with new data | Appropriate | Defer until stable version |
| Stable version for community use | Appropriate | Appropriate |

This matrix provides a clear decision rule: the RefSeq path requires contiguity at or above the organism median, no detectable contamination, complete annotation, and a stable assembly version. If any of these conditions are not met, the GenBank path is the appropriate choice, and you can pursue RefSeq inclusion after addressing the deficiencies.

### Step 6: Document the Decision

Record the rationale for your submission path choice in your laboratory notebook or project documentation. Include the assembly quality metrics, the contamination screening results, the annotation status, and the intended use of the data. This documentation will be useful if NCBI staff ask questions during the submission process, and it will help you or your colleagues make decisions about future updates to the assembly.

The documentation should also include the versions of all software used for assembly, quality assessment, and contamination screening. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility, and the nf-core documentation describes community pipeline standards for reproducible workflow context. Following these standards for your own analysis will make the submission process smoother and will support the documentation requirements.

## Record System for Submission Tracking

Maintaining a structured record system throughout the submission process reduces the risk of errors and provides a clear audit trail if problems arise. The following record system is designed to track the key information needed for a successful submission and for responding to NCBI staff inquiries.

### Submission Tracking Table

Create a table with the following columns and update it at each step of the submission process.

| Date | Step Completed | Accession or Identifier | Notes | Staff Contact |
| --- | --- | --- | --- | --- |
| 2025-01-15 | BioProject registered | PRJNA123456 | Project title confirmed | Not applicable |
| 2025-01-16 | BioSample registered | SAMN12345678 | Organism name verified | Not applicable |
| 2025-01-20 | Assembly files prepared | Not applicable | FASTA and AGP validated | Not applicable |
| 2025-01-22 | Assembly submitted | GCA_123456789.1 | Submission ticket opened | J. Smith |
| 2025-01-25 | Validation warning received | GCA_123456789.1 | Low complexity region flagged | J. Smith |
| 2025-01-28 | Warning addressed | GCA_123456789.1 | Documentation provided | J. Smith |
| 2025-02-01 | Assembly published | GCA_123456789.1 | GenBank path completed | Not applicable |

This table provides a chronological record of the submission process, including the dates, the steps completed, the accessions assigned, and any correspondence with NCBI staff. The notes column captures important details such as validation warnings and the actions taken to address them.

### Metadata Consistency Checklist

Before submitting the assembly, verify the following metadata elements are consistent across the BioProject, BioSample, and assembly submission.

| Metadata Element | BioProject | BioSample | Assembly Submission | Status |
| --- | --- | --- | --- | --- |
| Organism scientific name | Not applicable | Tribolium castaneum | Tribolium castaneum | Match |
| Strain or isolate identifier | Not applicable | Tcas5.2 | Tcas5.2 | Match |
| Project title | Genome assembly of Tcas5.2 | Not applicable | Not applicable | Consistent |
| BioProject accession | PRJNA123456 | PRJNA123456 | PRJNA123456 | Match |
| BioSample accession | Not applicable | SAMN12345678 | SAMN12345678 | Match |

Complete this checklist before starting the assembly submission. Any mismatch should be corrected before proceeding, as metadata inconsistencies are the most common cause of submission delays.

### File Preparation Log

Maintain a log of the file preparation steps, including the commands used, the software versions, and the output file names. This log supports reproducibility and helps you respond to questions about the assembly process.

| Step | Software and Version | Command or Parameter | Output File | Date |
| --- | --- | --- | --- | --- |
| Contamination screening | BBDuk 38.84 | bbduk.sh in=assembly.fa out=clean.fa ref=contaminants.fa | clean.fa | 2025-01-18 |
| AGP file generation | Custom script v1.2 | python make_agp.py clean.fa scaffolds.txt | assembly.agp | 2025-01-19 |
| FASTA validation | NCBI validation tool | validate_fasta clean.fa | validation_report.txt | 2025-01-20 |

The file preparation log should include enough detail that another researcher could reproduce the exact steps. The Carpentries Lessons provide foundational computing and data training that supports this level of documentation practice.

## Troubleshooting Method for Validation Failures

When the submission portal reports validation errors or warnings, use a systematic troubleshooting method to identify and resolve the underlying cause. The following method is designed to address the most common failure patterns without requiring repeated submission attempts.

### Step 1: Categorize the Error

Read the validation report carefully and categorize each error or warning into one of the following types.

| Error Type | Description | Example |
| --- | --- | --- |
| Format error | File structure does not meet specification | Duplicate sequence identifiers in FASTA |
| Metadata error | Information does not match across records | Organism name differs between BioSample and assembly |
| Sequence error | Sequence content fails quality checks | Contamination detected in a contig |
| Annotation error | Feature placement or content is inconsistent | Gene model extends beyond contig boundary |

Categorizing the error helps you identify the appropriate corrective action and prevents you from addressing the wrong aspect of the submission.

### Step 2: Identify the Root Cause

For each error, trace the root cause back to the source. A format error may result from incorrect file preparation, such as using a text editor that introduced special characters into the FASTA headers. A metadata error may result from a typo in the organism name or from using an outdated taxonomy name. A sequence error may result from contamination introduced during library preparation or from misassembly of repetitive regions.

The root cause analysis should consider the entire workflow from sequencing through assembly to file preparation. The EMBL-EBI Training program offers bioinformatics learning pathways and data-resource training that can help you develop the skills needed for thorough root cause analysis.

### Step 3: Apply the Corrective Action

Apply the corrective action specific to the error type.

| Error Type | Corrective Action |
| --- | --- |
| Format error | Regenerate the file using the NCBI validation tools to identify and correct the specific formatting issue |
| Metadata error | Correct the metadata in the BioProject or BioSample record before resubmitting the assembly |
| Sequence error | Remove contaminated contigs or reassemble with different parameters to resolve the sequence issue |
| Annotation error | Regenerate the annotation using the appropriate pipeline or correct the feature placement manually |

Do not resubmit the same files without addressing the identified errors. Repeated submission of unchanged files will result in the same validation failures and will delay the processing of your submission.

### Step 4: Document the Resolution

Record the error, the root cause, and the corrective action in your submission tracking table. This documentation provides a record of the troubleshooting process and helps you respond to questions from NCBI staff. It also serves as a reference for future submissions, allowing you to avoid the same errors.

### Step 5: Escalate When Necessary

If the validation error persists after corrective action, or if you cannot identify the root cause, escalate the issue to NCBI support. Provide the submission ticket number, the validation report, and a description of the corrective actions you have taken. The NCBI support team can provide guidance on the submission process and can help resolve technical issues that are not apparent from the validation report.

## Common Failure Patterns in Submission Path Selection

### Pattern 1: Requesting RefSeq Inclusion Prematurely

Researchers often request RefSeq inclusion for assemblies that do not meet the quality standards. This results in a longer processing time and a rejection notice that could have been avoided by choosing the GenBank path initially. The RefSeq project evaluates assemblies for contamination and completeness of annotation prior to acceptance, and assemblies that do not meet these standards remain in GenBank without RefSeq status.

The corrective action is to apply the decision framework before submission. Assess contiguity, contamination status, annotation readiness, and intended use before selecting the submission path. If the assembly does not meet RefSeq standards, submit through the GenBank path and pursue RefSeq inclusion after improving the assembly.

### Pattern 2: Submitting Without Contamination Screening

Assemblies submitted without contamination screening are likely to fail validation or to be rejected during RefSeq review. The RefSeq project has increased its scrutiny of contamination, and assemblies with any evidence of contamination are unlikely to be accepted. The corrective action is to screen the assembly against contaminant databases before submission and to document the screening process.

### Pattern 3: Inconsistent Metadata Across Records

Metadata inconsistencies between the BioProject, BioSample, and assembly submission are the most common cause of submission delays. The organism name, strain identifier, and project title must match across all three records. The corrective action is to complete the metadata consistency checklist before starting the assembly submission and to correct any mismatches before proceeding.

### Pattern 4: Ignoring Validation Warnings

Validation warnings do not necessarily block submission, but they should be reviewed carefully. A warning about low sequence complexity or unusual GC content may indicate a problem with the assembly or contamination. The corrective action is to investigate each warning and to document why the warning is acceptable for your data or to address the underlying issue before submission.

## Welfare and Safety Context for Data Submission

### Data Integrity and Reproducibility

The submission of accurate and complete assembly data supports the integrity of the scientific record and enables reproducibility of genomic analyses. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that supports rigorous data analysis practices. The nf-core documentation describes community pipeline standards for reproducible workflow context. Following these standards for your own analysis ensures that the data you submit are reliable and that other researchers can reproduce your results.

### Responsible Data Release

GenBank submissions become public upon processing, and the data are accessible to anyone through the NCBI website. Consider the implications of public data release before submission. If the data are subject to restrictions, such as those from material transfer agreements or institutional policies, resolve those issues before submission. The NCBI data resources are described at https://www.ncbi.nlm.nih.gov/, and the data release policies are documented on the NCBI website.

### Pathogen Data Considerations

For pathogen data, including SARS-CoV-2, rapid data release is important for public health. The NCBI automated pipeline for SARS-CoV-2 sequences ensures the collection of contextual information about the virus source, assesses sequence quality, and annotates descriptive biological features. The pipeline has processed and published hundreds of thousands of annotated SARS-CoV-2 sequences. If you are submitting pathogen data, use the appropriate submission path and provide complete contextual information to support public health activities.

### Ethical and Legal Considerations

For human data or data from endangered species, additional ethical and legal considerations apply. Human data may require controlled access or specific consent documentation. Data from endangered species may be subject to international regulations. Consult with your institution's ethics board or legal counsel before submitting such data. The decision framework should include a review of these considerations before you select a submission path.

## Professional Escalation Criteria for Submission Path Decisions

### When to Contact NCBI Support

Contact NCBI support if you are uncertain about which submission path is appropriate for your assembly, if the submission portal reports errors that you cannot resolve, or if the submission has been pending for an extended period without processing. The NCBI support team can provide guidance on the submission process and can help resolve technical issues.

### When to Seek Additional Sequencing

If your assembly fails quality checks due to low contiguity, excessive gaps, or evidence of missing sequence, consider whether additional sequencing is needed before resubmission. The Tribolium castaneum example shows that adding large-distance jumping library DNA sequencing can join scaffolds and fill small gaps, substantially improving assembly quality. If your assembly has similar limitations, additional sequencing may be more effective than repeated attempts to submit the current assembly.

### When to Revise the Assembly

If the assembly fails validation due to contamination or assembly errors, revise the assembly before resubmission. This may involve removing contaminated contigs, correcting assembly errors, or reassembling with different parameters. Do not submit the same assembly repeatedly without addressing the identified problems.

### When to Defer RefSeq Inclusion

If your assembly does not meet the quality standards for RefSeq inclusion, defer the RefSeq request until the assembly is improved. The GenBank path allows you to deposit the assembly and make it publicly available while you work on improvements. Once the assembly reaches the quality standards, you can request RefSeq inclusion as an update to the existing GenBank record.

## Frequently Asked Questions

### What is the difference between GenBank and RefSeq submission?

GenBank submission deposits your assembly into the public nucleotide database with your provided or pipeline-generated annotation. RefSeq inclusion is a separate process where NCBI evaluates the assembly for quality and completeness and, if accepted, provides a curated reference annotation. RefSeq assemblies receive additional quality checks for contamination and annotation completeness. The Tribolium castaneum assembly was submitted to GenBank and then accepted as a RefSeq genome after quality improvements.

### How long does the NCBI assembly submission process take?

The processing time varies depending on the assembly type, the completeness of the metadata, and the current workload at NCBI. Viral assemblies submitted through the automated pipeline can be processed within minutes. Bacterial and eukaryotic assemblies typically take longer, ranging from days to weeks. The automated SARS-CoV-2 pipeline creates and publishes fully processed GenBank files within minutes of deposition, but this speed is specific to that pipeline.

### Can I update my assembly after submission?

Yes, you can submit an updated version of your assembly. The Assembly database tracks changes to updated genome assemblies and provides an assembly accession and version to unambiguously identify the set of sequences that make up a particular version. The update history is reported in the Assembly database. When submitting an update, provide the previous assembly accession and describe the changes made.

### What should I do if my assembly is rejected?

If your assembly is rejected, review the rejection reason provided by NCBI. Common reasons include contamination, metadata inconsistencies, and file format errors. Address the identified problems and resubmit. If the rejection reason is not clear, contact NCBI support for clarification. Do not resubmit the same assembly without addressing the identified issues.

### Do I need to annotate my assembly before submission?

Annotation is not required for GenBank submission, but it is required for RefSeq inclusion. If you do not provide annotation, NCBI will run the appropriate annotation pipeline after submission. For prokaryotes, PGAP is used to annotate nearly all RefSeq assemblies and is available as a stand-alone tool able to produce GenBank-ready files. You can run PGAP locally to review the annotation before submission or allow NCBI to run it after submission.

### What information do I need to provide about the sequencing technology?

Provide the sequencing platform, read length, library type, insert size, and sequencing coverage. This information is used to assess assembly quality and to help users of the data understand the assembly. The Leishmania infantum study provides an example of the level of detail expected, including the platform, library preparation kit, read length, and assembly approach.

### How do I check if my organism has a valid NCBI Taxonomy name?

Use the NCBI Taxonomy database to search for your organism name. The taxonomy database provides the authoritative list of accepted names and the taxonomic hierarchy. If your organism name is not found, you may need to request a new taxonomy ID before submission. The organism name in the assembly submission must match the organism name in the BioSample.

### What are the most common reasons for submission delays?

The most common reasons for delays are inconsistent metadata across the BioProject, BioSample, and assembly submission, file format errors, and contamination. Review all metadata for consistency before submission, run the NCBI validation tools on your files, and screen your assembly for contamination. Addressing these issues before submission will reduce the likelihood of delays.

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Spatial Transcriptomics Workflow: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/spatial-transcriptomics-workflow-from-sample-preparation-to-data-analysis)
- [Single-Cell Sequencing Workflow: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/single-cell-sequencing-workflow-from-sample-preparation-to-data-analysis)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Assembly: a resource for assembled genomes at NCBI.](https://pubmed.ncbi.nlm.nih.gov/26578580). Nucleic acids research, 2016.
- [Enhanced genome assembly and a new official gene set for Tribolium castaneum.](https://pubmed.ncbi.nlm.nih.gov/31937263). BMC genomics, 2020.
- [RefSeq and the prokaryotic genome annotation pipeline in the age of metagenomes.](https://pubmed.ncbi.nlm.nih.gov/37962425). Nucleic acids research, 2024.
- [Genome Sequencing of Leishmania infantum Causing Cutaneous Leishmaniosis from a Turkish Isolate with Next-Generation Sequencing Technology.](https://pubmed.ncbi.nlm.nih.gov/32691361). Acta parasitologica, 2021.
- [Rapid automated validation, annotation and publication of SARS-CoV-2 sequences to GenBank.](https://pubmed.ncbi.nlm.nih.gov/35230423). Database : the journal of biological databases and curation, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.