How to Submit Metagenomic Data to Public Repositories: SRA, ENA, and DDBJ Best Practices

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Submit Metagenomic Data to Public Repositories: SRA, ENA, and DDBJ Best Practices

Key Takeaways

  • Raw FASTQ files are paramount: Deposit raw sequencing reads (FASTQ format) directly from the instrument, before any quality filtering, host removal, or assembly, to enable reanalysis with different methods and ensure long-term data value.
  • Metadata completeness is critical for interpretability: Comprehensive metadata, including host species, body site/environment, collection details, sequencing platform, and library preparation, is essential for secondary users to understand and verify findings.
  • Adherence to FAIR principles is a best practice: Ensure data is Findable (standardized fields, naming conventions), Accessible (public repositories), Interoperable (standard formats), and Reusable (complete metadata, documented analysis choices).
  • Proactive quality control prevents submission failures: Conduct thorough quality control for adapter contamination, low-quality bases, duplicate reads, and cross-contamination before submission to avoid validation issues and ensure data integrity.
  • Systematic auditing is essential for accuracy: Implement a pre-submission audit covering file inventory verification (with checksums), metadata completeness assessment, file-metadata cross-checks, and technical format verification to prevent common submission errors.
  • Choose repository based on practicalities: Select SRA, ENA, or DDBJ based on geographic location, funding requirements, journal policies, and familiarity with submission portals, recognizing that all INSDC repositories synchronize data.

Metagenomic data submission to public repositories is a mandatory step for publishing microbiome research and satisfying funding agency requirements. This article provides a practical workflow for depositing raw sequencing data and associated metadata to the Sequence Read Archive (SRA), the European Nucleotide Archive (ENA), and the DNA Data Bank of Japan (DDBJ). The guidance covers file formats, metadata requirements, quality control before submission, and common pitfalls specific to metagenomic projects. Researchers who follow these procedures will produce submissions that pass repository validation, support replication by other groups, and remain interpretable years after the original study concludes.

The Role of Public Repositories in Metagenomic Research

Public nucleotide sequence repositories serve as the primary infrastructure for sharing raw sequencing data across the global research community. The National Center for Biotechnology Information (NCBI) operates the Sequence Read Archive alongside its other sequence resources, search systems, and analysis services that researchers use daily [<a href="#ref-1">1</a>]. The European Bioinformatics Institute provides complementary training pathways and data resources that help researchers understand submission requirements and analysis options [<a href="#ref-2">2</a>]. Together with DDBJ, these three repositories form the International Nucleotide Sequence Database Collaboration, which synchronizes data across their systems.

For metagenomic projects specifically, the repository record becomes the permanent link between a published paper and the underlying evidence. When a study reports taxonomic profiles, functional annotations, or assembled genomes, reviewers and readers need access to the raw reads to verify those claims. A submission that lacks complete metadata or omits sequencing runs makes independent verification difficult or impossible. The practical consequence is that incomplete archiving reduces the long-term value of the research investment.

The scale of metagenomic data in public repositories is substantial. One community resource cataloging bee-associated microbiomes found over 33,000 Sequence Read Archive experiments spanning 278 host species, demonstrating both the volume of deposited data and the difficulty of locating relevant datasets within large generalist repositories [<a href="#ref-3">3</a>]. This example illustrates why careful metadata submission matters: without accurate descriptions, even deposited data becomes effectively invisible to the researchers who need it.

Core Principles of Metagenomic Data Submission

Raw Data Takes Priority Over Processed Results

The most important principle in data archiving is that raw sequencing reads must be deposited, beyond the results derived from them. A survey of ancient genomics studies found that half of the studies archived incomplete datasets, preventing accurate replication and representing a loss of data with potential future use [<a href="#ref-4">4</a>]. While ancient DNA research has particular constraints due to destructive sampling of finite material, the underlying principle applies equally to metagenomics: the raw reads are the primary evidence, and downstream analyses can always be rerun if the raw data exists.

For metagenomic projects, this means depositing the FASTQ files that come directly from the sequencing instrument, before any quality filtering, host removal, or assembly steps. Some researchers mistakenly deposit only the reads that aligned to a reference genome or only the assembled contigs. This practice prevents other groups from applying different analysis methods or re-examining the data with improved tools. The recommendation from the ancient genomics survey applies directly: archive all sequencing reads, beyond those that aligned to a reference genome [<a href="#ref-4">4</a>].

Metadata Completeness Determines Data Usability

Metadata is the descriptive information that makes raw sequence data interpretable. Without proper metadata, a FASTQ file is just a collection of nucleotide sequences with no context. The repository requirements for metadata include sample identifiers, organism or environment descriptions, sequencing platform, library preparation method, and experimental design details.

The ancient genomics survey identified two specific metadata failures: incorrect experiment metadata on samples, libraries, and sequencing runs, and uninformative sample metadata [<a href="#ref-4">4</a>]. Both problems appear in metagenomic submissions as well. A sample described only as "gut" without host species, body site, health status, or collection method provides limited value to secondary users. Conversely, metadata that misidentifies the sequencing platform or library preparation method can lead to incorrect downstream analysis choices.

FAIR Principles Guide Submission Practice

The FAIR data principles, which emphasize findability, accessibility, interoperability, and reusability, provide a useful framework for metagenomic data submission. Community resources such as the BlastoDB database for Blastocystis research explicitly describe themselves as FAIR-aligned hubs that integrate epidemiological data, microbiome profiles, multi-omics datasets, reference sequences, protocols, and related metadata [<a href="#ref-5">5</a>]. While BlastoDB is a specialized community resource instead of a general repository, its approach illustrates the standard that researchers should apply to their own submissions.

For general repository submission, FAIR principles translate into practical actions: use standardized metadata fields, follow repository naming conventions, provide persistent identifiers, and document all analysis choices in the associated publication. The BeeBiome data portal demonstrates the value of this approach by making thousands of bee microbiome experiments findable through a single interface with filtering on relevant criteria [<a href="#ref-3">3</a>]. Individual researchers who submit complete, well-described data contribute to the collective value of these resources.

At a Glance: Repository Comparison for Metagenomic Submission

RepositoryPrimary Geographic FocusSubmission PortalKey Strength for MetagenomicsData Synchronization
Sequence Read Archive (SRA)United States and internationalNCBI Submission PortalIntegration with NCBI analysis tools and reference databases [<a href="#ref-1">1</a>]Yes, via INSDC
European Nucleotide Archive (ENA)Europe and internationalENA Submission PortalStrong metadata validation and training resources [<a href="#ref-2">2</a>]Yes, via INSDC
DNA Data Bank of Japan (DDBJ)Japan and Asia-PacificDDBJ Submission PortalEfficient submission for researchers in the regionYes, via INSDC

All three repositories accept metagenomic data and synchronize submissions through the International Nucleotide Sequence Database Collaboration. Researchers should choose the repository that best fits their geographic location, funding requirements, and existing account infrastructure. Journal requirements may specify a particular repository, so authors should check target journal policies before beginning the submission process.

Preparing Metagenomic Data for Submission

File Format Requirements

The standard file format for raw sequencing data submission is FASTQ, which contains both nucleotide sequences and quality scores. Most sequencing facilities deliver data in this format, often compressed with gzip to reduce file size. Repository submission systems accept compressed FASTQ files and will validate that the files are properly formatted.

Before submission, researchers should verify that their FASTQ files meet the following criteria:

  • Paired-end reads are in separate files for forward and reverse reads
  • File names are descriptive and match the sample identifiers used in metadata
  • Quality scores are in the correct encoding format for the sequencing platform
  • Files are not truncated or corrupted during transfer

For metagenomic projects that include multiple sequencing runs per sample, each run should be submitted as a separate experiment with its own file set. This allows secondary users to understand the sequencing depth and technical replication structure of the study.

Metadata Organization

The metadata submission process requires researchers to organize information at multiple levels. The sample level includes biological information about the source material. The experiment level describes the sequencing library preparation. The run level identifies the specific sequencing output files.

For metagenomic samples, the following metadata fields are particularly important:

  • Host species and strain or breed
  • Body site or environmental source
  • Collection date and geographic location
  • Health status or disease condition
  • Sample processing method
  • DNA extraction protocol
  • Sequencing platform and instrument model
  • Library preparation kit and protocol
  • Read length and insert size
  • Sequencing depth or coverage target

The ancient genomics survey specifically recommended providing informative sample metadata and correct experiment metadata on samples, libraries, and sequencing runs [<a href="#ref-4">4</a>]. These recommendations apply directly to metagenomic submissions, where the biological interpretation depends heavily on understanding the sample context.

Quality Control Before Submission

Quality control should occur before submission, not after deposition. Researchers should assess their sequencing data for common issues that could affect downstream analysis or repository validation:

  • Adapter contamination that should be removed before submission
  • Low-quality bases that may fail repository quality checks
  • Duplicate reads that inflate apparent sequencing depth
  • Cross-contamination between samples that could be detected in metadata review

The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover quality control procedures for sequencing data [<a href="#ref-6">6</a>]. Researchers who are uncertain about their quality control steps can follow these established workflows to prepare their data properly.

Step-by-Step Submission Workflow

Step 1: Create Repository Accounts and Understand Requirements

Before beginning a submission, researchers need accounts at their chosen repository. The NCBI provides official descriptions of its databases, search systems, and sequence resources on its website [<a href="#ref-1">1</a>]. The EMBL-EBI training portal offers learning pathways that explain data-resource requirements and practical analysis education [<a href="#ref-2">2</a>]. Researchers should review the specific submission guides for their chosen repository, as requirements can change over time.

Step 2: Organize Sample Information

Create a spreadsheet or table that lists every sample in the project with all relevant metadata. This table will serve as the basis for the repository submission. Each sample should have a unique identifier that matches the file names of the sequencing data. The metadata should include both the biological information described above and any technical information about how the sample was processed.

For metagenomic projects with many samples, this organization step is critical. A well-organized sample table makes the submission process faster and reduces the risk of metadata errors. The Carpentries lessons provide foundational training in data organization and computing skills that are directly applicable to managing large sample tables [<a href="#ref-7">7</a>].

Step 3: Prepare Sequencing Files

Ensure that all FASTQ files are properly named, compressed, and organized. The file names should include the sample identifier and read pair information. For example, a paired-end sample might have files named "Sample01_R1.fastq.gz" and "Sample01_R2.fastq.gz". Verify that the files are not corrupted by checking their integrity before upload.

Step 4: Complete the Submission Form

The repository submission portal will guide researchers through a series of forms that collect metadata at the study, sample, experiment, and run levels. The process typically includes:

  • Study information: title, abstract, funding source, publication status
  • Sample information: biological source, collection details, processing methods
  • Experiment information: library preparation, sequencing platform
  • Run information: file names, file formats, file sizes

The submission system will validate the metadata and flag any missing or inconsistent fields. Researchers should address all validation warnings before finalizing the submission.

Step 5: Upload Files and Submit

File upload can be done through the repository web interface or through command-line tools for large datasets. Metagenomic projects often generate substantial data volumes, so researchers should plan for adequate upload time and bandwidth. After upload, the repository will perform final validation and assign accession numbers.

Step 6: Verify Accession Numbers and Update Records

After submission, the repository assigns accession numbers that should be included in the associated publication. Researchers should verify that all expected accessions are present and that the data is publicly accessible. If the study is under embargo until publication, researchers should confirm the embargo release date and ensure it aligns with the publication timeline.

Options and Tradeoffs in Repository Selection

Geographic and Funding Considerations

The choice of repository often depends on the researcher's location and funding source. Many funding agencies require data deposition in a specific repository or in a repository that meets certain standards. Researchers should check their funding agreements and institutional policies before choosing a repository.

The three INSDC repositories synchronize data, so a submission to any one repository will eventually appear in the others. However, the submission process and interface differ, and researchers may find one system more intuitive than another. The EMBL-EBI training resources can help researchers understand the ENA submission process specifically [<a href="#ref-2">2</a>].

Specialized Repositories and Community Resources

Some research communities maintain specialized databases that complement the general repositories. The BlastoDB resource for Blastocystis research integrates epidemiological data, microbiome profiles, multi-omics datasets, reference sequences, protocols, and related metadata [<a href="#ref-5">5</a>]. The BeeBiome data portal provides access to bee microbiome data from thousands of SRA experiments [<a href="#ref-3">3</a>]. These specialized resources do not replace the general repositories but add value by making data more discoverable and providing community-specific context.

Researchers working in areas with established community resources should consider whether their data should also be submitted to these databases. The general repository submission remains the primary requirement, but specialized resources can increase the visibility and impact of the data.

Data Volume and Upload Considerations

Metagenomic projects can generate terabytes of raw sequencing data. The upload process can take significant time depending on internet bandwidth and repository infrastructure. Researchers should plan for this time and consider using command-line upload tools that support parallel transfers and resume capabilities.

For very large projects, researchers may want to contact the repository help desk before submission to discuss the best upload strategy. Repository staff can provide guidance on file organization and transfer methods that will work efficiently for large datasets.

Records and Measurements for Submission Quality

Tracking Submission Status

Researchers should maintain records of their submission progress, including:

  • The date each sample was submitted
  • Accession numbers assigned to each sample
  • Validation warnings and how they were resolved
  • Embargo release dates
  • Publication status of associated papers

This tracking is important for responding to journal queries and for ensuring that data becomes publicly available at the appropriate time. The ancient genomics survey found that no studies met all criteria that could be considered best practice for data archiving [<a href="#ref-4">4</a>], which suggests that systematic tracking can help researchers improve their submission completeness.

Measuring Metadata Completeness

Before finalizing a submission, researchers should review their metadata against a checklist of required and recommended fields. The repository submission system will enforce required fields, but recommended fields may be optional. Researchers should aim to complete all recommended fields that are relevant to their study.

A practical approach is to create a metadata completeness score for each sample, counting the proportion of relevant fields that are filled. This score can be tracked across samples to identify any that are missing important information.

Documenting Analysis Choices

The ancient genomics survey recommended documenting archiving choices in papers and having these choices peer reviewed [<a href="#ref-4">4</a>]. For metagenomic studies, this documentation should include:

  • Which files were deposited and which were not
  • Any filtering or processing applied before submission
  • The software and parameters used for quality control
  • The version of the submission that corresponds to the published analysis

This documentation helps readers understand exactly what data supports the published findings and how to reproduce the analysis.

Common Failure Patterns in Metagenomic Submission

Incomplete Read Archiving

The most common failure pattern is depositing only a subset of the sequencing reads. Some researchers submit only the reads that passed quality filtering or only the reads that aligned to a reference database. This practice prevents reanalysis with different methods and reduces the long-term value of the data. The ancient genomics survey specifically identified this as a common problem, recommending that all sequencing reads be archived, beyond those that aligned to a reference genome [<a href="#ref-4">4</a>].

For metagenomic projects, the solution is to deposit the raw FASTQ files exactly as they came from the sequencing facility. Any quality filtering or host removal should be documented in the methods section of the paper but should not be applied to the deposited files.

Missing or Incorrect Experiment Metadata

Another common failure is providing incorrect experiment metadata on samples, libraries, and sequencing runs [<a href="#ref-4">4</a>]. This can happen when researchers copy metadata from a previous submission or when the same sample is sequenced in multiple batches with different library preparation methods.

The solution is to verify that each experiment record accurately describes the library preparation and sequencing run that produced the corresponding files. If a sample was sequenced on two different platforms, each platform should have its own experiment record with the correct platform information.

Uninformative Sample Metadata

Sample metadata that lacks biological context makes data interpretation difficult. A sample described only as "stool" without host species, health status, or collection method provides limited value. The ancient genomics survey recommended providing informative sample metadata [<a href="#ref-4">4</a>], and this applies equally to metagenomic studies.

Researchers should include all relevant biological information in the sample metadata, even if it seems obvious. What is obvious to the researcher who collected the sample may be completely unknown to a secondary user who encounters the data years later.

Delayed or Missing Data Release

Some researchers submit their data but set embargo periods that extend well beyond the publication date. This practice delays access to data that should be publicly available once the associated paper is published. The solution is to set the embargo release date to match the expected publication date and to update the release date if the publication is delayed.

Quality Controls and Validation

Repository Validation Checks

All three repositories perform automated validation of submitted data. These checks verify that files are properly formatted, that metadata is complete, and that the submission is internally consistent. Researchers should address all validation warnings before finalizing their submission.

Common validation issues include:

  • FASTQ files with inconsistent read lengths
  • Quality score encoding that does not match the declared platform
  • Sample identifiers that do not match the file names
  • Missing required metadata fields

Independent Verification

After submission, researchers should independently verify that their data is accessible and correctly described. This verification can include:

  • Downloading a small sample of the submitted files to confirm they are intact
  • Checking that the accession numbers resolve to the correct records
  • Confirming that the metadata appears as intended in the public record

The Galaxy Training Network provides tutorials that cover data retrieval and verification workflows [<a href="#ref-6">6</a>]. Researchers can use these workflows to confirm that their deposited data is accessible and usable.

Cross-Repository Consistency

Because the INSDC repositories synchronize data, researchers should verify that their submission appears correctly in all three repositories after synchronization. This verification is particularly important for researchers who submit to one repository but have collaborators or readers who access data through another.

Limitations and Interpretation Boundaries

What Repository Submission Does Not Guarantee

Depositing data in a public repository ensures that the raw reads are available, but it does not guarantee that the data is biologically meaningful or that the analysis is correct. Repository validation checks file format and metadata completeness, not biological validity. Researchers should be clear about this distinction when interpreting data from public repositories.

Data Quality Variability Across Studies

Metagenomic datasets in public repositories vary widely in quality. Some studies have extensive metadata and well-documented methods, while others have minimal information. The BeeBiome data portal was created in part because datasets are difficult to find within large generalist repositories and are often not readily accessible [<a href="#ref-3">3</a>]. This variability means that secondary users must carefully evaluate the metadata and methods of any dataset they plan to reuse.

The Role of Community Curation

Specialized community resources like BlastoDB and BeeBiome add value by curating and organizing data from general repositories [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>]. These resources can help researchers find relevant datasets and understand their context. However, the curation depends on the quality of the original submissions. Researchers who submit complete, well-described data contribute to the collective value of these community resources.

Safety and Regulatory Context

Data Privacy and Human Subject Considerations

Metagenomic studies involving human samples raise privacy concerns that must be addressed before data submission. Researchers should ensure that their study has appropriate ethical approval and that the data submission complies with relevant privacy regulations. Some repositories offer controlled access options for human data, but these options vary by repository and data type.

For human metagenomic studies, researchers should consider whether the data can be de-identified and whether the repository's data access policies are appropriate for the sensitivity of the data. The NCBI provides guidance on data submission for human studies [<a href="#ref-1">1</a>], and researchers should review this guidance before beginning the submission process.

Compliance with Funding and Journal Requirements

Many funding agencies and journals have specific data deposition requirements. Researchers should review these requirements early in the project to ensure that their submission will comply. Common requirements include:

  • Deposition in a specific repository
  • Data availability before publication
  • Inclusion of accession numbers in the manuscript
  • Compliance with data sharing timelines

The EMBL-EBI training resources can help researchers understand the data management requirements associated with European funding [<a href="#ref-2">2</a>]. Researchers should also check their institutional data management policies.

Professional Escalation Criteria

When to Contact Repository Help Desks

Researchers should contact the repository help desk when they encounter issues that they cannot resolve through the documentation or submission interface. Common situations that warrant help desk contact include:

  • Upload failures that persist after troubleshooting
  • Validation errors that are not explained by the error message
  • Questions about metadata fields that are unclear
  • Requests for large data transfers that exceed standard limits

Repository help desks are staffed by experienced professionals who can provide guidance on complex submissions. The NCBI and EMBL-EBI both provide support resources for data submitters [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

When to Seek Institutional Support

Researchers should involve their institutional data management or bioinformatics support when:

  • The submission is part of a large project with many samples
  • The data has complex access requirements
  • The project involves multiple institutions with different data policies
  • The researcher is uncertain about compliance with funding or journal requirements

Institutional support can help researchers navigate the submission process and ensure that their data management practices meet all requirements.

Practical Implementation Steps

Before Sequencing Begins

The best time to plan for data submission is before sequencing starts. Researchers should:

  1. Review the submission requirements of their target repository
  2. Create a sample naming convention that will work for both the sequencing facility and the repository
  3. Plan the metadata that will be collected for each sample
  4. Determine whether any specialized community resources are relevant to their research area

The Carpentries lessons provide foundational training in data organization that is directly applicable to planning a metagenomic project [<a href="#ref-7">7</a>]. Researchers who establish good data management practices early will find the submission process much easier.

During Data Generation

While sequencing is underway, researchers should:

  1. Track the files generated for each sample
  2. Document any deviations from the planned protocol
  3. Begin compiling the sample metadata table
  4. Perform preliminary quality checks on the sequencing output

The nf-core documentation provides standards for reproducible workflow usage and configuration [<a href="#ref-8">8</a>]. Researchers who use standardized workflows for their analysis will find it easier to document their methods and prepare their data for submission.

After Analysis Completion

Once the analysis is complete, researchers should:

  1. Finalize the sample metadata table
  2. Verify that all raw data files are available and intact
  3. Complete the repository submission
  4. Verify the accession numbers and update the manuscript

The Galaxy Training Network provides accessible workflow training that covers the full analysis pipeline from raw data to results [<a href="#ref-6">6</a>]. Researchers can use these workflows to ensure that their analysis is reproducible and that their data submission supports the published findings.

Building a Pre-Submission Data Audit System for Metagenomic Projects

A structured pre-submission audit system prevents the most common causes of repository rejection and data rework. instead of discovering missing files or inconsistent metadata during the submission process, researchers can implement a verification protocol that checks every component before the first upload attempt. This section provides a practical audit framework, a record-keeping system, and troubleshooting methods that complement the submission workflow described earlier.

The Case for Systematic Auditing

The survey of ancient genomics studies found that half of the studies archived incomplete datasets and that no studies met all criteria that could be considered best practice [<a href="#ref-4">4</a>]. These findings came from published studies, meaning the authors believed their archiving was sufficient. The gap between perceived completeness and actual completeness demonstrates that informal checking is not reliable. A formal audit system forces researchers to verify each component against a defined standard instead of relying on memory or assumption.

For metagenomic projects, the audit must cover both the sequencing files and the metadata that describes them. A project with 50 samples, each sequenced on two platforms with multiple runs per platform, can easily generate 200 or more files. Tracking these files manually without a structured system invites errors. The audit system described here converts this tracking burden into a repeatable process that can be completed in a few hours.

Audit Components and Verification Steps

The audit system has five components that should be checked in sequence before any submission begins.

Component 1: File Inventory Verification

Create a complete inventory of all sequencing files for the project. This inventory should list every FASTQ file with its sample identifier, read pair designation, file size, and checksum value. The checksum, typically an MD5 or SHA256 hash, provides a way to verify that files have not been corrupted during transfer or storage.

For each file in the inventory, verify the following:

  • The file name matches the sample identifier used in the metadata table
  • Paired-end files exist for both forward and reverse reads
  • The file size is consistent with expected sequencing output
  • The checksum matches the value recorded at the sequencing facility

The Carpentries lessons provide foundational training in shell commands and data organization that are directly applicable to generating file inventories and checksums [<a href="#ref-7">7</a>]. Researchers who are not comfortable with command-line tools can use the repository submission interfaces, which typically display file information during the upload process, but a pre-upload inventory is faster and catches problems before the upload begins.

Component 2: Metadata Completeness Assessment

Create a metadata completeness score for each sample in the project. This score is the proportion of relevant metadata fields that are filled, expressed as a percentage. The relevant fields depend on the sample type and study design, but for metagenomic projects the core fields include host species, body site or environmental source, collection date, geographic location, health status, DNA extraction method, sequencing platform, library preparation kit, read length, and insert size.

For each sample, count the number of filled fields and divide by the total number of relevant fields. A sample with 8 of 10 relevant fields filled has a completeness score of 80 percent. The target should be 100 percent for all samples. Any sample below this target requires metadata completion before submission.

The ancient genomics survey specifically recommended providing informative sample metadata [<a href="#ref-4">4</a>]. A completeness score makes this recommendation measurable. Instead of relying on a general sense that the metadata is adequate, researchers can calculate exactly which samples are missing which fields and address those gaps directly.

Component 3: File-Metadata Cross-Check

The file inventory and the metadata table must be consistent. Every sample in the metadata table should have corresponding sequencing files, and every sequencing file should correspond to a sample in the metadata table. This cross-check catches several common errors:

  • Samples that were sequenced but never added to the metadata table
  • Sequencing files that were renamed or moved after the metadata was created
  • Duplicate files that inflate the apparent sequencing depth
  • Files from a different project that were accidentally included

The cross-check can be performed by comparing the sample identifiers in the metadata table against the sample identifiers extracted from the file names. Any identifier that appears in only one list requires investigation.

Component 4: Technical Format Verification

The sequencing files must be in the correct format for repository submission. The standard format is FASTQ, typically compressed with gzip. Before submission, verify that:

  • The files open correctly and are not truncated
  • The quality score encoding matches the sequencing platform
  • The read lengths are consistent within each file
  • The number of reads in each paired-end file matches

The Galaxy Training Network provides accessible workflow training that covers quality control procedures for sequencing data [<a href="#ref-6">6</a>]. Running a quality control workflow before submission serves the dual purpose of verifying the technical format and identifying any data quality issues that should be documented in the submission.

Component 5: Repository Requirement Review

Each repository has specific requirements that change over time. Before submission, review the current requirements for the chosen repository. The NCBI provides official descriptions of its databases and submission processes [<a href="#ref-1">1</a>], and the EMBL-EBI training portal offers learning pathways that explain data-resource requirements [<a href="#ref-2">2</a>]. The review should confirm that the planned submission format and metadata structure match the current repository expectations.

Record-Keeping System for Audit Results

The audit produces records that should be maintained alongside the project documentation. These records serve multiple purposes: they document the verification process, provide evidence of data quality for reviewers, and create a reference for future submissions of related data.

Audit Log Structure

Create an audit log that records the following information for each audit run:

  • The date the audit was performed
  • The person who performed the audit
  • The version of the audit checklist used
  • The results for each component
  • Any issues identified and their resolution status

The audit log should be stored with the project files and referenced in the data availability statement of the associated publication. This documentation aligns with the recommendation to document archiving choices in papers and have these choices peer reviewed [<a href="#ref-4">4</a>].

Issue Tracking

When the audit identifies an issue, record it in a tracking table with the following fields:

  • Issue identifier
  • Sample or file affected
  • Description of the issue
  • Date identified
  • Resolution action
  • Date resolved
  • Verification that the resolution was successful

This tracking table ensures that no issue is forgotten or assumed to be resolved without verification. The ancient genomics survey found that incorrect experiment metadata on samples, libraries, and sequencing runs was a common problem [<a href="#ref-4">4</a>]. An issue tracking system catches these errors before submission instead of after publication.

Version Control for Metadata

The metadata table will likely change during the audit process as missing fields are filled and errors are corrected. Maintain version control for the metadata table so that changes can be tracked and reviewed. The Carpentries lessons provide foundational training in Git and version control that is directly applicable to managing metadata tables [<a href="#ref-7">7</a>]. Even a simple versioning system that saves dated copies of the metadata table is better than overwriting the original without a record.

Troubleshooting Common Audit Failures

The audit system will identify issues that require troubleshooting. The following are common failure patterns and their resolution approaches.

Mismatched Sample Identifiers

When the file inventory and metadata table do not match, the first step is to determine which list is correct. Check the original sequencing facility records to confirm the sample identifiers assigned at the time of sequencing. If the file names were changed after sequencing, the original names may be recoverable from the facility records or from the file headers. The resolution is to standardize on the correct identifiers and update either the file names or the metadata table to match.

Missing Metadata Fields

When the metadata completeness score identifies missing fields, the resolution depends on whether the information is still available. Collection date and location may be recorded in laboratory notebooks or electronic lab notebooks. Health status may require consultation with clinical collaborators. DNA extraction methods should be documented in the laboratory protocol. If the information is genuinely unavailable, document this in the audit log and consider whether the missing field affects the usability of the data.

Corrupted or Truncated Files

When a file fails the technical format verification, the first step is to check whether the file was corrupted during transfer or storage. Re-download the file from the original source and compare checksums. If the file is corrupted at the source, contact the sequencing facility to request a replacement. If the file is truncated, the sequencing run may need to be repeated or the file may need to be regenerated from the instrument output.

Inconsistent Read Counts

When paired-end files have different numbers of reads, the files may have been processed differently or one file may be corrupted. Check the processing history for the sample. If adapter trimming or quality filtering was applied, the same filtering should have been applied to both read pairs. If the files are raw output from the sequencing instrument, inconsistent read counts may indicate a problem with the sequencing run itself.

Integrating the Audit with the Submission Process

The audit should be completed before the submission process begins. The submission workflow described earlier assumes that the data and metadata are ready for upload. The audit verifies this readiness and identifies any issues that need resolution before the first upload attempt.

The audit also provides a natural checkpoint for the documentation that should accompany the submission. The audit log, issue tracking table, and metadata version history together constitute the documentation of archiving choices that the ancient genomics survey recommended [<a href="#ref-4">4</a>]. This documentation can be referenced in the data availability statement and made available to reviewers upon request.

Measuring Audit Effectiveness

The audit system itself should be evaluated periodically to ensure it is catching the issues that matter. Track the following metrics across projects:

  • The number of issues identified per audit run
  • The types of issues most frequently identified
  • The time required to complete the audit
  • The number of submission validation failures that occur despite the audit

If submission validation failures persist, the audit checklist should be updated to include the checks that would have caught the failures. The nf-core documentation provides standards for reproducible workflow usage and configuration [<a href="#ref-8">8</a>], and these standards can inform the audit checklist for projects that use standardized analysis pipelines.

The audit system is not a guarantee against all submission problems, but it substantially reduces the risk of incomplete or inconsistent submissions. The investment of a few hours per project is small compared to the cost of discovering archiving problems after publication, when the data may be difficult or impossible to correct.

Frequently Asked Questions

What is the difference between SRA, ENA, and DDBJ?

The Sequence Read Archive at NCBI, the European Nucleotide Archive at EMBL-EBI, and the DNA Data Bank of Japan are the three major public repositories for raw sequencing data. They form the International Nucleotide Sequence Database Collaboration and synchronize data across their systems. Researchers can submit to any of the three repositories, and the data will be shared with the others. The choice of repository often depends on geographic location, funding requirements, and journal policies. The NCBI provides official descriptions of its databases and resources [<a href="#ref-1">1</a>], and the EMBL-EBI offers training on its data resources [<a href="#ref-2">2</a>].

What file formats are accepted for metagenomic data submission?

The standard format for raw sequencing data is FASTQ, typically compressed with gzip. The repository submission systems accept compressed FASTQ files and validate that they are properly formatted. Some repositories also accept BAM files for aligned data, but the raw FASTQ files should always be submitted as the primary evidence. The ancient genomics survey recommended archiving read alignments as secondary analysis files, in addition to the raw reads [<a href="#ref-4">4</a>].

How much metadata is required for a metagenomic submission?

The required metadata includes sample identifiers, organism or environment descriptions, sequencing platform, library preparation method, and experimental design details. Recommended metadata includes collection date and location, health status, sample processing methods, and DNA extraction protocols. The ancient genomics survey found that half of the studies surveyed archived incomplete datasets and recommended providing informative sample metadata [<a href="#ref-4">4</a>]. Researchers should aim to complete all relevant metadata fields, beyond the required ones.

Can I submit metagenomic data to a specialized database instead of a general repository?

Specialized databases like BlastoDB for Blastocystis research and BeeBiome for bee microbiomes provide valuable community resources [<a href="#ref-5">5</a>][<a href="#ref-3">3</a>]. However, these databases do not replace the general repositories. Researchers should submit their raw data to SRA, ENA, or DDBJ and may additionally submit to specialized databases for increased visibility. The specialized databases typically curate data from the general repositories instead of accepting primary submissions.

What should I do if my sequencing files are very large?

Large metagenomic datasets can be challenging to upload. Researchers should use the command-line upload tools provided by the repositories, which support parallel transfers and resume capabilities. For very large projects, contact the repository help desk before submission to discuss the best upload strategy. The repository staff can provide guidance on file organization and transfer methods.

How do I set an embargo on my metagenomic data?

Most repositories allow researchers to set an embargo period during which the data is not publicly accessible. The embargo release date should be set to match the expected publication date. Researchers should update the release date if the publication is delayed. Once the embargo expires, the data becomes publicly accessible and cannot be made private again.

What happens if my submission fails validation?

Repository submission systems perform automated validation checks on the submitted data and metadata. If validation fails, the system will provide error messages that explain the issues. Researchers should address these issues and resubmit. Common validation failures include inconsistent file formats, missing metadata fields, and mismatched sample identifiers. If the validation errors are not clear, contact the repository help desk for assistance.

Do I need to submit data from negative or failed experiments?

The ancient genomics survey specifically recommended archiving data from low-coverage and negative experiments [<a href="#ref-4">4</a>]. This recommendation applies to metagenomic studies as well. Data from experiments that did not produce the expected results can still be valuable for understanding technical variation and for meta-analyses. Researchers should submit all sequencing data, including data from experiments that were not included in the final analysis.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [2] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [3] [The BeeBiome data portal provides easy access to bee microbiome information.](https://doi.org/10.1186/s12859-025-06229-7). 2025. [4] [Improving data archiving practices in ancient genomics.](https://doi.org/10.1038/s41597-024-03563-y). 2024. [5] [BlastoDB: first release of a community-driven multi-omics and epidemiological resource for <,i>,Blastocystis<,/i>, biology and subtyping.](https://doi.org/10.12688/openreseurope.23235.2). 2026. [6] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [8] [nf-core Documentation](https://nf-co.re/docs). nf-core.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.