# Storing and Archiving Assembly Data: Best Practices for Raw Reads, Assemblies, and Intermediate Files


## Key Takeaways

- Implement the 3-2-1 backup rule: maintain three total copies of critical data on two different media types, with one copy stored off-site, to ensure data integrity against hardware failures and accidental deletion.
- Utilize standardized file formats like FASTQ for raw reads and FASTA for assemblies to guarantee long-term accessibility across evolving software versions, and archive associated quality metrics (e.g., N50, completeness) alongside assemblies.
- Document all metadata meticulously, including organism/sample identifiers, sequencing platform, library preparation, coverage depth, and assembly software/parameters, as this information is integral to the scientific value of sequence data.
- Archive raw sequencing reads and final assemblies in public repositories like NCBI SRA or GenBank, ensuring permanent storage, unique accession numbers, and public accessibility, while defining specific retention periods (1-2 years post-publication) for intermediate files.
- Employ a structured decision framework using criteria such as regeneration cost, access frequency, legal/funding requirements, file size, and collaboration needs to assign appropriate storage tiers for each data category, preventing both data loss and unnecessary expenditure.
- Establish a comprehensive Assembly Data Registry tracking file name/pattern, data category, creation date, storage location, backup status, retention period, and archival/deletion date, with monthly reviews to manage storage costs and verify backup integrity.

---

Genome assembly projects generate large volumes of data across multiple stages, from raw sequencing reads through polished assemblies and annotation files. Researchers often underestimate the storage requirements and long-term preservation needs of these data types, leading to data loss, corrupted files, or unnecessary expenditure. This article provides a practical framework for managing sequencing data and assemblies cost-effectively while ensuring long-term integrity and accessibility. The guidance applies to laboratory professionals, biology students, and researchers working with whole-genome, metagenomic, or organellar genome projects who need concrete decisions about storage tiers, backup frequency, and data archival.

## The Scope of Assembly Data Management

Assembly projects produce distinct data categories with different storage requirements and reuse potential. Raw sequencing reads are the primary evidence for every downstream analysis and are typically the largest files in a project. Assembled contigs and scaffolds represent the derived product that researchers analyze and publish. Intermediate files, including error-corrected reads, assembly graphs, and polishing iterations, occupy substantial space but have variable long-term value.

The National Center for Biotechnology Information maintains public sequence databases that serve as the primary repository for raw reads and assembled genomes [<a href="#ref-1">1</a>]. Understanding what belongs in public repositories versus local storage is the first decision in any data management plan. The European Bioinformatics Institute provides training pathways that cover data submission and retrieval practices across these repositories [<a href="#ref-2">2</a>]. Researchers should consult these resources when planning submission timelines and file format requirements.

A metagenomic project workflow described in a 2008 review emphasizes that data management decisions accompany every stage of analysis, from presequencing considerations through assembly and annotation [<a href="#ref-3">3</a>]. The authors note that the chain of decisions includes data storage and organization alongside computational analysis. This perspective remains relevant for modern assembly projects, where storage costs can rival sequencing costs for large genomes or deep-coverage datasets.

## Core Principles for Assembly Data Storage

### Data Integrity Requires Redundancy

Single-copy storage on a laboratory workstation or external drive does not constitute a preservation strategy. Hardware failures, accidental deletion, and file corruption are common causes of data loss in academic settings. A robust storage plan maintains at least two independent copies of critical data, with one copy stored off-site or in a different physical location.

The 3-2-1 rule provides a useful baseline: three total copies of the data, on two different media types, with one copy off-site. For assembly projects, this translates to maintaining the primary working copy on a local server or workstation, a second copy on a network-attached storage device or institutional server, and a third copy in cloud storage or on tape media stored at a different facility.

### File Formats Affect Long-Term Accessibility

Proprietary file formats from sequencing instruments and analysis software may become inaccessible as software versions change. Standardized formats such as FASTQ for raw reads, FASTA for assemblies, and SAM or BAM for alignments ensure that data remains readable by multiple software packages. The Galaxy Training Network provides tutorials that demonstrate working with these standard formats across different analysis platforms [<a href="#ref-4">4</a>].

When archiving assemblies, include both the primary FASTA file and the associated quality metrics, such as N50 values, completeness scores, and coverage statistics. These metrics allow future researchers to assess assembly quality without re-running the entire analysis pipeline. The CAMI benchmarking toolkit described in a 2021 Nature Protocols tutorial provides standardized metrics for assessing assembly quality that can be recorded alongside archived data [<a href="#ref-5">5</a>].

### Metadata Is Part of the Data

Sequence files without associated metadata lose most of their scientific value. Record the organism or sample identifier, sequencing platform, library preparation method, coverage depth, assembly software and version, and parameter settings. This information should be stored in a machine-readable format alongside the sequence data.

Voucher specimens provide the physical evidence for taxonomic identification of genome assemblies, as emphasized in a 2021 eLife article [<a href="#ref-6">6</a>]. The authors note that most vertebrate genomes in public databases lack links to voucher specimens, creating potential for taxonomic errors. For projects involving non-model organisms, record the voucher specimen identifier and deposition location in the metadata file.

## At a Glance: Storage Tier Comparison

| Storage Tier | Typical Cost per TB | Access Speed | Recommended Use | Retention Period |
| --- | --- | --- | --- | --- |
| Local workstation or server | Low initial cost, ongoing maintenance | Fastest | Active analysis, working files | Duration of project |
| Institutional or network storage | Moderate, often subsidized | Fast | Primary backup, shared access | Project duration plus 1 to 2 years |
| Cloud object storage | Pay-per-use, variable | Moderate | Off-site backup, collaboration | 3 to 5 years or longer |
| Tape or cold archival | Lowest per TB | Slow, requires retrieval | Long-term archival, raw reads | 5 to 10 years or permanent |

The table above provides a starting point for storage planning. Actual costs vary by institution, cloud provider, and data volume. Researchers should consult their institutional information technology services for available storage allocations before purchasing external storage or cloud subscriptions.

## Practical Workflow for Assembly Data Management

### Step 1: Estimate Storage Requirements Before Sequencing

Calculate expected data volume before the sequencing run begins. Multiply the genome size by the planned coverage depth to estimate raw read volume. For example, a 5 Mb bacterial genome sequenced at 100x coverage produces approximately 500 Mb of raw sequence data, with FASTQ files typically requiring 2 to 3 times this amount due to quality scores and read identifiers.

Metagenomic projects require larger estimates because community complexity and the need for deep coverage of multiple organisms increase data requirements. The 2008 metagenomics review notes that sequencing strategies must account for community composition and the desired depth for downstream analyses [<a href="#ref-3">3</a>]. Plan for 2 to 3 times the estimated raw data volume to accommodate quality trimming and error correction steps.

### Step 2: Establish a Directory Structure

Create a consistent directory structure at the start of each project. A recommended structure includes separate directories for raw reads, trimmed reads, assemblies, polishing iterations, annotations, and metadata. Include the project identifier and date in directory names to prevent confusion between similar projects.

Record the directory structure and file naming conventions in a README file stored at the project root. This documentation ensures that anyone accessing the data, including future lab members or collaborators, can navigate the project without prior knowledge.

### Step 3: Implement Automated Backup Procedures

Manual backup procedures fail because researchers forget to run them or run them inconsistently. Use automated backup tools that run on a schedule, such as nightly incremental backups to a network drive and weekly full backups to cloud storage. The Carpentries lessons provide foundational instruction on using command-line tools and shell scripts that can automate backup procedures [<a href="#ref-7">7</a>].

Test backup restoration procedures regularly. A backup that cannot be restored is not a backup. Schedule quarterly restoration tests where a random file is retrieved from each backup tier and verified against the original.

### Step 4: Archive Raw Reads in Public Repositories

Raw sequencing reads should be deposited in public repositories such as the NCBI Sequence Read Archive or equivalent databases [<a href="#ref-1">1</a>]. These repositories provide permanent storage, unique accession numbers, and public accessibility. Submission typically occurs after quality assessment and before or concurrent with manuscript submission.

Deposit assembled genomes in public genome databases with associated metadata, including the sequencing platform, assembly method, and quality metrics. The NCBI provides multiple databases for different data types, and researchers should select the appropriate database for their organism and data type [<a href="#ref-1">1</a>].

### Step 5: Define Retention Periods for Intermediate Files

Intermediate files, including error-corrected reads, assembly graphs, and polishing iterations, have variable long-term value. Retain these files for the duration of the project and for a defined period after publication, typically 1 to 2 years, to allow for manuscript revisions and reader inquiries.

Assembly graphs and polishing iterations may be needed if the assembly is updated or if reviewers request additional analyses. However, these files are often large and can be regenerated from raw reads and the documented analysis pipeline. The nf-core documentation emphasizes that reproducible workflows allow regeneration of intermediate results from raw inputs [<a href="#ref-8">8</a>]. If the analysis pipeline is fully documented and version-controlled, intermediate files can be deleted after the project concludes.

## Storage Options and Tradeoffs

### Local Storage

Local storage on laboratory workstations or dedicated servers provides the fastest access for active analysis. Solid-state drives offer faster read and write speeds than traditional hard drives but cost more per terabyte. For most assembly projects, traditional hard drives provide sufficient performance at lower cost.

Local storage requires ongoing maintenance, including monitoring disk space, replacing failed drives, and managing user access. Institutional information technology services often provide managed storage with automated backups, which reduces the burden on individual researchers.

### Network and Institutional Storage

Most universities and research institutions provide network-attached storage for research data. This storage is typically backed up automatically and accessible from multiple computers. Institutional storage policies may include data retention requirements and access controls that align with funding agency mandates.

Check institutional storage quotas and policies before starting a large assembly project. Some institutions charge for storage above a base allocation, and understanding these costs helps with budget planning.

### Cloud Storage

Cloud storage providers offer scalable storage with pay-per-use pricing. Object storage services are well suited for archival data because they provide durability and accessibility from anywhere with internet access. Cloud storage also facilitates collaboration with researchers at other institutions.

Cloud storage costs include data transfer fees for uploading and downloading, which can be substantial for large datasets. Compare the total cost of cloud storage, including transfer fees, against institutional storage costs before committing to a cloud provider.

### Tape and Cold Archival

Tape storage provides the lowest cost per terabyte for large datasets that require long-term preservation. Tape media have a lifespan of 15 to 30 years when stored properly, making them suitable for permanent archival of raw reads and final assemblies.

Tape storage requires specialized hardware and retrieval procedures. Access times range from minutes to hours, making tape unsuitable for active analysis. Many institutions provide tape archival services for research data, and researchers should inquire about these services before purchasing their own tape hardware.

## Records and Measurements for Storage Management

### Track Storage Usage by Data Category

Maintain a spreadsheet or database that records storage usage by data category, including raw reads, assemblies, intermediate files, and backups. Update this record monthly or at each project milestone. The record should include file counts, total size, storage location, and backup status.

This record serves multiple purposes. It helps identify which data categories consume the most storage, informs future storage budget estimates, and provides documentation for funding agency reporting requirements.

### Record Assembly Quality Metrics

Assembly quality metrics provide the basis for deciding which assemblies warrant long-term preservation. Record N50 and L50 values, total assembly size, number of contigs, completeness scores from benchmarking tools, and coverage statistics. The CAMI benchmarking toolkit provides standardized metrics and procedures for assessing assembly quality [<a href="#ref-5">5</a>].

For organellar genomes, such as chloroplast genomes, standard data reporting practices include assembly verification and annotation quality checks, as described in a 2021 methods chapter [<a href="#ref-9">9</a>]. Record these verification results alongside the assembly files.

### Document Software Versions and Parameters

Reproducibility requires documentation of the exact software versions and parameters used for each analysis step. Record the assembly software, version number, and parameter settings in a configuration file or README. The nf-core documentation emphasizes that reproducible workflows require version-controlled pipelines and configuration files [<a href="#ref-8">8</a>].

Bioconductor provides reproducible genomic-analysis documentation and package management tools that help researchers track software versions [<a href="#ref-10">10</a>]. Use these tools or equivalent approaches to ensure that the analysis pipeline can be re-run if needed.

## Common Failure Patterns in Assembly Data Management

### Failure Pattern 1: Single-Copy Storage

Researchers who store data only on their laboratory workstation risk losing everything when the drive fails or the computer is stolen. This pattern is common in small laboratories without institutional storage policies. The solution is to implement automated backups to at least one additional location from the start of the project.

### Failure Pattern 2: Unlabeled or Poorly Labeled Files

Files named with generic identifiers such as "final_assembly.fasta" or "data_2.fastq" become impossible to distinguish as projects accumulate. This pattern leads to analysis errors and wasted time identifying the correct files. The solution is to establish naming conventions that include project identifiers, sample identifiers, and version numbers.

### Failure Pattern 3: Incomplete Metadata

Researchers who archive sequence files without associated metadata create data that cannot be interpreted by future researchers or even by the original researcher months later. This pattern is particularly problematic for public repository submissions, where incomplete metadata may result in data being rejected or misinterpreted. The solution is to create a metadata template at project initiation and complete it as data is generated.

### Failure Pattern 4: Ignoring Storage Costs Until They Become Critical

Researchers who do not track storage usage may exceed institutional quotas or incur unexpected cloud storage charges. This pattern leads to rushed decisions about data deletion that may result in loss of valuable intermediate files. The solution is to track storage usage from the start and plan for storage costs in grant budgets.

### Failure Pattern 5: Untested Backups

Researchers who create backups but never test restoration procedures discover that backups are corrupted or incomplete only when they need to restore data. This pattern is common when backup procedures are manual or when backup software is not monitored. The solution is to schedule regular restoration tests and verify backup logs.

## Limitations and Interpretation Boundaries

### Storage Decisions Cannot Compensate for Poor Sequencing

Storage and archival practices preserve the data that exists, but they cannot improve data quality. Poor-quality sequencing reads, inadequate coverage, or contaminated samples produce assemblies with limited scientific value regardless of how well the data is stored. The 2008 metagenomics review emphasizes that presequencing considerations, including community composition and sequencing strategy, greatly influence downstream analyses [<a href="#ref-3">3</a>].

### Public Repository Submission Does Not Guarantee Data Quality

Public repositories accept data submissions but do not independently verify assembly quality or metadata accuracy. Researchers are responsible for ensuring that submitted data meets quality standards and that metadata is complete and accurate. The eLife article on vouchers notes that taxonomic errors in public databases can propagate through the scientific literature [<a href="#ref-6">6</a>].

### Storage Media Have Finite Lifespans

All storage media degrade over time. Hard drives typically last 3 to 5 years, solid-state drives have limited write cycles, and optical media degrade over decades. Tape media have longer lifespans but require proper storage conditions. Regular data migration to new media is necessary for long-term preservation.

### Cloud Storage Does Not Eliminate Data Management Responsibilities

Cloud storage providers maintain infrastructure but do not manage research data. Researchers remain responsible for organizing files, maintaining metadata, and ensuring that data is accessible to authorized collaborators. Cloud storage also introduces security considerations, particularly for data subject to access restrictions.

## Professional Escalation Criteria

### When to Consult Institutional Information Technology Services

Consult institutional information technology services when storage requirements exceed available allocations, when specialized storage solutions such as tape archival are needed, or when data security requirements exceed standard institutional provisions. Information technology staff can provide guidance on institutional storage options and policies.

### When to Consult Bioinformatics Support

Consult bioinformatics support when assembly quality metrics indicate problems, when analysis pipelines fail to reproduce results, or when data format conversions are needed for repository submission. Bioinformatics support can also assist with documenting analysis pipelines for reproducibility.

### When to Consult Data Management Specialists

Consult data management specialists when planning large-scale projects with multiple data types, when funding agency data management requirements are unclear, or when data must be preserved for regulatory or legal reasons. Data management specialists can help develop data management plans that meet institutional and funding requirements.

## Preservation of Field-Collected Samples

### Sample Preservation Affects Data Quality

The preservation method used for field-collected samples affects the quality of extracted DNA and RNA, which in turn affects sequencing and assembly outcomes. A 2022 study comparing preservation methods for glacial snow and ice samples found that flash freezing of filters was the preferred method for preserving DNA and RNA from low-biomass samples [<a href="#ref-11">11</a>]. The study also found that preservation method had less impact on results when larger sample volumes were filtered.

For projects involving field-collected samples, document the preservation method and sample volume in the metadata. This information is essential for interpreting assembly quality and for comparing results across studies.

### Voucher Specimens Support Taxonomic Accuracy

Voucher specimens provide the physical evidence for taxonomic identification of genome assemblies [<a href="#ref-6">6</a>]. For projects involving non-model organisms, deposit voucher specimens in accessible, permanent research collections and link these vouchers to publications and public databases. This practice prevents taxonomic errors and supports reproducibility.

The eLife article recommends that researchers generating new genome assemblies deposit voucher specimens and recognize the work of local field biologists who collected the specimens [<a href="#ref-6">6</a>]. This practice also promotes a diverse and inclusive knowledge base by acknowledging the contributions of local researchers.

## Cost Considerations for Storage Planning

### Estimate Total Storage Costs

Total storage costs include the cost of storage media or cloud subscriptions, the cost of backup storage, and the labor cost of managing the storage system. For cloud storage, include data transfer fees in the estimate. For local storage, include the cost of replacement drives and the labor cost of monitoring and maintenance.

### Compare Storage Options Based on Total Cost

Compare storage options based on total cost over the expected retention period, beyond the initial purchase price. A local external drive may appear cheaper than cloud storage, but the cost of replacement drives, backup storage, and labor may exceed cloud storage costs over a 5-year period.

### Plan for Data Growth

Sequencing data volumes continue to increase as sequencing costs decrease and coverage depths increase. Plan for data growth by building in storage capacity for future projects and by regularly reviewing storage usage to identify data that can be archived or deleted.

## Quality Controls for Archived Data

### Verify File Integrity After Transfer

Verify file integrity after transferring data between storage systems. Use checksums or hash functions to confirm that files were transferred without corruption. Record checksums in the metadata file so that data integrity can be verified at any time.

### Verify Assembly Quality Before Archival

Verify assembly quality before archiving final assemblies. Use standardized metrics and benchmarking tools to assess assembly completeness and correctness. The CAMI benchmarking toolkit provides procedures for generating these metrics [<a href="#ref-5">5</a>].

### Verify Metadata Completeness Before Repository Submission

Verify that metadata is complete before submitting data to public repositories. Include the organism or sample identifier, sequencing platform, coverage depth, assembly software and version, and quality metrics. For projects involving field-collected samples, include preservation method and voucher specimen identifiers.

## A Practical Decision Framework for Assembly Data Storage Tiers

Choosing where to store each data type requires a structured decision process instead of intuition or habit. Researchers often default to storing everything on the fastest available storage or, conversely, deleting intermediate files to save space without considering future needs. A decision framework that evaluates each data category against explicit criteria helps balance cost, accessibility, and preservation requirements.

### Decision Criteria for Storage Placement

Five criteria determine the appropriate storage tier for any file in an assembly project: regeneration cost, access frequency, legal or funding requirements, file size, and collaboration needs. Regeneration cost asks whether the file can be recreated from raw reads and documented pipelines, and if so, at what computational expense. Access frequency distinguishes between files needed weekly for active analysis and files accessed once every few years. Legal or funding requirements may mandate specific retention periods or repository deposition. File size determines whether storage cost is a meaningful factor or a rounding error in the project budget. Collaboration needs identify whether multiple researchers at different institutions require simultaneous access.

Apply these criteria in sequence for each data category. Raw sequencing reads score high on regeneration cost because they cannot be recreated, high on legal requirements because most funding agencies and journals mandate deposition, and high on file size. This combination points to public repository deposition plus cold archival storage. Final assemblies score high on regeneration cost because they represent substantial computational investment, high on access frequency during manuscript preparation, and moderate on file size. This points to redundant local storage plus public repository deposition. Intermediate files score low on regeneration cost if the pipeline is documented, low on access frequency after project completion, and variable on file size. This points to deletion after the retention period instead of long-term archival.

### A Scoring System for Storage Decisions

A simple scoring system formalizes the decision process. For each file or data category, assign a score from 1 to 5 for each of the five criteria, where 5 represents high regeneration cost, frequent access, strong legal requirements, large file size, or high collaboration need. Sum the scores and use the total to guide storage placement.

Scores of 20 to 25 indicate that the data belongs in the highest protection tier with redundant copies and public repository deposition. Scores of 15 to 19 indicate that the data warrants institutional storage with automated backups and defined retention periods. Scores of 10 to 14 indicate that the data can be stored on local or network storage with periodic backups. Scores below 10 indicate that the data can be deleted after the project concludes, provided the analysis pipeline is fully documented.

This scoring system provides a defensible basis for storage decisions that can be documented in project records and explained to collaborators or funding agencies. It also prevents the common failure pattern of treating all data equally, which leads either to excessive storage costs or to premature deletion of valuable files.

### Applying the Framework to Specific Data Categories

Raw sequencing reads receive a regeneration cost score of 5 because they cannot be recreated, an access frequency score of 2 because they are rarely accessed after quality assessment, a legal requirements score of 5 because most funding agencies mandate deposition, a file size score of 5 because they dominate project storage, and a collaboration score of 3 because collaborators may request access. The total of 20 places raw reads in the highest protection tier with public repository deposition and cold archival storage.

Final assemblies receive a regeneration cost score of 4 because they can be regenerated from raw reads but at substantial computational expense, an access frequency score of 4 during active analysis and manuscript preparation, a legal requirements score of 4 because journals typically require accession numbers, a file size score of 3 because assemblies are smaller than raw reads, and a collaboration score of 4 because assemblies are the primary file shared with collaborators. The total of 19 places assemblies in the high protection tier with redundant local storage and public repository deposition.

Error-corrected reads receive a regeneration cost score of 2 because they can be regenerated from raw reads with documented trimming and correction steps, an access frequency score of 2 because they are rarely accessed after assembly, a legal requirements score of 1 because no repository accepts them, a file size score of 4 because they are nearly as large as raw reads, and a collaboration score of 1 because they are rarely shared. The total of 10 places error-corrected reads in the deletion category after the project retention period.

Assembly graphs receive a regeneration cost score of 3 because they require specific software versions and parameters to reproduce, an access frequency score of 1 because they are rarely accessed after assembly completion, a legal requirements score of 1, a file size score of 3 because they can be large for complex genomes, and a collaboration score of 1. The total of 9 places assembly graphs in the deletion category, provided the assembly software and parameters are documented.

Polishing iterations receive a regeneration cost score of 2, an access frequency score of 1, a legal requirements score of 1, a file size score of 2, and a collaboration score of 1. The total of 7 places polishing iterations in the deletion category after the project concludes.

### Documenting Storage Decisions

Record the scoring and resulting storage decision for each data category in the project README file. This documentation serves multiple purposes. It provides a rationale for storage choices that can be explained to collaborators or funding agencies. It creates a record that can be reviewed when storage costs are questioned or when data needs to be located years after project completion. It also establishes a template that can be reused for future projects, reducing the time needed to make storage decisions.

The documentation should include the date of the decision, the person responsible, the scores assigned to each criterion, and the resulting storage placement. Update the documentation when significant changes occur, such as a new collaboration that increases access frequency or a funding agency requirement that changes retention periods.

## A Record System for Assembly Data Lifecycle Tracking

Storage decisions are only useful if they are tracked over time. A record system that tracks each data file from creation through archival or deletion provides the visibility needed to manage storage costs, verify backup integrity, and locate files when needed. This record system complements the directory structure and backup procedures described in the existing article by adding a temporal dimension to data management.

### The Assembly Data Registry

Create a registry, implemented as a spreadsheet or simple database, that records each data file or data category in the project. The registry should include the file name or pattern, the data category, the date of creation, the storage location, the backup status, the retention period, and the archival or deletion date.

The registry serves as the authoritative record of what data exists, where it is stored, and what will happen to it. Without such a registry, researchers rely on memory or directory browsing to locate files, which becomes unreliable as projects accumulate and lab members change.

### Registry Fields and Their Purpose

The file name or pattern field identifies the data being tracked. For individual files, record the exact file name. For data categories that include many files, such as raw reads split across multiple FASTQ files, record the directory path and file pattern.

The data category field links each file to the storage decision framework. Categories include raw reads, trimmed reads, error-corrected reads, assembly graphs, draft assemblies, polished assemblies, final assemblies, annotations, and metadata.

The date of creation field records when the file was generated. This date supports retention period calculations and helps identify files that may be candidates for deletion or archival.

The storage location field records the current primary storage location, such as the local server path, network drive, or cloud bucket. This field must be updated when files are moved between storage tiers.

The backup status field records whether the file is included in automated backup procedures and the date of the last verified backup. This field supports the backup testing procedures described in the existing article.

The retention period field records the planned retention duration based on the storage decision framework. This field drives the review process that determines whether files are archived, retained, or deleted.

The archival or deletion date field records when the file was moved to archival storage or deleted. This field provides the audit trail needed to demonstrate compliance with data management policies.

### Monthly Registry Review

Review the registry monthly or at each project milestone. The review serves three purposes. First, it identifies files that have reached their retention period and require a decision about archival or deletion. Second, it identifies files whose storage location has changed and requires registry updates. Third, it identifies files whose backup status is outdated or unverified.

The monthly review should be a scheduled activity with a defined procedure. The researcher responsible for data management reviews each file that has reached its retention period, confirms that the analysis pipeline is documented for files proposed for deletion, and updates the registry accordingly.

### Registry Integration with Backup Verification

The registry provides the file list for backup verification procedures. When testing backup restoration, select files from the registry instead of from memory or directory browsing. This ensures that backup testing covers all data categories and storage tiers.

Record the results of backup verification in the registry, including the date of the test, the files tested, and the outcome. This record demonstrates that backup procedures are functioning and provides evidence for data management audits.

## Troubleshooting Storage and Archival Problems

Storage problems manifest in predictable ways, and recognizing the symptoms early prevents data loss and cost overruns. This troubleshooting method addresses the most common storage failures in assembly projects and provides concrete diagnostic steps.

### Symptom 1: Storage Quota Exceeded Without Warning

Researchers often discover that they have exceeded institutional storage quotas only when writes begin to fail. This symptom indicates that storage usage is not being tracked proactively. The registry described above provides the solution by tracking storage usage by data category.

Diagnostic steps include reviewing the registry to identify the largest data categories, checking whether intermediate files have reached their retention periods, and verifying that raw reads have been deposited in public repositories so that local copies can be moved to cold storage.

### Symptom 2: Backup Completion Reports Show Errors

Automated backup systems generate reports that researchers often ignore until a restoration attempt fails. Backup errors indicate either that files are locked or in use during backup, that the backup destination is full, or that the backup software is misconfigured.

Diagnostic steps include reviewing backup logs to identify the specific files that failed, checking whether those files are open in analysis software during the backup window, and verifying that the backup destination has sufficient free space. Adjust the backup schedule to avoid active analysis periods or configure the backup software to handle open files.

### Symptom 3: File Corruption Detected During Analysis

Analysis software may fail to read files that were previously accessible, indicating file corruption. This symptom requires immediate action to prevent further data loss.

Diagnostic steps include running checksum verification on the affected files, comparing checksums against the values recorded in the metadata file, and restoring corrupted files from backup copies. If no backup copy exists, assess whether the file can be regenerated from raw reads and documented pipelines. Document the corruption event in the registry and review storage hardware for signs of failure.

### Symptom 4: Cloud Storage Costs Exceed Budget

Cloud storage bills that exceed budget estimates indicate either that data volumes are larger than expected, that data transfer fees are higher than anticipated, or that the storage tier is more expensive than necessary.

Diagnostic steps include reviewing cloud storage usage reports to identify the largest data categories, checking whether infrequently accessed data can be moved to lower-cost storage tiers such as cold or archive classes, and reviewing data transfer logs to identify repeated downloads of the same files. Consider whether collaborators can access data through shared links instead of individual downloads.

### Symptom 5: Public Repository Submission Rejected

Repository submissions may be rejected for incomplete metadata, incorrect file formats, or missing quality metrics. This symptom indicates that submission requirements were not reviewed before data generation.

Diagnostic steps include reviewing the repository submission requirements, comparing the prepared files against the requirements, and consulting repository documentation or training materials. The NCBI provides documentation for its databases and submission systems [<a href="#ref-1">1</a>], and the EMBL-EBI provides training on data submission and retrieval [<a href="#ref-2">2</a>]. Correct the identified issues and resubmit.

### Symptom 6: Assembly Cannot Be Reproduced

When a researcher attempts to reproduce an assembly from raw reads and documented pipelines, the result may differ from the original assembly. This symptom indicates that the pipeline documentation is incomplete or that software versions have changed.

Diagnostic steps include reviewing the recorded software versions and parameters, checking whether the same software versions are still available, and comparing the original assembly quality metrics with the reproduced assembly quality metrics. The CAMI benchmarking toolkit provides standardized metrics for comparing assembly results [<a href="#ref-5">5</a>]. If the reproduced assembly differs substantially, the original assembly files should be retained instead of deleted.

## Welfare and Safety Context for Data Management

### Data Management as Research Integrity

Data management practices directly affect research integrity. Incomplete metadata, untracked files, and untested backups can lead to unreproducible results, retracted publications, and wasted public research funding. The eLife article on vouchers emphasizes that taxonomic errors in public databases can propagate through the scientific literature [<a href="#ref-6">6</a>], and incomplete metadata increases the risk of such errors.

Researchers have an obligation to ensure that the data they generate remains accessible and interpretable for the duration of its scientific value. This obligation extends beyond the active analysis period and includes the archival phase when data may be accessed by researchers who were not involved in the original project.

### Safety Considerations for Storage Hardware

Storage hardware presents physical safety considerations that are often overlooked. Hard drives generate heat and noise, and large arrays require adequate ventilation and cooling. Tape media require controlled temperature and humidity for long-term preservation. Researchers should follow institutional guidelines for equipment placement and environmental conditions.

Electrical safety is also relevant. Storage devices should be connected to surge protectors or uninterruptible power supplies to prevent data loss from power fluctuations. Institutional information technology services can provide guidance on safe equipment installation.

### Data Security and Access Control

Data management includes security considerations, particularly for data subject to access restrictions such as human sequence data or data covered by material transfer agreements. Storage decisions must account for access control requirements, and cloud storage may not be appropriate for restricted data.

Researchers should consult institutional data security policies before selecting storage options. The decision framework described above should include a security assessment as part of the legal requirements criterion.

## Implementation Checklist for Storage Decision Framework

Implement the storage decision framework with the following steps.

First, create the storage decision scoring template as a spreadsheet or document that can be reused across projects. Include the five criteria, the scoring scale, and the interpretation guide.

Second, score each data category for the current project using the template. Record the scores and the resulting storage decisions in the project README file.

Third, create the assembly data registry with the fields described above. Populate the registry with all existing data files and categories.

Fourth, schedule the monthly registry review and the quarterly backup verification. Add these activities to the project calendar.

Fifth, review the troubleshooting symptoms and ensure that the diagnostic steps are documented in the project records. Assign responsibility for responding to each symptom type.

Sixth, verify that the storage decision framework and registry are documented in the project README file and that all project members understand how to use them.

## Common Failure Patterns in Storage Decision Implementation

### Failure Pattern 1: Scoring All Data the Same

Researchers who assign the same scores to all data categories defeat the purpose of the decision framework. This pattern occurs when the scoring is done hastily or when researchers do not distinguish between data types. The result is either excessive storage costs or premature deletion of valuable files.

The solution is to score each data category separately and to review the scores with another researcher or the project supervisor. The review provides an opportunity to challenge assumptions about regeneration cost and access frequency.

### Failure Pattern 2: Registry Not Updated

A registry that is created but not maintained becomes inaccurate and loses its value. This pattern occurs when registry updates are not assigned to a specific person or scheduled at regular intervals.

The solution is to assign registry maintenance to a specific researcher and to include registry updates in the monthly review schedule. The registry should be updated whenever files are created, moved, or deleted, also during scheduled reviews.

### Failure Pattern 3: Retention Periods Set but Not Enforced

Researchers who set retention periods but never review files at the retention date accumulate intermediate files indefinitely. This pattern leads to storage costs that grow without bound and makes it difficult to locate active files among obsolete ones.

The solution is to enforce retention periods through the monthly registry review. Files that have reached their retention period should be archived or deleted according to the documented decision, not retained indefinitely because the review was skipped.

### Failure Pattern 4: Troubleshooting Without Documentation

Researchers who troubleshoot storage problems without documenting the diagnosis and resolution repeat the same problem-solving process each time the symptom recurs. This pattern wastes time and may lead to inconsistent solutions.

The solution is to document each troubleshooting event in the registry or a separate troubleshooting log. Include the symptom, the diagnostic steps taken, the root cause identified, and the resolution applied. This documentation becomes a reference for future troubleshooting.

### Failure Pattern 5: Framework Applied Only to New Projects

Researchers who apply the storage decision framework only to new projects leave existing data unmanaged. This pattern creates a divide between well-managed new data and poorly managed legacy data.

The solution is to apply the framework retroactively to existing projects. Score existing data categories, create registry entries for all files, and make storage decisions for files that have already exceeded their retention periods.

## Limitations of the Decision Framework

The storage decision framework provides structure but does not eliminate the need for judgment. Scores are subjective, and different researchers may assign different values to the same data category. The framework should be used as a discussion tool instead of an automated decision maker.

The framework does not account for all institutional policies and funding agency requirements. Researchers must review their specific obligations and adjust the framework accordingly. The framework also does not account for the cost of storage hardware failures, which can exceed the cost of the storage itself when data loss occurs.

The framework assumes that analysis pipelines are documented and version-controlled. For projects without such documentation, intermediate files have higher regeneration cost and should be retained longer. The nf-core documentation emphasizes that reproducible workflows allow regeneration of intermediate results from raw inputs [<a href="#ref-8">8</a>], and the Carpentries lessons provide foundational instruction on version control and reproducible computing [<a href="#ref-7">7</a>].

The framework does not address the specific requirements of different sequencing platforms and data types. Researchers working with long-read sequencing, single-cell sequencing, or other specialized data types should adapt the framework to their specific needs.

## Professional Escalation Criteria for Storage Problems

### When to Escalate Storage Problems to Information Technology Services

Escalate storage problems to institutional information technology services when storage hardware fails, when backup systems fail repeatedly, when storage quotas cannot accommodate project needs, or when data security requirements exceed standard provisions. Information technology staff can diagnose hardware failures, reconfigure backup systems, and provide guidance on institutional storage options.

### When to Escalate Storage Problems to Bioinformatics Support

Escalate storage problems to bioinformatics support when assembly files cannot be reproduced from documented pipelines, when file format conversions fail, or when repository submissions are rejected for technical reasons. Bioinformatics support can help diagnose pipeline issues, perform format conversions, and prepare files for repository submission.

### When to Escalate Storage Problems to Data Management Specialists

Escalate storage problems to data management specialists when funding agency requirements are unclear, when data must be preserved for regulatory or legal reasons, or when large-scale projects require coordinated data management across multiple institutions. Data management specialists can help develop data management plans that meet institutional and funding requirements.

## Frequently Asked Questions

### How much storage do I need for a typical bacterial genome assembly project?

A typical bacterial genome project with 100x coverage of a 5 Mb genome produces approximately 500 Mb of raw sequence data. FASTQ files typically require 2 to 3 times the raw sequence volume, so plan for 1 to 1.5 GB of raw data storage. Assembly files and intermediate files add another 0.5 to 1 GB. Plan for 2 to 3 GB total storage per bacterial genome project, plus backup storage.

### Should I store intermediate assembly files or only raw reads and final assemblies?

Store intermediate files for the duration of the project and for 1 to 2 years after publication. Error-corrected reads, assembly graphs, and polishing iterations may be needed for manuscript revisions or reviewer requests. After this period, intermediate files can be deleted if the analysis pipeline is fully documented and version-controlled, allowing regeneration from raw reads.

### What is the difference between backup and archival storage?

Backup storage maintains copies of active data for disaster recovery and is accessed regularly. Archival storage preserves data for long-term retention and is accessed rarely. Backup storage typically uses faster media with shorter retention periods, while archival storage uses slower, cheaper media with longer retention periods. A complete storage plan includes both backup and archival components.

### How do I choose between cloud storage and institutional storage?

Compare total costs, including data transfer fees for cloud storage and labor costs for managing institutional storage. Consider access requirements, collaboration needs, and institutional policies. Many researchers use institutional storage for primary backups and cloud storage for off-site backups. Consult institutional information technology services to understand available options and costs.

### What metadata should I record for each assembly project?

Record the organism or sample identifier, voucher specimen identifier if applicable, sequencing platform and instrument, library preparation method, coverage depth, assembly software and version, parameter settings, quality metrics including N50 and completeness scores, and the date of each analysis step. Store this information in a machine-readable format alongside the sequence data.

### How long should I retain raw sequencing reads?

Retain raw sequencing reads permanently or for the maximum period allowed by institutional policies. Raw reads are the primary evidence for all downstream analyses and cannot be regenerated. Deposit raw reads in public repositories such as the NCBI Sequence Read Archive to ensure permanent preservation [<a href="#ref-1">1</a>].

### What should I do if I discover corrupted files in my archived data?

Restore the corrupted files from backup copies and verify file integrity using checksums. If no backup copy exists, assess whether the corrupted files can be regenerated from raw reads or other sources. Document the corruption event and the restoration procedure in the project records. Review backup procedures to prevent future corruption.

### How do I prepare data for submission to public repositories?

Review the submission requirements for the target repository, including file formats, metadata requirements, and quality standards. The NCBI provides documentation for its databases and submission systems [<a href="#ref-1">1</a>]. The EMBL-EBI provides training on data submission and retrieval [<a href="#ref-2">2</a>]. Prepare files in the required formats, complete all metadata fields, and verify file integrity before submission.

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Genomic Data Processing: From Raw Sequencing to Analysis-Ready Files](/knowledge/bioinformatics/genomic-data-processing-from-raw-sequencing-to-analysis-ready-files)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [A bioinformatician's guide to metagenomics.](https://pubmed.ncbi.nlm.nih.gov/19052320). Microbiology and molecular biology reviews : MMBR, 2008.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Tutorial: assessing metagenomics software with the CAMI benchmarking toolkit.](https://pubmed.ncbi.nlm.nih.gov/33649565). Nature protocols, 2021.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [The critical importance of vouchers in genomics.](https://pubmed.ncbi.nlm.nih.gov/34061026). eLife, 2021.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [Sequencing of Complete Chloroplast Genomes.](https://pubmed.ncbi.nlm.nih.gov/33301089). Methods in molecular biology (Clifton, N.J.), 2021.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [DNA/RNA Preservation in Glacial Snow and Ice Samples.](https://pubmed.ncbi.nlm.nih.gov/35677909). Frontiers in microbiology, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.