# Automated vs. Manual Curation of Genome Assemblies: When to Invest in Human Review


## Key Takeaways

- Automated curation tools like `purge_dups` effectively reduce assembly size by removing haplotypic duplications based on read depth and sequence similarity, but can incorrectly purge genuine sequence in highly polymorphic species or fail to distinguish recent whole-genome duplications.
- Hi-C scaffolding tools leverage chromosomal interaction patterns to order contigs into pseudomolecules, but their accuracy is compromised by structural rearrangements, repetitive regions, or assembly errors that disrupt the assumed correlation between contact frequency and linear genomic distance.
- Manual curation is essential for resolving complex genomic regions such as tandem repeats and segmental duplications, where automated algorithms struggle due to ambiguous read mapping and coverage patterns, allowing human expertise to interpret evidence and make informed decisions.
- Biological plausibility checks, such as verifying chromosome counts against known karyotypes or confirming the representation of sex chromosomes, are critical manual curation steps that automated tools cannot perform, especially for species with unusual genomic features.
- Genome complexity, characterized by high heterozygosity, extensive repeat content, or polyploidy, significantly increases the need for manual curation to ensure accurate representation of all genomic sequences and structures, particularly for reference-quality assemblies.
- Project goals dictate the required curation rigor; draft assemblies for gene content analysis may suffice with automated tools and quality metrics, while chromosome-level or reference-quality assemblies necessitate extensive manual review to ensure contiguity, completeness, and correctness.

---

Genome assembly projects face a practical decision point after the initial assembly is produced: whether to rely on automated curation tools or to invest laboratory time in manual review. This article provides a decision framework for biology students, researchers, laboratory professionals, and life-science practitioners who need to balance cost and accuracy when curating genome assemblies. The framework compares automated tools such as purge_dups and Hi-C scaffolding with manual curation approaches, and it offers concrete criteria based on genome complexity and project goals.

## The Curation Problem in Modern Genome Projects

Genome assembly has become a routine step in many biological research programs, yet the quality of the final product depends heavily on what happens after the assembler finishes its work. Automated curation tools can identify and correct many common assembly errors, but they operate within limits that are not always obvious to the user. Manual curation, by contrast, requires substantial time and expertise but can resolve issues that automated methods miss.

The scale of modern genome projects makes this decision consequential. Reference genome initiatives such as the Darwin Tree of Life project produce assemblies for eukaryotic species found in Britain and Ireland, and each assembly requires curation decisions that affect downstream analyses [<a href="#ref-1">1</a>]. When a project produces assemblies for hundreds or thousands of species, the choice between automated and manual curation has major implications for both budget and data quality.

The core problem is that no single curation strategy works for every genome. A small bacterial genome with a single circular chromosome presents different challenges than a large plant genome with high repeat content and polyploidy. A project aiming for a chromosome-level reference genome has different quality requirements than a project that only needs a draft assembly for gene content analysis. The decision framework in this article helps readers match curation effort to project needs.

## What Automated Curation Tools Actually Do

Automated curation tools perform specific, well-defined operations on genome assemblies. Understanding what these tools can and cannot do is essential for deciding when to use them.

### Purge Haplotigs and Related Tools

Purge_dups and similar tools identify and remove haplotypic duplications from primary assemblies. When a genome is heterozygous, the assembler may produce two versions of the same genomic region, one from each haplotype. These tools use read depth and sequence similarity to distinguish true duplications from haplotypic variants.

The practical effect is that automated tools can reduce assembly size and improve contiguity by removing redundant sequence. However, the tools make decisions based on statistical thresholds that may not be appropriate for all genomes. A highly polymorphic species may have haplotypic divergence that exceeds the tool's expectations, leading to incorrect removal of genuine sequence. Conversely, a recent whole-genome duplication may produce similar sequences that the tool incorrectly classifies as haplotypic.

### Hi-C Scaffolding

Hi-C data provides information about the three-dimensional organization of chromosomes in the nucleus. Automated scaffolding tools use this information to order and orient contigs into chromosomal pseudomolecules. The approach works well for genomes with clear Hi-C interaction patterns, where the contact frequency between loci decreases with genomic distance.

The limitation of automated Hi-C scaffolding is that it relies on the assumption that interaction patterns reflect linear genome organization. Structural rearrangements, repetitive regions, and assembly errors can disrupt this assumption. When the Hi-C signal is ambiguous, automated tools may place contigs incorrectly or fail to resolve complex regions.

### Quality Assessment Tools

Automated quality assessment tools such as BUSCO and QUAST provide metrics that help researchers evaluate assembly completeness and contiguity. These tools compare the assembly against expected gene content or measure statistics such as N50 and L50. They are valuable for identifying assemblies that need further work, but they do not fix problems.

The distinction between assessment and curation is important. Assessment tools tell you that a problem exists. Curation tools attempt to fix the problem. Both are useful, but they serve different purposes in the workflow.

## What Manual Curation Adds

Manual curation involves a researcher examining the assembly in detail, often using visualization tools to inspect alignments, coverage patterns, and Hi-C contact maps. The curator makes decisions about whether to break contigs, join contigs, remove sequence, or adjust scaffolding based on evidence that automated tools may not fully consider.

### Resolving Complex Regions

Automated tools struggle with regions that have unusual characteristics. Tandem repeats, segmental duplications, and regions with extreme GC content can produce patterns that confuse automated algorithms. A human curator can examine the underlying evidence and make informed decisions about how to handle these regions.

For example, a curator might notice that a particular contig has coverage that is consistently half the genome average, suggesting that it represents a haplotypic variant instead of a true duplication. The curator can then decide whether to remove the contig or keep it, based on the project's goals.

### Evaluating Biological Plausibility

Manual curation allows the researcher to consider biological context that automated tools cannot access. A curator who knows the species' karyotype can check whether the number of chromosomal pseudomolecules matches expectations. A curator who knows about a particular genomic feature, such as a sex chromosome or a satellite array, can verify that the assembly represents it correctly.

This biological knowledge is particularly valuable for genomes with unusual features. The tub gurnard assembly, for example, places most of haplotype 1 into 24 chromosomal pseudomolecules [<a href="#ref-2">2</a>]. A curator familiar with fish karyotypes could verify that this number is plausible for the species.

### Correcting Automated Errors

Manual curation also serves as a check on automated tools. When automated scaffolding produces a questionable result, the curator can examine the evidence and decide whether to accept the automated decision or override it. This verification step is important because automated tools can produce confident but incorrect results.

The cost of manual curation is time. A curator must examine the assembly, interpret evidence, and document decisions. For large or complex genomes, this process can take weeks or months. The decision framework in this article helps researchers determine when this investment is justified.

## At a Glance: Curation Decision Framework

The following table summarizes the key factors that should influence the decision between automated and manual curation.

| Genome Complexity | Project Goal | Recommended Approach | Primary Considerations |
| --- | --- | --- | --- |
| Low complexity, haploid or highly inbred | Draft assembly for gene content analysis | Automated only | Use purge_dups and quality metrics, manual review only if metrics fail |
| Moderate complexity, some heterozygosity | Chromosome-level reference for a well-studied species | Automated with targeted manual review | Use Hi-C scaffolding, then manually inspect chromosome boundaries and sex chromosomes |
| High complexity, polyploid or highly repetitive | Reference-quality assembly for a new species | Automated plus extensive manual curation | Expect multiple rounds of curation, budget for significant human time |
| Any complexity | Comparative genomics across many species | Automated with standardized quality gates | Use consistent automated pipelines, manual curation only for outliers |

The decision framework rests on two main axes: genome complexity and project goals. Genome complexity includes factors such as heterozygosity, repeat content, polyploidy, and genome size. Project goals include the intended uses of the assembly, the quality standards that must be met, and the resources available for curation.

## Genome Complexity as a Curation Driver

Genome complexity is the most important factor in determining how much manual curation is needed. Complex genomes produce more assembly errors, and those errors are more difficult for automated tools to resolve.

### Heterozygosity

Heterozygous genomes present a fundamental challenge to assembly. When the two haplotypes differ substantially, the assembler may produce separate assemblies for each haplotype. Automated tools can identify and remove haplotypic duplications, but they may struggle when heterozygosity is high or when the haplotypes differ in complex ways.

The Parastichtis suspecta assembly illustrates this challenge. The assembly contains two haplotypes with total lengths of 868.84 megabases and 869.52 megabases [<a href="#ref-1">1</a>]. Haplotype 1 was scaffolded into 30 chromosomal pseudomolecules, while haplotype 2 was assembled only to scaffold level. This difference in assembly quality between haplotypes reflects the difficulty of assembling both haplotypes to the same standard.

For highly heterozygous species, manual curation may be necessary to ensure that the primary assembly represents a single haplotype consistently. The curator can examine coverage patterns and sequence alignments to identify regions where haplotypic variation has caused assembly problems.

### Repeat Content

Repetitive sequences create ambiguity in assembly because reads from different copies of a repeat may be identical. Automated tools use read depth and linkage information to resolve this ambiguity, but the resolution is not always correct.

Highly repetitive genomes, such as those of many plants and some animals, may require extensive manual curation. The curator can use Hi-C data and other evidence to determine the correct placement of repeat copies. This process is time-consuming but can substantially improve assembly quality.

### Polyploidy

Polyploid genomes present a special challenge because they contain more than two copies of each chromosome. Automated tools designed for diploid genomes may not handle polyploid genomes correctly. The tools may collapse homoeologous chromosomes or produce chimeric assemblies that combine sequence from different subgenomes.

Manual curation of polyploid genomes requires specialized knowledge and tools. The curator must understand the species' ploidy and be able to distinguish homoeologous sequences from paralogous sequences. This is one of the most difficult curation scenarios.

### Genome Size

Genome size affects curation effort in two ways. Larger genomes require more computational resources for automated curation, and they produce more data that must be examined during manual curation. A curator can examine a small bacterial genome in a few hours, but a large plant genome may require weeks of work.

The Phaonia angelicae assembly, with haplotypes of 1,593.88 megabases and 1,575.57 megabases, represents a moderately large insect genome [<a href="#ref-3">3</a>]. The assembly placed 97.48% of haplotype 1 into 5 chromosomal pseudomolecules. Curation of this assembly required balancing the need for completeness against the practical limits of manual review.

## Project Goals and Quality Standards

The intended use of the assembly should determine the curation standard. A draft assembly may be sufficient for some purposes, while other purposes require a reference-quality assembly.

### Draft Assemblies for Gene Content Analysis

If the goal is to identify the gene content of a species, a draft assembly may be sufficient. The assembly needs to contain most of the coding sequence, but the ordering and orientation of contigs may not be critical. Automated curation with quality assessment may be enough for this purpose.

The quality gate for this use case is gene completeness. The researcher should check whether the assembly contains expected single-copy orthologs and whether gene models can be predicted reliably. If these checks pass, manual curation may not be necessary.

### Chromosome-Level References

A chromosome-level reference genome requires that most of the sequence be placed into chromosomal pseudomolecules. This standard requires Hi-C scaffolding or equivalent methods, and it often requires manual curation to resolve problematic regions.

The quality gate for this use case is the proportion of sequence in chromosomal pseudomolecules and the correctness of the chromosome structures. The curator should verify that chromosome boundaries are correct and that no large-scale misjoins are present.

### Reference-Quality Assemblies

Reference-quality assemblies meet the highest standards of contiguity, completeness, and correctness. These assemblies are intended to serve as the definitive genomic reference for a species, and they require extensive curation.

The quality gate for this use case includes all the checks for chromosome-level assemblies plus additional verification of complex regions. The curator should examine repetitive regions, segmental duplications, and other difficult features to ensure that they are represented correctly.

### Comparative Genomics Projects

Projects that assemble many species for comparative analysis face a different challenge. The goal is consistency across species, not necessarily the highest quality for each individual assembly. Automated curation with standardized quality gates may be the most appropriate approach.

The quality gate for this use case is the consistency of the assembly metrics across species. The researcher should establish thresholds for completeness and contiguity and apply them uniformly. Manual curation may be reserved for assemblies that fail these thresholds.

## Practical Workflow for Curation Decisions

The following workflow provides a structured approach to deciding when to invest in manual curation.

### Step 1: Assess Genome Complexity Before Assembly

Genome complexity can be estimated before assembly from read data and biological knowledge. The researcher should consider the species' ploidy, heterozygosity, repeat content, and genome size. This assessment informs the choice of assembly strategy and the expected curation burden.

The researcher should also consider what is known about the species' karyotype and genomic features. This knowledge helps set expectations for the assembly and provides a basis for later curation decisions.

### Step 2: Run Automated Curation and Quality Assessment

After the initial assembly, run automated curation tools to remove haplotypic duplications and scaffold the assembly. Then run quality assessment tools to measure completeness and contiguity.

The results of these automated steps provide the first indication of whether manual curation is needed. If the quality metrics meet the project's standards, manual curation may be optional. If the metrics fall short, manual curation is likely necessary.

### Step 3: Examine Quality Metrics Against Project Standards

Compare the quality metrics against the standards established for the project. The comparison should consider the intended use of the assembly and the resources available for curation.

The researcher should document the quality metrics and the standards against which they are being compared. This documentation provides a record of the curation decision and its rationale.

### Step 4: Identify Problem Regions for Manual Review

If manual curation is needed, the first task is to identify the regions that require review. These regions may include assembly breaks, questionable joins, regions with abnormal coverage, and areas where the assembly conflicts with biological expectations.

The researcher should create a list of problem regions and prioritize them based on their likely impact on assembly quality. This prioritization helps allocate curation effort efficiently.

### Step 5: Perform Manual Curation and Document Decisions

Manual curation involves examining each problem region in detail and making decisions about how to resolve it. The curator should document each decision, including the evidence that supported it and the action taken.

This documentation is important for reproducibility. Other researchers should be able to understand why the assembly has its current structure and what evidence supports the curation decisions.

### Step 6: Reassess Quality After Curation

After manual curation, run the quality assessment tools again to measure the improvement. This reassessment confirms that the curation effort produced the expected benefits and identifies any remaining issues.

The reassessment also provides a record of the curation process. The before and after metrics document the value of the manual curation effort.

## Records and Measurements for Curation Decisions

Curation decisions should be based on measurable evidence, not intuition. The following records and measurements support informed decisions.

### Assembly Quality Metrics

Standard assembly quality metrics include N50, L50, and the proportion of sequence in chromosomal pseudomolecules. These metrics provide a quantitative basis for comparing assemblies and tracking improvement.

The researcher should record these metrics at each stage of the curation process. The records show how the assembly changed in response to curation actions.

### Completeness Scores

Completeness scores measure the proportion of expected genes or genomic features present in the assembly. These scores provide a functional measure of assembly quality that complements the structural metrics.

The researcher should record completeness scores before and after curation. The comparison shows whether curation improved the assembly's functional completeness.

### Coverage Analysis

Coverage analysis examines the depth of read coverage across the assembly. Regions with abnormal coverage may indicate assembly errors, such as collapsed repeats or haplotypic duplications.

The researcher should record coverage statistics for the assembly and for specific regions of interest. These records help identify problem regions and support curation decisions.

### Hi-C Contact Maps

Hi-C contact maps show the interaction frequency between genomic loci. These maps provide evidence for the linear organization of chromosomes and can identify misjoins and other structural errors.

The researcher should preserve the Hi-C contact maps used during curation. These maps document the evidence that supported scaffolding decisions.

### Curation Log

A curation log records each action taken during manual curation. The log should include the region examined, the evidence considered, the decision made, and the rationale for the decision.

The curation log is an essential record for reproducibility. It allows other researchers to understand the curation process and to revisit decisions if new evidence becomes available.

## Common Failure Patterns in Automated Curation

Automated curation tools fail in predictable ways. Recognizing these failure patterns helps researchers decide when manual review is necessary.

### Overpurge of Haplotypic Sequence

Purge tools may remove sequence that is not actually haplotypic. This failure occurs when the tool's thresholds are too aggressive for the genome being processed. The result is an assembly that is missing genuine sequence.

The warning sign for this failure is a completeness score that is lower than expected. If the assembly is missing genes that should be present, the purge may have removed too much sequence.

### Underpurge of Haplotypic Sequence

The opposite failure occurs when purge tools fail to remove haplotypic duplications. The result is an assembly that contains redundant sequence, inflating the assembly size and potentially confusing downstream analyses.

The warning sign for this failure is an assembly that is larger than expected and contains regions with half the average coverage. These regions likely represent haplotypic variants that should have been removed.

### Misjoins in Hi-C Scaffolding

Hi-C scaffolding can produce incorrect joins when the interaction data are ambiguous. The result is a chromosome structure that does not reflect the true genome organization.

The warning sign for this failure is a Hi-C contact map that shows abrupt changes in interaction patterns at the join point. The curator should examine these regions carefully to determine whether the join is correct.

### Collapsed Repeats

Automated assembly may collapse multiple copies of a repeat into a single sequence. The result is an assembly that underestimates the repeat content of the genome.

The warning sign for this failure is a region with abnormally high coverage. The high coverage suggests that reads from multiple repeat copies have been mapped to the same location.

### Chimeric Contigs

Contigs may contain sequence from different genomic locations, creating chimeric structures. These errors are difficult for automated tools to detect because the sequence may be internally consistent.

The warning sign for this failure is a contig that shows inconsistent coverage or that maps to multiple locations in the Hi-C contact map. Manual examination is often required to resolve these cases.

## Limitations of Automated Curation

Automated curation tools have inherent limitations that cannot be overcome by parameter tuning or better algorithms. Understanding these limitations is essential for deciding when manual curation is necessary.

### Lack of Biological Context

Automated tools do not know the biology of the species being assembled. They cannot use knowledge of the karyotype, sex chromosome system, or genomic features to guide their decisions.

This limitation is particularly important for species with unusual genomic features. A curator who knows that a species has a particular chromosome arrangement can verify that the assembly reflects this arrangement. An automated tool cannot make this check.

### Statistical Assumptions

Automated tools make statistical assumptions about the data. These assumptions may not hold for all genomes, leading to incorrect decisions.

For example, purge tools assume that haplotypic divergence follows a particular distribution. If the actual divergence is higher or lower than expected, the tool may make incorrect classifications. The researcher should be aware of these assumptions and their potential impact.

### Inability to Resolve Ambiguity

Automated tools must make a decision even when the evidence is ambiguous. The tools do not have the ability to defer a decision or to seek additional evidence.

A human curator can examine ambiguous regions in detail, consider multiple lines of evidence, and make a judgment about the most likely correct structure. This judgment is particularly valuable for complex regions where automated tools are likely to make errors.

### No Learning from Context

Automated tools process each assembly independently. They do not learn from the results of previous assemblies or from the curation decisions made by humans.

A human curator accumulates experience across projects. This experience informs decisions about new assemblies and helps the curator recognize patterns that automated tools miss.

## When Manual Curation Is Essential

Manual curation is not optional for certain types of projects. The following situations require human review.

### New Reference Genomes

A new reference genome for a species will be used by many researchers for many purposes. The quality of this assembly has long-term consequences for the research community. Manual curation is essential to ensure that the assembly meets the highest standards.

The cost of manual curation is justified by the value of the reference genome. A well-curated reference genome enables accurate downstream analyses for years or decades.

### Genomes with Known Complex Features

If the species is known to have complex genomic features, such as large segmental duplications, extensive heterozygosity, or unusual sex chromosomes, manual curation is likely necessary. These features create assembly challenges that automated tools may not resolve correctly.

The researcher should identify these features before assembly and plan for the curation effort they will require.

### Comparative Genomics with High Stakes

Comparative genomics projects that will be used for evolutionary inference or clinical applications require high-quality assemblies. Errors in individual assemblies can propagate through comparative analyses and produce incorrect conclusions.

The researcher should establish quality gates that trigger manual curation when assemblies fail to meet the required standards.

### Assemblies for Regulatory or Clinical Use

Assemblies used for regulatory or clinical purposes must meet rigorous quality standards. The consequences of errors are potentially severe, and the curation process must be documented thoroughly.

Manual curation provides the human oversight that regulatory and clinical applications require. The curation log provides the documentation that demonstrates the quality of the assembly.

## Cost Considerations for Curation Decisions

The cost of curation includes both the direct cost of researcher time and the indirect cost of delayed results. The decision framework should consider both types of cost.

### Direct Costs of Manual Curation

Manual curation requires skilled personnel who understand genome assembly and the species being studied. These personnel are expensive, and their time is limited.

The researcher should estimate the time required for manual curation based on the genome complexity and the quality standards. This estimate should include time for examining problem regions, making decisions, and documenting the process.

### Indirect Costs of Delayed Results

Manual curation delays the availability of the assembly. This delay may affect downstream analyses, publications, or applications that depend on the assembly.

The researcher should consider whether the delay is acceptable given the project timeline. In some cases, a draft assembly may be sufficient for immediate needs, with manual curation deferred until later.

### Costs of Inadequate Curation

The cost of inadequate curation includes the cost of incorrect downstream analyses and the cost of redoing the assembly. These costs can exceed the cost of manual curation.

The researcher should weigh the cost of manual curation against the potential cost of errors. For high-stakes projects, the balance favors manual curation.

### Budgeting for Curation

Curation should be included in the project budget from the beginning. The budget should account for both automated curation tools and manual curation time.

The researcher should document the curation budget and track actual expenditures. This record provides a basis for estimating curation costs in future projects.

## Training and Expertise for Manual Curation

Manual curation requires specific skills that are not always taught in standard bioinformatics training. Researchers who plan to perform manual curation should ensure that they have the necessary expertise.

### Foundational Bioinformatics Skills

Manual curation requires comfort with command-line tools, scripting, and data manipulation. These skills are taught in programs such as The Carpentries lessons, which provide foundational computing and data skills [<a href="#ref-4">4</a>].

The researcher should be able to work with common bioinformatics file formats, run analysis tools, and interpret the results. These skills are prerequisites for effective manual curation.

### Genome Assembly and Curation Training

Specialized training in genome assembly and curation is available through resources such as the Galaxy Training Network, which provides accessible workflow training and analysis tutorials [<a href="#ref-1">1</a>]. The EMBL-EBI Training program also offers bioinformatics learning pathways and practical analysis education [<a href="#ref-5">5</a>].

The researcher should seek training that covers the specific tools and approaches used in their curation workflow. This training should include hands-on practice with real datasets.

### Reproducible Workflow Skills

Manual curation should be conducted within a reproducible workflow. The researcher should use version control, document their commands, and maintain records of their decisions.

Resources such as the nf-core documentation provide guidance on community pipeline standards and reproducible workflow context [<a href="#ref-6">6</a>]. The Bioconductor project offers official package and workflow documentation for reproducible genomic analysis [<a href="#ref-7">7</a>].

### Domain Knowledge

Manual curation requires knowledge of the species being assembled. The curator should understand the species' biology, including its karyotype, reproductive system, and genomic features.

This domain knowledge is essential for evaluating the biological plausibility of the assembly. The curator should consult with experts in the species or taxonomic group when needed.

## Quality Control During Manual Curation

Manual curation should be subject to quality control, just like any other scientific process. The following practices help ensure that curation decisions are sound.

### Independent Verification

Curation decisions should be verified by an independent reviewer when possible. The reviewer can check the evidence and confirm that the decision is justified.

Independent verification is particularly important for decisions that have a major impact on the assembly, such as breaking a contig or removing sequence. The reviewer provides a check on the curator's judgment.

### Documentation of Evidence

Each curation decision should be documented with the evidence that supported it. The documentation should include the specific data examined and the reasoning behind the decision.

This documentation serves multiple purposes. It provides a record for reproducibility, it supports the quality of the assembly, and it helps other researchers understand the assembly structure.

### Consistency Checks

The curator should perform consistency checks across the assembly. For example, the curator should verify that the chromosome count matches the expected karyotype and that the assembly is consistent with known genomic features.

Consistency checks help identify errors that may not be apparent from examining individual regions. They provide a holistic view of assembly quality.

### Iterative Review

Manual curation is often an iterative process. The curator may need to revisit regions after examining other parts of the assembly or after obtaining new evidence.

The curator should plan for multiple rounds of review. Each round should focus on the most important remaining issues.

## Common Failure Patterns in Manual Curation

Manual curation is not infallible. The following failure patterns are common and should be recognized.

### Overcorrection

The curator may make changes that are not supported by the evidence. This overcorrection can introduce errors into the assembly.

The warning sign for this failure is a change that is not clearly justified by the evidence. The curator should be conservative and only make changes when the evidence is strong.

### Inconsistent Decisions

The curator may apply different standards to different regions of the assembly. This inconsistency can produce an assembly with uneven quality.

The warning sign for this failure is a pattern of decisions that varies across the assembly. The curator should apply consistent standards throughout.

### Failure to Document

The curator may make decisions without documenting the evidence and rationale. This failure undermines the reproducibility of the curation process.

The warning sign for this failure is a curation log with missing entries or vague descriptions. The curator should document every decision, no matter how minor.

### Anchoring on Initial Assembly

The curator may be influenced by the initial assembly structure and fail to consider alternative arrangements. This anchoring can prevent the curator from identifying errors in the initial assembly.

The warning sign for this failure is a curation process that makes only minor changes to the initial assembly. The curator should be willing to consider major structural changes when the evidence supports them.

## Safety and Regulatory Context for Curation Decisions

Genome assembly curation has implications for safety and regulatory compliance in certain contexts. The researcher should be aware of these implications.

### Data Quality for Clinical Applications

Assemblies used in clinical applications must meet rigorous quality standards. Errors in the assembly can lead to incorrect diagnoses or treatment decisions.

The researcher should ensure that the curation process meets the quality standards required for clinical use. This may require additional verification and documentation.

### Synthetic Biology Applications

Genome assemblies are used in synthetic biology to design and construct novel biological systems [<a href="#ref-8">8</a>]. The accuracy of these assemblies affects the safety and efficacy of engineered organisms.

The researcher should ensure that assemblies used in synthetic biology applications are curated to the highest standards. Errors in the assembly could lead to unintended consequences in engineered organisms.

### Data Sharing and Reproducibility

Many funding agencies and journals require that genome assemblies be deposited in public databases. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-9">9</a>].

The researcher should ensure that the assembly and its curation documentation meet the requirements for data sharing. This includes providing sufficient metadata and documentation for other researchers to understand and reproduce the curation process.

### Community Resources

Some genome projects contribute to community resources that require consistent curation standards. For example, BlastoDB is a community-driven resource that integrates multi-omics datasets and reference sequences for Blastocystis subtypes [<a href="#ref-10">10</a>].

The researcher should be aware of community standards for data curation and ensure that their assembly meets these standards. This may require following specific protocols for data submission and documentation.

## Professional Escalation Criteria

Researchers should know when to escalate curation decisions to more experienced colleagues or to seek additional expertise. The following situations warrant escalation.

### Unresolvable Assembly Conflicts

If the curator cannot resolve a conflict between different lines of evidence, the decision should be escalated. A more experienced curator may be able to interpret the evidence differently or may know of additional data that can resolve the conflict.

The escalation should include a clear description of the conflict and the evidence that has been considered. The more experienced curator can then make an informed decision.

### Unusual Genomic Features

If the assembly contains features that the curator does not understand, the decision should be escalated to a specialist. The specialist may have knowledge of similar features in related species.

The escalation should include the location of the feature and the evidence that has been gathered. The specialist can then determine whether the feature is real or an assembly artifact.

### High-Impact Decisions

Decisions that have a major impact on the assembly should be escalated for review. These decisions include breaking large contigs, removing substantial sequence, or making major changes to chromosome structures.

The escalation should include a description of the proposed change and the evidence supporting it. The reviewer can then confirm or reject the proposed change.

### Quality Metric Failures

If the assembly fails to meet quality standards after multiple rounds of curation, the decision should be escalated. The escalation may result in a decision to restart the assembly with different parameters or to accept a lower quality standard.

The escalation should include the quality metrics and the curation history. The decision maker can then determine the best course of action.

## Decision Documentation and Reporting

Curation decisions should be documented and reported in a way that supports reproducibility and transparency.

### Curation Report Structure

A curation report should include the assembly version, the curation actions taken, the evidence for each action, and the quality metrics before and after curation. The report should be written so that another researcher can understand the curation process.

The report should also include any unresolved issues and recommendations for future work. This information helps other researchers understand the limitations of the assembly.

### Data Availability

The curation report and supporting data should be made available to other researchers. This may include depositing the assembly in a public database and providing the curation log as supplementary material.

The NCBI provides resources for depositing and accessing sequence data [<a href="#ref-9">9</a>]. The researcher should follow the appropriate procedures for data submission.

### Publication Standards

Many journals require that genome assembly papers include information about the curation process. The researcher should be prepared to describe the curation methods and to justify the curation decisions.

The researcher should also be prepared to respond to reviewer questions about the curation process. A thorough curation log supports this response.

## Frequently Asked Questions

### What is the difference between automated and manual curation of genome assemblies?

Automated curation uses software tools to identify and correct assembly errors based on statistical models and algorithmic rules. These tools can remove haplotypic duplications, scaffold contigs using Hi-C data, and assess assembly quality. Manual curation involves a researcher examining the assembly in detail, interpreting evidence, and making decisions that require biological knowledge and judgment. Manual curation can resolve issues that automated tools miss, but it requires substantial time and expertise.

### How do I know if my genome assembly needs manual curation?

The need for manual curation depends on the complexity of the genome and the quality standards for the project. If the assembly quality metrics meet the project's standards and the genome is not highly complex, automated curation may be sufficient. If the quality metrics fall short, or if the genome has high heterozygosity, high repeat content, or polyploidy, manual curation is likely necessary. The decision should be based on measurable evidence, not intuition.

### What are the most common errors that automated curation tools miss?

Automated tools commonly miss collapsed repeats, chimeric contigs, and misjoins in repetitive regions. These errors occur because the tools rely on statistical patterns that break down in complex genomic regions. Manual curation can identify these errors by examining coverage patterns, Hi-C contact maps, and other evidence that automated tools do not fully consider.

### How much time does manual curation typically require?

The time required for manual curation varies widely depending on genome complexity and quality standards. A small bacterial genome may require only a few hours of review, while a large, complex eukaryotic genome may require weeks or months. The researcher should estimate the curation time based on the specific characteristics of the genome and the project's quality requirements.

### Can automated curation tools be improved by adjusting parameters?

Parameter adjustment can improve the performance of automated tools for specific genomes, but it cannot overcome the fundamental limitations of the tools. The tools still lack biological context and cannot resolve all ambiguity. Parameter adjustment should be seen as a way to optimize automated curation, not as a substitute for manual review when it is needed.

### What training do I need to perform manual curation?

Manual curation requires foundational bioinformatics skills, including command-line proficiency and data manipulation. Training resources such as The Carpentries lessons provide foundational computing and data skills [<a href="#ref-4">4</a>]. Specialized training in genome assembly and curation is available through resources such as the Galaxy Training Network [<a href="#ref-1">1</a>] and EMBL-EBI Training [<a href="#ref-5">5</a>]. Domain knowledge about the species being assembled is also essential.

### How should I document my curation decisions?

Each curation decision should be documented with the evidence that supported it and the action taken. The documentation should include the specific data examined, the reasoning behind the decision, and the expected impact on the assembly. This documentation supports reproducibility and allows other researchers to understand the assembly structure.

### When should I escalate a curation decision to a more experienced colleague?

Escalate a curation decision when you cannot resolve a conflict between lines of evidence, when the assembly contains features you do not understand, when a decision has a major impact on the assembly, or when the assembly fails to meet quality standards after multiple rounds of curation. The escalation should include a clear description of the issue and the evidence that has been considered.

## Related Bioinformatics Guides

- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data](/knowledge/bioinformatics/evaluating-metagenomic-assembly-tools-a-benchmarking-framework-for-short-read-and-long-read-data)
- [Human Reference Genomes: Navigating and Downloading hg38 and hg19 FASTA Assemblies](/knowledge/bioinformatics/human-reference-genomes-hg38-hg19-assemblies)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [The genome sequence of the Suspected, &lt,i&gt,Parastichtis suspecta&lt,/i&gt, (Hübner, 1809) (Lepidoptera: Noctuidae).](https://doi.org/10.12688/wellcomeopenres.26587.1). 2026.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [The genome sequence of the tub gurnard, &lt,i&gt,Chelidonichthys lucerna&lt,/i&gt, (Linnaeus, 1758) (Perciformes: Triglidae).](https://doi.org/10.12688/wellcomeopenres.25417.2). 2026.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [The genome sequence of a muscid fly, &lt,i&gt,Phaonia angelicae&lt,/i&gt, (Scopoli, 1763) (Diptera: Muscidae).](https://doi.org/10.12688/wellcomeopenres.26330.1). 2026.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Engineered bacteria in disease diagnosis and therapy: A synthetic biology perspective.](https://doi.org/10.1016/j.synbio.2026.04.028). 2026.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [BlastoDB: first release of a community-driven multi-omics and epidemiological resource for &lt,i&gt,Blastocystis&lt,/i&gt, biology and subtyping.](https://doi.org/10.12688/openreseurope.23235.2). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.