# Manual Curation of Genome Assemblies: A Step-by-Step Workflow for Resolving Misjoins and Improving Contiguity


## Key Takeaways

- Manual genome assembly curation is essential for resolving structural errors like misjoins and collapsed repeats, particularly when downstream applications demand high accuracy, such as comparative genomics or structural variation studies.
- Initial quality assessment involves computing assembly statistics (e.g., N50, L50) using tools like `gfastats` and evaluating gene completeness with BUSCO, establishing a baseline for measuring curation impact.
- Assembly graphs visualized with tools like Bandage are critical for identifying misjoins, characterized by anomalous node connectivity and coverage patterns, and breakpoints are confirmed by analyzing read pair orientations and long read alignments.
- Correction strategies include splitting misjoined contigs using tools like `gfastats` and rejoining sequences based on graph topology and read support, with Hi-C data providing long-range evidence for scaffolding validation.
- Collapsed repeats, indicated by abnormally high coverage, require reconstruction of individual copies, while false gene loss is addressed by identifying regions with read support but absent assembly sequence and integrating missing segments.
- Rigorous validation post-curation involves recomputing assembly statistics, remapping reads to assess coverage uniformity and mapping rates, and re-running BUSCO to confirm improvements in gene completeness.

---

Automated genome assembly pipelines produce draft sequences that frequently contain structural errors, including misjoins where unrelated genomic regions are incorrectly concatenated, collapsed repeats, and unnecessary fragmentation. Manual curation is the process of inspecting assembly graphs, validating contig boundaries, and applying targeted corrections to resolve these errors. This article provides a structured workflow for researchers who need to move beyond automated outputs and produce assemblies that accurately represent the underlying genome. The workflow covers initial quality assessment, graph visualization, breakpoint identification, sequence correction, and validation, with concrete decision criteria at each stage.

## Scope and Prerequisites for Manual Curation

Manual curation is warranted when downstream applications depend on accurate genome structure. Comparative genomics, gene family analysis, and studies of structural variation all require assemblies where contig order and orientation reflect biological reality. Automated pipelines often cannot unambiguously resolve a bacterial genome, for example due to the presence of sequence repeat structures on the chromosome or on plasmids, and a more sophisticated approach or manual curation is needed in those cases [<a href="#ref-1">1</a>]. The same principle applies to eukaryotic genomes, where repetitive content and polyploidy create assembly challenges that algorithms may not fully resolve.

Before beginning manual curation, confirm that you have the necessary inputs and skills. You need the draft assembly in FASTA or FASTQ format, the raw sequencing reads used for the assembly, and ideally an assembly graph file in GFA format. You also need familiarity with command-line tools and basic scripting. Foundational training in shell computing, data handling, and version control is available through The Carpentries lessons, which provide structured instruction for researchers who need to strengthen these skills [<a href="#ref-2">2</a>]. If you are new to genome assembly analysis, the Galaxy Training Network offers accessible workflow tutorials that cover quality assessment and assembly evaluation in a reproducible environment [<a href="#ref-3">3</a>].

The time investment for manual curation varies substantially with genome size and complexity. A small bacterial genome with a few contigs may require hours of work, while a large eukaryotic genome with hundreds of scaffolds can require days or weeks. Plan your curation effort according to the biological questions you need to answer. If you only need gene presence or absence calls, targeted curation of specific loci may suffice. If you need chromosome-level structure, comprehensive curation of all scaffolds is necessary.

## At a Glance

| Curation Stage | Primary Tools | Key Evidence Sources | Typical Duration |
| --- | --- | --- | --- |
| Initial quality assessment | gfastats, BUSCO, read mappers | Assembly statistics, completeness scores, coverage profiles | 1 to 3 hours |
| Graph visualization and misjoin detection | Bandage, IGV | Graph topology, node connectivity, coverage patterns | 2 to 8 hours |
| Breakpoint confirmation and correction | gfastats, read mappers, Hi-C tools | Read pair orientations, long read spans, contact maps | 4 to 16 hours per complex region |
| Validation and reporting | gfastats, BUSCO, version control | Updated statistics, remapping rates, completeness scores | 2 to 4 hours |

## Understanding Assembly Graphs and Misjoins

Assembly graphs represent the relationships between sequence contigs. In a correct assembly, the graph reflects the true path of the genome through sequence space. Misjoins occur when the graph contains connections that do not exist in the biological genome, often because repetitive sequences create ambiguous paths that the assembler resolves incorrectly.

### Graph Representations and File Formats

The Graphical Fragment Assembly format, commonly abbreviated as GFA, stores assembly graphs in a text-based format that includes sequence nodes and the edges connecting them. Tools that manipulate assembly graphs, including gfastats, store assembly sequences internally in a GFA-like format, which allows seamless conversion between FASTA, FASTQ, and GFA files [<a href="#ref-4">4</a>]. Understanding this format is essential because many curation operations involve converting between sequence-only files and graph-aware representations.

Bandage is the primary visualization tool for assembly graphs. It displays the graph as a network where nodes represent sequences and edges represent connections supported by read data. When you open a GFA file in Bandage, you can see the overall graph structure, identify regions of complexity, and inspect the sequence at specific nodes. This visual inspection is the first step in identifying potential misjoins.

### Common Structural Errors in Draft Assemblies

Misjoins typically appear in assembly graphs as nodes with unexpected connectivity patterns. A contig that connects to two regions that should be distant in the genome suggests a misjoin. Collapsed repeats appear as nodes with unusually high coverage relative to surrounding sequence, because reads from multiple genomic locations map to the same assembled sequence. False gene loss occurs when single-copy sequences are incorrectly removed during assembly, often because they are mistaken for repeats [<a href="#ref-1">1</a>].

Inter-plasmidic repeat collapse is a specific error type observed in bacterial hybrid assemblies where plasmids share repetitive sequences. The Unicycler assembly pipeline, which combines short and long reads, can collapse these inter-plasmidic repeats, producing a single sequence that represents multiple distinct plasmids [<a href="#ref-1">1</a>]. Detecting this error requires comparing the expected plasmid content with the assembled output and inspecting coverage patterns across the affected regions.

## Initial Quality Assessment Before Curation

Manual curation should begin with a systematic assessment of assembly quality. This assessment establishes a baseline against which you can measure the impact of your curation efforts and identifies the regions most likely to contain errors.

### Computing Assembly Statistics

Assembly statistics provide a quantitative summary of contiguity and completeness. The N50 statistic, which represents the contig length at which half the assembly is contained in contigs of that length or longer, is a standard metric. The L50 statistic indicates how many contigs are needed to reach that half-length threshold. The number of contigs and the total assembly length are also essential baseline measurements.

The gfastats tool computes assembly summary statistics and manipulates assembly sequences in FASTA, FASTQ, or GFA format [<a href="#ref-4">4</a>]. Run gfastats on your draft assembly before curation to establish baseline metrics. Record the total length, number of contigs, N50, L50, and the length of the longest contig. These values will serve as your reference point for evaluating whether curation improves the assembly.

### Evaluating Completeness with BUSCO

Completeness assessment measures whether the assembly contains expected conserved genes. BUSCO, which stands for Benchmarking Universal Single-Copy Orthologs, compares the assembly against a set of single-copy genes expected to be present in the target taxonomic group. The output reports the percentage of complete, fragmented, and missing BUSCO genes.

Genome completeness is a key quality indicator. In a study of nine insect genome assemblies, BUSCO scores ranged from 85.5 percent completeness for the largest genome to 98.8 percent completeness for the smallest genome [<a href="#ref-5">5</a>]. This range demonstrates that completeness varies with genome characteristics and assembly strategy. Low BUSCO completeness may indicate collapsed repeats or missing sequence, both of which manual curation can address.

### Mapping Reads Back to the Assembly

Read mapping provides evidence for assembly correctness. When you map the original sequencing reads back to the draft assembly, regions with abnormally high or low coverage indicate potential problems. High coverage may indicate collapsed repeats where reads from multiple genomic locations map to one assembled sequence. Low coverage may indicate misassembled regions where reads cannot map correctly.

Visualize read coverage along the assembly using a genome browser such as IGV. Coverage plots that show abrupt transitions from high to low coverage at a single position often indicate misjoins. Read pairs that map with unexpected orientation or insert size also signal assembly errors. For example, read pairs that should map in a forward-reverse orientation but map in a reverse-forward orientation suggest that the assembly has inverted a region relative to the true genome.

## Visualizing the Assembly Graph with Bandage

Bandage provides the graphical interface for inspecting assembly graphs and identifying regions that require manual intervention. The tool displays nodes as rectangles or circles, with edges connecting nodes that have read support. Node size can represent sequence length, and node color can represent coverage depth.

### Loading and Navigating the Graph

Load your assembly graph into Bandage by opening the GFA file. If your assembler produced only FASTA output, you may need to generate a graph file using a tool that can build an assembly graph from sequences. Gfastats can build an assembly graph that can be used to manipulate the underlying sequences following instructions provided by the user [<a href="#ref-4">4</a>]. This capability allows you to create graph representations even when your original assembly pipeline did not produce them.

Once the graph is loaded, navigate to regions of interest. Bandage allows you to search for specific node sequences, zoom into complex regions, and highlight nodes based on coverage or length thresholds. Develop a systematic approach to graph inspection. Start with the largest nodes, which represent the most significant portions of the assembly, and examine their connections. Then move to nodes with unusual coverage patterns, which may indicate collapsed repeats.

### Identifying Misjoin Candidates in the Graph

Misjoins produce characteristic graph patterns. A node that connects to two or more nodes that should be genomically distant suggests that the assembler incorrectly joined unrelated sequences. In a correct assembly graph, the path through the graph should be linear for haploid genomes, with branches only at repetitive regions.

Look for nodes where the graph branches and then rejoins. These structures may indicate alternative paths through repetitive sequence, where the assembler chose one path but the true genome follows another. Nodes with coverage that is substantially higher than the surrounding sequence are candidates for collapsed repeats. Nodes with coverage that drops to near zero at one end may indicate a misjoin where the assembler connected a real sequence to an artifact.

### Documenting Candidate Regions

Create a curation log that records each candidate region you identify. For each candidate, note the node identifier, the genomic coordinates, the coverage pattern, and your hypothesis about the error type. This documentation is essential for tracking your curation progress and for reporting the changes you make. The log also provides a record that other researchers can review to understand the curation decisions.

## Identifying Breakpoints with Read Support

Breakpoint identification is the process of determining the exact position where a misjoin occurs. This step requires integrating information from the assembly graph, read mapping, and sequence characteristics.

### Using Coverage Transitions to Locate Breakpoints

Coverage transitions provide the first indication of breakpoint locations. When you map reads to the assembly and visualize coverage, a misjoin often appears as an abrupt change in coverage depth. The position where coverage changes from high to low, or from uniform to variable, is a candidate breakpoint.

In IGV, zoom into the region around the coverage transition and examine the read alignments. Look for reads that span the transition point. If no reads span the position, the assembler likely joined two sequences that are not connected in the genome. If reads span the position but with mismatches or indels, the join may be correct but contain sequencing errors.

### Analyzing Read Pair Orientations

Read pair information provides powerful evidence for breakpoint identification. In a correct assembly, read pairs map with expected orientations and insert sizes. A cluster of read pairs that map with unexpected orientations at a specific position indicates a misjoin.

For example, if read pairs on the left side of a position map in the expected orientation but read pairs on the right side map with inverted orientation, the assembly likely contains an inversion at that position. If read pairs map with insert sizes that are much larger than expected, the assembly may have deleted sequence between the paired reads. Document the read pair patterns you observe for each candidate breakpoint.

### Confirming Breakpoints with Long Reads

Long reads provide the strongest evidence for breakpoint confirmation. A single long read that spans a candidate breakpoint with high identity confirms that the join is correct. Conversely, long reads that end at a candidate breakpoint or that map to different genomic regions on either side of the breakpoint confirm a misjoin.

If your assembly used long reads, map the raw long reads back to the draft assembly and examine the alignments at candidate breakpoints. Long reads that span the breakpoint with no mismatches provide definitive evidence. Long reads that split at the breakpoint, with one portion mapping to one side and the other portion mapping elsewhere, confirm the misjoin.

## Correcting Misjoins and Splitting Contigs

Once you have identified and confirmed misjoins, the next step is to correct them. The correction process involves splitting the misjoined contig at the breakpoint and then determining the correct connections for the resulting sequences.

### Splitting Contigs at Breakpoints

Splitting a contig at a breakpoint produces two sequences that correspond to the true genomic regions. The gfastats tool can manipulate assembly sequences following instructions provided by the user, which includes splitting sequences at specified positions [<a href="#ref-4">4</a>]. When you split a contig, you must decide whether to keep both resulting sequences or discard one.

If the misjoin connected two genuine genomic regions, keep both sequences after splitting. If the misjoin connected a genuine region to an artifact, such as a contaminant or an assembly error, discard the artifact sequence. Use coverage and read support to make this decision. Genuine regions have consistent read coverage and read pairs that support their existence. Artifacts often have low coverage or reads that map with poor identity.

### Rejoining Sequences with Correct Connections

After splitting, you need to determine the correct connections for the resulting sequences. This determination requires examining the assembly graph for alternative paths and using read support to choose the correct path. In some cases, the correct connection is already present in the graph but was not chosen by the assembler. In other cases, you need to use additional evidence, such as Hi-C contact data, to determine the correct order and orientation.

Hi-C data provides proximity information that can resolve scaffolding questions. The Puzzler pipeline automates Hi-C-based scaffolding and can generate input files for manual Hi-C curation [<a href="#ref-6">6</a>]. If you have Hi-C data available, use contact maps to validate that your corrected connections are consistent with the three-dimensional organization of the genome.

### Handling Ambiguous Cases

Some misjoins cannot be resolved with certainty using the available data. In these cases, document the ambiguity and make a conservative decision. If you cannot determine the correct connection, leave the sequence as a separate contig instead of joining it to an uncertain partner. An assembly with more contigs but correct sequence is preferable to an assembly with fewer contigs but incorrect joins.

Record ambiguous cases in your curation log with a clear description of the evidence and the decision you made. This documentation allows other researchers to revisit the case if additional data become available.

## Resolving Collapsed Repeats and False Gene Loss

Collapsed repeats and false gene loss are related error types that affect assembly completeness and accuracy. Both errors arise from the difficulty of assembling repetitive sequence, but they require different correction strategies.

### Detecting Collapsed Repeats

Collapsed repeats appear as regions with abnormally high coverage. When two or more genomic regions share identical or nearly identical sequence, the assembler may produce a single sequence that represents all copies. Reads from all copies map to this single sequence, producing coverage that is two or more times the average genome coverage.

To detect collapsed repeats, calculate the average coverage across the assembly and then identify regions with coverage substantially above this average. In a diploid genome, a collapsed repeat that represents two copies will have approximately twice the average coverage. In a genome with more copies, the coverage increase is correspondingly higher. The inter-plasmidic repeat collapse observed in Unicycler hybrid assemblies is a specific example where plasmids sharing repetitive sequence are collapsed into a single representation [<a href="#ref-1">1</a>].

### Reconstructing Collapsed Repeat Copies

Reconstructing collapsed repeats requires determining the number of copies and the flanking sequence for each copy. Long reads that span the repeat and extend into unique flanking sequence can resolve the copy number and the genomic context of each copy. If long reads are not available, you may need to use coverage estimates and comparative information to infer the copy number.

The correction process involves replacing the collapsed sequence with the appropriate number of copies, each with its correct flanking sequence. This process is complex and may require manual sequence construction. In some cases, it is more practical to leave the collapsed region as a single sequence and document the limitation in the assembly report.

### Addressing False Gene Loss

False gene loss occurs when single-copy sequences are incorrectly removed from the assembly. This error is particularly problematic because it leads to incorrect conclusions about gene content. The Galaxy workflow described for Unicycler hybrid assemblies addresses false loss of single-copy sequences by detecting regions that should be present based on read support but are absent from the assembly [<a href="#ref-1">1</a>].

To detect false gene loss, compare the read mapping results with the assembly sequence. Regions where reads map but no assembled sequence exists indicate potential false gene loss. You can also compare the predicted gene content with the expected gene content for the species. Missing genes that have strong read support are candidates for false gene loss.

Recovering falsely lost sequence requires extracting the sequence from the read data and adding it to the assembly. This process involves assembling the reads that map to the missing region and then integrating the resulting sequence at the correct position. The Galaxy workflow provides a standardized approach for this process in the context of Unicycler assemblies [<a href="#ref-1">1</a>].

## Gap Closing and Contiguity Improvement

Gap closing is the process of filling unknown sequence regions, typically represented as runs of N characters in the assembly. Improving contiguity involves reducing the number of contigs and scaffolds by establishing correct connections between sequences.

### Identifying Gap Locations and Sizes

Gaps in assemblies are represented as runs of N characters in FASTA files. The length of the gap is indicated by the number of N characters. Gap locations and sizes should be recorded in your curation log. Large gaps are more likely to contain biologically significant sequence and may warrant targeted gap-closing efforts.

Use read mapping to assess whether reads span the gap. If reads span the gap, you can potentially fill the gap by assembling the spanning reads. If no reads span the gap, the gap may represent a region that was not sequenced or that could not be assembled with the available data.

### Gap-Filling Strategies

Several strategies exist for gap filling. If you have long reads, map them to the assembly and identify reads that span gap boundaries. These reads can provide the sequence needed to fill the gap. If you have Hi-C data, the contact information can help determine the correct sequence to place in the gap, even if the exact sequence is not available.

The choice of gap-filling strategy depends on the data available and the size of the gap. Small gaps that are spanned by reads can often be filled with targeted assembly. Large gaps that are not spanned by reads may require additional sequencing or may need to remain as gaps in the final assembly.

### Evaluating Contiguity Improvements

After gap closing and contiguity improvement efforts, recompute the assembly statistics to quantify the changes. Compare the new N50, number of contigs, and total length with the baseline values recorded before curation. An increase in N50 and a decrease in contig count indicate improved contiguity.

Record the specific changes you made and the impact on assembly statistics. This information is essential for reporting the curation process and for evaluating whether the effort produced meaningful improvements.

## Using Hi-C Data for Scaffolding Validation

Hi-C data provides genome-wide proximity information that is valuable for validating scaffolding and for resolving ambiguous connections. The contact frequency between two genomic regions reflects their physical proximity in the nucleus, which correlates with their linear distance along the chromosome.

### Generating Hi-C Contact Maps

Hi-C contact maps visualize the frequency of contacts between genomic regions. The Puzzler pipeline integrates Hi-C data for chromosome-scale assembly and includes quality control with Hi-C contact maps [<a href="#ref-6">6</a>]. If you have Hi-C data, generate contact maps for your assembly before and after curation to evaluate whether your corrections improve the contact signal.

A correct assembly produces contact maps with a characteristic pattern. Contacts are strongest along the diagonal, representing nearby genomic regions, and decrease with distance from the diagonal. Misjoins produce visible disruptions in this pattern, such as strong off-diagonal contacts or breaks in the diagonal signal.

### Interpreting Contact Map Patterns

When you examine a contact map, look for regions where the expected pattern is disrupted. A misjoin that places two distant genomic regions adjacent to each other produces a contact map where the regions show weak contact across the join but strong contact with their true genomic neighbors. This pattern allows you to identify the correct connections even when the assembly graph is ambiguous.

The Puzzler pipeline can operate reference-free and can generate input files for manual Hi-C curation [<a href="#ref-6">6</a>]. This capability is valuable when you need to curate an assembly without a closely related reference genome. The manual curation input files provide the contact information you need to make informed scaffolding decisions.

### Integrating Hi-C Evidence with Other Data

Hi-C evidence should be integrated with read mapping and graph information when making curation decisions. Hi-C data provides long-range information that is complementary to the short-range information from read pairs and the local information from assembly graphs. When multiple lines of evidence support the same correction, you can make the correction with high confidence.

When evidence conflicts, document the conflict and consider whether additional data are needed. For example, if the assembly graph suggests one connection but Hi-C data suggest a different connection, you may need to examine the raw Hi-C reads or generate additional sequencing data to resolve the conflict.

## Quality Control and Validation After Curation

After completing curation, you must validate that the corrected assembly is accurate and complete. This validation involves recomputing quality metrics, mapping reads, and assessing biological plausibility.

### Recomputing Assembly Statistics

Recompute all assembly statistics after curation and compare them with the baseline values. The gfastats tool can compute assembly summary statistics for the corrected assembly [<a href="#ref-4">4</a>]. Record the updated values for total length, number of contigs, N50, L50, and longest contig. Also record the number of gaps and the total gap length.

Interpret the changes in statistics in the context of the corrections you made. Splitting misjoined contigs will increase the contig count and may decrease N50. Resolving collapsed repeats will increase the total assembly length. Gap closing will decrease the total gap length. These changes should be consistent with the corrections you documented.

### Remapping Reads to the Corrected Assembly

Map the original sequencing reads to the corrected assembly and examine the mapping statistics. The percentage of reads that map, the percentage that map uniquely, and the coverage uniformity are all informative. A well-curated assembly should have high read mapping rates and relatively uniform coverage, except in known repetitive regions.

Compare the read mapping results before and after curation. Regions where you corrected misjoins should now show consistent read support across the corrected join. Regions where you resolved collapsed repeats should show appropriate coverage for the copy number. Regions where you filled gaps should show reads spanning the formerly gapped region.

### Running Completeness Assessment

Run BUSCO on the corrected assembly and compare the results with the baseline assessment. Improvements in BUSCO completeness indicate that curation recovered missing sequence or corrected errors that affected gene prediction. The expected improvement depends on the types of errors you corrected. Resolving collapsed repeats and recovering falsely lost genes should increase completeness. Correcting misjoins may not change BUSCO scores if the misjoined regions contained complete genes.

Record the BUSCO results before and after curation in your curation log. This record provides evidence of the impact of your curation efforts and is useful for reporting to collaborators or reviewers.

## Common Failure Patterns in Manual Curation

Manual curation is a complex process with several common failure patterns. Recognizing these patterns can help you avoid them and can improve the quality of your curation efforts.

### Overcorrection and Introducing New Errors

The most significant risk in manual curation is introducing new errors while attempting to fix existing ones. Splitting a contig at a position that is not a true breakpoint creates two incorrect sequences. Joining sequences based on insufficient evidence creates new misjoins. To avoid overcorrection, require strong evidence before making changes. If you are uncertain about a correction, document the uncertainty and consider leaving the sequence unchanged.

### Incomplete Documentation

Curation decisions that are not documented cannot be reviewed or reproduced. Incomplete documentation is a common failure that undermines the value of curation. Maintain a detailed curation log that records every candidate region, the evidence for each correction, the exact changes made, and the rationale for the decisions. This log should be stored with the assembly files and made available to other researchers.

### Focusing on Statistics instead of Biology

Assembly statistics are useful indicators, but they do not capture all aspects of assembly quality. A curation effort that focuses solely on improving N50 may produce an assembly with fewer contigs but incorrect joins. Prioritize biological accuracy over statistical improvement. An assembly with more contigs but correct sequence is more valuable than an assembly with fewer contigs but errors.

### Ignoring Contamination

Contamination, where sequence from another organism is included in the assembly, is a serious error that manual curation should address. The Puzzler pipeline includes BlobTools contamination screening as part of its quality control [<a href="#ref-6">6</a>]. If you have not screened your assembly for contamination, do so before or during curation. Contaminant sequences often have unusual coverage or composition compared with the target genome.

## Limitations of Manual Curation

Manual curation has inherent limitations that you should understand before beginning the process. These limitations affect the outcomes you can achieve and the confidence you can place in the final assembly.

### Data Limitations

The accuracy of manual curation is limited by the quality and quantity of the available data. If the sequencing reads do not cover a region, you cannot determine the correct sequence for that region. If the reads contain systematic errors, your corrections may propagate those errors. If you lack Hi-C data, you may not be able to determine the correct scaffolding for large genomic regions.

### Expertise Requirements

Manual curation requires substantial expertise in genome biology, sequencing technology, and bioinformatics tools. The process is time consuming and often requires substantial expertise [<a href="#ref-7">7</a>]. Inexperienced curators may miss errors or introduce new errors. If you are new to genome assembly, consider working with an experienced curator or using automated curation tools as a first pass before manual inspection.

### Scalability Constraints

Manual curation does not scale efficiently to large numbers of genomes. The process is labor intensive and requires individual attention to each assembly. For projects that require dozens or hundreds of assemblies, automated pipelines with integrated quality control are more practical. The Puzzler pipeline was designed for scalable, high-throughput platinum-quality genome assembly, demonstrating that automation can achieve high quality for many genomes [<a href="#ref-6">6</a>].

### Incomplete Resolution of Complex Regions

Some genomic regions cannot be fully resolved with any available technology. Highly repetitive regions, such as centromeres and ribosomal RNA gene clusters, may remain as gaps or collapsed sequences even after extensive curation. Document these limitations in the assembly report and be transparent about the residual uncertainty.

## Reproducibility and Reporting Standards

Reproducibility is essential for scientific credibility. Your curation process should be documented in sufficient detail that another researcher could repeat the process and obtain similar results.

### Version Control for Assembly Files

Use version control to track changes to assembly files. Git is the standard tool for this purpose. The Carpentries lessons provide training in Git and version control for researchers [<a href="#ref-2">2</a>]. Maintain the draft assembly, each intermediate version, and the final curated assembly in a version-controlled repository. Commit messages should describe the curation changes made at each step.

### Documenting Curation Decisions

The curation log should include the date, the curator name, the assembly version, and a description of each change. For each change, record the evidence that supported the decision and any alternative interpretations that were considered. This documentation allows reviewers to evaluate the quality of the curation and to identify any decisions they might question.

### Reporting in Publications

When you publish results based on a curated assembly, report the curation process in the methods section. Describe the tools used, the types of errors corrected, and the impact of curation on assembly statistics. Provide access to the curation log and the version-controlled assembly files. This transparency allows other researchers to understand the quality of the assembly and to compare their results with yours.

### Using Workflow Management Systems

Workflow management systems can improve the reproducibility of the curation process. The nf-core documentation describes community standards for pipeline usage and configuration [<a href="#ref-8">8</a>]. While manual curation is inherently interactive, the surrounding steps of quality assessment, read mapping, and validation can be automated in reproducible workflows. The Galaxy platform provides a framework for reproducible analysis that can incorporate manual curation steps [<a href="#ref-3">3</a>].

## Professional Escalation Criteria

Some assembly problems require expertise beyond what a typical researcher can provide. Recognizing when to escalate is important for achieving the best possible assembly quality.

### When to Seek Specialized Assistance

If you encounter assembly problems that you cannot resolve with the available tools and expertise, seek assistance from specialized resources. The NCBI provides official descriptions of databases, search systems, and analysis services that can support genome assembly work [<a href="#ref-9">9</a>]. The EMBL-EBI Training program offers learning pathways and practical analysis education for bioinformatics [<a href="#ref-10">10</a>]. These resources can help you develop the skills needed to address challenging curation problems.

### Collaborating with Genome Centers

For complex genomes or challenging assembly problems, collaboration with a genome center may be appropriate. Genome centers have specialized expertise and computational resources that can address problems beyond the scope of individual laboratories. The Vertebrate Genomes Project, which developed gfastats, is an example of a large-scale effort that has developed and refined curation approaches [<a href="#ref-4">4</a>].

### Using Community Resources

Community resources can provide support for specific curation challenges. The Bioconductor project offers packages and workflows for genomic analysis that may include tools relevant to your curation needs [<a href="#ref-11">11</a>]. The Galaxy Training Network provides tutorials that cover specific analysis scenarios [<a href="#ref-3">3</a>]. These resources can help you find solutions to problems you encounter during curation.

## Records and Measurements for Curation

Maintaining systematic records is essential for effective curation. The following measurements should be recorded for every curation project.

### Baseline Measurements

Record the following measurements before beginning curation: total assembly length, number of contigs, N50, L50, longest contig, number of gaps, total gap length, BUSCO completeness percentage, and read mapping rate. These baseline values provide the reference for evaluating curation impact.

### Per-Region Measurements

For each region you curate, record the node identifier, genomic coordinates, coverage statistics, read support evidence, and the type of error identified. Record the correction applied and the evidence supporting the correction. Record any alternative interpretations that were considered and rejected.

### Final Measurements

After curation, record the same measurements as the baseline assessment. Calculate the changes in each metric and interpret these changes in the context of the corrections applied. Record any regions that remain unresolved and the reasons for the residual uncertainty.

### Curation Log Format

A curation log should be a structured document that can be reviewed by other researchers. Use a table format with columns for date, curator, assembly version, region identifier, error type, evidence, correction applied, and notes. Store the log with the assembly files in the version-controlled repository.

## Safety and Ethical Context

Genome assembly and curation have safety and ethical considerations that researchers should understand.

### Data Handling and Privacy

Genome sequence data may include information about human subjects or endangered species. Handle all sequence data in accordance with applicable regulations and institutional policies. The NCBI provides information about data submission and access policies [<a href="#ref-9">9</a>]. Ensure that your data handling practices comply with these policies.

### Responsible Reporting

Report assembly quality accurately and transparently. Do not overstate the quality of your assembly or the completeness of your curation. Acknowledge limitations and residual uncertainty in your reports. This transparency is essential for scientific integrity and for the appropriate use of your assembly by other researchers.

### Ethical Use of Computational Resources

Genome assembly and curation require substantial computational resources. Use these resources responsibly and efficiently. Avoid unnecessary recomputation and share intermediate results when appropriate. The nf-core documentation describes standards for efficient and reproducible pipeline usage [<a href="#ref-8">8</a>].

## Frequently Asked Questions

### What is the difference between assembly polishing and manual curation?

Assembly polishing corrects base-level errors, such as single nucleotide substitutions and small indels, typically using read alignment information. Manual curation addresses structural errors, such as misjoins, collapsed repeats, and incorrect scaffolding. Polishing improves sequence accuracy within contigs, while curation improves the overall structure of the assembly. Both processes are often needed to produce a high-quality genome assembly.

### How do I know if my assembly needs manual curation?

Run a quality assessment that includes BUSCO completeness, read mapping statistics, and assembly graph inspection. If BUSCO completeness is low, if read mapping reveals coverage anomalies, or if the assembly graph shows complex branching patterns, manual curation is likely needed. Automated pipelines often cannot unambiguously resolve a genome when repeat structures are present, and manual curation is required in those cases [<a href="#ref-1">1</a>].

### What tools do I need for manual curation?

The essential tools are Bandage for graph visualization, IGV for read mapping visualization, and gfastats for computing statistics and manipulating sequences [<a href="#ref-4">4</a>]. You also need the tools used in your assembly pipeline for read mapping and sequence extraction. The Bioconductor project offers additional packages that may be useful for specific analysis tasks [<a href="#ref-11">11</a>].

### How long does manual curation take?

The time required varies substantially with genome size and complexity. A small bacterial genome may require hours, while a large eukaryotic genome may require days or weeks. The number of errors in the draft assembly and the availability of supporting data, such as Hi-C, also affect the time required. Plan your curation effort according to the biological questions you need to answer.

### Can manual curation introduce new errors?

Yes, manual curation can introduce new errors if corrections are made without sufficient evidence. Splitting a contig at a position that is not a true breakpoint creates incorrect sequences. Joining sequences based on insufficient evidence creates new misjoins. To minimize this risk, require strong evidence before making changes and document the evidence for each correction.

### What should I do if I cannot resolve a region?

If you cannot resolve a region with the available data, document the ambiguity and make a conservative decision. Leave the sequence as a separate contig instead of joining it to an uncertain partner. Record the ambiguity in your curation log so that other researchers can revisit the case if additional data become available.

### How should I report manual curation in my publications?

Describe the curation process in the methods section, including the tools used, the types of errors corrected, and the impact on assembly statistics. Provide access to the curation log and the version-controlled assembly files. This transparency allows other researchers to understand the quality of the assembly and to compare their results with yours.

### What are the limitations of manual curation?

Manual curation is limited by the quality and quantity of the available data, the expertise of the curator, and the inherent complexity of some genomic regions. Highly repetitive regions may remain unresolved even after extensive curation. Manual curation also does not scale efficiently to large numbers of genomes, making automated approaches more practical for large-scale projects [<a href="#ref-6">6</a>].

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [A practical guide and Galaxy workflow to avoid inter-plasmidic repeat collapse and false gene loss in Unicycler's hybrid assemblies.](https://pubmed.ncbi.nlm.nih.gov/38197876). Microbial genomics, 2024.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Gfastats: conversion, evaluation and manipulation of genome sequences using assembly graphs.](https://pubmed.ncbi.nlm.nih.gov/35799367). Bioinformatics (Oxford, England), 2022.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [High-quality genome assemblies for nine non-model North American insect species representing six orders (Insecta: Coleoptera, Diptera, Hemiptera, Hymenoptera, Lepidoptera, Neuroptera).](https://pubmed.ncbi.nlm.nih.gov/39155537). Molecular ecology resources, 2024.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [Puzzler: scalable one-command platinum-quality genome assembly from HiFi and Hi-C.](https://pubmed.ncbi.nlm.nih.gov/41573168). Bioinformatics advances, 2026.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [A quick guide for student-driven community genome annotation.](https://pubmed.ncbi.nlm.nih.gov/30943207). PLoS computational biology, 2019.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.