How to Use Hi-C Data for Manual Curation of Genome Assemblies: A Guide to Scaffolding and Error Correction

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Use Hi-C Data for Manual Curation of Genome Assemblies: A Guide to Scaffolding and Error Correction

Key Takeaways

  • Manual curation of genome assemblies using Hi-C contact maps involves visually inspecting chromatin interaction patterns to identify and correct scaffolding errors, primarily misjoins, by applying targeted breaks and joins to achieve chromosome-level contiguity.
  • Characteristic square-shaped off-diagonal blocks in Hi-C contact maps at multiple resolutions serve as the primary visual diagnostic for identifying misjoins, which occur when non-adjacent genomic sequences are incorrectly joined in the assembly.
  • While automated scaffolding tools are essential for initial large-scale ordering and orientation, manual curation remains critical for resolving complex structural errors, particularly in challenging genomes like those of birds with microchromosomes, or for achieving reference-quality assemblies.
  • Input data preparation involves aligning Hi-C reads to a draft assembly and generating normalized contact maps in formats compatible with visualization tools like Juicebox, with bin sizes ranging from 500 kb to 1 Mb for initial assessment and 10 kb to 50 kb for breakpoint localization.
  • Curation is an iterative process requiring meticulous record-keeping of every break and join, including coordinates and rationale, to ensure reproducibility and facilitate verification through independent metrics like BUSCO and gene annotation.
  • Distinguishing true misjoins from artifacts like those arising from repetitive sequences or segmental duplications necessitates examining patterns at multiple resolutions and, when ambiguity persists, seeking convergent evidence from other genomic data types such as long reads or optical mapping.

Direct Answer and Scope

Manual curation of genome assemblies using Hi-C contact maps is the process of visually inspecting chromatin interaction patterns to identify and correct scaffolding errors, then applying targeted breaks and joins to produce chromosome-level assemblies. This guide provides a practical protocol for researchers who have generated a draft assembly and Hi-C data and need to improve contiguity and correctness through manual intervention. The workflow centers on generating contact maps, loading them into visualization tools such as Juicebox, identifying misjoins and misassemblies, and executing curated corrections. This approach has been applied successfully across diverse taxa, from mammals to birds, and remains essential even as automated scaffolding tools improve. The intended reader is a biology student, researcher, or laboratory professional who has Hi-C data in hand and needs a concrete path from raw contact information to a curated chromosome-level assembly.

Why Manual Curation Remains Necessary

Automated Hi-C scaffolding tools produce chromosome-scale assemblies, but they do not always resolve every structural error. The Etruscan shrew genome assembly required manual curation to identify 22 chromosomes including sex chromosomes, despite the use of PacBio long reads, 10X Genomics linked reads, optical mapping, and Hi-C linked reads. This example demonstrates that even multi-platform assemblies benefit from expert visual inspection of contact maps. Similarly, the LT1 human reference genome used 72 GB of Hi-C chromosomal mapping data for scaffolding, and the final assembly quality depended on manual curation after automated scaffolding raised the NG50 from 12.0 Mbp to 137 Mbp. The jump in contiguity came from both automated scaffolding and subsequent manual correction of remaining errors.

Bird genomes present a particularly difficult challenge because microchromosomes are often fragmented or missing entirely from draft assemblies. Expert manual curation through manipulation of genome-wide Hi-C contact maps is required to resolve these small chromosomes, and many published bird assemblies still do not reflect the known karyotype. The MicroFinder pipeline was developed specifically to accelerate this process by identifying conserved dot chromosome proteins that anchor manual curation efforts. This illustrates that manual curation is not a legacy step but an active area of method development.

The Human Pangenome Reference Consortium compared assembly approaches and found that highly accurate long reads combined with graph-based haplotype phasing reduced the need for manual curation, but did not eliminate it. Their first high-quality diploid reference assembly still contained approximately four gaps per chromosome on average. For most research projects, manual curation remains the step that transforms a good draft assembly into a reference-quality chromosome-level assembly.

At a Glance

Decision PointWhat to DoWhy It MattersCommon Mistake
Input data preparationAlign Hi-C reads to the draft assembly and generate contact maps in Juicebox formatContact maps are the visual substrate for all curation decisionsAligning to an unpolished assembly that contains many base-level errors
Initial scaffoldingRun automated Hi-C scaffolding before manual curationAutomated tools resolve most large-scale ordering and orientation problemsSkipping automated scaffolding and attempting to build chromosomes entirely by hand
Misjoin identificationLook for square-shaped off-diagonal patterns in contact mapsThese patterns indicate sequences that are joined incorrectly in the assemblyConfusing local noise with true misjoins, especially in repeat-rich regions
Break and join executionApply breaks at misjoin boundaries and join corrected scaffoldsManual breaks and joins fix errors that automated tools missBreaking at the wrong position because the contact map resolution is too low
Quality verificationRegenerate contact maps after each round of curation and check for new errorsCuration can introduce new misjoins that must be caughtAssuming one round of curation is sufficient for a complete assembly
Record keepingDocument every break and join with coordinates and rationaleReproducibility and troubleshooting depend on complete recordsMaking changes without notes, then being unable to reconstruct the curation history

Core Principles of Hi-C Contact Map Interpretation

How Hi-C Data Reflects Chromatin Structure

Hi-C sequencing captures physical proximity between genomic loci by crosslinking DNA within the nucleus, digesting the crosslinked DNA, ligating the fragments, and sequencing the resulting chimeric molecules. The number of read pairs linking two genomic regions reflects how close those regions are in three-dimensional nuclear space. In a correctly assembled genome, loci on the same chromosome show strong interaction signals, while loci on different chromosomes show weaker signals. This property makes Hi-C data useful for both scaffolding and validation.

The contact map is a matrix where each axis represents the linear order of the assembly. The intensity of each cell represents the number of Hi-C read pairs linking the two positions. A well-assembled chromosome displays a characteristic pattern: strong signal along the diagonal, representing local interactions, and a plaid or checkerboard pattern of off-diagonal blocks representing the spatial organization of chromatin into compartments and topologically associating domains. These patterns are the visual language of manual curation.

Expected Patterns in a Correct Assembly

A correctly assembled chromosome shows a continuous strong diagonal with gradually decreasing signal as distance from the diagonal increases. The off-diagonal signal forms a roughly triangular shape that reflects the polymer physics of chromatin folding. Large structural features such as centromeres and telomeres may show distinct interaction patterns. Centromeres often appear as regions of reduced signal or as interaction hubs, while telomeres show a gradual signal decrease at chromosome ends.

In a genome with multiple chromosomes, the contact map shows clear boundaries between chromosomes. Interchromosomal interactions are present but much weaker than intrachromosomal interactions. The transition between chromosomes appears as a sharp drop in signal intensity. When this pattern is disrupted, it indicates a scaffolding error.

Identifying Misjoins Through Contact Map Patterns

A misjoin occurs when two sequences that are not adjacent in the true genome are joined together in the assembly. In a contact map, a misjoin produces a characteristic square-shaped off-diagonal block. This pattern arises because the sequences on either side of the misjoin belong to different genomic contexts, and their interaction patterns do not match the expected linear arrangement.

For example, if a scaffold contains the end of chromosome 1 joined to the middle of chromosome 2, the contact map shows strong signal between the chromosome 1 portion and other chromosome 1 sequences, and strong signal between the chromosome 2 portion and other chromosome 2 sequences. The junction between these two regions appears as a sharp transition, often with a square block of signal that extends away from the diagonal. This square pattern is the primary visual cue for manual curation.

Distinguishing Real Misjoins from Artifacts

Not every off-diagonal block indicates a misjoin. Repetitive sequences, segmental duplications, and centromeric regions can produce interaction patterns that resemble misjoins. The key is to examine the pattern at multiple resolutions. A true misjoin shows a consistent square pattern across multiple resolutions, while an artifact may appear only at certain resolutions or in specific genomic contexts.

The Etruscan shrew assembly identified segmental duplications as part of its characterization, and these regions can complicate contact map interpretation. When a suspected misjoin falls in a repeat-rich region, additional evidence from long reads or optical mapping may be needed to confirm the error. The decision to break a scaffold should be based on convergent evidence from multiple sources, not solely on the contact map pattern.

Preparing Input Data for Manual Curation

Required Data Types

Manual curation requires three inputs: a draft genome assembly in FASTA format, Hi-C sequencing reads in FASTQ format, and a reference genome or karyotype information when available. The draft assembly should be as contiguous as possible before Hi-C scaffolding. Long-read assemblies from PacBio or Oxford Nanopore platforms provide a good starting point because they resolve repetitive regions better than short-read assemblies. The LT1 assembly used 57x nanopore long reads and was polished with 47x short paired-end reads before Hi-C scaffolding, demonstrating the value of polishing before scaffolding.

The Hi-C data should have sufficient sequencing depth to produce a clear contact map. The LT1 project used 72 GB of Hi-C data for scaffolding a 2.73 Gbp genome. The required depth depends on genome size, complexity, and the resolution needed for curation. Higher resolution contact maps allow finer breakpoint identification but require more sequencing.

Aligning Hi-C Reads to the Draft Assembly

The first computational step is aligning Hi-C reads to the draft assembly. This alignment is performed with a read aligner that can handle chimeric reads, because Hi-C library preparation produces fragments that ligate two distant genomic loci. The alignment output is then processed to retain only valid Hi-C pairs, which are read pairs where each read maps uniquely to the assembly and the two reads map to different restriction fragments.

Alignment parameters must be chosen carefully. The insert size distribution for Hi-C libraries is broad, and the aligner must be configured to allow for this. Reads that map to multiple locations should be handled according to the analysis goals. For manual curation, uniquely mapping reads provide the clearest signal, but multi-mapping reads can be informative in repetitive regions when used with caution.

Generating Contact Maps

Contact maps are generated from the aligned Hi-C pairs by binning the genome into intervals and counting the number of read pairs linking each pair of bins. The bin size determines the resolution of the map. Coarse bins of 500 kb to 1 Mb are useful for identifying chromosome-level patterns and large misjoins. Finer bins of 10 kb to 50 kb are needed to localize breakpoints precisely.

The contact map must be normalized to account for biases in Hi-C data, including restriction enzyme cutting frequency, GC content, and mappability. Normalization methods such as iterative correction and eigenvector decomposition produce maps where the interaction signal reflects true chromatin structure instead of technical artifacts. The normalized map is then converted to a format compatible with visualization tools.

Visualization Tools for Manual Curation

Juicebox is the most widely used tool for manual curation of Hi-C contact maps. It provides an interactive interface for viewing contact maps at multiple resolutions, navigating the genome, and marking breaks and joins. The tool was used in the curation of the Etruscan shrew assembly and is the standard for this workflow.

The input to Juicebox is a Hi-C map file in the .hic format, which stores contact matrices at multiple resolutions for efficient visualization. The .hic file is generated from aligned Hi-C pairs using the Juicer pipeline. An assembly file in .assembly format describes the current scaffold structure and allows Juicebox to display the assembly along the axes of the contact map.

Practical Workflow for Manual Curation

Step 1: Generate an Initial Contact Map

Begin by aligning Hi-C reads to the draft assembly and generating a contact map at coarse resolution. The goal of this first map is to assess the overall quality of the assembly and identify large-scale problems. Load the map into Juicebox and examine each chromosome-scale scaffold. Look for the expected plaid pattern of intrachromosomal interactions and sharp boundaries between scaffolds.

At this stage, the most common problems are scaffolds that contain sequences from multiple chromosomes, scaffolds that are oriented incorrectly relative to each other, and scaffolds that are joined in the wrong order. These problems appear as square off-diagonal blocks, inverted patterns, or discontinuities in the diagonal signal.

Step 2: Identify Misjoins at Coarse Resolution

Work through the contact map systematically, chromosome by chromosome. For each scaffold, trace the diagonal and note any position where the signal pattern changes abruptly. A misjoin is indicated by a square block of signal that extends away from the diagonal and does not match the surrounding pattern.

When a suspected misjoin is found, zoom in to higher resolution to confirm the breakpoint. The breakpoint is the position where the interaction pattern changes from one genomic context to another. In many cases, the breakpoint is sharp and unambiguous. In other cases, especially in repeat-rich regions, the breakpoint may be spread over a range of positions.

Step 3: Mark Breaks and Joins in Juicebox

Juicebox provides tools for marking breaks and joins in the assembly. A break splits a scaffold at a specified position. A join connects two scaffolds in a specified orientation. The assembly file is updated to reflect these changes, and the contact map is regenerated to visualize the result.

The decision to break a scaffold should be based on the strength of the evidence. A clear square pattern at multiple resolutions is strong evidence for a misjoin. A pattern that appears only at one resolution or that is confounded by repeats requires additional evidence. In these cases, examine the underlying Hi-C read pairs to confirm that the interaction pattern is real and not an artifact of alignment or normalization.

Step 4: Iterate Between Breaking and Joining

Manual curation is an iterative process. After breaking a misjoined scaffold, the resulting fragments may need to be joined to other scaffolds. The contact map provides the evidence for these joins: fragments from the same chromosome show strong interaction signal with each other and with other fragments from that chromosome.

The order and orientation of fragments within a chromosome can be determined from the contact map pattern. Adjacent fragments show strong local interactions, and the diagonal signal is continuous across the join. Fragments that are inverted relative to their neighbors show a characteristic pattern where the interaction signal is strongest between positions that are far apart in the assembly coordinate system.

Step 5: Verify the Curated Assembly

After each round of curation, regenerate the contact map and examine the result. The curated assembly should show clean chromosome-level patterns with no remaining square blocks or discontinuities. Each chromosome should have a continuous diagonal and a clear boundary with other chromosomes.

Verification should also include independent quality metrics. The LT1 assembly was assessed using BUSCO, which identified 89.3% of single-copy orthologous genes. The Etruscan shrew assembly used the NCBI genome annotation pipeline to identify 39,091 genes. These external validations confirm that the curated assembly supports accurate gene prediction and annotation.

Options and Tradeoffs in Curation Approaches

Fully Manual versus Semi-Automated Curation

Manual curation can be performed entirely by visual inspection of contact maps, or it can be assisted by automated tools that identify candidate misjoins. The fully manual approach gives the curator complete control but is time-consuming and requires substantial expertise. The semi-automated approach uses computational methods to flag likely errors, which the curator then verifies and corrects.

The Human Pangenome Reference Consortium found that assembly approaches using highly accurate long reads and graph-based haplotype phasing required minimal manual curation. This finding suggests that improvements in assembly algorithms can reduce but not eliminate the need for manual curation. For most projects, a combination of automated scaffolding and manual curation provides the best balance of speed and accuracy.

Reference-Based versus Reference-Free Curation

When a reference genome is available for the same species or a closely related species, the reference can guide curation by providing expected chromosome structure and gene order. The Puzzler pipeline includes an optional module for chromosome assignment via synteny, which uses the reference to order and orient scaffolds. This approach is particularly useful for genomes with conserved karyotypes.

Reference-free curation relies entirely on the Hi-C contact map signal. This approach is necessary for species without a close reference, but it requires more expertise to interpret. The contact map signal alone can resolve chromosome boundaries and order, but it cannot distinguish between different chromosome arrangements that produce similar interaction patterns.

Manual Curation Aids for Difficult Genomes

Some genomes present specific challenges that require specialized curation aids. Bird genomes have tiny microchromosomes that are often fragmented or missing from draft assemblies. The MicroFinder pipeline identifies conserved dot chromosome proteins in draft assemblies and uses these as anchors for manual curation. This approach dramatically speeds up curation and improves the sequence content of dot microchromosomes.

The MicroFinder-enabled re-curation of 12 previously released bird genome assemblies increased the sequence content of dot microchromosome models. This result demonstrates that even assemblies that have undergone expert curation can be improved with targeted tools. For researchers working on difficult genomes, investing in specialized curation aids can substantially improve assembly quality.

Observations and Measurements During Curation

Recording Curation Decisions

Every break and join should be recorded with the scaffold name, position, orientation, and rationale. This record serves multiple purposes. It allows the curation to be reproduced and verified by others. It provides a basis for troubleshooting if problems are found later. It documents the evidence for each decision, which is valuable for publications and data releases.

The record should also include the resolution at which each decision was made and the evidence that supported it. For example, a break might be recorded as "scaffold_12, position 45,320,000, square pattern at 100 kb resolution, confirmed at 25 kb resolution." This level of detail allows the decision to be revisited if new data become available.

Measuring Assembly Quality Before and After Curation

Assembly quality should be measured before, during, and after manual curation. The primary metrics are contiguity measures such as N50 and NG50, completeness measures such as BUSCO, and correctness measures such as the number of misjoins identified by independent methods. The LT1 assembly showed a dramatic improvement in NG50 from 12.0 Mbp to 137 Mbp after Hi-C scaffolding and manual curation.

Contiguity metrics alone do not capture assembly correctness. An assembly can be highly contiguous but contain misjoins that are invisible to N50 calculations. The contact map provides a direct measure of correctness by showing whether the assembly structure matches the expected chromatin interaction pattern. This is why manual curation is essential even when contiguity metrics look good.

Tracking Curation Progress

A curation log should track the number of breaks and joins made, the number of scaffolds before and after curation, and the number of chromosomes resolved. This log provides a quantitative measure of curation progress and helps identify when the assembly has reached a stable state. The log also provides the data needed to describe the curation process in publications.

For the Etruscan shrew assembly, manual curation identified 22 chromosomes including X and Y sex chromosomes. This outcome required multiple rounds of breaking and joining, and the curation log would document the path from the initial draft to the final chromosome-level assembly.

Quality Controls and Verification

Independent Validation of Curated Assemblies

The curated assembly should be validated using methods that do not depend on Hi-C data. Long-read alignments can confirm that the assembly structure is consistent with the underlying sequence data. Optical mapping provides an independent measure of genome structure. Genetic maps, when available, can confirm the order and orientation of markers.

The Etruscan shrew assembly used PacBio long reads, 10X Genomics linked reads, optical mapping, and Hi-C linked reads. The combination of these platforms provided multiple lines of evidence for the final assembly structure. For projects with limited resources, at least one independent validation method should be used.

BUSCO and Gene Annotation as Quality Metrics

BUSCO assesses assembly completeness by searching for a set of single-copy orthologous genes that are expected to be present in the genome. The LT1 assembly achieved 89.3% BUSCO completeness, indicating that most expected genes were present and complete. Gene annotation provides a complementary assessment by showing that the assembly supports accurate gene prediction.

The NCBI genome annotation pipeline identified 39,091 genes in the Etruscan shrew assembly, 19,819 of them protein-coding. This annotation provides strong evidence that the assembly is complete enough to support functional genomics. For species without close references, gene annotation may be less informative, but BUSCO remains a useful universal metric.

Checking for Curation-Induced Errors

Manual curation can introduce new errors if breaks are placed at incorrect positions or joins are made in the wrong orientation. The contact map should be regenerated after each round of curation and examined for new square patterns or discontinuities. This verification step is essential because curation errors can be as damaging as the original assembly errors.

The iterative nature of curation means that errors introduced in one round can be caught in the next round. The key is to maintain a complete record of all changes so that any error can be traced to its source. If a new misjoin is detected, the curation log should identify the break and join that caused it.

Common Failure Patterns in Manual Curation

Misinterpreting Local Interaction Patterns

The most common failure in manual curation is misinterpreting local interaction patterns as evidence for misjoins. Hi-C contact maps contain substantial noise, especially in repeat-rich regions and at coarse resolutions. A curator who is too aggressive in breaking scaffolds can fragment an assembly that was largely correct.

The solution is to require convergent evidence before making a break. A suspected misjoin should show a clear square pattern at multiple resolutions and should be supported by the underlying read pairs. If the evidence is ambiguous, the safer choice is to leave the scaffold intact and note the ambiguity for future investigation.

Breaking at the Wrong Position

When a misjoin is confirmed, the break must be placed at the correct position. Breaking too far from the true junction leaves sequence from the wrong chromosome attached to the scaffold. Breaking too close to the junction may split a sequence that should remain intact.

The breakpoint should be localized using high-resolution contact maps. The transition between interaction patterns marks the junction, and the break should be placed at this transition. In some cases, the transition is gradual, and the breakpoint is uncertain. In these cases, the break should be placed conservatively, and the resulting fragments should be examined in the next round of curation.

Joining Fragments in the Wrong Orientation

Joining two fragments in the wrong orientation produces an assembly that appears contiguous but has an inverted region. This error is detected by examining the contact map pattern across the join. A correct join shows a continuous diagonal with gradually decreasing signal. An inverted join shows a discontinuity in the diagonal and a characteristic off-diagonal pattern.

The orientation of a fragment can be determined from the interaction pattern with neighboring fragments. The strand of the fragment that shows stronger interaction with the upstream neighbor is the correct orientation for that join. This determination requires careful examination of the contact map at the join site.

Overlooking Small Misjoins

Small misjoins, involving sequences shorter than the resolution of the contact map, can be missed during manual curation. These errors are invisible at coarse resolution and may only appear at fine resolution. The curator should examine the assembly at multiple resolutions, focusing on regions where the interaction pattern is ambiguous.

Small misjoins are more likely in genomes with recent duplications or other complex structural features. The Etruscan shrew assembly identified segmental duplications, which are regions that can produce confusing interaction patterns. For these regions, additional evidence from long reads or other platforms may be needed to resolve the structure.

Limitations of Manual Curation

Resolution Limits of Hi-C Data

The resolution of the contact map limits the precision of breakpoint identification. At 10 kb resolution, a breakpoint can be localized to a 10 kb interval. At 1 Mb resolution, the same breakpoint is localized to a 1 Mb interval. The required resolution depends on the size of the misjoin and the complexity of the surrounding sequence.

Higher resolution requires more sequencing depth. The relationship between depth and resolution is not linear, and the cost of sequencing increases rapidly as resolution increases. For most curation purposes, 25 to 100 kb resolution provides a good balance between precision and cost.

Difficulty with Repetitive Regions

Repetitive regions present the greatest challenge for manual curation. These regions produce ambiguous interaction patterns because reads from different copies of a repeat cannot be uniquely mapped. The contact map signal in these regions reflects the average interaction pattern of all repeat copies, which may not match the true structure of any individual copy.

The Etruscan shrew assembly identified segmental duplications, and these regions required careful interpretation during curation. In some cases, the structure of a repetitive region cannot be resolved with Hi-C data alone, and additional data from long reads or optical mapping is needed.

Expertise Required for Reliable Curation

Manual curation requires substantial expertise in both Hi-C data interpretation and genome biology. A curator must be able to distinguish real interaction patterns from artifacts, recognize the signatures of different types of structural variation, and make decisions that balance contiguity against correctness.

The Human Pangenome Reference Consortium found that assembly approaches using highly accurate long reads and graph-based haplotype phasing required minimal manual curation. This finding suggests that improvements in assembly algorithms can reduce the expertise required for curation. However, for most projects, manual curation remains a skilled task that benefits from training and experience.

Safety and Reproducibility Context

Reproducibility of Curation Decisions

Manual curation is inherently subjective, and different curators may make different decisions when presented with the same contact map. This subjectivity is a limitation for reproducibility. The curation log mitigates this limitation by documenting every decision and the evidence that supported it.

The Puzzler pipeline includes a checkpointing system that ensures previously completed tasks are not re-executed. This approach supports reproducibility by providing a clear record of the assembly process. For manual curation, the curation log serves a similar purpose.

Data Management for Curation Projects

Hi-C data and contact maps are large files that require substantial storage. The LT1 project used 72 GB of Hi-C data for scaffolding. The .hic files generated from this data are also large, especially when stored at multiple resolutions. Data management plans should account for these storage requirements.

The assembly file, curation log, and contact maps should be versioned and backed up. The curation process can span weeks or months, and the ability to return to a previous state is essential for troubleshooting. Version control systems such as Git provide a framework for managing these files.

Training Resources for Manual Curation

Manual curation skills are learned through practice and training. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover Hi-C data processing and visualization. The Carpentries lessons provide foundational computing and data skills that are useful for managing curation projects. EMBL-EBI Training offers bioinformatics learning pathways that include data-resource training and practical analysis education.

These training resources are valuable for researchers who are new to manual curation. The skills required include command-line computing, data visualization, and genome biology. The Carpentries lessons cover shell, Git, and programming fundamentals that are prerequisites for efficient curation work.

Professional Escalation Criteria

When to Seek Additional Expertise

Manual curation can reach a point where the available data and expertise are insufficient to resolve remaining problems. This situation is indicated by persistent ambiguous contact map patterns, conflicting evidence from different data types, or an inability to resolve chromosome structure despite multiple rounds of curation.

In these cases, seeking additional expertise is appropriate. Collaborators with experience in difficult genomes, such as the bird microchromosome problem addressed by MicroFinder, can provide guidance. Specialized services may also be available for particularly challenging assemblies.

When to Generate Additional Data

If the contact map resolution is insufficient to resolve breakpoints, generating additional Hi-C data may be necessary. If repetitive regions cannot be resolved, additional long-read data or optical mapping may be needed. The decision to generate additional data should be based on the value of resolving the remaining problems relative to the cost of sequencing.

The Etruscan shrew assembly used multiple platforms, including PacBio long reads, 10X Genomics linked reads, optical mapping, and Hi-C linked reads. This multi-platform approach provided the data needed to resolve a complex genome. For most projects, a similar approach is warranted when the genome is particularly complex or when the assembly will serve as a long-term reference.

When to Consider Alternative Assembly Strategies

If manual curation is unable to produce a satisfactory assembly, alternative assembly strategies should be considered. The Human Pangenome Reference Consortium found that assembly approaches using highly accurate long reads and graph-based haplotype phasing outperformed other approaches. Switching to a different assembly strategy may resolve problems that cannot be fixed by curation alone.

The Puzzler pipeline provides an integrated approach that automates contig assembly, duplicate purging, Hi-C-based scaffolding, and chromosome assignment via synteny. This pipeline has been validated on genomes ranging from 24 Mbp to 6.5 Gbp and delivers highly contiguous assemblies with minimal user input. For projects that require many chromosome-scale assemblies, such as pan-genomic efforts, this type of pipeline may be more appropriate than manual curation for every genome.

Frequently Asked Questions

What is the minimum Hi-C sequencing depth needed for manual curation?

The required depth depends on genome size, complexity, and the resolution needed for curation. The LT1 human genome assembly used 72 GB of Hi-C data for a 2.73 Gbp genome. Smaller genomes may require less data, while complex genomes with many repeats may require more. The appropriate depth is best determined empirically by generating a contact map and assessing whether the interaction patterns are clear enough for curation decisions.

How long does manual curation take for a typical genome?

The time required varies widely depending on genome size, complexity, and curator experience. Simple genomes with few misjoins may be curated in days. Complex genomes with many small chromosomes, such as bird genomes, may require weeks or months. The MicroFinder pipeline was developed to speed up bird genome curation, demonstrating that specialized tools can substantially reduce curation time.

Can automated tools replace manual curation entirely?

Automated scaffolding tools can resolve most large-scale ordering and orientation problems, but they do not catch every error. The Human Pangenome Reference Consortium found that assembly approaches using highly accurate long reads and graph-based haplotype phasing required minimal manual curation, but even these approaches did not eliminate the need for curation. For most projects, a combination of automated scaffolding and manual curation provides the best results.

What is the difference between a misjoin and a misassembly?

A misjoin is an error in the order or orientation of sequences within a scaffold. A misassembly is a broader term that includes any error in the assembly structure, including base-level errors, collapsed repeats, and incorrect haplotype representation. Manual curation using Hi-C contact maps primarily addresses misjoins, while other methods are needed for other types of misassembly.

How do I know if a square pattern in the contact map is a real misjoin?

A real misjoin shows a consistent square pattern at multiple resolutions and is supported by the underlying Hi-C read pairs. An artifact may appear only at certain resolutions or in specific genomic contexts. When the evidence is ambiguous, additional data from long reads or optical mapping can confirm whether a misjoin is present.

What should I do if my genome has many small chromosomes that are difficult to resolve?

Genomes with many small chromosomes, such as bird genomes, require specialized curation approaches. The MicroFinder pipeline identifies conserved dot chromosome proteins in draft assemblies and uses these as anchors for manual curation. This approach dramatically speeds up curation and improves the sequence content of dot microchromosomes.

How should I document my curation decisions for publication?

Each break and join should be recorded with the scaffold name, position, orientation, and rationale. The record should include the resolution at which each decision was made and the evidence that supported it. This documentation allows the curation to be reproduced and verified by others and provides the data needed to describe the curation process in publications.

What quality metrics should I report for a curated assembly?

Report contiguity metrics such as N50 and NG50, completeness metrics such as BUSCO, and the number of chromosomes resolved. The LT1 assembly reported an NG50 of 137 Mbp and 89.3% BUSCO completeness after curation. The Etruscan shrew assembly reported 22 chromosomes and 39,091 genes identified by the NCBI genome annotation pipeline. These metrics provide a comprehensive picture of assembly quality.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.