Using IGV to Manually Inspect Genome Assemblies: A Practical Guide for Spotting Misassemblies

By Dr. Zubair Khalid, DVM, MS, PhD ·

Using IGV to Manually Inspect Genome Assemblies: A Practical Guide for Spotting Misassemblies

Key Takeaways

  • Manual inspection of genome assemblies in IGV is crucial for identifying misassemblies missed by automated metrics (e.g., N50, BUSCO), particularly in repetitive regions or those with structural variants.
  • Read alignment quality and coverage depth are foundational; inconsistent coverage across contig boundaries or elevated coverage in extended regions strongly suggests misjoins or collapsed repeats, respectively.
  • Characteristic visual signatures in IGV include abrupt coverage changes at contig boundaries indicating misjoins, and elevated coverage proportional to copy number in collapsed repeat regions.
  • Paired-end read orientation and insert size anomalies, alongside read termination patterns at contig junctions, are critical diagnostic indicators for chimeric joins and misassemblies.
  • Base-level accuracy verification in IGV, by examining mismatches and indels against aligned reads, is essential for validating critical genomic regions like coding sequences.
  • Documentation of inspection findings, including coverage statistics and read alignment patterns for prioritized regions (contig boundaries, repeats, genes of interest), is vital for reproducibility and downstream decision-making.

Genome assembly quality directly affects every downstream analysis, from variant calling to comparative genomics. The Integrative Genomics Viewer (IGV) provides a visual interface for examining assembled contigs against aligned sequencing reads, allowing researchers to detect misjoins, collapsed repeats, and structural errors that automated quality metrics often miss. This guide presents a systematic protocol for loading assembly FASTA files into IGV, navigating coverage and alignment tracks, and identifying common assembly error patterns with concrete decision criteria for when to trust or reject a genomic region.

Why Manual Assembly Inspection Matters

Automated assembly quality metrics such as N50, BUSCO completeness scores, and quality values provide useful summary statistics, but they cannot reveal every structural error. A scaffold N50 of 33.48 Mb and 95.5% BUSCO completeness, as reported for the bighead catfish genome assembly, indicates high continuity and completeness at the aggregate level, yet individual regions may still contain misjoins or collapsed repeats that only visual inspection can expose. The gap between summary statistics and regional accuracy is where manual inspection becomes essential.

Long-read sequencing technologies including Pacific Biosciences HiFi and Oxford Nanopore produce reads that span repetitive elements and structural variants more effectively than short reads alone. However, assembly algorithms still make errors when confronted with highly repetitive sequence, segmental duplications, or regions of extreme base composition. The KRAB-zinc finger protein gene clusters described in the mouse genome illustrate this challenge: these loci contain rapidly evolving gene families with high sequence similarity and endogenous retrovirus insertions that create complex repeat structures. Assemblies of such regions frequently contain collapsed paralogs or misjoined contigs that would escape detection without visual inspection.

Manual inspection serves three distinct purposes in the assembly validation workflow. First, it confirms that aligned reads show consistent coverage and mapping quality across contig boundaries. Second, it reveals whether repetitive regions have been collapsed into fewer copies than actually exist in the genome. Third, it identifies chimeric joins where sequences from different genomic locations have been incorrectly fused into a single contig. Each of these error types produces characteristic visual patterns in IGV that trained researchers can recognize reliably.

Core Principles of Visual Assembly Validation

Read Alignment as the Foundation for Inspection

Visual inspection of an assembly requires aligned sequencing reads that map back to the assembled contigs. The alignment step transforms raw sequencing data into a format that IGV can display, creating the coverage tracks and read pileups that reveal assembly errors. Without this alignment step, the assembly FASTA file alone provides no information about whether the assembled sequence is supported by the underlying read data.

The choice of aligner depends on the sequencing technology used for the original assembly. Long reads from Pacific Biosciences or Oxford Nanopore require aligners designed for high error rates and long read lengths, while Illumina short reads use different alignment algorithms optimized for high accuracy at short lengths. The alignment parameters must be appropriate for the read type, since incorrect settings can produce spurious alignments that obscure genuine assembly errors or create false positives that waste inspection time.

Coverage depth directly influences the reliability of visual inspection. Regions with very low coverage provide insufficient evidence to distinguish true assembly errors from alignment artifacts. Regions with extremely high coverage may indicate collapsed repeats, where reads from multiple genomic copies align to a single assembled locus. Understanding the expected coverage distribution for the sequencing run provides the baseline against which anomalies are judged.

The IGV Interface for Assembly Inspection

IGV displays genomic data as horizontal tracks aligned along a reference coordinate system. For assembly inspection, the assembled contig serves as the reference sequence, and aligned reads appear as individual horizontal bars beneath the reference. Coverage is typically shown as a separate track with a histogram or line plot indicating read depth at each position.

The viewer supports navigation to specific contigs and coordinates, zooming from chromosome-level views down to individual base pairs. This multi-scale navigation is essential for assembly inspection because different error types manifest at different scales. Misjoins often appear as abrupt changes in coverage or read orientation at a specific breakpoint, while collapsed repeats show elevated coverage across an extended region. Base-level zoom reveals mismatches and indels that indicate sequence errors.

IGV also provides features for examining read-level details, including base quality scores, mapping qualities, and read orientations. Paired-end reads display as connected pairs, allowing the inspector to identify pairs that map to unexpected locations or orientations. These read-level details distinguish true assembly errors from alignment artifacts and provide evidence for the nature of the underlying problem.

Preparing Assembly Data for IGV Inspection

Required Input Files and Their Formats

The inspection workflow requires three categories of input files. The assembly FASTA file contains the contig or scaffold sequences to be inspected. The aligned reads file contains sequencing reads mapped to the assembly, typically in BAM format with its associated index file. A genome annotation file in GFF or BED format is optional but useful for correlating assembly errors with gene content.

The assembly FASTA file must be indexed before IGV can load it efficiently. The index file allows IGV to access specific regions of the assembly without loading the entire file into memory. For large assemblies such as the 880 Mb bighead catfish genome, indexing is essential for responsive navigation. The index file is created using standard bioinformatics tools and must be located in the same directory as the FASTA file with a matching filename.

The BAM file containing aligned reads must also be sorted and indexed. Sorting places reads in genomic order, which is required for efficient random access. The index file enables IGV to retrieve reads from specific genomic regions quickly. Both the sorted BAM file and its index must be present for IGV to display read alignments properly.

Loading the Assembly and Alignments into IGV

The standard workflow begins by loading the indexed assembly FASTA file as the reference genome. IGV treats the assembly contigs as chromosomes, displaying each contig as a separate sequence in the navigation dropdown. The contig names from the FASTA file appear in the selection menu, allowing direct navigation to any assembled sequence.

After loading the reference, the aligned reads BAM file is loaded as a data track. IGV automatically displays the coverage track and read alignments for the currently viewed region. The user can then navigate to specific contigs and genomic coordinates to begin the inspection process.

For assemblies with many contigs, prioritizing which regions to inspect first requires a systematic approach. The highest priority regions include contig boundaries, where misjoins are most likely to occur, and regions with known repetitive content, where collapses are common. Regions containing genes of interest or structural variants identified by automated tools also warrant careful inspection.

Creating a Genome Annotation Track for Context

A genome annotation track provides essential context for interpreting assembly errors. Gene models, repeat annotations, and other genomic features help the inspector distinguish genuine assembly errors from biologically unusual but correct sequence. For example, a region annotated as containing a tandem gene family may legitimately show elevated coverage, while the same coverage pattern in a single-copy region suggests a collapsed repeat.

Annotation files in GFF or BED format can be loaded into IGV as additional tracks. The annotation track displays gene models, exons, and other features aligned to the assembly coordinates. This context is particularly valuable when inspecting regions where assembly errors affect gene structure, since a misjoin that disrupts a gene model is immediately visible.

Systematic Workflow for Inspecting Assembly Quality

Step 1: Establish the Inspection Scope and Priorities

Before beginning visual inspection, define which regions of the assembly require examination. A complete manual inspection of a large genome is impractical, so prioritization is essential. The inspection scope should be guided by the research questions the assembly will support and the known risk factors for assembly errors.

Contig boundaries deserve the highest priority because misjoins most frequently occur at the ends of assembled sequences. When two sequences from different genomic locations are incorrectly joined, the breakpoint typically falls at a contig boundary. Examining the coverage and read alignment patterns at each boundary reveals whether the join is supported by consistent read evidence.

Regions with repetitive content constitute the second priority tier. Tandem repeats, transposable elements, and segmental duplications create assembly challenges that often result in collapsed or expanded copy numbers. The KRAB-zinc finger clusters in mice exemplify this risk, with their rapidly evolving gene families and ERV insertions creating complex repeat structures that challenge assembly algorithms.

Genes and genomic features of biological interest form the third priority tier. If the assembly will be used to study specific gene families or genomic regions, those regions warrant direct inspection regardless of their repeat content. This targeted inspection ensures that the most biologically important regions meet quality standards.

Step 2: Examine Coverage Consistency Across Contigs

Coverage analysis provides the first line of evidence for assembly errors. The expected coverage for a region depends on the sequencing depth and the genome size relative to the sequencing yield. For a diploid genome assembled from long reads, coverage should be relatively uniform across the genome, with deviations indicating potential problems.

Abrupt coverage changes at a specific position suggest a misjoin, where reads from two different genomic regions with different coverage levels have been fused. A gradual coverage increase across an extended region suggests a collapsed repeat, where reads from multiple genomic copies align to a single assembled locus. Coverage that drops to zero at a contig boundary may indicate an unassembled gap or a misjoin that prevented read alignment.

The coverage track in IGV provides a visual representation of read depth across the viewed region. The inspector should examine coverage at multiple scales, from broad chromosome-level views to detailed base-level views. Broad views reveal large-scale coverage anomalies, while detailed views localize the exact position of coverage transitions.

Step 3: Inspect Read Alignment Patterns at Contig Boundaries

Contig boundaries require careful examination of read alignment patterns. In a correctly assembled genome, reads should align continuously across the boundary, with read pairs spanning the junction and supporting the join. Reads that terminate abruptly at the boundary or map to unexpected locations indicate a potential misjoin.

The orientation of aligned reads provides additional evidence. In a correctly assembled region, reads align in the expected orientation relative to the reference. Reads that switch orientation at a specific position suggest that sequences from opposite strands have been incorrectly joined. Paired-end reads that map with unexpected insert sizes or orientations also indicate assembly problems.

For long-read assemblies, the read alignment pattern at contig boundaries should show reads spanning the junction with high mapping quality. Reads that partially align to one side of the boundary and partially to the other, with a soft-clipped or hard-clipped segment, suggest that the assembled sequence does not match the underlying read data at that position.

Step 4: Identify Collapsed Repeats Through Coverage and Read Patterns

Collapsed repeats occur when an assembly algorithm merges multiple similar genomic copies into a single assembled sequence. The resulting contig shows elevated coverage, since reads from all copies align to the single assembled locus. The coverage elevation is approximately proportional to the number of collapsed copies.

The visual signature of a collapsed repeat in IGV includes elevated coverage across the repeated region, reads with high mapping quality that span the region, and a sharp transition to normal coverage at the repeat boundaries. The boundaries of the collapsed region often show characteristic read patterns, including reads that partially align to the collapsed region and partially to flanking unique sequence.

Distinguishing collapsed repeats from genuine tandem duplications requires additional evidence. A genuine tandem duplication in the genome will show elevated coverage that matches the expected copy number, with read pairs spanning the duplication junction. A collapsed repeat shows coverage consistent with the number of collapsed copies, but the read pattern may reveal inconsistencies at the boundaries.

Step 5: Detect Misjoins Through Read Orientation and Pairing Anomalies

Misjoins produce characteristic read alignment patterns that differ from collapsed repeats. A misjoin fuses sequences from different genomic locations, creating a chimeric contig. Reads from each contributing location align to their respective portions of the contig, but no reads span the junction between the two portions.

The visual signature of a misjoin includes a sharp transition in read coverage, read orientation, or read pairing at the junction position. Reads to one side of the junction align with high mapping quality, while reads to the other side also align with high mapping quality, but no reads connect the two sides. Paired-end reads may show pairs that map to distant locations, indicating that the two sides of the junction originate from different genomic regions.

For long-read assemblies, misjoins may be less common than in short-read assemblies, but they still occur, particularly in regions with complex repeat structure. The KRAB-zinc finger clusters with their ERV insertions and recombination events create conditions where misjoins are more likely. Careful inspection of these regions is essential for detecting chimeric joins.

Step 6: Verify Base-Level Accuracy in Regions of Interest

Base-level inspection focuses on the accuracy of individual nucleotides within the assembled sequence. The read alignment at base level reveals mismatches, insertions, and deletions that indicate sequence errors. High-quality assemblies should show few mismatches between the assembled sequence and the aligned reads.

The base-level view in IGV displays the reference sequence at the top and the aligned reads below, with mismatches highlighted in color. Insertions and deletions appear as gaps in the read alignment or as extra bases in the read sequence. The inspector should examine these patterns to distinguish genuine sequence variants from assembly errors.

For regions that will be used for downstream analysis, base-level accuracy is critical. A single base error in a coding region can alter the predicted protein sequence, while errors in regulatory regions can affect the interpretation of gene expression data. The inspection should verify that the assembled sequence matches the consensus of the aligned reads at every position in these critical regions.

Common Assembly Error Patterns and Their Visual Signatures

Misjoined Contigs and Chimeric Sequences

Misjoined contigs represent one of the most serious assembly errors because they create chimeric sequences that do not exist in the actual genome. These errors typically occur when an assembly algorithm incorrectly connects two sequences from different genomic locations. The resulting contig contains sequence from both locations, with no biological relationship between the two portions.

The visual signature of a misjoin in IGV includes a sharp transition in read coverage at the junction, reads that align to one side of the junction but not the other, and paired-end reads that map to unexpected locations. The coverage on each side of the junction reflects the coverage of the respective genomic locations, which may differ substantially. The absence of reads spanning the junction is the most reliable indicator of a misjoin.

Misjoins can occur at any scale, from small insertions of a few hundred bases to large rearrangements involving entire chromosome arms. The severity of the error depends on the size of the misjoined region and its biological significance. Small misjoins in non-coding regions may have minimal impact, while large misjoins that disrupt gene structure can invalidate downstream analyses.

Collapsed Repeats and Copy Number Errors

Collapsed repeats occur when an assembly algorithm merges multiple similar genomic copies into a single sequence. This error is particularly common in regions with tandem gene families, segmental duplications, or transposable element insertions. The resulting assembly has fewer copies of the repeated sequence than the actual genome, leading to copy number errors.

The visual signature of a collapsed repeat includes elevated coverage across the repeated region, with the coverage level approximately proportional to the number of collapsed copies. The boundaries of the collapsed region show a transition from elevated to normal coverage. Reads within the collapsed region may show higher mapping quality than expected, since they align perfectly to the single assembled copy.

Collapsed repeats are particularly problematic for gene family analysis because they obscure the true copy number and sequence diversity of the gene family. The KRAB-zinc finger clusters in mice illustrate this challenge, with their rapidly evolving gene families and ERV insertions creating conditions where collapse is likely. Detecting collapsed repeats requires careful comparison of coverage levels with expected copy numbers.

Expanded Repeats and False Duplications

Expanded repeats represent the opposite error from collapsed repeats, where an assembly algorithm creates multiple copies of a sequence that exists only once in the genome. This error is less common than collapse but can occur in regions with complex repeat structure. The resulting assembly has more copies of the repeated sequence than the actual genome.

The visual signature of an expanded repeat includes reduced coverage across the repeated region, since reads from the single genomic copy are distributed across multiple assembled copies. The boundaries of the expanded region show a transition from normal to reduced coverage. Reads may show lower mapping quality within the expanded region, since they align to multiple similar assembled copies.

Expanded repeats are particularly problematic for genome size estimation and gene copy number analysis. The false duplication inflates the apparent genome size and creates artificial gene family members that do not exist in the actual genome. Detecting expanded repeats requires careful comparison of coverage levels with expected copy numbers.

Unresolved Gaps and Missing Sequence

Unresolved gaps represent regions where the assembly contains no sequence, typically represented by runs of N characters in the FASTA file. These gaps occur when the assembly algorithm cannot determine the sequence in a region, often due to repetitive content or insufficient read coverage. The gaps may represent genuine missing sequence or may indicate regions where the assembly failed.

The visual signature of an unresolved gap in IGV includes a region with no reference sequence, no aligned reads, and a gap in the coverage track. The size of the gap indicates the amount of missing sequence. Flanking regions may show normal coverage and read alignment, with the gap representing an island of missing data.

Unresolved gaps are problematic for downstream analysis because they create regions where no sequence information is available. Genes that span gaps cannot be fully annotated, and structural variants that fall within gaps cannot be detected. The inspection should document the location and size of all gaps to inform downstream analysis decisions.

Base Errors and Small Indels

Base errors and small indels represent the smallest scale of assembly errors, affecting individual nucleotides or short sequence stretches. These errors can arise from sequencing errors that were not corrected during assembly, or from alignment errors during the polishing process. While individual base errors may have minimal impact, they can be significant in coding regions or regulatory elements.

The visual signature of base errors in IGV includes mismatches between the assembled reference sequence and the aligned reads. The mismatches appear as colored bases in the read alignment track, with the color indicating the alternate base. Small indels appear as gaps in the read alignment or as extra bases in the read sequence.

The frequency of base errors varies across the assembly, with some regions showing high error rates and others showing perfect accuracy. The inspection should document the error rate in regions of interest and determine whether the errors are concentrated in specific sequence contexts, such as homopolymer runs or repetitive regions.

At a Glance: Assembly Error Patterns and IGV Signatures

Error TypePrimary Visual Signature in IGVTypical Genomic ContextRecommended Action
Misjoined contigSharp coverage transition, no reads spanning junction, discordant read pairsContig boundaries, repeat-flanked regionsConfirm with flanking analysis, consider reassembly or masking
Collapsed repeatElevated coverage proportional to copy number, high mapping quality readsTandem gene families, segmental duplications, ERV-rich lociCompare coverage to expected copy number, consider targeted reassembly
Expanded repeatReduced coverage across region, lower mapping qualityComplex repeat structures, recent duplicationsVerify copy number with orthogonal methods, correct if confirmed
Unresolved gapNo reference sequence, no aligned reads, coverage gapRepetitive content, low coverage regionsDocument gap coordinates, assess impact on downstream analysis
Base errors and small indelsMismatches or gaps in read alignment at specific positionsHomopolymer runs, low complexity sequenceEvaluate error rate, polish if concentrated in critical regions

Practical Implementation Steps for Assembly Inspection

Setting Up the Inspection Environment

The inspection environment requires IGV installed on a computer with sufficient memory to load the assembly and alignment files. The memory requirements depend on the size of the assembly and the number of aligned reads. Large genomes such as the 880 Mb bighead catfish assembly require substantial memory for efficient navigation.

The IGV software can be downloaded from the official distribution site and runs on Windows, macOS, and Linux operating systems. The software requires Java runtime environment and benefits from a computer with multiple processor cores and substantial RAM. For very large assemblies, a dedicated workstation or server may be necessary.

The input files must be prepared before starting the inspection. The assembly FASTA file must be indexed, the BAM file must be sorted and indexed, and any annotation files must be in the correct format. The file preparation steps are standard bioinformatics operations that can be performed using command-line tools. Training in these foundational computing skills is available through The Carpentries lessons, which cover shell, data, and programming fundamentals that support bioinformatics workflows.

Organizing the Inspection Workflow

A systematic inspection workflow ensures that all priority regions are examined and that the results are documented consistently. The workflow should include a checklist of regions to inspect, a standard set of observations to record for each region, and a system for tracking the inspection status of each region.

The inspection checklist should include all contig boundaries, all regions with known repetitive content, all genes of interest, and any regions flagged by automated quality assessment tools. The checklist should be organized by contig and genomic coordinate, allowing the inspector to work through the regions systematically.

The observation record for each region should include the contig name, start and end coordinates, coverage statistics, read alignment patterns, and any anomalies detected. The record should also include a quality assessment for the region, indicating whether the assembly appears correct or contains errors that require attention.

Documenting Inspection Results

Documentation of inspection results is essential for reproducibility and for communicating findings to collaborators. The documentation should include the inspection date, the software version used, the input file versions, and the specific observations made for each region.

The documentation format should be consistent across all inspected regions, allowing easy comparison and aggregation of results. A spreadsheet or table format works well for recording coverage statistics and quality assessments. Screenshots of IGV views provide visual evidence for specific findings and can be included in the documentation.

The documentation should distinguish between confirmed assembly errors, suspected errors that require additional investigation, and regions that appear correct. This classification guides follow-up actions, with confirmed errors requiring correction or masking and suspected errors requiring additional analysis.

Records and Measurements for Assembly Quality Assessment

Coverage Statistics as Quality Indicators

Coverage statistics provide quantitative measures of assembly quality that complement visual inspection. The mean coverage across the assembly indicates the overall sequencing depth, while the coverage distribution reveals regions with unusually high or low depth. The coefficient of variation provides a measure of coverage uniformity.

For a diploid genome, the expected coverage follows a distribution centered on the mean sequencing depth. Regions with coverage substantially above the mean suggest collapsed repeats, while regions with coverage substantially below the mean suggest expanded repeats or missing sequence. The inspection should document coverage statistics for each inspected region and compare them with the genome-wide distribution.

The coverage statistics should be calculated separately for different sequence contexts, such as coding regions, repetitive regions, and intergenic regions. This context-specific analysis reveals whether coverage anomalies are concentrated in particular sequence types, which informs the interpretation of the anomalies.

Read Alignment Statistics for Error Detection

Read alignment statistics provide additional evidence for assembly quality assessment. The mapping rate indicates the proportion of reads that align to the assembly, with low mapping rates suggesting assembly errors or contamination. The mapping quality distribution reveals the confidence of read alignments, with low mapping qualities indicating ambiguous alignments.

The insert size distribution for paired-end reads provides information about the consistency of read pairs. Reads with insert sizes that deviate substantially from the expected distribution may indicate misjoins or structural errors. The orientation of read pairs also provides evidence, with unexpected orientations suggesting assembly problems.

The alignment statistics should be calculated for the entire assembly and for individual contigs. Contigs with unusually low mapping rates or unusual alignment patterns warrant closer inspection. The statistics provide a quantitative basis for prioritizing regions for visual inspection.

Quality Value and Completeness Metrics

Quality value (QV) and completeness metrics provide aggregate measures of assembly accuracy. The QV represents the base-level accuracy of the assembly, with higher values indicating fewer errors. A QV of 50, as reported for the bighead catfish assembly, corresponds to an estimated error rate of 1 in 100,000 bases.

BUSCO completeness scores measure the proportion of conserved single-copy orthologs present in the assembly. A completeness score of 95.5% for the bighead catfish assembly indicates that most conserved genes are present and complete. Missing or fragmented BUSCO genes may indicate assembly errors in those regions.

These aggregate metrics provide context for the visual inspection but do not replace it. A high QV and completeness score do not guarantee that every region is correctly assembled, and a low score in a specific region may indicate a localized problem that requires investigation.

Common Failure Patterns in Assembly Inspection

Overlooking Low-Complexity Regions

Low-complexity regions, including homopolymer runs, simple repeats, and regions with extreme base composition, present special challenges for assembly inspection. These regions often show unusual coverage patterns and read alignment characteristics that can be mistaken for assembly errors. The inspector must distinguish genuine assembly errors from the expected behavior of low-complexity sequence.

Homopolymer runs, for example, often show reduced read coverage because sequencing technologies have difficulty reading long stretches of the same base. The reduced coverage does not necessarily indicate an assembly error. Similarly, simple repeats may show variable coverage because reads align ambiguously to multiple repeat copies.

The inspection should account for the expected behavior of low-complexity regions when interpreting coverage and alignment patterns. Regions with known low-complexity content should be flagged for special consideration, and the inspection should document the expected coverage patterns for these regions.

Misinterpreting Biological Variation as Assembly Errors

Biological variation can produce coverage and alignment patterns that resemble assembly errors. Tandem gene families, segmental duplications, and transposable element insertions all create elevated coverage that could be mistaken for collapsed repeats. The inspector must distinguish genuine biological variation from assembly artifacts.

The KRAB-zinc finger clusters in mice provide an example of biological variation that complicates assembly inspection. These clusters contain rapidly evolving gene families with high sequence similarity, and their ERV insertions create complex repeat structures. The elevated coverage in these regions reflects genuine biological copy number variation, not assembly collapse.

The inspection should use annotation data and comparative genomics to distinguish biological variation from assembly errors. Regions with known biological copy number variation should be interpreted differently from regions where the copy number is unexpected. The inspection should document the evidence for biological variation when interpreting coverage anomalies.

Failing to Document Inspection Decisions

The failure to document inspection decisions represents a common failure pattern that undermines the value of manual inspection. Without documentation, the inspection results cannot be reproduced, verified, or communicated to collaborators. The documentation should record the observations and the reasoning behind quality assessments.

The documentation should include the specific evidence that led to each quality assessment, including coverage statistics, read alignment patterns, and annotation context. This evidence allows other researchers to verify the assessment and to understand the basis for the decision. The documentation should also record any uncertainties or ambiguities in the assessment.

The documentation format should be standardized across all inspected regions, allowing easy comparison and aggregation. The documentation should be stored with the assembly files and made available to all researchers using the assembly.

Limitations of Manual Assembly Inspection

Time and Resource Constraints

Manual assembly inspection is time-intensive, particularly for large genomes. A complete inspection of an 880 Mb genome at base-level resolution is impractical, requiring the inspector to prioritize regions for examination. The time required for each region depends on the complexity of the region and the experience of the inspector.

The time constraints of manual inspection mean that some assembly errors will inevitably be missed. The inspection provides a sampling of assembly quality instead of a complete assessment. The sampling strategy should prioritize regions where errors are most likely and where errors would have the greatest impact on downstream analysis.

The resource constraints of manual inspection also include the computational requirements for loading and navigating large assembly files. IGV requires substantial memory for large genomes, and the navigation can be slow on less powerful computers. The inspection environment should be configured to provide adequate performance for the assembly size.

Subjectivity in Quality Assessment

Manual inspection involves subjective judgments about assembly quality. Different inspectors may reach different conclusions about the same region, depending on their experience and interpretation of the evidence. The subjectivity of the assessment should be acknowledged and addressed through standardized protocols and documentation.

The standardization of inspection protocols reduces subjectivity by defining specific criteria for quality assessment. The criteria should include quantitative measures such as coverage thresholds and qualitative measures such as read alignment patterns. The criteria should be documented and applied consistently across all inspected regions.

The subjectivity of manual inspection also means that the results should be interpreted with appropriate caution. The inspection provides evidence about assembly quality but does not provide a definitive assessment. The results should be combined with automated quality metrics and other evidence to form a comprehensive assessment.

Incomplete Coverage of the Assembly

Manual inspection cannot examine every region of a large assembly, leaving some regions uninspected. The uninspected regions may contain assembly errors that would be detected by a complete inspection. The inspection results therefore provide a lower bound on the number of assembly errors, with the true number potentially higher.

The incomplete coverage of the assembly should be acknowledged in the inspection documentation. The documentation should specify which regions were inspected and which were not, allowing users of the assembly to understand the scope of the quality assessment. The documentation should also specify the criteria used to prioritize regions for inspection.

The incomplete coverage of manual inspection can be partially addressed by combining it with automated quality assessment tools. Automated tools can flag regions for inspection, and the manual inspection can focus on the flagged regions. This combined approach provides more complete coverage than manual inspection alone.

Professional Escalation Criteria for Assembly Issues

When to Seek Additional Expertise

Certain assembly issues require expertise beyond the scope of routine manual inspection. Complex structural rearrangements, unusual repeat structures, and regions with conflicting evidence may require specialized analysis. The inspector should recognize when the available evidence is insufficient to reach a confident assessment.

The escalation criteria should include situations where the assembly error pattern is unfamiliar, where the evidence is contradictory, or where the downstream impact of the error is severe. The inspector should document the evidence and the reason for escalation, providing the specialist with the context needed to address the issue.

The escalation process should be defined in advance, with clear criteria for when to seek additional expertise. The process should include the documentation requirements and the expected timeline for resolution. The escalation should be initiated promptly when the criteria are met, avoiding delays in the assembly validation process.

When to Consider Reassembly or Polishing

Some assembly issues cannot be resolved through manual inspection alone and require reassembly or polishing. The decision to reassemble or polish should be based on the severity and extent of the identified errors. The decision should consider the cost of reassembly in terms of time and computational resources.

The criteria for reassembly include extensive misjoins, widespread collapsed repeats, or base error rates that exceed acceptable thresholds. The criteria for polishing include localized base errors or small indels that can be corrected without full reassembly. The decision should be documented with the evidence supporting the need for reassembly or polishing.

The reassembly or polishing process should be followed by a new round of manual inspection to verify that the identified errors have been corrected. The verification inspection should focus on the regions where errors were previously identified, confirming that the corrections are effective.

When to Mask or Exclude Problematic Regions

Some assembly issues cannot be resolved through reassembly or polishing, requiring the problematic regions to be masked or excluded from downstream analysis. The decision to mask or exclude should be based on the severity of the errors and the impact on downstream analysis.

The criteria for masking or excluding include unresolvable gaps, regions with persistent misjoins, and regions with collapsed repeats that cannot be resolved. The masking or exclusion should be documented, with the specific regions and the reason for the decision recorded. The documentation should be available to all users of the assembly.

The masking or exclusion decision should consider the biological significance of the affected regions. Regions containing genes of interest may warrant additional effort to resolve, while regions with no known biological significance may be masked or excluded more readily. The decision should balance the cost of resolution against the impact of exclusion.

Training and Reproducibility Considerations

Building Inspection Skills Through Structured Training

Manual assembly inspection requires a combination of conceptual knowledge and practical skills. Structured training programs provide the foundation for both. The EMBL-EBI training portal offers learning pathways covering bioinformatics data resources and practical analysis education, including topics relevant to genome assembly and visualization. These training resources help researchers develop the skills needed to interpret assembly quality evidence effectively.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover genome assembly and quality assessment. These tutorials offer hands-on experience with assembly tools and visualization approaches, allowing researchers to practice inspection techniques in a guided environment. The practical experience gained through these tutorials translates directly to independent inspection work.

Foundational computing skills are equally important for assembly inspection. The Carpentries lessons cover shell, data, and programming fundamentals that support bioinformatics workflows. These skills enable researchers to prepare input files, run alignment tools, and manage inspection documentation efficiently.

Reproducibility in Assembly Inspection

Reproducibility in manual assembly inspection requires careful attention to version control and documentation. The software versions used for alignment, indexing, and visualization should be recorded, since different versions may produce different results. The input file versions should also be documented, including the assembly FASTA file, the read data, and any annotation files.

The inspection workflow itself should be documented in sufficient detail that another researcher could repeat the process. The documentation should include the specific regions inspected, the order of inspection, and the criteria used for quality assessment. This level of detail allows the inspection to be verified and extended by other researchers.

Reproducible workflow standards from the bioinformatics community provide guidance for structuring assembly inspection as part of a larger analysis pipeline. The nf-core documentation describes community pipeline standards for usage, configuration, and reproducible workflow context. These standards emphasize version control, containerization, and automated documentation, which can be applied to assembly inspection workflows.

Frequently Asked Questions

What is the minimum coverage needed for reliable assembly inspection in IGV?

The minimum coverage for reliable inspection depends on the sequencing technology and the error type being investigated. Higher coverage provides more evidence for assessing assembly quality, while lower coverage makes it difficult to distinguish genuine errors from alignment artifacts. For long-read assemblies, coverage of at least 30x provides sufficient read depth for most inspection purposes, while lower coverage may still allow detection of large-scale errors such as misjoins. The inspection should document the coverage level for each region and interpret the evidence accordingly.

How do I distinguish a collapsed repeat from a genuine tandem duplication?

A collapsed repeat shows elevated coverage that is approximately proportional to the number of collapsed copies, with reads from all copies aligning to the single assembled locus. A genuine tandem duplication also shows elevated coverage, but the coverage level matches the expected copy number and read pairs span the duplication junction. The distinction requires careful examination of read alignment patterns at the boundaries of the elevated coverage region. Annotation data and comparative genomics can provide additional evidence for distinguishing these two possibilities.

Can IGV detect all types of assembly errors?

IGV can detect assembly errors that produce visible patterns in coverage and read alignment, including misjoins, collapsed repeats, expanded repeats, and base errors. However, some errors may not produce visible patterns, particularly in regions with low coverage or complex repeat structure. IGV inspection should be combined with automated quality metrics and other validation approaches for a comprehensive assessment. The limitations of visual inspection should be acknowledged in the interpretation of results.

What should I do when I find a suspected misjoin in my assembly?

When a suspected misjoin is identified, the first step is to document the evidence, including the contig name, coordinates, coverage statistics, and read alignment patterns. The next step is to examine the flanking regions to determine whether the misjoin is supported by additional evidence. If the misjoin is confirmed, the decision to correct, mask, or exclude the region should be based on its severity and impact on downstream analysis. The documentation should be shared with collaborators and included in the assembly quality assessment.

How long does a manual assembly inspection take?

The time required for manual assembly inspection depends on the size of the assembly, the number of priority regions, and the experience of the inspector. A focused inspection of priority regions in a large genome can take several days, while a complete inspection of a small genome may take only a few hours. The time investment should be balanced against the value of the inspection for downstream analysis. The inspection should be prioritized to focus on regions where errors are most likely and most impactful.

What are the most common mistakes in manual assembly inspection?

The most common mistakes include overlooking low-complexity regions, misinterpreting biological variation as assembly errors, and failing to document inspection decisions. These mistakes can lead to incorrect quality assessments and missed assembly errors. The mistakes can be avoided through standardized inspection protocols, careful documentation, and appropriate interpretation of evidence. The inspection should be conducted systematically, with clear criteria for quality assessment and documentation.

How do I prepare my assembly files for IGV inspection?

The assembly FASTA file must be indexed, and the aligned reads BAM file must be sorted and indexed. The index files allow IGV to access specific regions efficiently. Annotation files in GFF or BED format can be loaded as additional tracks for context. The file preparation steps are standard bioinformatics operations that can be performed using command-line tools. The prepared files should be verified before starting the inspection.

What qualifications do I need to perform manual assembly inspection?

Manual assembly inspection requires familiarity with genome assembly concepts, read alignment, and the IGV software. Training in bioinformatics fundamentals provides the foundation for understanding the evidence and interpreting the patterns. The Carpentries lessons offer foundational computing and data skills that support bioinformatics analysis. EMBL-EBI training provides learning pathways for bioinformatics data resources and practical analysis education. The inspection skills develop with practice and experience.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.