# Bandage vs. IGV: Choosing the Right Visualization Tool for Your Genome Assembly


## Key Takeaways

- Bandage visualizes the assembly graph structure (GFA format), revealing contig connectivity, branching patterns indicative of repeats or heterozygosity, and graph completeness via connectivity analysis. This is crucial for initial de novo assembly inspection and troubleshooting complex repetitive regions.
- IGV visualizes linear alignments of sequencing reads and assembled contigs against a reference genome, enabling base-pair resolution inspection of mismatches, gaps, and structural variations. This is essential for post-assembly validation and for informing polishing decisions by examining read support for sequence errors.
- Bandage is indispensable for understanding the inherent structure of a de novo assembly, particularly for long-read data where graph complexity reflects biological variation (e.g., tandem arrays) and assembly challenges. It directly addresses how contigs are connected without requiring a reference.
- IGV is critical for validating assembled contigs against a known reference, identifying misassemblies as disruptions in linear alignment, and assessing the quality of read support for polishing. It answers questions about how the assembly conforms to established genomic coordinates.
- A comprehensive workflow necessitates using both Bandage and IGV iteratively: Bandage for initial graph assessment and repeat analysis, followed by IGV for linear validation and read-level error inspection, potentially returning to Bandage for graph-level investigation of issues identified in linear view.
- Quantitative metrics from Bandage (node count, component number) and coverage analysis within the graph are vital for assessing assembly fragmentation and repeat resolution, complementing visual inspection. IGV provides alignment statistics (coverage, mismatch rates) for quantitative validation against a reference.

---

Genome assembly visualization serves two distinct purposes that require different tools. Bandage displays the assembly graph structure itself, showing how contigs connect through branching and repetitive regions. IGV displays aligned sequencing reads and assembled contigs against a linear reference coordinate system. Researchers working with de novo assemblies, particularly those generated from long-read sequencing, need both tools at different stages of the assembly and polishing workflow. This article provides a practical comparison of Bandage and IGV, focusing on graph-level versus linear visualization, with concrete guidance on when each tool is appropriate for assembly inspection, quality assessment, and polishing decisions.

## Understanding the Two Visualization Paradigms

### Assembly Graphs as the Primary Output of Modern Assemblers

Modern genome assemblers do not produce a simple linear sequence as their direct output. Instead, they generate a graph structure that represents the relationships between sequence fragments. The graphical fragment assembly (GFA) format stores this graph information, including nodes that represent sequence segments and edges that represent connections between those segments. This graph structure is particularly important for long-read assemblies, where repetitive regions create complex branching patterns that must be resolved during the assembly process.

Recent work on phased genomes of polar fishes demonstrates that assembly uncertainty is common in repetitive regions, particularly in tandem arrays of genes that have undergone copy number expansion. The researchers developed gfa_parser, a tool that computes and extracts all possible contiguous sequences from GFA files, revealing that standard processing of graphical fragment assemblies can bias measurements of haplotype copy number variation. This finding underscores the importance of examining the assembly graph directly instead of relying solely on linearized output.

Bandage was designed specifically to visualize these assembly graphs. The tool reads GFA files and renders the graph structure as a network diagram, allowing researchers to see how contigs connect, where branches occur, and which regions of the assembly are unresolved. This visualization is essential for understanding whether an assembly is complete, whether repetitive regions have been properly collapsed or expanded, and whether misassemblies have created incorrect connections between unrelated genomic regions.

### Linear Visualization for Read Alignment and Variant Inspection

IGV operates on a fundamentally different principle. The tool displays genomic data along a linear coordinate axis, typically using a reference genome as the backbone. Researchers load aligned sequencing reads, annotation tracks, and variant calls onto this linear framework, allowing them to inspect specific genomic regions at base-pair resolution.

For genome assembly work, IGV serves several important functions. After an assembly is complete, researchers can align the assembled contigs back to a reference genome and visually inspect the alignment quality. This inspection reveals structural variations, misassembled regions, and areas where the assembly may have introduced errors. IGV also supports the visualization of RNA-seq data aligned to an assembly, which is useful for gene annotation and expression analysis.

The distinction between these two visualization paradigms is fundamental. Bandage answers the question of how the assembly is structured as a graph. IGV answers the question of how sequence data aligns to a linear reference. Both questions are relevant during the assembly process, but they arise at different stages and require different analytical approaches.

## Core Principles of Assembly Graph Inspection

### Identifying Branching Structures and Repeat-Induced Complexity

Assembly graphs contain branching structures that indicate either genuine genomic complexity or assembly errors. In a perfect assembly of a haploid genome, the graph would be a simple linear path with no branches. In practice, repetitive regions create branches because identical or nearly identical sequence repeats cannot be unambiguously placed in the assembly.

When inspecting an assembly graph in Bandage, researchers should look for several structural features. Bubbles in the graph represent regions where two alternative paths exist between the same nodes, often indicating heterozygous variants or assembly uncertainty. Complex tangles of interconnected nodes typically indicate highly repetitive regions that the assembler could not resolve. Long unbranched paths represent confidently assembled regions that are ready for downstream analysis.

The polar fish study provides a concrete example of why graph inspection matters. The researchers found that assembly uncertainty was ubiquitous across antifreeze protein gene arrays, which are repetitive tandem arrays that vary in copy number between haplotypes. Without examining the graph structure directly, measurements of copy number variation would be biased by the way standard assembly processing handles these ambiguous regions.

### Evaluating Assembly Completeness Through Graph Connectivity

Assembly completeness can be assessed by examining whether the graph contains a single connected component or multiple disconnected components. A complete assembly of a single chromosome should produce a graph that is connected, with all nodes reachable from any starting point. Disconnected components indicate gaps in the assembly where the assembler could not bridge between adjacent genomic regions.

Bandage provides tools for exploring graph connectivity, including the ability to highlight paths between specified nodes and to identify dead ends where the graph terminates without connecting to other nodes. These dead ends may represent telomeres, assembly gaps, or regions where sequencing coverage was insufficient to establish connections.

For phased assemblies, the graph structure becomes more complex because each haplotype must be represented separately. The spinach downy mildew pangenome study used haplotype-resolved pangenome graph analysis to examine effector candidate genes across 19 pathogen isolates. This analysis required careful inspection of graph structures to distinguish genuine haplotype variation from assembly artifacts.

## Core Principles of Linear Alignment Visualization

### Reference-Based Inspection of Assembled Contigs

Once an assembly is complete, researchers often want to compare it to a reference genome or to examine specific genomic regions in detail. This comparison requires aligning the assembled contigs to a reference and visualizing the alignment in a linear format. IGV provides this functionality through its support for BAM files containing aligned reads or contigs.

When inspecting an assembly aligned to a reference in IGV, researchers should examine several features. Mismatches between the assembly and reference may indicate genuine sequence differences or assembly errors. Gaps in the alignment may indicate missing sequence or misassembled regions. Breakpoints where the alignment jumps between distant reference positions may indicate structural rearrangements in the assembly.

The RPE-1 cell line genome project provides an example of reference-based assembly validation. The researchers generated a near-complete diploid assembly of the hTERT RPE-1 cell line and compared both haplotypes with the CHM13 human reference genome. This comparison detected haplotype-specific genomic variations, including a translocation between chromosome 10 and chromosome X that is characteristic of RPE-1 cells. Visual inspection of these structural variants in a linear viewer was essential for validating the assembly and understanding its relationship to the reference.

### Examining Read Alignments for Polishing Decisions

Assembly polishing requires examining how sequencing reads align to the assembled contigs. This examination reveals sequencing errors in the assembly that can be corrected by incorporating information from multiple reads covering the same genomic position. IGV is well suited for this task because it displays individual reads aligned to a linear reference, showing mismatches, insertions, and deletions at base-pair resolution.

When inspecting read alignments for polishing decisions, researchers should look for consistent patterns of errors. A position where many reads show the same mismatch may indicate a systematic sequencing error or a genuine sequence difference. A position where reads show conflicting bases may indicate a heterozygous variant or a region of low sequencing quality. Understanding these patterns helps researchers decide which polishing tools to use and how to interpret their results.

The Galaxy Training Network provides accessible tutorials on assembly and polishing workflows that incorporate both graph-based and linear visualization approaches. These tutorials emphasize the importance of visual inspection at multiple stages of the assembly process, from initial graph construction through final polishing and validation.

## At a Glance: Bandage versus IGV for Assembly Work

| Feature | Bandage | IGV |
|---------|---------|-----|
| Primary visualization | Assembly graph structure from GFA files | Linear alignments of reads and contigs to a reference |
| Input formats | GFA files from assemblers | BAM, VCF, GFF, FASTA, and related genomic formats |
| Best use stage | Initial assembly inspection, graph troubleshooting, repeat resolution | Post-assembly validation, polishing decisions, variant inspection |
| Key questions answered | How are contigs connected? Where are branches and repeats? | How do reads align to the assembly? Where are errors and variants? |
| Reference requirement | None, works directly on assembly graph | Requires a reference sequence for coordinate-based display |
| Typical user | Assembly specialists, genome project teams | Broad genomics community, including variant analysis |
| Learning curve | Moderate, requires understanding of graph concepts | Lower, familiar genome browser interface |
| Output for records | Graph screenshots, path extractions, connectivity statistics | Alignment screenshots, coverage tracks, variant views |

## Practical Workflow for Assembly Visualization

### Step 1: Initial Graph Inspection with Bandage

Begin the assembly visualization workflow by loading the GFA file produced by your assembler into Bandage. This initial inspection provides an overview of the assembly structure and reveals major issues that require attention before proceeding to detailed analysis.

During this first inspection, record the following observations. Note the number of connected components in the graph and whether they correspond to expected chromosome or scaffold numbers. Identify any large branching structures that may indicate unresolved repeats. Check for dead ends that may represent assembly gaps. Capture screenshots of the overall graph and of specific regions that show unusual structure.

For phased assemblies, examine whether the graph contains separate paths for each haplotype. The polar fish study demonstrated that phased assemblies of repetitive regions require careful graph inspection to avoid bias in copy number measurements. If the graph shows unexpected complexity in a region that should be simple, investigate whether the complexity represents genuine biological variation or assembly error.

### Step 2: Targeted Graph Exploration for Problem Regions

After the initial overview, focus on specific regions that show problematic structure. Use Bandage's navigation tools to zoom into these regions and examine the connections between nodes in detail. Identify whether branches represent alternative paths through the same sequence or connections between different genomic regions.

For each problem region, record the following information. Describe the graph structure, including the number of nodes involved and the pattern of connections. Note whether the region contains sequences that match known repeats or transposable elements. Document the coverage information associated with each node, as coverage differences can indicate collapsed repeats or haplotype-specific sequences.

The pangenome representation review emphasizes that visualization formats must effectively convey complex genomic structures and variations. When examining problem regions in your assembly graph, consider whether the structure you observe represents genuine variation that should be preserved in the final assembly or an artifact that should be resolved through additional sequencing or assembly parameter adjustment.

### Step 3: Linear Alignment for Validation

Once the assembly graph appears structurally sound, align the assembled contigs to an appropriate reference genome and load the alignment into IGV. This step provides a different perspective on assembly quality, revealing how the assembly relates to known genomic coordinates.

When examining the linear alignment, record the following observations. Note regions where the assembly aligns cleanly to the reference with few mismatches or gaps. Identify regions where the alignment breaks down, showing large insertions, deletions, or rearrangements. Check whether the assembly covers the expected genomic regions completely or whether gaps exist that require additional sequencing.

For projects without a close reference genome, consider using a related species or a pangenome representation as the alignment target. The spinach downy mildew study constructed a pangenome graph from 19 isolates, providing a reference framework that captured diversity across the species. This approach allowed the researchers to examine effector variation in the context of the broader species diversity instead of against a single reference isolate.

### Step 4: Read Alignment Inspection for Polishing

Before running polishing tools, inspect how the original sequencing reads align to the assembled contigs. This inspection provides baseline information about sequencing quality and reveals systematic errors that polishing should address.

Load the read alignments into IGV and examine several regions across the assembly. Look for consistent mismatch patterns that indicate systematic sequencing errors. Check whether coverage is uniform across the assembly or whether some regions have unusually high or low coverage. Identify any regions where reads show evidence of misassembly, such as reads that span unexpected junctions or that have inconsistent insert sizes.

Record the specific positions where polishing is likely to make corrections. Note whether errors appear to be random or clustered in specific sequence contexts. This information helps in selecting appropriate polishing tools and in evaluating whether polishing has been successful after the fact.

### Step 5: Post-Polishing Validation

After running polishing tools, repeat the linear alignment inspection to verify that polishing improved assembly quality. Compare the pre-polishing and post-polishing alignments to identify positions where errors were corrected and to check whether polishing introduced any new errors.

For each polished region, record the following information. Document the number of corrections made and their locations. Note whether any positions show new mismatches that were not present before polishing. Check whether polishing affected the assembly graph structure, particularly in repetitive regions where aggressive polishing can sometimes introduce errors.

The nf-core documentation emphasizes the importance of reproducible workflow standards in genomic analysis. When documenting your polishing validation, record the exact parameters used for polishing tools and the versions of all software involved. This documentation ensures that the polishing process can be reproduced and evaluated by other researchers.

## Options and Tradeoffs in Visualization Tool Selection

### When Bandage Is the Appropriate Choice

Bandage is the appropriate visualization tool when the primary question concerns the structure of the assembly graph itself. This situation arises during initial assembly inspection, when troubleshooting assembly failures, and when examining repetitive regions that create complex graph structures.

Choose Bandage when you need to understand how contigs connect to each other. The graph visualization reveals whether the assembly is fragmented, whether repeats have been properly resolved, and whether misassemblies have created incorrect connections. This information is essential for deciding whether to accept an assembly, to adjust assembly parameters and rerun, or to generate additional sequencing data to resolve problematic regions.

Bandage is also the appropriate choice when working with assemblies that lack a close reference genome. Because Bandage works directly on the GFA file without requiring a reference, it can be used for any assembly project regardless of whether a suitable reference exists. This makes it particularly valuable for non-model organisms and for projects exploring novel genomic diversity.

### When IGV Is the Appropriate Choice

IGV is the appropriate visualization tool when the primary question concerns how sequence data aligns to a linear reference. This situation arises during post-assembly validation, when examining specific genomic regions in detail, and when integrating assembly data with other genomic datasets.

Choose IGV when you need to examine base-pair level details of the assembly. The linear visualization shows individual mismatches, insertions, and deletions, providing the resolution needed for polishing decisions and variant identification. This level of detail is not available in graph-based visualization, which focuses on the overall structure instead of individual sequence positions.

IGV is also the appropriate choice when working with multiple data types that need to be integrated. The tool supports the simultaneous display of read alignments, variant calls, gene annotations, and other genomic features. This integration is essential for understanding how the assembly relates to functional genomic elements and for identifying variants that may have biological significance.

### Combining Both Tools in a Complete Workflow

Most assembly projects benefit from using both Bandage and IGV at different stages of the workflow. The tools address different questions and provide complementary information about assembly quality and structure.

A typical workflow begins with Bandage for initial graph inspection, moves to IGV for detailed alignment analysis, and returns to Bandage if structural issues are identified that require graph-level investigation. This iterative approach ensures that both the overall structure and the detailed sequence are examined and validated.

The Bioconductor project provides packages that support both graph-based and linear genomic analysis within the R programming environment. These packages allow researchers to move between graph and linear representations programmatically, supporting reproducible analysis workflows that combine both visualization approaches.

## Observations and Measurements for Assembly Assessment

### Quantitative Metrics from Graph Analysis

Bandage provides several quantitative measurements that support assembly assessment. The number of nodes in the graph indicates the level of fragmentation, with more nodes generally indicating a more fragmented assembly. The number of connected components indicates whether the assembly is complete or contains gaps. The length distribution of nodes provides information about the contiguity of the assembly.

Record these metrics for each assembly you evaluate. Compare the metrics across different assembly runs to determine whether parameter changes improved assembly quality. Document the metrics alongside the assembly parameters so that the relationship between parameters and outcomes can be examined.

The polar fish study used gfa_parser to extract all possible contiguous sequences from GFA files, providing a systematic approach to measuring assembly uncertainty in repetitive regions. This approach demonstrates that quantitative analysis of graph structure can reveal biases that are not apparent from visual inspection alone.

### Coverage-Based Measurements for Repeat Resolution

Coverage information stored in the assembly graph provides important clues about repeat structure. Regions with approximately double the average coverage may indicate collapsed repeats where two identical copies were assembled as one. Regions with coverage lower than average may indicate haplotype-specific sequences in a heterozygous diploid assembly.

When examining coverage in Bandage, record the coverage values for nodes in regions of interest. Compare these values to the genome-wide average coverage. Note any regions where coverage deviates substantially from the expected value, as these regions may require additional investigation.

The RPE-1 cell line project used high-coverage Pacific Biosciences and Oxford Nanopore Technologies long-read sequencing to generate a high-quality de novo assembly. The high coverage provided confidence in the assembly and allowed the researchers to validate the assembly through multiple methods, including coverage-based assessments of repeat resolution.

### Alignment Statistics for Assembly Validation

After aligning the assembly to a reference, record alignment statistics that support quality assessment. The percentage of the assembly that aligns to the reference indicates the completeness of the assembly relative to the reference. The percentage of the reference that is covered by the assembly indicates whether the assembly spans the expected genomic regions. The number and size of alignment gaps indicate regions where the assembly may be incomplete or misassembled.

These statistics provide a quantitative complement to visual inspection. While visual inspection reveals the specific regions that require attention, alignment statistics provide an overall assessment of assembly quality that can be compared across different assembly runs and across different projects.

The EMBL-EBI Training program provides educational materials on sequence analysis that include guidance on interpreting alignment statistics and using them to assess assembly quality. These materials emphasize the importance of combining quantitative metrics with visual inspection for a complete assessment.

## Records and Documentation for Assembly Visualization

### Maintaining a Visualization Log

Keep a detailed log of all visualization inspections performed during the assembly process. This log should include the date of each inspection, the software versions used, the assembly version examined, and the specific regions or features that were inspected. Record any issues identified and the actions taken to address them.

The visualization log serves several purposes. It provides a record of the assembly assessment process that can be reviewed by other researchers. It documents the rationale for assembly decisions, supporting the reproducibility of the project. It also provides a basis for comparing assembly versions and for evaluating whether changes to the assembly process improved quality.

The Carpentries lessons emphasize the importance of documentation and reproducible practices in computational research. Following these principles in your visualization work ensures that your assembly assessment can be understood and reproduced by others.

### Capturing and Storing Visualization Images

Capture screenshots of important visualization results and store them with descriptive filenames that include the assembly version, the region examined, and the date. Organize these images in a directory structure that mirrors the assembly workflow, making it easy to locate images from specific stages of the process.

For Bandage images, include the graph structure and any annotations that help interpret the structure. For IGV images, include the genomic coordinates displayed and the tracks that are visible. These images provide visual evidence of the assembly assessment that can be referenced in publications and reports.

The Galaxy Training Network provides guidance on documenting analysis workflows and results, including the capture and storage of visualization outputs. Following these practices ensures that your visualization records are complete and useful for future reference.

### Recording Assembly Parameters and Versions

Document the exact parameters used for each assembly run and the versions of all software involved. This documentation includes the assembler version, the parameters for read trimming and error correction, and the settings for any polishing tools. Record the versions of Bandage and IGV used for visualization, as different versions may display data differently.

This parameter documentation is essential for reproducing the assembly and for understanding how parameter choices affected the final result. When assembly issues are identified through visualization, the parameter documentation allows researchers to determine which parameters should be adjusted and to test whether the adjustments resolve the issues.

The nf-core documentation provides standards for documenting workflow parameters and versions, supporting reproducible analysis. Applying these standards to your assembly project ensures that the relationship between parameters, assembly outcomes, and visualization results is clearly documented.

## Common Failure Patterns in Assembly Visualization

### Misinterpreting Graph Branches as Assembly Errors

A common failure pattern is interpreting all graph branches as assembly errors when some branches represent genuine biological variation. Heterozygous variants in diploid organisms create bubbles in the assembly graph that are biologically meaningful. Repetitive regions with copy number variation create complex graph structures that reflect real genomic diversity.

To avoid this failure, examine the coverage of each branch in the graph. Branches with similar coverage likely represent genuine alternative sequences, such as haplotypes. Branches with very different coverage may indicate collapsed repeats or assembly artifacts. Consider the biological context of the region when interpreting graph structure.

The spinach downy mildew study demonstrated that pangenome graph analysis can distinguish isolates that emerged from recent sexual recombination from those that evolved via prolonged asexual reproduction and loss of heterozygosity. This distinction required careful interpretation of graph structure in the context of the organism's reproductive biology.

### Overlooking Misassemblies in Linear Alignments

Another common failure pattern is overlooking misassemblies when examining linear alignments in IGV. Misassemblies can appear as subtle disruptions in the alignment that are easy to miss during visual inspection. Reads that span unexpected junctions, inconsistent insert sizes, and abrupt changes in coverage can all indicate misassembly.

To avoid this failure, systematically examine the alignment across the entire assembly instead of focusing only on regions of interest. Use the IGV navigation tools to scroll through the assembly at a consistent zoom level, looking for any disruptions in the alignment pattern. Compare the alignment across multiple regions to identify patterns that may indicate systematic misassembly.

The polar fish study developed switch_error_screen to flag potential switch errors, which are changes from one parental haplotype to another in a contiguous assembly. This tool demonstrates that automated detection of misassembly patterns can complement visual inspection, catching errors that might be missed during manual review.

### Relying on a Single Visualization Approach

A third common failure pattern is relying exclusively on either graph-based or linear visualization. Each approach reveals different aspects of assembly quality, and neither is sufficient on its own. Graph visualization reveals structural issues but does not provide base-pair resolution. Linear visualization reveals sequence details but does not show the overall graph structure.

To avoid this failure, integrate both visualization approaches into the assembly workflow. Use Bandage for initial structural assessment and for investigating regions that show unusual graph structure. Use IGV for detailed sequence inspection and for validating the assembly against a reference. Move between the two tools as needed to fully understand the assembly.

The pangenome representation review emphasizes that different representation formats and visualization techniques are needed to convey the complex genomic structures captured in pangenome assemblies. This principle applies equally to individual genome assemblies, where multiple visualization approaches are needed for a complete assessment.

## Limitations of Bandage and IGV for Assembly Work

### Scale Limitations for Large Genomes

Both Bandage and IGV have limitations when working with very large genomes. Bandage may struggle to render the complete graph of a large genome with millions of nodes, requiring the user to focus on specific regions or to use simplified graph representations. IGV may have difficulty loading and displaying alignments for very large genomes, particularly when working with many samples or high coverage data.

For large genome projects, consider using subsampled data for initial visualization and then focusing on specific regions for detailed inspection. The RPE-1 cell line project, which assembled a near-complete diploid human genome, required careful management of visualization resources given the size and complexity of the data.

### Interpretation Challenges in Complex Repetitive Regions

Both tools face interpretation challenges in complex repetitive regions. Graph visualization of highly repetitive regions can produce dense tangles that are difficult to interpret visually. Linear visualization of these regions can show ambiguous alignments where reads map to multiple locations.

In these challenging regions, combine visualization with quantitative analysis. Use coverage statistics and sequence composition analysis to complement visual inspection. Consider using specialized tools designed for repeat analysis, which may provide additional insights beyond what Bandage and IGV can offer.

The polar fish study demonstrated that repetitive tandem arrays of antifreeze protein genes require specialized analysis approaches. The researchers developed gfa_parser specifically to handle the complexity of these regions, showing that standard visualization and analysis tools may be insufficient for highly repetitive genomic regions.

### Reference Bias in Linear Visualization

IGV and other linear visualization tools are inherently reference-based, which introduces potential bias when the reference does not accurately represent the sample being studied. Regions where the sample differs substantially from the reference may show poor alignment that is difficult to interpret. Structural variants that are not present in the reference may be invisible in the linear view.

For projects where the reference is distantly related to the sample, consider whether graph-based visualization should play a larger role in the analysis. The spinach downy mildew study used a pangenome graph approach to avoid reference bias, allowing the researchers to examine effector variation across the full diversity of the species instead of against a single reference isolate.

## Safety and Reproducibility Context for Assembly Visualization

### Ensuring Reproducible Visualization Results

Visualization results should be reproducible, meaning that the same input data and parameters should produce the same visual output. This reproducibility requires documenting the exact versions of Bandage and IGV used, the parameters applied, and the input files examined.

For Bandage, record the GFA file version and any filtering or simplification options applied. For IGV, record the reference genome version, the alignment files loaded, and the display settings used. This documentation allows other researchers to reproduce the visualization and to verify the conclusions drawn from it.

The Bioconductor project provides reproducible analysis frameworks that can be applied to visualization workflows. By integrating visualization into scripted analysis pipelines, researchers can ensure that visualization results are reproducible and can be regenerated as needed.

### Data Management for Visualization Inputs

Proper data management is essential for effective visualization. Keep assembly files organized with clear naming conventions that include the assembly version and date. Store alignment files in a consistent directory structure that makes it easy to locate the files needed for visualization. Maintain backups of all visualization inputs to prevent data loss.

For large projects, consider using a data management system that tracks file versions and relationships. The nf-core documentation provides guidance on data management for reproducible workflows, including standards for file naming and organization.

### Professional Escalation Criteria for Assembly Issues

Some assembly issues identified through visualization require escalation to more specialized expertise. Escalate when the assembly graph shows unexpected complexity that cannot be resolved through parameter adjustment. Escalate when linear alignments reveal structural variants that may indicate biological significance or assembly error. Escalate when repetitive regions show patterns that are not explained by known biology.

When escalating, provide the visualization records and documentation to the specialist. Include the assembly parameters, the visualization log, and the specific regions of concern. This information allows the specialist to understand the issue and to provide informed guidance.

The EMBL-EBI Training program provides pathways for developing specialized bioinformatics skills, including advanced assembly and visualization techniques. Researchers who encounter assembly issues beyond their expertise should consider seeking additional training or consulting with specialists.

## Frequently Asked Questions

### What is the main difference between Bandage and IGV for assembly visualization?

Bandage displays the assembly graph structure directly from GFA files, showing how contigs connect through nodes and edges. IGV displays sequence data aligned to a linear reference coordinate system, showing reads and contigs at base-pair resolution. Bandage answers structural questions about the assembly graph, while IGV answers sequence-level questions about alignments and variants.

### Can I use Bandage to examine read alignments for polishing decisions?

Bandage is not designed for read alignment visualization. The tool focuses on the assembly graph structure and does not display individual read alignments. For polishing decisions that require examining how reads align to assembled contigs, use IGV or another linear alignment viewer that displays read-level data.

### Do I need a reference genome to use Bandage?

No, Bandage works directly on the GFA file produced by the assembler and does not require a reference genome. This makes Bandage suitable for any assembly project, including those for non-model organisms without a close reference. IGV, in contrast, requires a reference sequence for coordinate-based display.

### How do I identify collapsed repeats in my assembly graph?

Collapsed repeats typically appear as nodes with approximately double the average coverage, because two identical copies are represented as a single sequence. Examine coverage values in Bandage and compare them to the genome-wide average. Regions with substantially elevated coverage may indicate collapsed repeats that require additional analysis.

### What should I do if my assembly graph shows many disconnected components?

Disconnected components indicate gaps in the assembly where the assembler could not bridge between adjacent genomic regions. Record the number and size of the components, then investigate whether additional sequencing data or adjusted assembly parameters could resolve the gaps. Consider whether the disconnected components correspond to expected features such as separate chromosomes or organelles.

### How can I verify that polishing improved my assembly?

Compare read alignments before and after polishing in IGV. Look for positions where mismatches were corrected and check whether polishing introduced any new errors. Record the number of corrections made and their locations. Also examine the assembly graph in Bandage to verify that polishing did not disrupt the overall graph structure.

### What are the limitations of using a single reference genome for assembly validation?

A single reference genome may not represent the diversity of the sample being studied, leading to poor alignments in regions where the sample differs from the reference. Structural variants not present in the reference may be invisible in linear visualization. Consider using a pangenome representation or multiple references to reduce reference bias in assembly validation.

### When should I seek help from a specialist for assembly visualization issues?

Seek help when the assembly graph shows unexpected complexity that you cannot resolve through parameter adjustment, when linear alignments reveal structural variants that may indicate biological significance or assembly error, or when repetitive regions show patterns not explained by known biology. Provide the specialist with your visualization records, assembly parameters, and documentation of the specific regions of concern.

## Related Bioinformatics Guides

- [RNA-Seq Alignment: Choosing the Right Tool and Parameters](/knowledge/bioinformatics/rna-seq-alignment-choosing-the-right-tool-and-parameters)
- [Metagenomics vs Metabarcoding: Choosing the Right Approach for Your Study](/knowledge/bioinformatics/metagenomics-vs-metabarcoding-choosing-the-right-approach-for-your-study)
- [Genomic Data Visualization Tools: Choosing and Using Them Effectively](/knowledge/bioinformatics/genomic-data-visualization-tools-choosing-and-using-them-effectively)
- [Metagenomics vs Metatranscriptomics: Choosing the Right Approach for Functional Profiling](/knowledge/bioinformatics/metagenomics-vs-metatranscriptomics-choosing-the-right-approach-for-functional-profiling)
- [Spatial Proteomics vs. Single-Cell Proteomics: Choosing the Right Approach](/knowledge/bioinformatics/spatial-proteomics-vs-single-cell-proteomics-choosing-the-right-approach)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Mitigating assembly and switch errors in phased genomes of polar fishes reveals haplotype diversity in copy number of antifreeze protein genes.](https://doi.org/10.1038/s41437-025-00803-8). 2026.
- [Pangenome graph analysis reveals evolution of resistance breaking in spinach downy mildew.](https://doi.org/10.1371/journal.pbio.3003596). 2026.
- [The reference genome of the human diploid cell line RPE-1.](https://doi.org/10.1038/s41467-025-62428-z). 2025.
- [Developing pangenomes for large and complex plant genomes and their representation formats.](https://doi.org/10.1016/j.jare.2025.01.052). 2025.
- [Reply to: The genomic structure of complex chromosomal rearrangement at the Fm locus in black-bone Silkie chicken.](https://doi.org/10.1038/s42003-025-07826-1). 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.