# How to Visualize Assembly Graphs with Bandage: A Step-by-Step Tutorial for Beginners


## Key Takeaways

- Bandage visualizes genome assembly graphs (GFA/FASTG formats) to assess assembly quality by revealing structural issues like repeats, errors, and contamination through node and edge topology.
- Graph navigation features (zoom, pan, rotate) and configurable color schemes (coverage, GC content) are crucial for identifying high/low coverage regions and potential misassemblies.
- Integrated BLAST searches within Bandage allow direct localization of specific genes or sequences within the assembly graph, aiding in functional annotation and contamination detection.
- Exporting graphs as PNG or SVG formats enables the creation of publication-ready figures, essential for documenting assembly quality and presenting findings.
- A systematic decision framework, assessing component structure, node length distribution, branching patterns, and coverage thresholds, guides interpretation of graph features for informed assembly improvement strategies.
- Troubleshooting common issues like empty graph displays or slow performance involves verifying file formats, updating graphics drivers, and adjusting display settings or considering computational resources for large datasets.

---

Assembly graph visualization is a critical step in genome assembly quality assessment, yet new users often struggle with loading, navigating, and interpreting the graphical output produced by assemblers. Bandage (a Bioinformatics Application for Navigating De novo Assembly Graphs Easily) provides a practical solution for viewing assembly graphs generated by tools like SPAdes, Flye, and Canu. This tutorial walks through the complete workflow from raw reads to polished assemblies, covering graph loading, color schemes, BLAST searches, and image export, with troubleshooting guidance for common problems.

## Understanding Assembly Graphs and Their Role in Genome Assembly

Genome assembly reconstructs complete sequences from fragmented DNA reads. Modern assemblers do not simply output contiguous sequences. They first construct assembly graphs that represent the relationships between sequence fragments. These graphs contain nodes, which represent sequence contigs or unitigs, and edges, which represent connections between those sequences. Understanding this structure is essential before attempting visualization.

De Bruijn graph assemblers, such as SPAdes, construct graphs by splitting reads into k-mers of fixed length. Each k-mer becomes a node or part of a node, and edges connect k-mers that overlap by k-1 bases. Overlap graph assemblers, such as Canu and Flye, instead build graphs where nodes represent reads or read overlaps. Both graph types convey the same fundamental information: how sequence fragments connect to form larger genomic structures.

The practical value of viewing assembly graphs lies in diagnosing assembly problems. A well-assembled genome typically produces a graph with long, unbranched paths. A poorly assembled genome produces graphs with many branches, loops, and disconnected components. These visual patterns reveal issues such as repetitive regions, sequencing errors, heterozygous variants, or contamination. Researchers can use this information to decide whether to adjust assembly parameters, perform additional sequencing, or proceed with downstream analysis.

Bandage was designed specifically to make assembly graph exploration accessible. The software renders graphs in a two-dimensional layout and allows users to zoom, pan, and search within the graph structure. It supports graphs from multiple assemblers and can display additional information such as read depth, GC content, and BLAST hit locations. The program runs on Windows, macOS, and Linux systems and requires no programming expertise to operate.

## At a Glance: Bandage Workflow Overview

The table below summarizes the core steps covered in this tutorial, the primary actions required, and the expected outcomes at each stage.

| Workflow Stage | Primary Actions | Expected Outcome |
| --- | --- | --- |
| Graph Loading | Open Bandage, load assembly graph file (GFA or FASTG format), verify node and edge counts | Graph renders in main window with visible nodes and connections |
| Graph Navigation | Zoom, pan, rotate, and adjust node spacing to inspect structure | Clear view of graph topology including branches and loops |
| Color and Label Configuration | Apply depth-based coloring, set node labels, adjust display thresholds | Visual differentiation of high and low coverage regions |
| BLAST Search Integration | Run BLAST search within Bandage or import BLAST results | Query sequences highlighted on the graph with alignment information |
| Image Export | Configure output settings, export as PNG or SVG | Publication-ready figure for reports or manuscripts |

## Preparing Your Input Files for Bandage

Bandage requires an assembly graph file in one of several supported formats. The most common formats are GFA (Graphical Fragment Assembly) and FASTG. Most modern assemblers produce GFA files by default or can be configured to output them. Before loading a graph into Bandage, verify that the file exists and contains the expected content.

SPAdes produces a file named assembly_graph.fastg in its output directory. This file contains the assembly graph in FASTG format. Newer versions of SPAdes also produce assembly_graph.gfa. Flye produces assembly_graph.gfa in its output directory. Canu produces a file with a .gfa extension in the assembly output folder. Check the documentation for your specific assembler version to confirm the exact file name and location.

The GFA format stores graph information as lines beginning with segment (S), link (L), and path (P) records. Segment lines define nodes and contain the sequence data. Link lines define edges between nodes. Path lines define known traversals through the graph. Bandage parses these records to reconstruct the graph structure. FASTG format stores similar information but uses a different syntax that some older assemblers prefer.

Before loading a graph, confirm that the file is not empty and that it contains sequence data. A graph file with zero bytes indicates that the assembly failed or that the output path was incorrect. Check the assembler log files for error messages if the graph file is missing or empty. Some assemblers produce multiple graph files at different stages of the assembly process. Use the final graph file, not intermediate files, for visualization.

For users working with public datasets, the [National Center for Biotechnology Information](https://www.ncbi.nlm.nih.gov/) provides access to assembled genomes and raw sequencing data that can be used to generate assembly graphs. NCBI databases include assembly records with associated graph files for many organisms. Downloading a public assembly graph can be useful for practicing Bandage skills before working with your own data.

## Installing and Launching Bandage

Bandage is distributed as a standalone application for Windows, macOS, and Linux. The software is available from the official Bandage website and requires no installation dependencies beyond the operating system. Download the appropriate version for your platform, extract the archive if necessary, and run the executable.

On Windows systems, Bandage is distributed as a ZIP archive containing an executable file. Extract the archive to a folder of your choice and double-click the executable to launch the program. On macOS, Bandage is distributed as a DMG file. Open the DMG and drag the application to your Applications folder. On Linux, Bandage is distributed as a tarball. Extract the archive and run the executable from the command line or file manager.

Bandage requires a graphics card that supports OpenGL 2.0 or higher. Most modern computers meet this requirement. If the program fails to launch or displays a blank window, update your graphics drivers. On Linux systems, you may need to install additional OpenGL libraries. Check the Bandage documentation for platform-specific troubleshooting guidance.

The Bandage graphical user interface consists of a main graph display window, a menu bar, and a control panel. The control panel contains tabs for graph drawing settings, node labels, BLAST search, and other features. Familiarize yourself with these interface elements before loading a graph. The menu bar provides access to file operations, graph drawing commands, and help resources.

## Loading an Assembly Graph into Bandage

To load an assembly graph, click File in the menu bar and select Load Graph. Navigate to the location of your graph file and select it. Bandage will parse the file and display the graph in the main window. The loading process may take a few seconds for large graphs. A status bar at the bottom of the window shows progress information.

After loading, the graph appears as a collection of nodes and edges. Nodes are drawn as rectangles or lines, depending on the drawing style selected. Edges are drawn as lines connecting nodes. The default layout places nodes in a circular arrangement. You can change the layout algorithm using the Graph Drawing tab in the control panel.

If the graph does not appear after loading, check the following items. First, confirm that the file format is supported. Bandage supports GFA version 1 and version 2, as well as FASTG. Some assemblers produce graph files in formats that Bandage does not recognize. Convert the file to a supported format if necessary. Second, check that the file contains nodes. A graph file with only link records and no segment records will produce an empty display. Third, verify that the file path contains no special characters that might confuse the parser.

The graph drawing tab contains options for node spacing, edge width, and layout style. The default settings work well for most graphs. For large graphs, increase node spacing to reduce visual clutter. For small graphs, decrease node spacing to bring nodes closer together. The layout algorithm options include circular, force-directed, and stacked layouts. Experiment with different layouts to find the one that best reveals the graph structure.

## Navigating the Assembly Graph Display

Once a graph is loaded, you can navigate the display using mouse and keyboard controls. Left-click and drag to pan the view. Use the mouse wheel to zoom in and out. Right-click and drag to rotate the graph in three-dimensional space. These controls allow you to inspect different regions of the graph at different magnifications.

The graph display window includes a toolbar with buttons for common navigation actions. The zoom-to-fit button adjusts the view to show the entire graph. The zoom-in and zoom-out buttons change magnification by a fixed factor. The reset view button returns the display to the default orientation. Use these buttons to quickly adjust the view when navigating large graphs.

For graphs with many nodes, zooming in to inspect individual nodes is essential. At high magnification, you can see the sequence associated with each node by hovering the mouse over the node. The node information panel displays the node name, length, coverage, and other attributes. This information helps you identify nodes of interest for further analysis.

Bandage also supports searching for specific nodes by name or sequence. Use the search box in the toolbar to enter a node name or partial sequence. The program highlights matching nodes in the display. This feature is useful when you know the identity of a particular contig and want to locate it within the graph structure.

The graph drawing tab includes an option to show node labels. Labels display the node name or sequence length next to each node. For large graphs, labels can clutter the display. Toggle labels on only when you need to identify specific nodes. The label size and font can be adjusted in the settings.

## Understanding Node and Edge Information

Each node in an assembly graph represents a sequence fragment. The length of the sequence is stored as an attribute of the node. Coverage, also called depth, represents the number of reads that map to that sequence. High coverage nodes typically represent genuine genomic sequence. Low coverage nodes may represent sequencing errors or contamination.

Bandage displays node information in the node info panel when you click on a node. The panel shows the node name, sequence length, coverage, and GC content. This information helps you assess the quality of individual contigs. Nodes with unusually high or low coverage compared to the rest of the graph may indicate problems.

Edges in the graph represent connections between nodes. In a de Bruijn graph, an edge indicates that the sequences of the two nodes overlap by k-1 bases. In an overlap graph, an edge indicates that the reads represented by the nodes overlap. The edge direction matters. A directed edge from node A to node B means that the sequence of A is followed by the sequence of B in the assembly.

The graph drawing tab includes options for edge display. You can show edge arrows to indicate direction. You can also color edges by the type of connection. These visual cues help you understand the graph topology and identify regions where the assembly is ambiguous.

Bandage can display coverage information as a color gradient. Nodes with high coverage appear in one color, while nodes with low coverage appear in another. The default color scheme uses a blue-to-red gradient, where blue represents low coverage and red represents high coverage. You can customize the color scheme in the graph drawing tab.

## Using Color Schemes to Interpret Graph Structure

Color is a powerful tool for interpreting assembly graphs. Bandage provides several color schemes that highlight different aspects of the graph. The default scheme colors nodes by coverage. This scheme helps you identify high and low coverage regions at a glance.

The coverage color scheme uses a logarithmic scale by default. This scale compresses the range of coverage values so that both high and low coverage nodes are visible. You can switch to a linear scale if you prefer. The color gradient can be customized by setting minimum and maximum coverage values. Nodes with coverage below the minimum appear in the minimum color. Nodes with coverage above the maximum appear in the maximum color.

Bandage also provides a color scheme based on GC content. This scheme colors nodes according to the proportion of guanine and cytosine bases in the sequence. GC content varies between organisms and between genomic regions. This scheme can help you identify contamination from different species or detect regions with unusual base composition.

The random color scheme assigns colors to nodes based on graph components. Each connected component of the graph receives a distinct color. This scheme helps you see the overall structure of the graph and identify disconnected components. Disconnected components may represent separate chromosomes, plasmids, or contamination.

For BLAST search results, Bandage colors nodes based on the position and orientation of BLAST hits. Nodes with hits appear in distinct colors, while nodes without hits appear in gray. This scheme helps you locate specific genes or sequences within the graph structure.

## Performing BLAST Searches Within Bandage

BLAST (Basic Local Alignment Search Tool) searches identify sequences in the assembly graph that match a query sequence. Bandage integrates BLAST search functionality, allowing you to search for specific genes or sequences directly within the graph. This feature is valuable for locating genes of interest, verifying assembly completeness, and identifying contamination.

To perform a BLAST search, click the BLAST tab in the control panel. Enter a query sequence in FASTA format or load a FASTA file containing the query. Select the BLAST program to use. Bandage supports BLASTN for nucleotide queries and BLASTP for protein queries. Choose the appropriate program based on your query type.

Bandage uses the BLAST+ suite for sequence alignment. The software must locate the BLAST executables on your system. On Windows, BLAST+ may need to be installed separately. On macOS and Linux, BLAST+ may already be installed or can be installed through package managers. Configure the BLAST path in the Bandage settings if the program cannot find the executables.

After running a BLAST search, Bandage displays the results in the BLAST tab. Each hit shows the query name, the target node, the alignment position, and the E-value. Click on a hit to highlight the corresponding region in the graph display. The graph view zooms to the hit location and colors the matching nodes.

BLAST search results can be saved to a file for later use. Bandage supports exporting BLAST results in tabular format. This file can be imported into other analysis tools or used for reporting. The export function is available in the BLAST tab.

The choice of alignment tool affects the accuracy and efficiency of sequence identification in assembly graphs. A 2024 study in [Microorganisms](https://doi.org/10.3390/microorganisms12112168) compared Bandage, SPAligner, and GraphAligner for detecting antimicrobial resistance gene sequences in assembly graphs. The study found that Bandage offered the most precise and efficient identification of antimicrobial resistance gene sequences among the tools tested. This result supports the use of Bandage for gene detection and genomic context analysis in bacterial populations.

## Exporting High-Quality Images from Bandage

Bandage can export the graph display as an image file for use in reports, presentations, and publications. The export function supports PNG and SVG formats. PNG is a raster format suitable for most purposes. SVG is a vector format that scales without loss of quality, making it ideal for publication figures.

To export an image, click File in the menu bar and select Export Graphic. Choose the output format and specify the file name and location. The export dialog includes options for image size and resolution. For publication-quality images, set the resolution to at least 300 dots per inch.

Before exporting, adjust the graph display to show the region of interest. Zoom and pan to frame the desired view. The exported image captures the current view, so take time to compose the figure carefully. Consider which nodes and edges are important to show and which can be cropped out.

The export dialog includes an option to include a legend. The legend explains the color scheme used in the graph. Including a legend is essential for figures that use coverage or GC content coloring. Without a legend, readers cannot interpret the colors in the image.

For figures that will be edited in other software, export as SVG. SVG files can be opened in vector graphics editors such as Inkscape or Adobe Illustrator. These editors allow you to modify the figure, add annotations, and adjust colors. For figures that will be inserted directly into documents, export as PNG.

## Practical Workflow: From Raw Reads to Polished Assemblies

The complete workflow from raw reads to a polished assembly involves multiple steps. Bandage visualization fits into this workflow at several points. Understanding where and how to use Bandage improves the efficiency and quality of the assembly process.

The first step is quality control of raw reads. Tools like FastQC assess read quality and identify problems such as adapter contamination and low-quality bases. Trimming tools remove low-quality bases and adapters. This step is essential because poor-quality reads produce poor assemblies. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on quality control and read processing that cover these steps in detail.

The second step is assembly. Choose an assembler appropriate for your data type. Short-read assemblers like SPAdes work well for Illumina data. Long-read assemblers like Flye and Canu work well for Oxford Nanopore and PacBio data. Each assembler produces an assembly graph file that can be loaded into Bandage.

After assembly, load the graph into Bandage to assess the assembly quality. Look for the following features. Long, unbranched paths indicate well-assembled regions. Short, branched nodes may indicate repeats or errors. Disconnected components may indicate separate chromosomes or contamination. Use the coverage color scheme to identify low-coverage nodes that may represent errors.

If the assembly graph shows problems, consider adjusting assembly parameters. Increasing the k-mer size in SPAdes may resolve some repeat structures. Increasing the minimum coverage threshold may remove low-coverage nodes. The [nf-core documentation](https://nf-co.re/docs) provides guidance on configuring assembly pipelines and parameters for reproducible results.

Assembly polishing is the next step. Polishing tools use read alignments to correct errors in the assembled sequence. For long-read assemblies, tools like Medaka and Racon correct errors using the original reads. For short-read assemblies, tools like Pilon correct errors using Illumina reads. After polishing, reload the graph into Bandage to verify that the corrections improved the assembly.

The final step is assembly evaluation. Tools like QUAST and BUSCO assess assembly completeness and accuracy. These tools provide quantitative metrics that complement the visual assessment from Bandage. The combination of visual and quantitative assessment provides a complete picture of assembly quality.

## Recording Observations and Measurements

Systematic recording of assembly graph observations supports reproducible research and informed decision-making. Maintain a laboratory notebook or electronic record that documents each assembly attempt and the corresponding graph characteristics.

Record the following information for each assembly. The assembler name and version. The input data type and quantity. The assembly parameters used. The graph file name and location. The number of nodes and edges in the graph. The total length of all nodes. The N50 length of the assembly. The number of connected components.

When examining the graph in Bandage, record qualitative observations. Note the presence of long unbranched paths. Note the presence of branched structures that may indicate repeats. Note the presence of disconnected components. Note any nodes with unusually high or low coverage. These observations provide context for interpreting quantitative metrics.

Capture screenshots of the graph at each stage of the workflow. Save these images with descriptive file names that include the sample name, assembly stage, and date. These images serve as a visual record of the assembly process and can be referenced in reports and publications.

The [European Bioinformatics Institute](https://www.ebi.ac.uk/training) provides training resources on data management and reproducible analysis. These resources emphasize the importance of documenting analysis steps and maintaining organized records. Following these practices ensures that assembly results can be reproduced and verified by other researchers.

## Common Failure Patterns and Troubleshooting

New users often encounter specific problems when working with Bandage. Recognizing these common failure patterns helps you diagnose and resolve issues quickly.

**Empty graph display after loading.** This problem occurs when the graph file contains no nodes or when the file format is not recognized. Check the file size and content. A file with zero bytes contains no data. A file with only link records and no segment records produces an empty display. Convert the file to a supported format if necessary.

**Graph appears as a single blob.** This problem occurs when the graph contains many short nodes that are densely connected. Increase node spacing in the graph drawing tab to separate the nodes. Alternatively, apply a coverage threshold to hide low-coverage nodes. The threshold can be set in the graph drawing tab.

**Graph contains many disconnected components.** This problem may indicate that the assembly failed to connect related sequences. Check the assembly log for error messages. Consider adjusting assembly parameters to improve connectivity. For metagenomic samples, disconnected components may represent different species in the community.

**BLAST search returns no results.** This problem occurs when the query sequence does not match any sequence in the graph. Check that the query is in the correct format and that the BLAST program is appropriate for the query type. Verify that the BLAST executables are correctly configured in Bandage settings.

**Program crashes or displays a blank window.** This problem usually indicates a graphics driver issue. Update your graphics drivers to the latest version. On Linux systems, install the required OpenGL libraries. Check the Bandage documentation for platform-specific troubleshooting guidance.

**Graph loads slowly or becomes unresponsive.** This problem occurs with very large graphs containing millions of nodes. Reduce the graph size by applying a coverage threshold or by loading only a subset of the graph. Consider using a computer with more memory for large assemblies.

## Limitations of Assembly Graph Visualization

Bandage is a powerful tool, but it has limitations that users should understand. The software visualizes the graph structure but does not perform assembly or polishing. It cannot fix assembly errors. It can only help you identify problems that may require reassembly with different parameters.

The graph layout algorithms used by Bandage are heuristic. Different layout algorithms may produce visually different arrangements of the same graph. The layout does not reflect the true genomic arrangement of the sequences. It only provides a convenient way to view the graph structure.

Bandage displays the graph as constructed by the assembler. If the assembler produced an incorrect graph, Bandage will display that incorrect graph. The accuracy of the visualization depends on the accuracy of the assembly. Always verify assembly quality using multiple methods, including quantitative metrics and read mapping.

The coverage values displayed by Bandage are computed from the graph structure. These values may differ from read mapping coverage computed by other tools. The graph coverage represents the number of reads that contributed to each node during assembly. This value is a useful approximation but may not reflect the true genomic coverage.

For metagenomic samples, assembly graphs can be extremely complex. The graphs may contain thousands of components representing different species. Bandage can display these graphs, but interpreting them requires careful analysis. The [HiMT toolkit](https://doi.org/10.1016/j.xplc.2025.101467), described in a 2025 study in Plant Communications, provides an alternative approach for organelle genome assembly with a graphical user interface and interactive reports. This tool may be more appropriate for chloroplast and mitochondrial genome projects.

## Quality Assessment Using Assembly Graphs

Assembly graphs provide a visual basis for assessing assembly quality. The [Human Pangenome Reference Consortium](https://doi.org/10.1038/s41586-023-05896-x) demonstrated that high-quality assemblies can be generated from diverse individuals and that these assemblies capture known variants and haplotypes. The quality of these assemblies was verified using multiple methods, including graph-based analysis.

When assessing assembly quality with Bandage, consider the following criteria. The proportion of the graph contained in long nodes. A high-quality assembly has most of its sequence in long nodes. The proportion of the graph contained in the largest connected component. A complete assembly has most of its sequence in one component. The presence of circular structures that may indicate complete circular chromosomes or plasmids.

The [Jorg method](https://doi.org/10.1371/journal.pcbi.1008972), described in a 2021 study in PLOS Computational Biology, uses iterative assembly, binning, and read mapping to circularize small bacterial, archaeal, and viral genomes. This method exposes potential misassemblies from k-mer based assemblies. Bandage can be used to visualize the circular structures produced by this method and to verify that circularization was successful.

Circularized genomes are important for several reasons. They provide a reference collection for future assemblies. They provide complete gene content of a genome. They confirm little or no contamination. They allow study of genomic context and synteny of genes. They link protein coding genes to ribosomal RNA genes for metabolic inference. Bandage visualization supports these applications by making circular structures visible.

For organelle genomes, assembly quality assessment requires attention to specific features. Chloroplast and mitochondrial genomes contain repetitive sequences that complicate assembly. The [HiMT toolkit](https://doi.org/10.1016/j.xplc.2025.101467) addresses these challenges with a fixed k-mer prefix strategy and automatic coverage estimation. Bandage can complement this approach by providing visual confirmation of organelle genome structure.

## Professional Escalation Criteria

Knowing when to seek additional help or escalate a problem is important for efficient workflow. The following situations warrant consultation with a bioinformatics specialist or the Bandage user community.

**Persistent graph loading failures.** If a graph file repeatedly fails to load despite troubleshooting, the file may be corrupted or in an unsupported format. Consult the assembler documentation to verify the output format. Seek help from the assembler user community or the Bandage GitHub issues page.

**Unexpected graph structures.** If the assembly graph shows structures that you cannot interpret, such as complex loops or unusual branching patterns, consult a specialist. These structures may indicate biological features such as repeats or may indicate assembly errors that require parameter adjustment.

**BLAST search failures.** If BLAST searches consistently fail to return expected results, the query sequence may not be present in the assembly. This situation may indicate that the gene of interest was not assembled or that the assembly is incomplete. Consult a specialist to determine the appropriate next steps.

**Large-scale assembly problems.** If the assembly graph indicates widespread problems, such as many short nodes or excessive branching, the assembly strategy may need revision. Consult a specialist to review the assembly parameters and data quality.

**Metagenomic assembly complexity.** If the assembly graph for a metagenomic sample is too complex to interpret, specialized tools may be needed. The [nf-core community](https://nf-co.re/docs) provides pipelines for metagenomic analysis that include assembly and binning steps. Consult the nf-core documentation for guidance on configuring these pipelines.

## Reproducibility and Documentation Practices

Reproducible analysis requires careful documentation of all steps in the assembly and visualization workflow. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training on computing skills that support reproducible research. These lessons cover shell scripting, version control with Git, and data management practices.

Document the exact version of Bandage used for visualization. Bandage versions may differ in features and behavior. Record the version number in your laboratory notebook or analysis documentation. This information allows others to reproduce your visualization results.

Document the settings used for graph drawing. Node spacing, color scheme, and threshold values affect the visual appearance of the graph. Record these settings when you export images for publication. This documentation ensures that others can reproduce your figures.

Store graph files and exported images in organized directories. Use descriptive file names that include the sample name, assembly stage, and date. Maintain a README file that describes the contents of each directory. This practice supports data sharing and collaboration.

The [Bioconductor project](https://bioconductor.org/) provides tools for reproducible genomic analysis in the R programming language. While Bandage is a standalone application, Bioconductor tools can complement Bandage for downstream analysis. The Bioconductor documentation provides guidance on installing and using these tools in reproducible workflows.

## Building a Graph Interpretation Decision Framework

A systematic decision framework helps you move from passive graph viewing to active assembly troubleshooting. The framework below converts visual observations into concrete management decisions. It is organized around five graph features that you can assess in under five minutes after loading any assembly graph.

### Step 1: Assess Component Structure

Start by counting the connected components in the graph. The component count appears in the graph statistics panel after loading. For a single chromosome assembly, you expect one dominant component containing most of the sequence. Multiple large components may indicate separate chromosomes, plasmids, or contamination from other organisms.

Record the number of components and the total sequence length in each. If one component contains less than 90 percent of the total assembled sequence, investigate whether the smaller components represent genuine biological entities or assembly artifacts. For bacterial genomes, small circular components often represent plasmids. For eukaryotic genomes, multiple components may represent different chromosomes that failed to connect due to repetitive regions.

### Step 2: Evaluate Node Length Distribution

The node length distribution reveals assembly continuity. Select the node info panel and sort nodes by length. A healthy assembly has most of its sequence concentrated in a small number of long nodes. A fragmented assembly has sequence distributed across many short nodes.

Apply the following thresholds as starting points for assessment. For short-read assemblies, nodes shorter than 500 base pairs often represent repetitive or low-complexity sequence. For long-read assemblies, nodes shorter than 5,000 base pairs may indicate assembly problems. These thresholds vary by organism and sequencing technology, so treat them as initial screening values instead of absolute rules.

Calculate the proportion of total sequence contained in nodes above your chosen threshold. If this proportion falls below 80 percent, consider whether assembly parameters need adjustment. The [nf-core documentation](https://nf-co.re/docs) provides guidance on configuring assembly pipelines that can improve node length distributions through parameter optimization.

### Step 3: Identify Branching Patterns

Branching in assembly graphs indicates ambiguous connections. A node with two outgoing edges represents a repeat that the assembler could not resolve. A node with many outgoing edges may represent a repetitive family or a misassembly.

Classify branches into three categories. Simple branches have two or three outgoing edges and often represent repeats that are shorter than the read length. Complex branches have more than three outgoing edges and may indicate collapsed repeats or contamination. Circular branches form loops and may represent complete circular genomes or tandem repeats.

For each branch type, record the node names and the coverage values of the branching nodes. Compare coverage between the branch point and the surrounding nodes. A branch point with coverage roughly double the surrounding nodes suggests a collapsed repeat. A branch point with coverage similar to surrounding nodes may indicate a genuine genomic rearrangement.

### Step 4: Apply Coverage Thresholds Systematically

Coverage thresholds filter the graph display to reveal high-confidence sequence. The graph drawing tab includes minimum and maximum coverage settings. Start with a minimum coverage threshold that removes nodes below the expected coverage for your sequencing depth.

To determine an appropriate threshold, examine the coverage histogram in the node info panel. Identify the main coverage peak. Set the minimum threshold to approximately 20 percent of this peak value. Nodes below this threshold often represent sequencing errors or very low abundance contamination.

Apply the threshold and observe which nodes disappear. If removing low-coverage nodes eliminates branches and simplifies the graph, those branches likely represented errors. If the graph structure remains complex after thresholding, the complexity reflects genuine biological features such as repeats or heterozygous regions.

### Step 5: Document Decisions and Outcomes

Create a structured record for each assembly graph you examine. The [European Bioinformatics Institute](https://www.ebi.ac.uk/training) provides training on data management practices that support systematic documentation. Include the following fields in your record.

| Assessment Field | Observation | Decision | Outcome |
| --- | --- | --- | --- |
| Component count | Number and sizes of connected components | Accept or investigate | Action taken |
| Node length distribution | N50 and proportion of sequence in long nodes | Accept or reassemble | Parameter changes |
| Branching patterns | Location and type of branches | Accept or resolve | Resolution method |
| Coverage distribution | Main peak and outlier nodes | Accept or filter | Threshold applied |
| Circular structures | Presence of circular paths | Verify completeness | Validation method |

This table format allows you to compare assemblies across samples and conditions. It also provides a record that other researchers can review when you share your assembly results.

## Implementing the Framework in Practice

Apply the framework to a test assembly before working with your own data. The [National Center for Biotechnology Information](https://www.ncbi.nlm.nih.gov/) provides access to public sequencing data and assembled genomes. Download a small bacterial genome assembly graph and practice the five assessment steps.

Start with a graph from a well-characterized organism with a known genome size. This allows you to compare your graph observations against the expected genome structure. Work through the five steps in order, recording your observations in the table format. Compare your assessments against published descriptions of the genome to validate your interpretation skills.

For metagenomic samples, the framework requires modification. The [Jorg method](https://doi.org/10.1371/journal.pcbi.1008972), described in a 2021 study in PLOS Computational Biology, demonstrates that circularized genomes from metagenomic data provide important validation of assembly completeness. When examining metagenomic graphs, focus on identifying circular structures that may represent complete genomes. The component count will be much higher than for single-organism assemblies, so adjust your expectations accordingly.

For organelle genomes, the framework needs additional attention to repetitive regions. The [HiMT toolkit](https://doi.org/10.1016/j.xplc.2025.101467), described in a 2025 study in Plant Communications, addresses the challenges of organelle genome assembly with a fixed k-mer prefix strategy. When examining chloroplast or mitochondrial graphs, pay particular attention to branches that may represent the inverted repeats common in these genomes.

## Common Interpretation Errors

New users often make predictable mistakes when interpreting assembly graphs. Recognizing these errors improves the accuracy of your assessments.

**Confusing graph layout with genomic arrangement.** The visual layout produced by Bandage is heuristic and does not reflect the true genomic order of sequences. Two nodes that appear close together in the display may be far apart in the genome. Always verify genomic relationships using sequence alignment or read mapping instead of visual proximity.

**Overinterpreting coverage differences.** Coverage values in the graph represent the number of reads that contributed to each node during assembly. These values can vary due to sequencing bias, GC content, and copy number variation. Do not conclude that a node with lower coverage represents contamination without additional evidence.

**Ignoring graph orientation.** Assembly graphs are directed. The orientation of edges matters for interpreting sequence connections. A node can be traversed in either direction, and the graph display may not make orientation obvious. Check edge directions before drawing conclusions about sequence order.

**Assuming all branches are errors.** Some branches represent genuine biological features. Heterozygous variants in diploid organisms create alternative paths through the graph. Repetitive regions create branches that reflect real genomic structure. Use coverage and sequence information to distinguish biological branches from assembly errors.

**Neglecting to verify with read mapping.** The graph structure reflects the assembler's interpretation of the reads. This interpretation can be incorrect. Always verify important graph features by mapping reads back to the assembly and checking that the read alignments support the graph structure.

## Validation Steps for Framework Decisions

Before acting on a graph interpretation decision, validate your observation using independent methods. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on assembly validation that complement visual graph assessment.

For suspected misassemblies, map the original reads to the assembled contigs. Look for read pairs that map to distant locations in the assembly, which indicates a misjoin. For suspected contamination, compare the GC content and coverage of the suspect nodes against the rest of the assembly. For suspected repeats, check whether the repeat sequence appears in multiple locations in the assembled genome.

The [Human Pangenome Reference Consortium](https://doi.org/10.1038/s41586-023-05896-x) demonstrated that high-quality assemblies require validation at multiple levels. Their assemblies were verified for accuracy at the structural and base pair levels using multiple methods. Apply similar rigor to your own assemblies by combining graph visualization with quantitative validation tools.

For antimicrobial resistance gene detection, the choice of alignment tool affects results. A 2024 study in [Microorganisms](https://doi.org/10.3390/microorganisms12112168) found that Bandage offered the most precise identification of antimicrobial resistance gene sequences compared to SPAligner and GraphAligner. When using the decision framework to locate specific genes, verify BLAST hits by examining the surrounding graph context and confirming the gene is in the expected genomic location.

## When to Escalate to Professional Support

The decision framework helps you identify problems, but some situations require specialized expertise. Escalate to a bioinformatics specialist when you encounter the following conditions.

**Persistent graph complexity after parameter adjustment.** If multiple assembly attempts with different parameters produce similarly complex graphs, the underlying data may have issues. A specialist can assess whether additional sequencing or different assembly strategies are needed.

**Unexpected circular structures in unexpected organisms.** Circular structures in bacterial assemblies often represent plasmids or complete chromosomes. Circular structures in eukaryotic assemblies may indicate contamination or assembly artifacts. A specialist can help determine the biological significance of these structures.

**Conflicting evidence from different assessment methods.** If the graph structure suggests one interpretation but quantitative metrics suggest another, specialist input can resolve the discrepancy. This situation often arises with complex genomes containing many repeats.

**Metagenomic assemblies requiring binning.** The [nf-core community](https://nf-co.re/docs) provides pipelines for metagenomic analysis that include binning steps. If your metagenomic graph is too complex to interpret manually, these pipelines can automate the process of separating genomes from the assembly graph.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training that can help you develop the skills needed to resolve common assembly problems independently. These lessons cover shell scripting, data management, and reproducible analysis practices that support effective assembly troubleshooting.

## Frequently Asked Questions

### What file formats does Bandage support for assembly graphs?

Bandage supports GFA version 1 and version 2 files, as well as FASTG files. Most modern assemblers produce GFA files by default. SPAdes produces assembly_graph.fastg and assembly_graph.gfa. Flye produces assembly_graph.gfa. Canu produces .gfa files. Check your assembler documentation to confirm the output format and file location.

### How do I interpret the colors in a Bandage graph?

The default color scheme colors nodes by coverage. Blue represents low coverage and red represents high coverage. You can change the color scheme in the graph drawing tab. Other schemes include GC content coloring and random coloring by connected component. BLAST search results are shown with distinct colors for hit regions.

### Why does my graph appear as a single dense blob?

A dense blob appearance usually indicates that the graph contains many short nodes that are densely connected. Increase node spacing in the graph drawing tab to separate the nodes. You can also apply a coverage threshold to hide low-coverage nodes. Adjust these settings until the graph structure becomes visible.

### Can Bandage perform BLAST searches automatically?

Bandage integrates BLAST search functionality. You can enter a query sequence in FASTA format and run BLASTN or BLASTP searches within the program. Bandage requires the BLAST+ executables to be installed on your system. Configure the BLAST path in the Bandage settings if the program cannot find the executables.

### How do I export a publication-quality image from Bandage?

Use the File menu and select Export Graphic. Choose PNG for raster output or SVG for vector output. Set the resolution to at least 300 dots per inch for publication quality. Adjust the graph display to frame the region of interest before exporting. Include a legend to explain the color scheme.

### What should I do if my assembly graph shows many disconnected components?

Disconnected components may represent separate chromosomes, plasmids, or contamination. For metagenomic samples, they may represent different species. Check the assembly log for errors and consider adjusting assembly parameters. For bacterial genomes, circular structures may indicate complete chromosomes or plasmids.

### How does Bandage compare to other graph visualization tools?

A 2024 study in [Microorganisms](https://doi.org/10.3390/microorganisms12112168) compared Bandage, SPAligner, and GraphAligner for detecting antimicrobial resistance gene sequences in assembly graphs. The study found that Bandage offered the most precise and efficient identification of antimicrobial resistance gene sequences among the tools tested. Bandage is particularly well suited for gene detection and genomic context analysis.

### Can Bandage help with organelle genome assembly?

Bandage can visualize organelle genome assembly graphs, but specialized tools may be more appropriate. The [HiMT toolkit](https://doi.org/10.1016/j.xplc.2025.101467), described in a 2025 study in Plant Communications, provides one-click chloroplast and mitochondrial genome assembly with a graphical user interface and interactive reports. HiMT is designed for use on standard laptops and is freely available for non-commercial use.

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Metagenomic Binning with Assembly Graph Embeddings: A New Frontier](/knowledge/bioinformatics/metagenomic-binning-with-assembly-graph-embeddings-a-new-frontier)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Metagenomics Pipeline: From Raw Reads to Taxonomic and Functional Profiles](/knowledge/bioinformatics/metagenomics-pipeline-from-raw-reads-to-taxonomic-and-functional-profiles)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [HiMT: An integrative toolkit for assembling organelle genomes using HiFi reads.](https://doi.org/10.1016/j.xplc.2025.101467). 2025.
- [A draft human pangenome reference.](https://doi.org/10.1038/s41586-023-05896-x). 2023.
- [A method for achieving complete microbial genomes and improving bins from metagenomics data.](https://doi.org/10.1371/journal.pcbi.1008972). 2021.
- [Reference-quality genome assembly created for widely used RPE-1 human cell line](https://www.semanticscholar.org/paper/db0bce8138deeb5ee7a6b503f40f41dfe971935d)
- [Evaluating Sequence Alignment Tools for Antimicrobial Resistance Gene Detection in Assembly Graphs](https://doi.org/10.3390/microorganisms12112168). Microorganisms, 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.