A Practical Guide to QUAST: Generating Comprehensive Assembly Quality Reports

By Dr. Zubair Khalid, DVM, MS, PhD ·

A Practical Guide to QUAST: Generating Comprehensive Assembly Quality Reports

Key Takeaways

  • QUAST provides essential reference-based and reference-free metrics for genome assembly quality assessment, including N50, genome fraction, misassembly counts, and duplication ratio, which are critical for evaluating contiguity, completeness, structural integrity, and redundancy.
  • Reference-based evaluation, when a closely related reference genome is available, is paramount for identifying structural errors like relocations, translocations, and inversions, and for quantifying base-level accuracy through mismatch and indel rates per 100 kbp.
  • Reference-free assessment, while limited in detecting structural errors, offers crucial insights into assembly contiguity (N50, L50), total length, and potential contamination through GC content distribution analysis, particularly vital for novel organisms.
  • QUAST metrics directly inform downstream decisions, such as initiating polishing for high mismatch rates (e.g., using Illumina reads to correct ONT assemblies) or restarting assembly with different parameters or assemblers for high misassembly counts or low genome fraction.
  • Integrating QUAST into reproducible workflows using tools like Nextflow or Snakemake, coupled with containerization (Docker/Singularity) and meticulous record-keeping of versions and parameters, ensures the reliability and defensibility of assembly quality reports for publications and database submissions.
  • Complementary tools like BUSCO for gene-space completeness and read mapping for independent validation are crucial for a comprehensive assessment, especially when a close reference genome is unavailable or when specific biological features are under investigation.

Genome assembly quality assessment is a mandatory step in any whole-genome sequencing project, yet many researchers treat it as a formality instead of a diagnostic process. QUAST, the Quality Assessment Tool for genome assemblies, provides a standardized framework for evaluating how well an assembly represents the true genome sequence. This guide explains how to run QUAST effectively, interpret its core metrics, and use the results to make informed decisions about assembly improvement, polishing, and downstream analysis.

The practical outcome of this guide is straightforward: you will learn to generate QUAST reports that tell you precisely where your assembly succeeds and where it fails, and you will know what actions to take based on those findings. The target audience includes biology students, researchers, laboratory professionals, and life-science practitioners who need to produce defensible assembly quality metrics for publications, submissions to public databases, or internal quality control.

What QUAST Measures and Why It Matters

QUAST evaluates genome assemblies using two distinct modes: reference-based evaluation and reference-free evaluation. The reference-based mode compares your assembly against a known genome sequence, while the reference-free mode assesses assembly characteristics without external comparison. Both modes produce metrics that inform different aspects of assembly quality.

The original QUAST publication describes the tool as an improvement over existing assembly comparison software, introducing new ideas and quality metrics for evaluating assemblies both with and without a reference genome. The authors note that genome sequencing techniques have limitations that have led to dozens of assembly algorithms, none of which is perfect, and that most existing methods for comparing assemblies were only applicable to new assemblies of finished genomes. QUAST was designed to address the problem of evaluating assemblies of previously unsequenced species, which had not been adequately considered at the time.

The core value of QUAST lies in its ability to produce many reports, summary tables, and plots that help scientists in their research and publications. For a researcher preparing a genome announcement or a comparative genomics study, QUAST provides the standardized metrics that reviewers and database curators expect to see.

The Role of QUAST in Modern Genome Projects

QUAST does not operate in isolation. Modern genome projects typically integrate QUAST into larger workflows that include assembly, annotation, and completeness assessment. The AquaaG pipeline, for example, integrates genome assembly retrieval from NCBI, assembly quality assessment using QUAST, organism-specific annotation using Prokka for prokaryotes and BRAKER3 for eukaryotes, gene-space completeness evaluation using BUSCO, and functional annotation using EggNOG-mapper. This pipeline demonstrates that QUAST serves as one component of a reproducible genome annotation and assessment framework.

In clinical microbiology, QUAST is used alongside taxonomic classification tools to assess assembly quality before downstream analysis. A benchmarking study comparing Illumina and Oxford Nanopore Technologies sequencing platforms for bacterial whole-genome sequencing used QUAST and GTDB-Tk to assess assembly quality, then proceeded to identify antimicrobial resistance genes, plasmids, and perform core genome MLST analysis. The study found that Illumina-based assemblies generated fewer genes annotated as disrupted, while for ONT assemblies, the base-caller affected assembly annotation accuracy, with High accuracy and Super accuracy base-calling models performing better than the FAST model.

These examples illustrate that QUAST results directly influence downstream analytical decisions. A poor QUAST report should trigger assembly improvement efforts before proceeding to annotation or variant calling.

Understanding QUAST Input Requirements

Before running QUAST, you must understand what input files are required and how to prepare them. QUAST accepts FASTA files containing contig sequences, which are the primary output of genome assemblers. The tool also accepts optional inputs that enable additional analyses.

Required Input Files

The minimum input for QUAST is one or more FASTA files containing assembled contigs or scaffolds. Each FASTA file represents one assembly to be evaluated. QUAST can compare multiple assemblies simultaneously, which is useful for benchmarking different assemblers or different parameter settings on the same dataset.

For reference-based evaluation, you must provide a reference genome in FASTA format. The reference should be closely related to the sequenced organism. For bacterial genomes, a reference from the same species is ideal. For eukaryotic genomes, a reference from the same genus may be acceptable, but you should expect lower alignment rates if the reference is distantly related.

Optional Input Files

QUAST accepts several optional inputs that enhance its analysis capabilities. A reference annotation file in GFF or BED format enables QUAST to report how many genes are fully assembled, partially assembled, or missing. This information is valuable for assessing whether the assembly captures the complete gene space of the organism.

For metagenomic assemblies or assemblies of complex samples, you can provide a list of known genes or operons to check for their presence in the assembly. This feature is particularly useful for clinical microbiology applications where specific resistance genes or virulence factors must be detected.

Preparing Input Files for QUAST

FASTA files from assemblers often contain contig names with special characters or spaces that can cause issues in downstream tools. Before running QUAST, rename contigs to simple identifiers such as contig_1, contig_2, and so on. This practice prevents parsing errors and makes report interpretation easier.

Check that your FASTA files are properly formatted with sequence lines of consistent length. Some assemblers produce files with very long lines that can slow down processing. You can reformat these files using standard bioinformatics tools available through The Carpentries lessons, which provide foundational training in shell and data manipulation skills.

Installing and Running QUAST

QUAST is available through multiple installation methods, and the choice of method depends on your computing environment and expertise level.

Installation Options

The simplest installation method is through the Bioconda package manager, which handles dependencies automatically. Bioconda is part of the Bioconductor ecosystem, which provides official package, workflow, installation, and reproducible genomic-analysis documentation. If you use conda, the command conda install -c bioconda quast installs QUAST and its dependencies.

For users who prefer containerized workflows, QUAST is available as a Docker image. This approach ensures that the software environment is identical across different machines, which is important for reproducibility. The nf-core documentation describes community pipeline standards and usage patterns that emphasize containerization for reproducible workflow execution.

For users who want to run QUAST within a graphical interface, the Galaxy Training Network provides accessible workflow training and analysis tutorials. Galaxy offers QUAST as a tool within its web-based platform, allowing users to run quality assessment without command-line expertise.

Basic QUAST Command

The basic QUAST command for reference-free evaluation is:

quast assembly.fasta -o quast_output

This command evaluates the assembly in assembly.fasta and writes all reports to the quast_output directory. QUAST generates multiple output files, including a report text file, HTML report, and various plots.

For reference-based evaluation, add the reference genome:

quast assembly.fasta -r reference.fasta -o quast_output

When a reference is provided, QUAST aligns the assembly to the reference and computes additional metrics such as genome fraction, misassembly counts, and structural variation statistics.

Running QUAST on Multiple Assemblies

To compare multiple assemblies, provide all FASTA files as positional arguments:

quast assembly1.fasta assembly2.fasta assembly3.fasta -r reference.fasta -o quast_comparison

QUAST generates comparative tables and plots that allow side-by-side evaluation of assembly quality. This feature is particularly useful when benchmarking different assemblers or testing different parameter settings.

Advanced QUAST Options

QUAST includes several advanced options that address specific analysis needs. The --gene-finding option predicts genes in the assembly and reports statistics about gene completeness. The --conserved-genes-finding option searches for a set of conserved genes that are expected to be present in the organism.

For large genomes, the --large option adjusts QUAST behavior to handle assemblies with many contigs. This option is recommended for eukaryotic genomes that may contain hundreds of thousands of contigs.

The --min-contig option sets a minimum contig length threshold. Contigs shorter than this threshold are excluded from some analyses. The default threshold is 500 base pairs, which is appropriate for most bacterial assemblies.

At a Glance: QUAST Metrics and Their Interpretation

The table below summarizes the most important QUAST metrics, what they measure, and how to interpret them in the context of a bacterial genome assembly project.

MetricWhat It MeasuresInterpretation Guidance
N50Contig length at which half of the total assembly is in contigs of that length or longerHigher values indicate more contiguous assemblies. Short-read bacterial assemblies typically show lower N50 than long-read assemblies. Always interpret alongside misassembly counts.
Genome fractionPercentage of the reference genome covered by the assemblyValues near 100 percent indicate complete capture of the reference sequence. Values below 95 percent for bacterial assemblies warrant investigation of missing regions.
Misassembly countNumber of structural errors where contigs are joined incorrectlyZero misassemblies is the expected standard for complete bacterial genomes. High counts indicate the assembly strategy needs revision.
Duplication ratioTotal assembly length divided by aligned assembly lengthValues near 1.0 indicate minimal redundant sequence. Values above 1.05 for haploid bacterial genomes suggest over-assembly of repeats or separate haplotype assembly.
Mismatches per 100 kbpBase-level substitution errors per 100 kilobases of aligned sequenceLower values indicate higher base accuracy. Elevated rates suggest polishing is needed, particularly for assemblies from error-prone sequencing platforms.
GC contentPercentage of guanine and cytosine bases in the assemblyCompare to the expected value for the organism. Multiple peaks in the GC distribution plot may indicate contamination.

Interpreting Core QUAST Metrics

QUAST reports numerous metrics, but a subset of these metrics provides the most informative assessment of assembly quality. Understanding these metrics is essential for making decisions about assembly improvement.

N50 and Related Contiguity Metrics

N50 is the contig length such that contigs of this length or longer account for at least half of the total assembly length. A higher N50 indicates a more contiguous assembly. For bacterial genomes assembled from short reads, N50 values in the hundreds of thousands of base pairs are typical. For long-read assemblies, N50 values often exceed one million base pairs.

The N50 metric has limitations. It does not account for misassemblies, and a high N50 can be achieved by an assembly with significant structural errors. Therefore, N50 should always be interpreted alongside misassembly metrics.

QUAST also reports L50, which is the number of contigs needed to reach half of the assembly length. A lower L50 indicates a more contiguous assembly. The NG50 metric is similar to N50 but is calculated relative to the reference genome size instead of the assembly size. NG50 is more informative when the assembly size differs substantially from the expected genome size.

Genome Fraction and Duplication Ratio

Genome fraction is the percentage of the reference genome that is covered by the assembly. A genome fraction close to 100 percent indicates that the assembly captures nearly all of the reference sequence. Low genome fraction values suggest that significant portions of the genome are missing from the assembly.

The duplication ratio is the total assembly length divided by the aligned assembly length. A duplication ratio close to 1.0 indicates that the assembly contains little redundant sequence. Higher duplication ratios suggest that some regions of the genome are represented multiple times in the assembly, which can occur when haplotypes are assembled separately or when repetitive regions are over-assembled.

Misassembly Metrics

Misassemblies are structural errors in the assembly where contigs are joined incorrectly. QUAST reports several types of misassemblies, including relocations, translocations, and inversions. A relocation occurs when two regions that are adjacent in the reference are placed far apart in the assembly. A translocation occurs when regions from different chromosomes are joined together. An inversion occurs when a region is assembled in the reverse orientation relative to the reference.

The number of misassemblies is a critical quality indicator. For bacterial genomes, a high-quality assembly should have zero or very few misassemblies. The presence of many misassemblies indicates that the assembly strategy needs revision, possibly through different assembler parameters, additional sequencing data, or hybrid assembly approaches.

Mismatches and Indels

QUAST reports the number of mismatches and indels per 100 kilobases of aligned sequence. These metrics reflect the base-level accuracy of the assembly. High mismatch rates suggest sequencing errors that have not been corrected during assembly. Indels can indicate homopolymer errors, which are common in certain sequencing technologies.

The benchmarking study comparing Illumina and ONT sequencing found that Illumina sequencing provided consistently high-quality reads with a median Q-score of 35, while ONT R10.4.1 with the SUP model showed a higher median quality score of 15.3 compared to R9.4.1 at 13.9. These sequencing quality differences directly affect assembly base accuracy, and QUAST metrics capture these differences.

Reference-Free Quality Assessment

When no reference genome is available, QUAST still provides useful metrics for assessing assembly quality. This situation is common for novel organisms or metagenomic samples where no closely related reference exists.

Contig Statistics Without a Reference

Without a reference, QUAST reports contig count, total length, N50, L50, and GC content. These metrics provide a basic description of the assembly but do not indicate structural correctness. A highly fragmented assembly with many small contigs suggests that the assembly process encountered difficulties, possibly due to repetitive regions or insufficient sequencing depth.

GC Content Analysis

QUAST plots GC content distribution across the assembly. Deviations from the expected GC content for the organism can indicate contamination or assembly errors. For example, the Neobacillus sedimentimangrovi UE25 genome was reported to have 37.75 percent GC content, and this value was confirmed through QUAST quality parameters. If your assembly shows a GC distribution with multiple peaks, this may indicate that sequences from different organisms were assembled together.

Limitations of Reference-Free Assessment

Reference-free assessment cannot detect misassemblies or structural errors. A contig that joins sequences from different parts of the genome will appear identical to a correctly assembled contig in reference-free metrics. Therefore, reference-free QUAST reports should be interpreted as preliminary quality indicators, not as evidence of assembly correctness.

For novel organisms, additional validation methods are needed. BUSCO analysis assesses gene-space completeness by searching for conserved single-copy orthologs. The AquaaG pipeline integrates BUSCO alongside QUAST to provide complementary quality information. Gene-space completeness is a strong indicator of assembly quality because conserved genes are expected to be present in any complete assembly.

Using QUAST with Reference Genomes

Reference-based evaluation is the most informative mode of QUAST analysis. When a high-quality reference genome is available, QUAST can identify structural errors and base-level inaccuracies that would otherwise go undetected.

Selecting an Appropriate Reference

The reference genome should be from the same species as the sequenced organism whenever possible. For bacterial species with multiple sequenced strains, choose the reference that is most closely related to your strain based on phylogenetic analysis or average nucleotide identity.

For novel species, a reference from the same genus may be acceptable, but you should expect lower genome fraction values due to genuine genomic differences between species. The Kalamiella piersonii genome report describes the first complete genome of this species, which was submitted to GenBank. For such novel species, the choice of reference is constrained by the availability of closely related genomes.

Reference Annotation Files

Providing a reference annotation file enables QUAST to report gene-level statistics. The annotation file should be in GFF or BED format and should correspond to the reference genome. QUAST uses this information to determine how many genes are fully assembled, partially assembled, or missing from your assembly.

Gene-level statistics are particularly valuable for clinical microbiology applications where the presence of specific genes, such as antimicrobial resistance genes, is of interest. The Kalamiella piersonii study used QUAST v5.0.2 to verify genome assembly quality before proceeding to annotation with PROKKA and identification of resistance genes and virulence factors using Abricate.

Interpreting Reference-Based Reports

The QUAST report for reference-based evaluation includes a section on genomic features. This section reports the number of genes, operons, and other features that are fully or partially covered by the assembly. A high percentage of fully covered genes indicates that the assembly captures the gene space of the organism.

The report also includes a structural variation section that lists large-scale differences between the assembly and the reference. These differences may represent genuine genomic variation between strains or assembly errors. Distinguishing between these possibilities requires additional analysis, such as read mapping to validate the assembly structure.

Practical Workflow for QUAST Analysis

A systematic workflow ensures that QUAST analysis is reproducible and that results are properly interpreted. The following workflow is suitable for bacterial genome projects and can be adapted for larger genomes.

Step 1: Prepare Input Files

Rename contigs to simple identifiers and ensure that FASTA files are properly formatted. If you have multiple assemblies to compare, place each assembly in a separate FASTA file with a descriptive name that indicates the assembler and parameters used.

Download the reference genome and annotation file from NCBI. The NCBI Data Resources provide official descriptions of database search systems and sequence resources. Ensure that the reference genome and annotation file are from the same assembly version to avoid inconsistencies.

Step 2: Run QUAST

Run QUAST with the reference genome and annotation file. Use the --gene-finding option to enable gene prediction in the assembly. Save the output to a directory with a descriptive name that includes the date and assembly version.

For large genomes, add the --large option to prevent memory issues. For metagenomic assemblies, consider using the --metagenome option, which adjusts some analyses for metagenomic data.

Step 3: Review the Report

Open the report text file and review the key metrics. Start with genome fraction and misassembly counts. If genome fraction is below 95 percent for a bacterial assembly, investigate whether the missing sequence is concentrated in specific regions or distributed throughout the genome.

Check the N50 value and compare it to expectations for your sequencing technology. Short-read assemblies typically have lower N50 values than long-read assemblies. The benchmarking study comparing Illumina and ONT found that ONT assemblies resolved rRNA operons better than Illumina assemblies, which is relevant for repetitive regions.

Step 4: Examine the Plots

QUAST generates cumulative length plots and GC content plots. The cumulative length plot shows how contig lengths accumulate, providing a visual representation of assembly contiguity. The GC content plot can reveal contamination or systematic biases in the assembly.

Step 5: Document Results

Record the QUAST metrics in your project documentation. Include the QUAST version number, the command used, and the date of analysis. This documentation is essential for reproducibility and for preparing methods sections for publications.

Common Failure Patterns in QUAST Reports

Recognizing common failure patterns helps you diagnose assembly problems quickly and take corrective action.

Low Genome Fraction with High N50

This pattern indicates that the assembly is contiguous but missing significant portions of the genome. The missing sequence may be in repetitive regions that are difficult to assemble, or it may indicate that some sequencing data was excluded during assembly. Check whether the missing regions correspond to known repetitive elements or ribosomal RNA operons.

High Misassembly Count with High N50

This pattern indicates that the assembler produced long contigs by joining sequences that are not adjacent in the true genome. This problem is common when assemblers attempt to resolve repetitive regions without sufficient evidence. Consider using a different assembler, adjusting repeat resolution parameters, or incorporating long-read data to resolve the structure.

High Duplication Ratio

A duplication ratio significantly above 1.0 suggests that the assembly contains redundant sequence. This can occur when haplotypes are assembled separately in a diploid organism or when the assembler fails to collapse repetitive regions. For haploid bacterial genomes, a duplication ratio above 1.05 warrants investigation.

Elevated Mismatch Rate

A high mismatch rate indicates base-level errors in the assembly. This problem is common in assemblies from sequencing platforms with higher error rates. The benchmarking study found that ONT assemblies had more disrupted genes when using the FAST base-calling model compared to HAC and SUP models. Polishing with additional sequencing data or improved base-calling models can reduce mismatch rates.

GC Content Deviation

If the GC content of the assembly differs substantially from the expected value for the organism, contamination may be present. Check the GC content distribution plot for multiple peaks, which indicate the presence of sequences from different organisms. The Neobacillus sedimentimangrovi study reported a GC content of 37.75 percent, and deviations from this value would warrant investigation.

Using QUAST Results for Assembly Improvement Decisions

QUAST results should drive concrete decisions about whether to accept an assembly, polish it, or restart assembly with different parameters.

Accepting an Assembly

An assembly is suitable for downstream analysis when it meets quality thresholds appropriate for the intended use. For bacterial genome submissions to public databases, a complete genome should have zero misassemblies, a genome fraction above 99 percent, and a mismatch rate below 1 per 100 kilobases. For draft genomes, higher misassembly counts may be acceptable, but the limitations should be documented.

Polishing Decisions

When base-level errors are detected, polishing is appropriate. Polishing tools use read mapping information to correct errors in the assembly. The choice of polishing tool depends on the sequencing technology used. For Illumina-based assemblies, polishing with additional Illumina reads can correct residual errors. For ONT assemblies, polishing with Illumina reads is often necessary to achieve high base accuracy.

The benchmarking study found that hybrid assemblies using both Illumina and ONT data produced better results than either technology alone for some applications. If your QUAST report shows elevated mismatch rates, consider generating additional sequencing data for polishing.

Restarting Assembly

When misassembly counts are high or genome fraction is low, restarting assembly with different parameters or a different assembler may be necessary. The QUAST comparative report, which evaluates multiple assemblies side by side, is particularly useful for this purpose. Run several assembly strategies and compare their QUAST metrics to identify the best approach.

Long-Read Assembly Considerations

For long-read assemblies, QUAST metrics should be interpreted with the understanding that long-read assemblers produce different error profiles than short-read assemblers. The benchmarking study found that ONT assemblies resolved rRNA operons better than Illumina assemblies, but ONT assemblies had more disrupted genes when using lower-quality base-calling models. These tradeoffs should be considered when choosing an assembly strategy.

Records and Documentation for QUAST Analysis

Maintaining detailed records of QUAST analysis is essential for publication, database submission, and internal quality control.

Essential Records

For each QUAST analysis, record the following information:

  • QUAST version number
  • Command line arguments used
  • Input file names and versions
  • Reference genome and annotation file versions
  • Date of analysis
  • Computing environment and resource usage

This information should be stored alongside the QUAST output files in a structured directory.

Publication Reporting

When reporting QUAST metrics in publications, include the QUAST version number and the parameters used. The original QUAST publication recommends reporting the metrics that are most relevant to the study. For genome announcements, report genome size, N50, GC content, and misassembly counts. For comparative studies, report the full set of metrics for all assemblies evaluated.

Database Submission Requirements

Public databases such as NCBI have specific requirements for genome assembly quality documentation. The NCBI Data Resources provide official descriptions of submission requirements. QUAST reports provide the metrics needed to complete submission forms and to respond to curator inquiries about assembly quality.

Integrating QUAST into Reproducible Workflows

Reproducibility is a core requirement for bioinformatics analysis. QUAST should be integrated into workflows that can be rerun with consistent results.

Workflow Management Systems

Workflow management systems such as Nextflow and Snakemake provide frameworks for reproducible analysis. The nf-core documentation describes community pipeline standards that emphasize reproducibility through containerization and version control. Integrating QUAST into a workflow management system ensures that the same version of QUAST is used across analyses and that results can be regenerated.

Containerization

Containerization through Docker or Singularity ensures that the software environment is identical across different computing systems. This approach eliminates the variability introduced by different software versions and dependencies. The nf-core documentation emphasizes containerization as a standard practice for reproducible workflows.

Version Control

Record the QUAST version number in your analysis documentation. QUAST has evolved significantly since its initial release, and metrics may differ between versions. The original QUAST publication from 2013 describes the initial feature set, while later versions have added new metrics and improved performance.

Training and Skill Development

Developing the skills needed for effective QUAST analysis requires training in bioinformatics fundamentals. The EMBL-EBI Training program provides learning pathways for bioinformatics data resources and practical analysis education. The Carpentries lessons offer foundational training in shell, Git, and programming skills that are essential for reproducible analysis. The Galaxy Training Network provides accessible workflow training that includes genome assembly quality assessment.

Limitations of QUAST and Complementary Tools

QUAST provides valuable assembly quality metrics, but it has limitations that should be understood.

Reference Bias

Reference-based QUAST analysis is biased toward the reference genome. If the reference is distantly related to the sequenced organism, genuine genomic differences will be reported as misassemblies or mismatches. This limitation is particularly relevant for novel species where no close reference exists.

Inability to Detect All Errors

QUAST cannot detect all types of assembly errors. Errors in repetitive regions may be missed if the reference has similar repetitive structure. Base-level errors in regions with low sequencing coverage may not be detected if the reference has the same errors.

Complementarity with BUSCO

BUSCO analysis provides complementary information to QUAST by assessing gene-space completeness. The AquaaG pipeline integrates QUAST and BUSCO to provide a more complete picture of assembly quality. BUSCO searches for conserved single-copy orthologs that are expected to be present in the organism. A high BUSCO completeness score indicates that the assembly captures the gene space, even if reference-based metrics are limited by the absence of a close reference.

Complementarity with Read Mapping

Read mapping provides an independent assessment of assembly quality. Mapping the original sequencing reads back to the assembly can identify regions with low coverage, which may indicate assembly errors. Tools such as Minimap2 and BWA can be used for this purpose. The benchmarking study used read mapping and SNP calling tools including Minimap2, Pilon, and Snippy to validate assemblies.

Safety and Ethical Considerations

Genome assembly quality assessment has implications for biosafety and biosecurity. Assemblies of pathogenic organisms should be handled according to institutional biosafety guidelines. The Kalamiella piersonii study describes a multidrug-resistant strain of a novel human pathogen, and such data should be handled with appropriate security measures.

When working with clinical isolates, ensure that sample collection and sequencing comply with ethical guidelines and institutional review board requirements. The benchmarking study used clinical isolates of ESKAPE pathogens and ATCC strains, and such work requires appropriate approvals.

Data sharing should follow the principles of responsible genomic data sharing. When submitting assemblies to public databases, ensure that you have the right to share the data and that any ethical or legal restrictions are documented.

Professional Escalation Criteria

Knowing when to seek expert assistance is important for efficient problem-solving in genome assembly projects.

When to Consult a Bioinformatics Specialist

Consult a bioinformatics specialist when QUAST reports show persistent problems that you cannot resolve through parameter adjustment. Examples include:

  • Misassembly counts that remain high across multiple assemblers and parameter sets
  • Genome fraction values below 90 percent for bacterial assemblies
  • Unexplained GC content deviations that suggest contamination
  • Assembly failures that prevent generation of any usable contigs

When to Consult a Sequencing Facility

Consult your sequencing facility when assembly problems suggest issues with the sequencing data itself. Examples include:

  • Consistently low read quality scores across multiple runs
  • Unexpected GC bias in the sequencing data
  • Contamination in the sequencing data that appears in the assembly

When to Consult a Domain Expert

Consult a domain expert in the biology of your organism when QUAST metrics suggest genuine biological variation instead of assembly errors. Examples include:

  • Structural variants that are consistently detected across multiple assembly strategies
  • Genomic rearrangements that are supported by read mapping evidence
  • Presence of mobile genetic elements that may explain assembly complexity

The Kalamiella piersonii study identified three plasmids of 513,647 bp, 261,771 bp, and 106,029 bp, and such findings require domain expertise to interpret correctly.

Frequently Asked Questions

What is the difference between N50 and NG50 in QUAST reports?

N50 is calculated based on the assembly length, while NG50 is calculated based on the reference genome length. N50 is the contig length at which contigs of that length or longer account for at least half of the total assembly length. NG50 is the contig length at which contigs of that length or longer account for at least half of the reference genome length. NG50 is more informative when the assembly size differs from the expected genome size, because it indicates how much of the expected genome is captured in contigs of a given length.

How many misassemblies are acceptable in a bacterial genome assembly?

For a complete bacterial genome, zero misassemblies is the expected standard. For draft genomes, some misassemblies may be acceptable, but the number should be documented and the regions affected should be identified. The acceptable threshold depends on the intended use of the assembly. For clinical applications where structural variants are being analyzed, even a single misassembly can lead to incorrect conclusions. For phylogenetic analysis based on core genome SNPs, a small number of misassemblies may have minimal impact.

Can QUAST evaluate assemblies without a reference genome?

Yes, QUAST can evaluate assemblies without a reference genome. In reference-free mode, QUAST reports contig statistics such as contig count, total length, N50, L50, and GC content. However, reference-free mode cannot detect misassemblies or structural errors. For novel organisms without a close reference, reference-free QUAST metrics should be complemented with BUSCO analysis to assess gene-space completeness.

What is the difference between QUAST and BUSCO?

QUAST assesses assembly quality through metrics such as contiguity, genome fraction, and misassembly counts. BUSCO assesses gene-space completeness by searching for conserved single-copy orthologs that are expected to be present in the organism. The two tools provide complementary information. QUAST tells you about the structural properties of the assembly, while BUSCO tells you whether the assembly captures the expected gene content. The AquaaG pipeline integrates both tools for comprehensive assembly assessment.

How should I choose a reference genome for QUAST analysis?

Choose a reference genome from the same species as your sequenced organism whenever possible. For bacterial species with multiple sequenced strains, select the reference that is most closely related to your strain based on phylogenetic analysis or average nucleotide identity. For novel species, a reference from the same genus may be acceptable, but you should expect lower genome fraction values due to genuine genomic differences. The reference genome and annotation file should be from the same assembly version to ensure consistency.

What does the duplication ratio mean in QUAST reports?

The duplication ratio is the total assembly length divided by the aligned assembly length. A ratio close to 1.0 indicates that the assembly contains little redundant sequence. Higher ratios suggest that some regions of the genome are represented multiple times in the assembly. This can occur when haplotypes are assembled separately in diploid organisms or when repetitive regions are over-assembled. For haploid bacterial genomes, a duplication ratio above 1.05 warrants investigation.

How can I reduce the number of misassemblies in my assembly?

Reducing misassemblies requires addressing the underlying causes, which are often related to repetitive regions or insufficient sequencing data. Consider using a different assembler that handles repeats differently, adjusting repeat resolution parameters, or incorporating long-read data to resolve complex regions. The benchmarking study comparing Illumina and ONT found that ONT assemblies resolved rRNA operons better than Illumina assemblies, suggesting that long-read data can help resolve repetitive regions. Run multiple assembly strategies and compare their QUAST metrics to identify the best approach.

What should I do if my QUAST report shows low genome fraction?

Low genome fraction indicates that significant portions of the genome are missing from the assembly. Investigate whether the missing sequence is concentrated in specific regions or distributed throughout the genome. Check whether the missing regions correspond to known repetitive elements or ribosomal RNA operons. Consider whether additional sequencing data is needed to cover these regions. If the missing sequence is distributed throughout the genome, this may indicate that the sequencing depth was insufficient or that the assembler discarded reads that did not assemble.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.