Evaluating Long-Read Assemblies: Key Metrics and Tools for Assessing Contiguity and Accuracy

By Dr. Zubair Khalid, DVM, MS, PhD ·

Evaluating Long-Read Assemblies: Key Metrics and Tools for Assessing Contiguity and Accuracy

Key Takeaways

  • Assembly quality is a function of the entire pipeline, not just raw read data, as demonstrated by studies showing polishing improves accuracy and continuity across different platforms and assemblers.
  • Contiguity metrics like N50 and NG50 are insufficient alone; they can be inflated by misjoins, necessitating complementary assessments of base-level accuracy (QV) and gene content completeness (BUSCO).
  • Reference-free evaluation using k-mer analysis (e.g., Merqury) is crucial for novel genomes, providing insights into completeness, consensus accuracy, and haplotype structure without relying on existing genomic references.
  • BUSCO assesses gene content completeness by identifying conserved single-copy orthologs, with interpretation requiring knowledge of the organism's expected gene repertoire and potential for whole-genome duplication.
  • Haplotype representation is critical for heterozygous genomes, with tools like Merqury's k-mer copy number spectrum helping to distinguish between haplotype collapse and legitimate allelic duplication.
  • A structured decision framework, defining quality tiers before analysis and employing a metric conflict resolution protocol, is essential for objectively evaluating assemblies and guiding iterative improvement or acceptance.

Long-read sequencing platforms such as Pacific Biosciences (PacBio) HiFi and Oxford Nanopore Technologies (ONT) produce reads that can span repetitive regions and resolve complex genomic structures, but the assemblies built from these reads require systematic evaluation before they can support downstream biological conclusions. This article explains which assembly quality metrics matter most for long-read projects, how to compute them with QUAST, BUSCO, and Merqury, and how to interpret results within the context of sequencing depth, platform choice, and research objectives. The practical outcome is a decision framework for determining whether an assembly meets quality thresholds for publication, comparative genomics, or clinical applications.

The Problem With Judging Assemblies by Read Quality Alone

Raw sequencing reads can look excellent on the instrument while the resulting assembly contains misjoins, collapsed repeats, or missing genes. The error profiles of long-read platforms differ substantially from short-read sequencing, and these differences propagate into assembly artifacts that are not visible when you inspect reads individually. ONT reads historically carried higher single-base error rates than PacBio HiFi reads, and while polishing tools can correct many of these errors, the correction process itself can introduce new problems if applied incorrectly. A 2022 benchmarking study of yeast genome assemblies compared 455 ONT assemblies and 88 HiFi assemblies across multiple depths and assembler combinations, finding that the choice of assembler played an essential role in genome construction, especially for low-depth datasets, and that polishing improved accuracy and continuity in most quality metrics [<a href="#ref-1">1</a>]. This finding underscores a central principle: assembly quality is a property of the entire pipeline, not of the sequencing run alone.

Researchers often ask whether their assembly is good enough without defining what good enough means for their specific application. A draft assembly intended for gene discovery has different quality requirements than a telomere-to-telomere reference genome intended for structural variant analysis. The metrics you choose must align with the biological questions you plan to answer. If you need to identify antimicrobial resistance genes in a bacterial pathogen, you care about gene content completeness and base-level accuracy in coding regions. If you are assembling a highly heterozygous insect genome, you care about haplotype separation and contiguity across repetitive regions. The evaluation strategy should be designed before you commit to a sequencing budget, because depth requirements differ substantially between platforms and organisms.

Core Quality Metrics for Long-Read Assemblies

Contiguity Metrics: N50, L50, and NG50

Contiguity metrics describe how much of the genome is contained in the largest assembled sequences. N50 is the length value at which half of the total assembly length is contained in sequences of that size or larger. If your assembly has an N50 of 5 megabases, it means that half of all assembled bases reside in contigs or scaffolds of at least 5 megabases. L50 is the number of sequences required to reach that halfway point. NG50 is the same calculation applied against the estimated genome size instead of the assembly size, which prevents inflated contiguity values in assemblies that overrepresent repetitive content.

These metrics are useful for comparing assemblies of the same genome produced by different pipelines, but they have limitations. N50 does not tell you whether the assembly is correct. A misjoin that fuses two unrelated regions can inflate N50 while making the assembly biologically wrong. The 2022 lepidopteran genome study used BUSCO, Inspector, and EagleC to evaluate assemblies of silkworm strains, noting that quality assessment is essential for genome assembly and can provide better and more accurate results [<a href="#ref-2">2</a>]. The authors did not rely on contiguity alone, because highly contiguous but structurally incorrect assemblies would have passed a superficial review.

For practical interpretation, consider the expected genome size and complexity. A bacterial genome of 5 megabases should assemble into a single circular chromosome with long-read data, so an N50 below 1 megabase suggests a problem. A plant genome of 1 gigabase with extensive repeat content may have an N50 of 10 megabases and still be considered highly contiguous. The relevant comparison is against assemblies of similar genomes produced with similar data types.

Base-Level Accuracy: QV and Error Rate

Quality value (QV) expresses the probability that a base in the assembly is correct. A QV of 40 corresponds to one error per 10,000 bases, or 99.99 percent accuracy. QV50 corresponds to one error per 100,000 bases. For reference-quality genomes, QV40 is often cited as a minimum threshold, though the specific requirement depends on the downstream application.

The 2020 study on genome-scale models from nanopore assemblies demonstrated that nanopore-derived models of Escherichia coli K-12 were more than 99 percent complete even at sequencing depths below 10x coverage, and that these models could identify canonical antimicrobial resistance content and simulate strain-specific growth [<a href="#ref-3">3</a>]. This finding shows that for some applications, modest base-level accuracy is sufficient. However, the same study emphasized that sequencing errors inherent to the nanopore technique may negatively affect the quality and utility of downstream products [<a href="#ref-3">3</a>]. If your goal is to identify single nucleotide polymorphisms that confer drug resistance, you need higher accuracy than if you are assembling a genome for taxonomic classification.

QV is typically estimated by comparing the assembly to a reference genome when one exists, or by using k-mer based methods that do not require a reference. Merqury uses k-mers from the raw reads to evaluate consensus accuracy, completeness, and haplotype structure without requiring a reference genome. This approach is particularly valuable for non-model organisms where no closely related reference exists.

Completeness: BUSCO and Gene Content

Benchmarking Universal Single-Copy Orthologs (BUSCO) assesses completeness by searching the assembly for a set of genes that are expected to be present as single copies in the target lineage. The output categories are complete single copy, complete duplicated, fragmented, and missing. A high-quality assembly should have a high percentage of complete BUSCO genes and a low percentage of missing genes.

The yeast benchmarking study used QUAST, BUSCO, and a newly proposed Comprehensive_score to evaluate assembly quality across 455 ONT and 88 HiFi assemblies [<a href="#ref-1">1</a>]. The finding that Flye was superior to other tools for ONT datasets through Comprehensive_score evaluation, and that Flye and NextDenovo performed better for HiFi datasets, illustrates how BUSCO results can differentiate assembler performance in ways that contiguity metrics alone cannot [<a href="#ref-1">1</a>]. A fragmented assembly may still have a respectable N50 if the fragmentation occurs in gene-rich regions, but BUSCO will detect the missing gene content.

BUSCO results must be interpreted with knowledge of the expected gene content for your organism. A bacterial genome should have nearly all BUSCO genes complete and single copy. A plant genome with recent whole-genome duplication may legitimately have many complete duplicated BUSCO genes. The BUSCO lineage dataset you select must match your organism, and the version of the lineage dataset affects comparability across studies.

Haplotype Representation: Duplication Ratio and Assembly Phasing

Heterozygous genomes present a special challenge for assembly. If the assembler fails to separate haplotypes, it may collapse divergent alleles into a single sequence, creating chimeric contigs. If it separates haplotypes too aggressively, it may produce two copies of every region, inflating the assembly size and creating false duplications.

The silkworm study reported the first nearly complete telomere-to-telomere reference genome of Bombyx mori produced by PacBio HiFi sequencing, along with highly contiguous assemblies of two other silkworm strains produced by ONT or PacBio continuous long reads [<a href="#ref-2">2</a>]. The authors emphasized that Lepidoptera includes a wide variety of insects with high genetic diversity and heterozygosity, making the selection of an appropriate sequencing and assembly strategy critical [<a href="#ref-2">2</a>]. For HiFi data, hifiasm was the preferred assembler, while NextDenovo was superior for CLR and ONT data [<a href="#ref-2">2</a>]. These recommendations emerged from systematic comparison, not from vendor claims.

Merqury provides a copy number spectrum that shows how many k-mers appear at each copy number in the assembly compared to the expected distribution from the reads. A clean haploid assembly shows a dominant peak at copy number one. A collapsed assembly shows an excess of single-copy k-mers. A duplicated assembly shows an excess of multi-copy k-mers. This diagnostic is essential for detecting haplotype collapse or over-separation.

At a Glance: Metric Selection by Application

ApplicationPrimary MetricsSecondary MetricsMinimum ThresholdsRecommended Tools
Bacterial pathogen characterizationBUSCO completeness, QV, circular chromosome statusN50, gene content for AMR markersBUSCO > 99 percent complete, QV > 40 for variant callingQUAST, BUSCO, Merqury
Eukaryotic draft genome for gene discoveryBUSCO completeness, N50, assembly size vs expected genome sizeDuplication ratio, GC content distributionBUSCO > 90 percent complete, N50 > 1 Mb for most eukaryotesBUSCO, QUAST, Merqury
Telomere-to-telomere reference genomeChromosome-level scaffolds, telomere repeats, centromere representationQV, BUSCO, k-mer completenessAll chromosomes represented, QV > 50, BUSCO > 98 percent completeQUAST, BUSCO, Merqury, Inspector
Metagenome-assembled genome from long readsCompleteness, contamination, strain heterogeneityN50, read mapping rateCompleteness > 90 percent, contamination < 5 percentBUSCO, CheckM, QUAST

The QUAST Workflow for Assembly Evaluation

Preparing Input Files

QUAST accepts FASTA files of assembled contigs or scaffolds. For long-read assemblies, you should provide the estimated genome size using the -g parameter so that NG50 and other genome-normalized metrics are calculated correctly. If you have a reference genome, you can provide it for comparative evaluation, but QUAST also operates in reference-free mode.

The Galaxy Training Network provides accessible workflow training for genome assembly evaluation, including practical tutorials that walk through QUAST usage in a reproducible environment [<a href="#ref-4">4</a>]. For researchers who prefer command-line workflows, the nf-core documentation describes community pipeline standards for assembly and quality control, including configuration guidance for running these tools at scale [<a href="#ref-5">5</a>]. These resources are useful for standardizing your evaluation across multiple samples or projects.

Interpreting QUAST Output

QUAST generates a report with multiple sections. The contig statistics section includes N50, L50, total length, and number of contigs. The misassembly section reports misjoins, relocations, and translocations when a reference is provided. The GC content section can reveal contamination if you observe unexpected GC distributions. The Nx plot shows the cumulative length distribution, which helps you visualize how contiguity changes across the assembly.

For long-read assemblies, pay attention to the number of contigs relative to the expected chromosome count. A bacterial genome should produce one contig per replicon. A fungal genome should produce approximately one contig per chromosome, though some chromosomes may be missing telomeres or centromeres. The silkworm study achieved a nearly complete telomere-to-telomere assembly, which represents the upper bound of what is currently possible [<a href="#ref-2">2</a>]. Most projects will fall short of this standard, and the gap between your assembly and a telomere-to-telomere reference should be quantified instead of ignored.

Common QUAST Pitfalls

QUAST can mislead you if you use the wrong genome size estimate. If you underestimate the genome size, NG50 will appear artificially high. If you overestimate it, NG50 will appear artificially low. Use the best available estimate from flow cytometry, k-mer analysis, or related species. For genomes with high heterozygosity, the assembly size may exceed the haploid genome size because both haplotypes are represented. QUAST does not correct for this, so you must interpret the results with knowledge of your organism's biology.

BUSCO Analysis for Completeness Assessment

Selecting the Correct Lineage Dataset

BUSCO requires a lineage dataset that matches your organism. The available lineages include bacteria, fungi, arthropods, vertebrates, plants, and many others. Using the wrong lineage produces meaningless results. If you are assembling a bacterial genome, use the bacteria lineage. If you are assembling an insect, use the insecta or arthropod lineage. The EMBL-EBI Training portal offers learning pathways for bioinformatics data resources, including guidance on selecting appropriate reference datasets for comparative analysis [<a href="#ref-6">6</a>].

The version of the lineage dataset matters for cross-study comparisons. BUSCO updates lineage datasets periodically, and scores from different versions are not directly comparable. Record the BUSCO version and lineage dataset version in your methods so that others can reproduce your evaluation.

Running BUSCO in Genome Mode

BUSCO can run in genome mode, transcriptome mode, or protein mode. For genome assemblies, use genome mode with the appropriate lineage. The output includes a summary table with counts for complete single copy, complete duplicated, fragmented, and missing BUSCO genes. The complete percentage is the sum of complete single copy and complete duplicated divided by the total BUSCO genes in the lineage.

The yeast benchmarking study used BUSCO as one of three evaluation tools and found that polishing improved accuracy and continuity of preassemblies, with the combination of Pilon and Medaka working well in most quality metrics [<a href="#ref-1">1</a>]. This finding demonstrates that BUSCO scores can improve after polishing, so you should run BUSCO before and after polishing to quantify the improvement. If BUSCO scores do not improve after polishing, either the polishing was ineffective or the assembly has structural problems that base-level correction cannot fix.

Interpreting Duplicated BUSCO Genes

Complete duplicated BUSCO genes indicate that both haplotypes are represented in the assembly. For a haploid assembly, this is a problem. For a diploid assembly where haplotypes are intentionally separated, this is expected. The interpretation depends on your assembly strategy. If you used hifiasm with default settings, you may have a primary assembly and an alternate assembly. The primary assembly should have mostly single-copy BUSCO genes. If the primary assembly has many duplicated BUSCO genes, the assembler may have failed to collapse haplotypes properly.

The lepidopteran study used BUSCO, Inspector, and EagleC to evaluate assembly quality and recommended 3D-DNA with EagleC evaluation for chromosome-level high-quality genome construction [<a href="#ref-2">2</a>]. This recommendation reflects the need for multiple evaluation tools because no single metric captures all aspects of assembly quality.

Merqury for Reference-Free Quality Assessment

K-mer Based Completeness and Accuracy

Merqury uses k-mers from the raw sequencing reads to evaluate assembly quality without a reference genome. The method works by counting k-mers in the read set, then checking which k-mers are present in the assembly. The completeness metric is the fraction of read k-mers that are present in the assembly. The consensus accuracy metric is derived from the k-mer copy number spectrum.

This approach is particularly valuable for non-model organisms where no reference genome exists. The NCBI Data Resources provide access to sequence databases and analysis services that can support assembly evaluation, including tools for comparing your assembly to related sequences [<a href="#ref-7">7</a>]. However, for a truly novel genome, Merqury provides the most reliable reference-free quality estimates.

The K-mer Copy Number Spectrum

Merqury produces a spectrum plot showing the number of distinct k-mers at each copy number in the assembly. A high-quality haploid assembly shows a dominant peak at copy number one. A diploid assembly with separated haplotypes shows peaks at copy number one and two. The spectrum also reveals the proportion of k-mers that are absent from the assembly, which corresponds to missing sequence.

The copy number spectrum can detect haplotype collapse. If the assembly has collapsed two haplotypes into one sequence, the k-mer copy number for heterozygous regions will be lower than expected. Merqury quantifies this effect and provides a haplotype-specific completeness estimate.

Choosing K-mer Size

The choice of k-mer size affects Merqury results. Smaller k-mers are more sensitive to sequencing errors but provide more coverage. Larger k-mers are more specific but require higher sequencing depth. A common choice is k = 21 for bacterial genomes and k = 21 or 31 for eukaryotic genomes. The optimal k-mer size depends on the sequencing error rate and depth. For HiFi data with low error rates, larger k-mers work well. For ONT data with higher error rates, smaller k-mers may be necessary.

The 2022 benchmarking study on yeast genomes used QUAST, BUSCO, and Comprehensive_score to evaluate assemblies, and the authors noted that enough data depth is required for high-quality genome construction by ONT (greater than 80x) and HiFi (greater than 20x) datasets [<a href="#ref-1">1</a>]. These depth recommendations provide context for interpreting Merqury results. If your sequencing depth is below these thresholds, the k-mer spectrum may show incomplete representation even for a correctly assembled genome.

Practical Workflow for Assembly Evaluation

Step 1: Run QUAST for Contiguity and Structural Assessment

Run QUAST on your assembly with the estimated genome size. Record N50, L50, NG50, total length, and number of contigs. If you have a reference genome, provide it for misassembly detection. Save the report and plots for your records.

Step 2: Run BUSCO for Gene Content Completeness

Select the appropriate lineage dataset and run BUSCO in genome mode. Record the counts for complete single copy, complete duplicated, fragmented, and missing genes. Calculate the complete percentage. Compare your results to assemblies of related species published in the literature.

Step 3: Run Merqury for Reference-Free Accuracy

Count k-mers in your raw reads using meryl or a similar tool. Run Merqury with your assembly and the k-mer database. Record the completeness, consensus accuracy (QV), and copy number spectrum. Examine the spectrum for signs of haplotype collapse or duplication.

Step 4: Polish and Re-evaluate

If BUSCO completeness is below your target or Merqury QV is below 40, consider polishing the assembly. The yeast benchmarking study found that polishing by Pilon and Medaka improved accuracy and continuity of preassemblies, and their combination pipeline worked well in most quality metrics [<a href="#ref-1">1</a>]. After polishing, rerun BUSCO and Merqury to quantify the improvement. If polishing does not improve the metrics, the problem may be structural instead of base-level.

Step 5: Document and Report

Record all software versions, parameters, and lineage dataset versions. The nf-core documentation emphasizes reproducibility standards for bioinformatics pipelines, including version pinning and configuration tracking [<a href="#ref-5">5</a>]. The Carpentries lessons provide foundational training in version control with Git, which is essential for tracking changes to your analysis scripts [<a href="#ref-8">8</a>]. These practices ensure that your evaluation can be reproduced by others.

Records and Measurements for Assembly Projects

Maintain a laboratory notebook or electronic record that includes the following information for each assembly project:

Record TypeSpecific DataPurpose
Sequencing metadataPlatform, chemistry version, flow cell type, read length distribution, estimated depthContext for interpreting assembly quality and comparing across runs
Assembly parametersAssembler name and version, all non-default parameters, input read filtering stepsReproducibility and troubleshooting
Quality metricsQUAST N50, L50, NG50, BUSCO complete percentage, Merqury QV and completenessBaseline for comparison after polishing or parameter changes
Polishing recordsPolisher name and version, number of rounds, changes in QV and BUSCODocumentation of improvement or lack thereof
Decision logRationale for accepting or rejecting assemblies, thresholds usedTransparency for collaborators and reviewers

The Bioconductor project provides official documentation for reproducible genomic analysis workflows, including package installation and version management [<a href="#ref-9">9</a>]. While Bioconductor focuses on R packages, the principles of reproducible analysis apply equally to assembly evaluation pipelines.

Common Failure Patterns in Long-Read Assembly Evaluation

Overestimating Quality From Contiguity Alone

A high N50 can mask serious structural errors. Misjoins that fuse unrelated genomic regions increase contiguity while making the assembly biologically incorrect. The lepidopteran study used multiple evaluation tools because no single metric captures all aspects of assembly quality [<a href="#ref-2">2</a>]. If you report only N50, you may miss collapsed repeats, chimeric contigs, or missing genes.

Ignoring the Effects of Sequencing Depth

The yeast benchmarking study found that the assembler plays an essential role in genome construction, especially for low-depth datasets, and that enough data depth is required for high-quality genome construction by ONT (greater than 80x) and HiFi (greater than 20x) datasets [<a href="#ref-1">1</a>]. If you sequence below these depths, your assembly quality will suffer regardless of which assembler you use. Evaluate your depth before concluding that a particular assembler or polishing tool is inadequate.

Applying Bacterial Quality Standards to Eukaryotic Genomes

Bacterial genomes are small, haploid, and relatively simple. Eukaryotic genomes are larger, often repetitive, and may be highly heterozygous. The quality thresholds that work for bacteria do not transfer to eukaryotes. A BUSCO completeness of 95 percent may be excellent for a plant genome but unacceptable for a bacterial genome. The silkworm study highlighted the challenges of high genetic diversity and heterozygosity in Lepidoptera, which required careful assembler selection [<a href="#ref-2">2</a>].

Failing to Distinguish Between Haplotype Collapse and True Duplication

Merqury's copy number spectrum distinguishes between collapsed and duplicated regions. If you see an excess of single-copy k-mers in a diploid assembly, haplotypes may be collapsed. If you see an excess of multi-copy k-mers, the assembly may contain false duplications. The interpretation depends on your assembly strategy and the expected ploidy of your organism.

Using Outdated Lineage Datasets for BUSCO

BUSCO lineage datasets are updated periodically to include more species and better gene models. Using an outdated dataset can produce misleading completeness scores. Record the lineage dataset version and check for updates before starting a new project. The EMBL-EBI Training portal provides guidance on using bioinformatics data resources effectively, including understanding versioning and updates [<a href="#ref-6">6</a>].

Limitations of Current Evaluation Methods

Reference Bias in Comparative Evaluation

When you evaluate an assembly against a reference genome, you inherit the reference's errors and biases. The reference may have its own misassemblies, missing regions, or haplotype artifacts. QUAST misassembly detection compares your assembly to the reference, but if the reference is wrong, the comparison may flag correct assemblies as misassembled or miss real errors.

K-mer Based Methods Require Sufficient Depth

Merqury and similar k-mer based methods require enough sequencing depth to distinguish true k-mers from sequencing errors. At low depth, the k-mer spectrum becomes noisy, and completeness estimates become unreliable. The yeast benchmarking study recommended greater than 80x depth for ONT and greater than 20x for HiFi [<a href="#ref-1">1</a>]. Below these depths, k-mer based evaluation should be interpreted with caution.

BUSCO Measures Gene Content, Not Structural Correctness

BUSCO completeness tells you whether expected genes are present, but it does not tell you whether the assembly is structurally correct. A misjoined assembly can contain all BUSCO genes while having incorrect chromosome structure. The silkworm study used Inspector and EagleC in addition to BUSCO to evaluate structural correctness [<a href="#ref-2">2</a>]. For chromosome-level assemblies, structural evaluation tools are essential.

QV Estimates Vary by Method

QV estimated by Merqury may differ from QV estimated by reference-based comparison. The two methods measure different things. Merqury measures k-mer consensus accuracy, while reference-based QV measures agreement with a specific reference genome. Report the method used to estimate QV and interpret the value accordingly.

Professional Escalation Criteria

When to Seek Expert Consultation

If your assembly fails to improve after multiple polishing rounds, or if BUSCO completeness remains below 90 percent for a bacterial genome or below 80 percent for a eukaryotic genome, consult with a bioinformatics specialist or core facility. The problem may be in the sequencing library preparation, the assembler parameters, or the evaluation approach itself.

When to Consider Additional Sequencing

If Merqury completeness is below 95 percent or if the k-mer spectrum shows substantial missing k-mers, additional sequencing may be necessary. The yeast benchmarking study showed that depth requirements differ by platform, with ONT requiring greater than 80x and HiFi requiring greater than 20x [<a href="#ref-1">1</a>]. If you are below these thresholds, adding sequencing data is likely to improve assembly quality more than changing assemblers or polishing tools.

When to Reject an Assembly

Reject an assembly if it fails to meet the minimum quality thresholds for your intended application. For clinical or diagnostic applications, the quality requirements are higher than for exploratory research. The 2020 study on nanopore-derived genome-scale models showed that models could be more than 99 percent complete even at depths below 10x coverage, but the authors also noted that sequencing errors may negatively affect the quality and utility of downstream products [<a href="#ref-3">3</a>]. Define your quality thresholds before starting the project, and do not lower them to accommodate a poor assembly.

Building a Decision Framework for Assembly Acceptance and Iteration

Running QUAST, BUSCO, and Merqury produces a collection of metrics, but the harder task is deciding what to do when those metrics conflict or fall short of expectations. A genome assembly project rarely produces a single clean pass or fail result. More often, you will see strong contiguity alongside mediocre completeness, or excellent base-level accuracy with signs of haplotype collapse. Without a structured decision framework, researchers tend to either accept an assembly that is not fit for purpose or spend weeks polishing an assembly that has structural problems no amount of base correction can fix. This section provides a practical framework for interpreting conflicting metrics, deciding when to iterate, and knowing when to stop.

Defining Quality Tiers Before You Start

The most common mistake in assembly evaluation is defining quality thresholds after seeing the results. This creates a conflict of interest where the assembly you produced becomes the standard you judge it against. Instead, define three quality tiers before you run your first assembler, based on your downstream application and the biological questions you need to answer.

Tier 1 is the minimum acceptable quality for your stated research objective. For a bacterial pathogen study where you need to identify antimicrobial resistance genes, Tier 1 might require BUSCO completeness above 99 percent, a single circular chromosome, and QV above 40 for reliable variant detection. The 2020 study on nanopore-derived genome-scale models demonstrated that models could be more than 99 percent complete even at sequencing depths below 10x coverage, which shows that modest sequencing investment can meet some research needs [<a href="#ref-3">3</a>]. However, the same study cautioned that sequencing errors inherent to the nanopore technique may negatively affect the quality and utility of downstream products [<a href="#ref-3">3</a>]. Your Tier 1 threshold must account for the error sensitivity of your specific downstream analysis.

Tier 2 is the quality level you would be satisfied with for publication and most comparative analyses. This tier typically requires BUSCO completeness above 95 percent for eukaryotes, QV above 40, and contiguity that approaches chromosome-scale for organisms with known chromosome numbers. Tier 3 is the reference-quality standard, which for most projects means QV above 50, BUSCO completeness above 98 percent, and near-complete representation of chromosomes including telomeres and centromeres where feasible.

Write these thresholds into your project notebook before assembly begins. Record the rationale for each threshold, citing the requirements of your downstream analysis. When you later face the decision of whether to accept an assembly, you can compare against your precommitted standards instead of negotiating with yourself about what is good enough.

The Metric Conflict Resolution Protocol

Assembly metrics frequently disagree with each other. A common scenario is high N50 with low BUSCO completeness. This pattern suggests that the assembler produced long contigs by joining regions that may not belong together, or that repetitive regions collapsed during assembly, removing duplicated gene copies. Another common pattern is excellent BUSCO completeness with poor QV, which indicates that the gene content is present but the base-level accuracy is insufficient for variant calling.

When metrics conflict, use the following protocol to determine which metric takes priority for your application. First, identify your primary metric based on the biological question. If you are identifying single nucleotide polymorphisms associated with drug resistance, QV is your primary metric and contiguity is secondary. If you are studying gene family evolution, BUSCO duplication ratio and completeness are primary. If you are building a reference genome for a community resource, all three metric categories matter, and you should aim for Tier 3 across the board.

Second, determine whether the conflict indicates a correctable problem or a fundamental limitation. Base-level errors are correctable with polishing. The yeast benchmarking study found that polishing by Pilon and Medaka improved accuracy and continuity of preassemblies, and their combination pipeline worked well in most quality metrics [<a href="#ref-1">1</a>]. Structural problems such as misjoins, collapsed repeats, or haplotype chimeras are not correctable by polishing and may require reassembly with different parameters or additional data.

Third, use Merqury to adjudicate between contiguity and completeness conflicts. The k-mer copy number spectrum reveals whether long contigs are biologically real or artifacts of misassembly. If the spectrum shows a clean dominant peak at the expected ploidy, the contiguity is likely genuine. If the spectrum shows abnormal copy number distributions, the long contigs may be masking structural errors. The lepidopteran genome study used BUSCO, Inspector, and EagleC in addition to contiguity metrics because quality assessment is essential for genome assembly and can provide better and more accurate results [<a href="#ref-2">2</a>]. Multiple tools provide independent evidence that helps resolve metric conflicts.

The Iteration Decision Matrix

When your assembly falls short of your precommitted thresholds, you face a choice between polishing, reassembling with different parameters, adding sequencing data, or accepting a lower quality tier. The following matrix guides this decision based on the pattern of metric failures.

If BUSCO completeness is below target but QV is acceptable, the problem is likely missing sequence instead of incorrect sequence. This pattern suggests the assembler failed to incorporate some reads, possibly due to coverage gaps or repetitive regions that were not bridged. Try reassembling with different parameters, particularly those that affect repeat resolution, or add sequencing data to increase coverage. The yeast benchmarking study found that the assembler plays an essential role in genome construction, especially for low-depth datasets [<a href="#ref-1">1</a>]. If you are below the recommended depth thresholds of greater than 80x for ONT and greater than 20x for HiFi, additional sequencing is likely to help more than parameter changes [<a href="#ref-1">1</a>].

If QV is below target but BUSCO completeness is acceptable, the problem is base-level accuracy. This pattern is correctable with polishing. Run one or two rounds of polishing and re-evaluate. The yeast study found that polishing improved accuracy and continuity of preassemblies, with the combination of Pilon and Medaka working well in most quality metrics [<a href="#ref-1">1</a>]. If polishing does not improve QV, check whether your read set contains systematic errors that the polisher cannot correct, such as methylation-induced base modifications in ONT data.

If both BUSCO completeness and QV are below target, the problem may be insufficient sequencing depth or poor library quality. Evaluate your depth against the recommended thresholds and consider whether the library preparation introduced biases. The 2020 nanopore study showed that genome-scale models could be constructed from assemblies at depths below 10x coverage, but the authors noted that sequencing errors may negatively affect downstream utility [<a href="#ref-3">3</a>]. Low depth affects both completeness and accuracy simultaneously, and adding data is usually the most effective remedy.

If contiguity is below target but completeness and QV are acceptable, the assembly is fragmented but correct. This pattern is common in genomes with long repetitive regions that long reads cannot span. Consider whether your application requires higher contiguity. If you need chromosome-scale scaffolds, you may need ultralong reads or additional sequencing to bridge gaps. The silkworm study achieved a nearly complete telomere-to-telomere reference genome using PacBio HiFi sequencing, demonstrating what is possible with sufficient data and appropriate assembler choice [<a href="#ref-2">2</a>].

The Stop-Go-Change Decision Record

For each assembly project, maintain a decision record that documents the evaluation results, the decision made, and the rationale. This record serves multiple purposes. It provides transparency for collaborators and reviewers, it creates a baseline for comparing future assemblies, and it prevents you from repeating unsuccessful strategies.

The decision record should include the date, the assembly version, all quality metrics from QUAST, BUSCO, and Merqury, the precommitted thresholds for your quality tiers, the specific decision made (accept, polish, reassemble, add data, or reject), and the rationale referencing the metric patterns described above. The nf-core documentation emphasizes reproducibility standards for bioinformatics pipelines, including version pinning and configuration tracking [<a href="#ref-5">5</a>]. Apply the same rigor to your decision record by documenting software versions and parameters for every evaluation run.

The Carpentries lessons provide foundational training in version control with Git, which is essential for tracking changes to your analysis scripts and decision records [<a href="#ref-8">8</a>]. If you maintain your decision records in a version-controlled repository, you can trace how your evaluation strategy evolved over the course of a project and identify which decisions led to successful outcomes.

Escalation Criteria for Persistent Quality Failures

If your assembly fails to meet Tier 1 thresholds after two rounds of polishing and one reassembly attempt, escalate the problem instead of continuing to iterate blindly. Persistent failures across multiple assemblers and parameter sets indicate a problem that is not correctable by assembly optimization alone.

First, evaluate your sequencing data quality. Check read length distributions, estimated depth, and error profiles. The yeast benchmarking study found that enough data depth is required for high-quality genome construction by ONT (greater than 80x) and HiFi (greater than 20x) datasets [<a href="#ref-1">1</a>]. If your depth is below these thresholds, additional sequencing is the most direct remedy. If your depth is adequate but read quality is poor, the problem may be in library preparation or sequencing chemistry.

Second, consult with a bioinformatics specialist or core facility. The EMBL-EBI Training portal offers learning pathways for bioinformatics data resources, including practical analysis education that can help you identify problems in your evaluation approach [<a href="#ref-6">6</a>]. A specialist may identify issues in your assembler parameters, BUSCO lineage selection, or k-mer size choice that you have overlooked.

Third, consider whether your quality thresholds are appropriate for your organism and data type. The lepidopteran study highlighted the challenges of high genetic diversity and heterozygosity in insects, which required careful assembler selection [<a href="#ref-2">2</a>]. If you are working with a highly heterozygous or polyploid organism, your thresholds may need adjustment to reflect biological reality instead of assembly failure. Document any threshold adjustments in your decision record with clear justification.

When to Accept a Lower Quality Tier

Sometimes the cost of achieving Tier 2 or Tier 3 quality exceeds the value it provides for your research question. The 2020 nanopore study demonstrated that genome-scale models could be constructed from assemblies at depths below 10x coverage, and these models identified canonical antimicrobial resistance content and enabled simulations of strain-specific microbial growth [<a href="#ref-3">3</a>]. For this application, Tier 1 quality was sufficient, and investing in additional sequencing would have added cost without improving the biological conclusions.

Accepting a lower quality tier is a legitimate decision when you have documented the limitations and confirmed that your downstream analysis is robust to them. The decision record should state which metrics fall short of your original thresholds, why the assembly is still acceptable for your application, and what analyses might be compromised by the quality shortfall. This documentation protects you from overinterpreting results that the assembly quality cannot support.

The Bioconductor project provides official documentation for reproducible genomic analysis workflows, including package installation and version management [<a href="#ref-9">9</a>]. Apply the same reproducibility principles to your quality acceptance decisions by recording them in a format that collaborators and reviewers can examine. A well-documented decision to accept a lower quality tier is more scientifically defensible than an undocumented claim that an assembly is high quality when the metrics do not support it.

Building a Reusable Evaluation Pipeline

Once you have established your quality tiers, decision framework, and record system, codify them into a reusable pipeline. The Galaxy Training Network provides accessible workflow training for genome assembly evaluation, including practical tutorials that walk through analysis in a reproducible environment [<a href="#ref-4">4</a>]. The nf-core documentation describes community pipeline standards for assembly and quality control, including configuration guidance for running these tools at scale [<a href="#ref-5">5</a>]. These resources help you standardize your evaluation across multiple samples or projects.

A reusable pipeline should accept raw reads and an assembly as input, run QUAST, BUSCO, and Merqury with your standard parameters, generate a summary report, and compare the results against your precommitted quality tiers. The pipeline should also generate the decision record template so that every assembly project produces consistent documentation. This investment in pipeline development pays off when you assemble multiple genomes, because the evaluation step becomes automated and the decision framework becomes consistent across projects.

The NCBI Data Resources provide access to sequence databases and analysis services that can support assembly evaluation, including tools for comparing your assembly to related sequences [<a href="#ref-7">7</a>]. Integrating these resources into your pipeline can provide additional context for interpreting your quality metrics, such as expected genome sizes and gene content for related species.

Frequently Asked Questions

What is the difference between N50 and NG50?

N50 is calculated against the total assembly length, while NG50 is calculated against the estimated genome size. If the assembly size exceeds the genome size due to haplotype duplication or contamination, N50 will be higher than NG50. NG50 is the more conservative metric and is preferred when you have a reliable genome size estimate. For bacterial genomes, the two values should be similar because the assembly should match the genome size closely.

How many BUSCO genes should be complete for a high-quality assembly?

The target depends on the organism and the application. For bacterial genomes, more than 99 percent complete BUSCO genes is expected for a high-quality assembly. For eukaryotic genomes, more than 90 percent complete is often considered good, and more than 95 percent is considered excellent. The lepidopteran study achieved a nearly complete telomere-to-telomere assembly, which represents the upper bound of what is currently possible [<a href="#ref-2">2</a>]. Compare your results to published assemblies of related species to calibrate expectations.

What QV value indicates a reference-quality assembly?

QV40, corresponding to one error per 10,000 bases, is often cited as a minimum for reference-quality assemblies. QV50 corresponds to one error per 100,000 bases and is preferred for applications that require high base-level accuracy, such as variant calling. The yeast benchmarking study showed that polishing improves accuracy, so you should polish before estimating the final QV [<a href="#ref-1">1</a>]. Merqury provides a reference-free QV estimate that is useful when no reference genome is available.

Can I use QUAST without a reference genome?

Yes, QUAST operates in reference-free mode and reports contiguity metrics, GC content, and Nx plots. However, misassembly detection requires a reference genome. For reference-free structural evaluation, use Inspector or EagleC as described in the lepidopteran study [<a href="#ref-2">2</a>]. Merqury provides reference-free completeness and accuracy estimates based on k-mers.

How does sequencing depth affect assembly quality metrics?

The yeast benchmarking study found that enough data depth is required for high-quality genome construction by ONT (greater than 80x) and HiFi (greater than 20x) datasets [<a href="#ref-1">1</a>]. Below these depths, contiguity and completeness metrics decline, and polishing may not fully compensate. If your depth is below these thresholds, consider additional sequencing before investing time in assembly optimization.

What does a high duplication ratio in BUSCO indicate?

A high duplication ratio indicates that many BUSCO genes are present in two or more copies. This can result from haplotype separation in a diploid assembly, from a recent whole-genome duplication in the organism, or from assembly artifacts that duplicate regions. Merqury's copy number spectrum can help distinguish between these possibilities. If you are assembling a haploid genome and see high duplication, the assembly likely has structural errors.

Should I polish my assembly before or after running quality metrics?

Run quality metrics before and after polishing. The yeast benchmarking study found that polishing by Pilon and Medaka improved accuracy and continuity of preassemblies, and their combination pipeline worked well in most quality metrics [<a href="#ref-1">1</a>]. Running metrics before polishing establishes a baseline, and running them after quantifies the improvement. If polishing does not improve the metrics, the assembly may have structural problems that base-level correction cannot fix.

How do I choose between PacBio HiFi and ONT for my assembly project?

The choice depends on your budget, timeline, and quality requirements. The yeast benchmarking study found that HiFi required greater than 20x depth while ONT required greater than 80x depth for high-quality assemblies [<a href="#ref-1">1</a>]. HiFi reads have lower error rates and are well suited for base-level accuracy, while ONT can produce ultralong reads that span difficult repeats. The silkworm study used both platforms and found that hifiasm was better for HiFi data while NextDenovo was superior for CLR and ONT data [<a href="#ref-2">2</a>]. Consider the total cost of sequencing and the quality requirements of your downstream applications.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Benchmarking of long-read sequencing, assemblers and polishers for yeast genome.](https://pubmed.ncbi.nlm.nih.gov/35511110). Briefings in bioinformatics, 2022. [2] [Comparison of Long-Read Methods for Sequencing and Assembly of Lepidopteran Pest Genomes.](https://pubmed.ncbi.nlm.nih.gov/36614092). International journal of molecular sciences, 2022. [3] [High-Quality Genome-Scale Models From Error-Prone, Long-Read Assemblies.](https://pubmed.ncbi.nlm.nih.gov/33281796). Frontiers in microbiology, 2020. [4] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [5] [nf-core Documentation](https://nf-co.re/docs). nf-core. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.