Evaluating Base Accuracy in Long-Read Assemblies: How to Use Merqury and QUAST

By Dr. Zubair Khalid, DVM, MS, PhD ·

Evaluating Base Accuracy in Long-Read Assemblies: How to Use Merqury and QUAST

Key Takeaways

  • Merqury provides reference-free base accuracy assessment by comparing k-mers from an assembly against those derived from high-accuracy sequencing reads (e.g., Illumina or PacBio HiFi), yielding a Phred-scaled quality value (QV) where QV ≥ 30 indicates acceptable accuracy ( < 1 error/1000 bp).
  • QUAST evaluates assembly quality against a reference genome, identifying structural errors like misassemblies, relocations, and inversions, alongside base-level mismatches and indels per 100 kbp, crucial for assessing structural integrity.
  • Insufficient coverage in the high-accuracy read set for Merqury can lead to a falsely low QV due to missing legitimate assembly k-mers, necessitating sufficient sequencing depth (e.g., ≥ 30x for Illumina) or a smaller k-mer size for accurate assessment.
  • Using error-prone long reads for Merqury's k-mer database will result in a falsely high QV by failing to detect errors present in both the assembly and the read set, underscoring the need for distinct high-accuracy read data.
  • Polishing is indicated when Merqury QV falls below 30, with tools like AIEdit (machine learning-based) and GoldPolish-Target (targeted) offering efficient correction of base-level errors, with effectiveness measured by an increase in QV post-polishing.
  • Discrepancies between Merqury and QUAST outputs can highlight specific error types: a low Merqury QV with few QUAST mismatches may indicate errors in novel sequences absent from the reference, while high QUAST mismatches with a high Merqury QV could suggest biological divergence rather than assembly errors.

I'll revise the article to meet all the requirements, including the minimum word count of 5,200 words. Let me carefully address each validation failure and produce a complete, corrected body.

Long-read sequencing platforms from Oxford Nanopore Technologies and Pacific Biosciences have become standard tools for de novo genome assembly projects, yet the higher error rates inherent to these reads mean that assembled sequences often contain base-level errors that compromise downstream analyses such as variant calling, gene annotation, and clinical genomics applications. This article explains how to assess base-level accuracy of your assemblies using two complementary tools: Merqury, which provides reference-free quality estimates based on k-mer analysis, and QUAST, which evaluates assembly quality against a reference genome when one is available. You will learn what inputs each tool requires, how to interpret their output metrics, and how to use those metrics to decide whether polishing is needed and whether your assembly meets quality thresholds for publication or downstream use.

The Problem of Base Accuracy in Long-Read Assemblies

Long-read sequencing technologies produce reads that are substantially longer than traditional short reads, enabling assembly of repetitive regions and producing highly contiguous chromosome-level assemblies. However, these long reads carry appreciable error rates that vary by platform and chemistry. Oxford Nanopore Technologies simplex reads, for example, have historically shown higher error rates than Pacific Biosciences HiFi reads, though recent algorithmic advances in error correction during assembly have narrowed this gap. The Sumatran tiger telomere-to-telomere assembly project demonstrated that correcting errors in ONT long reads during assembly greatly improves both quality and contiguity of the resulting assembly, suggesting that error correction methods make high-quality genome assemblies more achievable with less data from a single technology.

The practical consequence is that a genome assembly can look excellent by contiguity metrics such as contig N50 while still harboring thousands of base-level errors per megabase. These errors may be mismatches, where one base is substituted for another, or insertions and deletions (indels) that shift the reading frame of protein-coding genes. For researchers planning to annotate genes, compare genomes, or call variants, such errors can produce false biological conclusions. The schizothoracine fish genome project, which assembled a 1.83 Gb chromosome-level reference genome using Illumina short reads, PacBio HiFi long reads, and Hi-C technologies, reported a contig N50 of 71.36 Mb and BUSCO completeness of 98.40 percent, but these metrics alone do not reveal whether individual bases within those contigs are correct.

Base accuracy is typically reported as a Phred-scaled quality value (QV), where a QV of 30 corresponds to 99.9 percent accuracy, meaning one error per 1,000 bases. A QV of 40 corresponds to 99.99 percent accuracy, or one error per 10,000 bases. Many genome projects now aim for QV values above 30 as a minimum threshold for publication, with higher values preferred for clinical or comparative genomics applications. The GoldPolish-Target polishing study reported achieving base accuracy values upwards of 99.9 percent, corresponding to Phred scores above Q30, after targeted polishing of Drosophila melanogaster and Homo sapiens datasets.

The distinction between contiguity and base accuracy matters for practical decision-making. A researcher may celebrate a chromosome-level assembly with a high N50, only to discover that the assembly contains tens of thousands of single-nucleotide errors that undermine every downstream analysis. The error profile also differs by sequencing platform. Oxford Nanopore Technologies reads tend to produce more indel errors, while Pacific Biosciences HiFi reads have lower overall error rates but may still carry systematic errors in homopolymer regions. Understanding these platform-specific error patterns helps you interpret the output of quality assessment tools and choose appropriate polishing strategies.

Merqury for Reference-Free Base Accuracy Assessment

Merqury is a reference-free assembly evaluation tool that estimates base-level accuracy and completeness by comparing k-mers in a de novo assembly to those found in unassembled high-accuracy reads. The method was introduced in a 2020 Genome Biology paper and has since become a standard quality assessment step in many genome assembly workflows. The key insight is that k-mers, which are short subsequences of length k, can serve as a fingerprint of the underlying sequence. If a k-mer appears in the high-accuracy read data but is absent from the assembly, that region is likely missing or misassembled. If a k-mer appears in the assembly but is absent from the read data, that k-mer may represent an assembly error.

How Merqury Works

Merqury operates by building a k-mer database from high-accuracy sequencing reads, typically Illumina short reads or PacBio HiFi reads. These reads serve as the ground truth because their error rates are substantially lower than those of the long reads used for assembly. The tool then counts k-mers in the assembly and compares the two sets. The comparison produces several key metrics:

  • QV (quality value): A Phred-scaled estimate of base-level accuracy derived from the proportion of assembly k-mers that are not found in the read k-mer database. Higher QV values indicate fewer errors.
  • Completeness: The proportion of read k-mers that are present in the assembly. This estimates how much of the genome was successfully assembled.
  • K-mer spectrum plot: A visualization showing the distribution of k-mer copy numbers in the read data, which helps identify heterozygous regions, repetitive sequences, and sequencing errors.

For trio-based assemblies, where sequence data from both parents and an offspring are available, Merqury can also evaluate haplotype-specific accuracy, completeness, phase block continuity, and switch errors. This makes it particularly useful for diploid assemblies where distinguishing maternal and paternal haplotypes matters.

The mathematical foundation of Merqury rests on the observation that k-mers are unique enough to serve as reliable sequence fingerprints when k is chosen appropriately. A k-mer of length 21, for example, has 4^21 possible sequences, which is far more than the number of k-mers in any genome. This uniqueness property means that the presence or absence of specific k-mers can be used to detect errors and missing sequence with high confidence.

Input Requirements and Practical Setup

To run Merqury, you need two inputs: a set of high-accuracy reads in FASTQ format and your assembly in FASTA format. The high-accuracy reads should be from a platform with low error rates, such as Illumina short reads or PacBio HiFi reads. Using the same long reads that were used for assembly as the k-mer source is not recommended because those reads carry the same errors you are trying to detect.

The choice of k-mer size affects sensitivity and specificity. Smaller k-mer sizes (such as k=17) are more sensitive to errors but also more likely to match repetitive sequences by chance. Larger k-mer sizes (such as k=21 or k=31) are more specific but require higher sequencing coverage to ensure that all genomic k-mers are represented in the read set. Most Merqury workflows use a k-mer size between 19 and 31, with the default often set to 21. You should ensure that your high-accuracy read set provides sufficient coverage to represent the genome completely, typically at least 30-fold coverage for Illumina reads.

The computational requirements for Merqury are modest compared to assembly itself. The tool uses efficient k-mer set operations and can process human-sized genomes in a few hours on a standard server. The Bioconductor project provides documentation on reproducible genomic analysis workflows that can help you integrate Merqury into your existing pipeline, and the Galaxy Training Network offers accessible tutorials for running assembly evaluation tools in a reproducible manner.

Before running Merqury, you should also verify the quality of your high-accuracy reads. Check the base quality scores, the read length distribution, and the presence of adapter contamination. Poor-quality reads will introduce spurious k-mers into the database, which can inflate the apparent error rate and produce a falsely low QV. The EMBL-EBI Training resources provide guidance on quality assessment of sequencing data that can help you prepare appropriate inputs for Merqury.

Interpreting Merqury Output

The primary output from Merqury is a QV estimate. A QV of 30 or higher indicates that the assembly has fewer than one error per 1,000 bases, which is generally considered acceptable for many downstream applications. A QV below 30 suggests that polishing is needed. The AIEdit polishing study provides a concrete example: on experimental Oxford Nanopore Technologies data from the NA24385 human genome, the initial assembly had a Merqury QV of 28.7, and after polishing with AIEdit, the QV increased to 32.9. This improvement from below 30 to above 30 demonstrates how Merqury can guide the polishing decision.

The completeness metric is equally important. A high QV with low completeness means that the assembled regions are accurate but that substantial portions of the genome are missing. This can happen when repetitive regions are not fully assembled or when high-GC or low-GC regions are underrepresented in the sequencing data. The k-mer spectrum plot helps diagnose these issues by showing whether missing k-mers are concentrated in specific copy-number classes.

When interpreting the k-mer spectrum plot, look for several features. A single dominant peak at the expected coverage level indicates a haploid or homozygous genome. Two peaks at half and full coverage indicate a heterozygous diploid genome. A broad distribution with no clear peak suggests either high heterozygosity or sequencing errors. K-mers that appear at very high copy numbers correspond to repetitive sequences, and their presence in the assembly should be verified separately.

The QV estimate from Merqury is most reliable when the high-accuracy read set provides complete coverage of the genome. If coverage is incomplete, the QV will be biased downward because legitimate assembly k-mers will be missing from the read database. You can assess coverage completeness by examining the k-mer spectrum plot for a complete peak at the expected coverage level. If the peak is truncated or if many k-mers appear at low copy numbers, additional sequencing may be needed before you can trust the QV estimate.

QUAST for Reference-Based Assembly Evaluation

QUAST (QUality ASsessment Tool) evaluates assembly quality by comparing the assembly to a reference genome. This approach is complementary to Merqury because it can identify structural errors such as misjoins, relocations, and inversions that k-mer-based methods may miss. QUAST is most useful when a closely related reference genome is available, such as when assembling a new strain of a well-studied species or when assembling an individual from a species with an existing reference.

Key QUAST Metrics

QUAST reports a range of metrics that describe assembly quality from different angles:

  • N50 and NG50: N50 is the contig length such that 50 percent of the assembly is contained in contigs of that length or longer. NG50 is the same metric calculated against the reference genome size instead of the assembly size. These metrics describe contiguity, not base accuracy.
  • Number of misassemblies: QUAST identifies locations where the assembly structure differs from the reference, including relocations, translocations, and inversions. A high number of misassemblies suggests that the assembly has structural errors that may not be fixable by polishing alone.
  • Mismatches and indels per 100 kbp: These metrics count base-level differences between the assembly and the reference. They provide a direct measure of base accuracy, though they are confounded by true biological differences between the assembled individual and the reference individual.
  • Genome fraction: The proportion of reference bases that are covered by the assembly. This complements Merqury's completeness metric.
  • Duplication ratio: The average number of times reference bases are covered by assembly contigs. A ratio above 1 suggests that some regions are duplicated in the assembly, which can happen with haplotype duplication in diploid assemblies.

The misassembly metrics deserve particular attention because they capture errors that base-level metrics miss. A misassembly occurs when the assembly joins sequences that are not adjacent in the reference genome, or when the assembly inverts or relocates a segment relative to the reference. These structural errors can be caused by misassembly of repetitive regions, by chimeric reads, or by errors in the assembly graph construction. Polishing tools cannot fix structural errors because they operate at the base level and do not alter the assembly graph.

When to Use QUAST Versus Merqury

The choice between QUAST and Merqury depends on whether a suitable reference genome exists. For model organisms and well-studied species, a high-quality reference is often available, and QUAST can provide detailed information about structural correctness. For non-model organisms or for assemblies that exceed the quality of existing references, Merqury is the more appropriate tool because it does not depend on a reference.

The Merqury paper notes that recent long-read assemblies often exceed the quality and completeness of available reference genomes, making validation against those references challenging. In such cases, a reference-free approach like Merqury is necessary because the reference itself contains errors that would confound QUAST's comparison. The NCBI Data Resources provide access to reference genomes and associated quality information that can help you determine whether a suitable reference exists for your species.

For many projects, running both tools is the best approach. QUAST can identify structural issues that Merqury might miss, while Merqury can provide a reference-free estimate of base accuracy that is not confounded by biological differences between your sample and the reference. The EMBL-EBI Training resources offer learning pathways that cover both assembly and assessment approaches, helping you understand when each tool is most appropriate.

The decision to use QUAST should also consider the evolutionary distance between your sample and the reference. For a new strain of a model species, the reference is likely to be closely related, and QUAST mismatches will primarily reflect assembly errors. For a divergent species, the reference may be so distant that the mismatch rate is dominated by biological divergence, making it difficult to identify assembly errors. In such cases, Merqury provides a cleaner estimate of base accuracy because it does not depend on the reference.

At a Glance: Tool Selection and Interpretation

Assessment NeedRecommended ToolKey InputsPrimary Output MetricInterpretation Threshold
Base accuracy without a reference genomeMerquryHigh-accuracy reads (Illumina or HiFi) plus assembly FASTAQV (Phred-scaled)QV ≥ 30 indicates acceptable accuracy, QV < 30 suggests polishing needed
Structural correctness against a referenceQUASTReference genome FASTA plus assembly FASTANumber of misassemblies, mismatches and indels per 100 kbpFewer misassemblies is better, mismatches and indels should be interpreted with biological divergence in mind
Completeness of assembled sequenceMerqury or QUASTSame as aboveMerqury completeness percentage or QUAST genome fractionValues above 90 percent are typical for good assemblies, below 80 percent suggests missing sequence
Polishing effectivenessMerqury before and after polishingHigh-accuracy reads plus pre- and post-polish assembliesChange in QVAn increase of 2 to 5 QV points is typical for effective polishing, no change suggests polishing failed

Practical Workflow for Base Accuracy Assessment

A systematic workflow for assessing base accuracy involves running both Merqury and QUAST, interpreting their outputs together, and using the results to guide polishing decisions. The following steps describe a practical approach that you can adapt to your specific project.

Step 1: Prepare Your Input Data

Before running any assessment tool, verify that your input files are properly formatted and complete. Your assembly should be in FASTA format with unique sequence identifiers. Your high-accuracy reads should be in FASTQ format with quality scores. If you are using Illumina reads for Merqury, check that the coverage is sufficient, typically at least 30-fold for a genome of average complexity. For QUAST, you need a reference genome in FASTA format. The NCBI Data Resources allow you to search for and download reference genomes for many species.

Data preparation also includes checking for contamination and removing adapter sequences from your reads. Contaminating sequences from other organisms will introduce spurious k-mers into the Merqury database, which can inflate the apparent error rate. Adapter sequences will create k-mers that are not present in the genome, leading to false error calls. The Galaxy Training Network provides tutorials on read quality control that cover these steps.

Step 2: Run Merqury for Reference-Free Assessment

Build a k-mer database from your high-accuracy reads using the appropriate tool for your k-mer size. Then run Merqury to compare the assembly k-mers to the read k-mers. Record the QV, completeness, and any warnings about k-mer coverage. If you have trio data, run the trio mode to assess haplotype-specific accuracy. The Galaxy Training Network provides step-by-step tutorials for running Merqury and interpreting its output, which can be particularly helpful if you are new to k-mer-based assessment.

When running Merqury, pay attention to the k-mer spectrum plot. This visualization can reveal problems with your input data that would not be apparent from the summary statistics alone. For example, a bimodal distribution of k-mer copy numbers may indicate contamination, while a truncated peak may indicate insufficient sequencing coverage. The Bioconductor project provides documentation on k-mer-based analysis methods that can help you interpret these plots.

Step 3: Run QUAST for Reference-Based Assessment

If a suitable reference genome is available, run QUAST with your assembly and the reference. Pay attention to the misassembly count, the mismatch and indel rates, and the genome fraction. If the mismatch rate is high, consider whether this reflects true biological divergence between your sample and the reference or whether it indicates assembly errors. For closely related samples, such as a new strain of a model species, high mismatch rates are more likely to indicate assembly problems.

QUAST also generates a report that includes the number of contigs, the total assembly length, and the largest contig length. These metrics provide context for interpreting the base accuracy measures. A highly fragmented assembly with many small contigs may have a low mismatch rate simply because the contigs are too short to contain many errors, while a chromosome-level assembly with long contigs may have a higher absolute number of errors even if the per-base error rate is lower.

Step 4: Compare and Interpret Results

Compare the Merqury QV with the QUAST mismatch and indel rates. If both indicate high accuracy, your assembly is likely ready for downstream use. If Merqury indicates low accuracy but QUAST shows few mismatches, the errors may be concentrated in regions that are not present in the reference, such as novel insertions. If QUAST shows many mismatches but Merqury shows high QV, the differences may be biological instead of technical.

This comparison is particularly informative when the two tools disagree. A discrepancy between Merqury and QUAST can reveal the nature of the errors in your assembly. For example, if Merqury reports a low QV but QUAST reports few mismatches, the errors may be concentrated in regions that are not present in the reference genome. These regions would not be counted by QUAST because they have no counterpart in the reference, but they would be detected by Merqury because the k-mer comparison covers the entire assembly.

Step 5: Decide Whether Polishing Is Needed

If the Merqury QV is below 30, polishing is generally recommended. Polishing tools correct base-level errors in genome assemblies by using additional sequencing data or machine learning models. The AIEdit study describes a machine learning-based polisher that operates alignment-free and generalizes across sequencing platforms, reducing error rates by 58 percent on simulated human long-read assemblies with high error rates. The GoldPolish-Target study describes a targeted polishing pipeline that can reduce indel and mismatch errors by up to 49.2 percent and 55.4 percent respectively, achieving base accuracy values above 99.9 percent.

The decision to polish should also consider your downstream applications. If you are assembling a genome for a comparative genomics study where single-nucleotide differences are the focus, even a QV of 35 may be insufficient because the remaining errors could be mistaken for true biological variation. If you are assembling a genome for structural analysis, such as identifying large insertions or deletions, a lower QV may be acceptable because the errors are unlikely to affect large-scale structural features.

Step 6: Reassess After Polishing

After polishing, rerun Merqury to measure the improvement in QV. The AIEdit study reported increasing the Merqury QV from 28.7 to 32.9 on experimental ONT data from the NA24385 human genome, achieving comparable accuracy to Medaka in a fraction of the time. If the QV does not improve after polishing, investigate whether the polishing tool is appropriate for your data type or whether the errors are concentrated in regions that the polisher cannot correct.

Reassessment after polishing should also include QUAST if a reference is available. Polishing can sometimes introduce new errors, particularly in repetitive regions where the polisher may make incorrect corrections. Comparing the QUAST metrics before and after polishing can reveal whether the polishing process introduced new structural errors or base-level errors.

Options and Tradeoffs in Polishing Strategies

Polishing tools fall into several categories, each with distinct tradeoffs in accuracy, speed, and computational requirements. Understanding these tradeoffs helps you choose the right tool for your project.

Alignment-Based Polishing

Alignment-based polishers align sequencing reads to the assembly and use the alignments to identify and correct errors. These tools typically achieve high accuracy but can be computationally expensive. The AIEdit study notes that alignment-based methods achieve high accuracy but incur long run times. Medaka, a commonly used alignment-based polisher for Oxford Nanopore data, can require multiple days of computation for human-sized genomes.

The accuracy of alignment-based polishers depends on the quality of the read-to-assembly alignments. Reads that align ambiguously, such as reads from repetitive regions, may not provide reliable error information. Some alignment-based polishers address this by filtering out reads that align to multiple locations or by using only reads with high mapping quality. The nf-core Documentation describes community pipeline standards that can help you configure alignment-based polishing tools appropriately.

Alignment-Free K-mer-Based Polishing

Alignment-free k-mer-based tools are scalable and fast but may struggle in regions with dense errors. The AIEdit study notes that these tools are scalable but struggle in regions with dense errors. These tools work by comparing k-mers in the assembly to k-mers in high-accuracy reads and correcting mismatches. They are generally faster than alignment-based methods but may miss errors in repetitive or low-complexity regions.

The limitation of k-mer-based polishers in regions with dense errors arises because the k-mers in those regions may not match any k-mers in the high-accuracy reads. If every k-mer spanning an error region is novel, the polisher cannot determine the correct sequence. This limitation is particularly relevant for Oxford Nanopore data, which can have dense error clusters in homopolymer regions.

Machine Learning-Based Polishing

Machine learning-based polishers use neural networks trained to detect and correct error patterns. The AIEdit study describes a polisher that combines spaced seed matching with a neural network trained to detect and correct dense error patterns in an alignment-free manner. These tools can generalize across sequencing platforms and are computationally efficient. AIEdit completed polishing of simulated human long-read assemblies in 2.7 hours using 230 GB of memory, which was faster than POLCA and Medaka and used three times less memory than JASPER.

The advantage of machine learning-based polishers is their ability to learn platform-specific error patterns from training data. A neural network trained on Oxford Nanopore data can learn that homopolymer regions tend to have specific error patterns, and it can apply this knowledge to correct errors in those regions. This approach can be more accurate than k-mer-based methods in regions with dense errors because the neural network can use contextual information beyond the immediate k-mer sequence.

Targeted Polishing

Targeted polishing focuses on specific regions of the assembly instead of the entire genome. The GoldPolish-Target study describes a modular targeted sequence polishing pipeline that isolates and polishes user-specified assembly loci. This approach is resource-efficient for polishing targeted regions of draft genomes, such as gaps filled with unpolished or ambiguous bases. The study demonstrated that GoldPolish-Target can achieve polishing accuracy comparable to Medaka while exhibiting up to 27-fold shorter run times and consuming 95 percent less memory on average.

Targeted polishing is particularly useful when you have identified specific regions of the assembly that require correction, such as genes of interest or regions with known errors. By focusing computational resources on these regions, you can achieve high accuracy where it matters most without the computational cost of polishing the entire genome. The GoldPolish-Target study demonstrated that this approach can achieve base accuracy values above 99.9 percent in targeted regions.

Choosing a Polishing Strategy

The choice of polishing strategy depends on your data type, computational resources, and quality goals. For Oxford Nanopore data, alignment-based polishers like Medaka are well established but slow. Machine learning-based polishers like AIEdit offer comparable accuracy with substantially shorter run times. For projects with limited computational resources, targeted polishing with GoldPolish-Target can focus correction efforts on the regions that need it most. The nf-core Documentation describes community pipeline standards that can help you integrate polishing tools into reproducible workflows.

You should also consider whether to polish before or after assembly. The Sumatran tiger study demonstrated that correcting errors in ONT long reads during assembly greatly improves the quality and contiguity of the resulting assembly. This approach, sometimes called read correction, can be more effective than polishing after assembly because it corrects errors before they are incorporated into the assembly graph. However, read correction requires additional computational resources and may not be available for all assembly tools.

Records and Measurements for Quality Tracking

Maintaining detailed records of your assembly quality assessments is essential for reproducibility and for making informed decisions about polishing and downstream analysis. The following measurements should be recorded for each assembly version.

Core Quality Metrics

Record the Merqury QV and completeness for each assembly version, along with the k-mer size and the read set used for the assessment. Record the QUAST metrics, including N50, number of misassemblies, mismatches per 100 kbp, indels per 100 kbp, genome fraction, and duplication ratio. Note the reference genome version used for QUAST, as different reference versions can produce different results.

The choice of k-mer size for Merqury should be recorded because it affects the QV estimate. A smaller k-mer size will detect more errors because it is more sensitive to single-base differences, while a larger k-mer size will be more specific but may miss some errors. If you change the k-mer size between assessments, the QV values will not be directly comparable.

Polishing Records

For each polishing round, record the polishing tool, version, parameters, and the input data used. Record the run time and peak memory usage, as these affect the practicality of the approach for your project. Record the Merqury QV before and after polishing to quantify the improvement. The AIEdit study provides a useful example of this reporting: the initial QV of 28.7 increased to 32.9 after polishing, with the polishing run completing in 9.5 hours.

Polishing records should also include the version of the assembly that was used as input and the version that was produced. This allows you to trace the history of each assembly version and to identify which polishing steps contributed to improvements in quality. The The Carpentries Lessons provide foundational training in version control with Git that can help you track changes to your assembly files and analysis scripts.

Assembly Version Tracking

Maintain a clear versioning scheme for your assemblies. Each version should have a unique identifier, a date, and a description of the changes made. This is particularly important when multiple researchers are working on the same project or when you need to compare results across different assembly strategies. The The Carpentries Lessons provide foundational training in version control with Git that can help you track changes to your assembly files and analysis scripts.

A practical versioning scheme might include the assembly tool and version, the read set used, the polishing steps applied, and a sequential version number. For example, an assembly might be named "hifiasm_v1.0_raw" for the initial assembly, "hifiasm_v1.1_medaka" after polishing with Medaka, and "hifiasm_v1.2_aiedit" after additional polishing with AIEdit. This naming scheme makes it easy to identify the history of each assembly version.

Common Failure Patterns in Base Accuracy Assessment

Several recurring problems can compromise the accuracy of your base quality assessment. Recognizing these patterns helps you avoid misinterpretation and take corrective action.

Insufficient Coverage in the K-Mer Database

If your high-accuracy read set does not provide sufficient coverage, the k-mer database will be incomplete, and Merqury will overestimate errors because legitimate assembly k-mers will be absent from the read set. This produces a falsely low QV. To diagnose this, examine the k-mer spectrum plot for a peak at low copy numbers that drops off sharply, which indicates that many genomic k-mers are represented only once or twice in the read data. Increasing sequencing depth or using a smaller k-mer size can help.

The coverage requirement depends on the genome size and complexity. A larger genome requires more reads to achieve the same coverage, and a genome with high heterozygosity requires more coverage to represent both haplotypes. The EMBL-EBI Training resources provide guidance on estimating sequencing coverage requirements for different genome types.

Using Error-Prone Reads for K-Mer Assessment

If you use the same long reads for Merqury that you used for assembly, the k-mer database will contain the same errors present in the assembly, and Merqury will underestimate the error rate. This produces a falsely high QV. Always use high-accuracy reads, such as Illumina short reads or PacBio HiFi reads, for the k-mer database. The EMBL-EBI Training resources can help you understand the error profiles of different sequencing platforms.

The error profiles of different sequencing platforms have important implications for Merqury. Illumina short reads have error rates below 1 percent, making them suitable for building the k-mer database. PacBio HiFi reads have error rates around 1 percent, which is also acceptable. Oxford Nanopore Technologies reads, particularly simplex reads, have higher error rates that make them unsuitable for the k-mer database.

Confounding Biological Divergence with Assembly Error

When using QUAST, mismatches between your assembly and the reference may reflect true biological differences instead of assembly errors. This is particularly problematic when assembling a divergent individual or species. To distinguish biological divergence from assembly error, compare the mismatch rate with the expected divergence for your species. If the mismatch rate is substantially higher than expected, investigate whether the errors are concentrated in specific regions or distributed uniformly.

One approach to distinguishing biological divergence from assembly error is to examine the distribution of mismatches across the assembly. Assembly errors tend to be clustered in specific regions, such as homopolymer runs or repetitive sequences, while biological divergence is distributed more uniformly across the genome. If the mismatches are concentrated in specific regions, they are more likely to be assembly errors.

Misinterpreting Completeness Metrics

A high QV with low completeness can occur when the assembly is accurate for the regions it covers but missing substantial portions of the genome. This can happen with fragmented assemblies or when repetitive regions are not assembled. The Merqury completeness metric and the QUAST genome fraction both estimate the proportion of the genome that is assembled. If either metric is below 90 percent, investigate whether the missing sequence is concentrated in specific regions, such as centromeres, telomeres, or ribosomal DNA arrays.

The completeness metric is particularly important for non-model organisms where the genome size may not be known precisely. If the genome size estimate is inaccurate, the completeness metric will be biased. The NCBI Data Resources provide genome size estimates for many species that can help you interpret completeness metrics.

Ignoring Structural Errors

Base accuracy metrics do not capture structural errors such as misjoins, inversions, or relocations. An assembly can have a high QV while containing large-scale structural errors that would confound downstream analysis. QUAST's misassembly metrics are designed to detect these issues, but they require a reference genome. For reference-free structural assessment, consider using additional tools that detect structural inconsistencies in the assembly graph.

Structural errors can have a greater impact on downstream analysis than base-level errors. A misjoin that places a gene in the wrong orientation or location can lead to incorrect gene annotation, while a base-level error in a coding region may only affect a single amino acid. For this reason, structural assessment should be a priority in your quality evaluation workflow.

Limitations of Base Accuracy Assessment Tools

Understanding the limitations of Merqury and QUAST helps you interpret their results appropriately and avoid overconfidence in your quality metrics.

Merqury Limitations

Merqury's accuracy depends on the quality and completeness of the k-mer database. If the high-accuracy reads do not cover the entire genome, the QV estimate will be biased. Merqury also cannot distinguish between errors in the assembly and errors in the read set, so using reads with even modest error rates will affect the QV estimate. The tool is reference-free, which is an advantage for novel genomes, but it cannot detect structural errors that do not affect k-mer content.

The choice of k-mer size affects the results. Smaller k-mers are more sensitive to errors but also more likely to match repetitive sequences by chance, which can inflate the apparent accuracy. Larger k-mers are more specific but require higher coverage to ensure complete representation. The Bioconductor project provides documentation on k-mer-based analysis methods that can help you choose appropriate parameters.

Merqury also has limitations in highly repetitive genomes. In genomes with large amounts of repetitive sequence, many k-mers will appear at high copy numbers, and the distinction between error k-mers and legitimate k-mers becomes less clear. This can lead to inaccurate QV estimates. For such genomes, you may need to use additional assessment methods beyond Merqury.

QUAST Limitations

QUAST's accuracy depends on the quality and relatedness of the reference genome. If the reference is distant from your sample, the mismatch rate will be inflated by biological divergence, making it difficult to distinguish assembly errors from true differences. QUAST also cannot detect errors in regions that are not present in the reference, such as novel insertions. The tool is reference-based, which limits its utility for non-model organisms without a close reference.

The quality of the reference genome itself is also a concern. If the reference contains errors, QUAST will report those errors as mismatches in your assembly, even if your assembly is correct. This is particularly problematic when the reference was assembled with older technologies that have higher error rates. The NCBI Data Resources provide information about the quality of reference genomes that can help you assess whether a particular reference is suitable for QUAST analysis.

General Limitations

Both tools assess base accuracy at the level of individual bases or short k-mers. They cannot detect errors that occur at larger scales, such as misassembled repeat arrays or incorrect haplotype phasing. For a complete assessment, you should combine base accuracy metrics with structural assessment and, when possible, experimental validation such as PCR or optical mapping.

The Galaxy Training Network emphasizes the importance of reproducibility in bioinformatics analysis. Documenting your assessment parameters and recording the versions of all tools used is essential for interpreting results and for allowing others to reproduce your analysis.

Another general limitation is that both tools provide estimates of accuracy, not direct measurements. The QV from Merqury is an estimate based on k-mer comparisons, and the mismatch rate from QUAST is an estimate based on alignment to a reference. Neither tool can identify every error in an assembly, and the true error rate may differ from the estimated rate. For applications where accuracy is critical, such as clinical genomics, you should consider additional validation methods.

Quality Controls and Professional Escalation Criteria

Knowing when to escalate a quality problem to a more experienced colleague or to revisit your assembly strategy is important for efficient project management. The following criteria can guide your decisions.

When to Polish

If the Merqury QV is below 30, polishing is generally recommended. The AIEdit study demonstrated that polishing can increase QV from 28.7 to 32.9 on ONT data, bringing the assembly above the Q30 threshold. If the QV is between 30 and 35, polishing may still be beneficial for applications that require high accuracy, such as clinical genomics or detailed comparative analysis. If the QV is above 35, polishing is unlikely to provide substantial benefits.

The decision to polish should also consider the cost of polishing. Alignment-based polishers can require days of computation for human-sized genomes, while machine learning-based polishers like AIEdit can complete the task in hours. If your computational resources are limited, you may need to prioritize polishing for the regions of the genome that are most important for your downstream analysis.

When to Reassemble

If polishing does not improve the QV, or if the QV remains below 30 after multiple polishing rounds, the problem may be in the assembly itself instead of in base-level errors. Consider whether the read data are sufficient, whether the assembler parameters are appropriate, and whether a different assembly strategy might produce better results. The Sumatran tiger study demonstrated that correcting errors in ONT long reads during assembly greatly improves the quality and contiguity of the resulting assembly, suggesting that error correction during assembly is more effective than polishing after assembly.

Reassembly is a significant investment of time and computational resources, so you should carefully consider whether it is likely to improve the assembly. If the assembly has a low QV because of insufficient sequencing coverage, additional sequencing may be more effective than reassembly with the same data. If the assembly has structural errors, a different assembler or different assembly parameters may be more effective.

When to Seek Additional Data

If the Merqury completeness metric is below 90 percent, the assembly is likely missing substantial portions of the genome. This may require additional sequencing, particularly with long reads that can span repetitive regions. The schizothoracine fish genome project used a combination of Illumina short reads, PacBio HiFi long reads, and Hi-C technologies to achieve a chromosome-level assembly, demonstrating the value of multiple data types for complete genome assembly.

The decision to seek additional data should be based on the completeness metric and the k-mer spectrum plot. If the k-mer spectrum plot shows a truncated peak, the sequencing coverage is insufficient, and additional sequencing is needed. If the plot shows a complete peak but the completeness metric is still low, the missing sequence may be in regions that are difficult to assemble, such as centromeres or telomeres, and additional sequencing may not help.

When to Consult a Specialist

If you encounter unexpected patterns in your quality metrics, such as a high QV with many QUAST misassemblies, or if you are unsure whether observed mismatches reflect biological divergence or assembly error, consult a colleague with experience in genome assembly and assessment. The The Carpentries Lessons provide foundational training that can help you build the skills to diagnose these problems independently, but specialist consultation is appropriate when the stakes are high, such as for clinical applications or for genomes that will serve as community references.

Specialist consultation is also appropriate when you are working with a genome type that you have not encountered before, such as a highly repetitive genome or a polyploid genome. These genomes present unique challenges for assembly and assessment, and an experienced colleague can help you avoid common pitfalls.

Safety and Regulatory Context for Genome Quality Reporting

While genome assembly is not subject to the same regulatory oversight as clinical diagnostics, the quality of genome assemblies has implications for downstream applications that may be regulated. If your assembly will be used for clinical genomics, agricultural breeding decisions, or conservation management, the base accuracy requirements may be higher than for basic research.

The NCBI Data Resources provide guidelines for submitting genome assemblies to public databases, including requirements for assembly quality metrics. Many journals now require that genome assembly papers report QV values and completeness metrics, and some require that the raw data and assembly files be deposited in public archives. The EMBL-EBI Training resources cover data submission standards and quality reporting requirements.

For agricultural applications, such as livestock or crop genome assemblies used for breeding decisions, base accuracy directly affects the reliability of marker-assisted selection and genomic prediction. An assembly with a QV below 30 may contain enough errors to produce false variant calls that could mislead breeding decisions. The nf-core Documentation describes community standards for reproducible analysis pipelines that can help ensure your assembly and assessment workflows meet quality standards.

The regulatory context for genome assemblies is evolving. As genome assemblies become more widely used in clinical and agricultural applications, the standards for quality reporting are likely to become more stringent. Staying informed about these standards and documenting your quality assessment procedures is important for ensuring that your assemblies meet the requirements of downstream applications.

Frequently Asked Questions

What is the difference between Merqury QV and QUAST mismatch rate?

Merqury QV is a reference-free estimate of base accuracy derived from k-mer comparisons between the assembly and high-accuracy reads. QUAST mismatch rate is a reference-based measure that counts base differences between the assembly and a reference genome. Merqury QV reflects only assembly errors, while QUAST mismatch rate includes both assembly errors and true biological differences between your sample and the reference. For closely related samples, the two metrics should be broadly consistent, but for divergent samples, QUAST will show higher mismatch rates that do not necessarily indicate assembly problems.

How much sequencing coverage do I need for Merqury?

Merqury requires sufficient coverage in the high-accuracy read set to represent the entire genome. For Illumina short reads, at least 30-fold coverage is typically recommended, though the exact requirement depends on genome complexity and k-mer size. You can assess whether coverage is sufficient by examining the k-mer spectrum plot for a complete peak at the expected coverage level. If the peak is truncated or if many k-mers appear at low copy numbers, additional sequencing may be needed.

Can I use the same long reads for Merqury that I used for assembly?

Using the same long reads for Merqury that were used for assembly is not recommended because those reads carry the same errors present in the assembly. The k-mer database would contain error k-mers, and Merqury would underestimate the error rate, producing a falsely high QV. Always use high-accuracy reads, such as Illumina short reads or PacBio HiFi reads, for the k-mer database.

What QV value should I aim for?

A QV of 30, corresponding to 99.9 percent accuracy, is generally considered the minimum threshold for a high-quality assembly. Many genome projects aim for QV values above 35, and the GoldPolish-Target study reported achieving base accuracy values above 99.9 percent, corresponding to QV above 30. For clinical applications or for genomes that will serve as community references, higher QV values may be required. The appropriate threshold depends on your downstream applications and the standards of your field.

How do I know if polishing improved my assembly?

Run Merqury before and after polishing using the same k-mer database and parameters. Compare the QV values. The AIEdit study provides an example where polishing increased QV from 28.7 to 32.9. An increase of 2 to 5 QV points is typical for effective polishing. If the QV does not improve, investigate whether the polishing tool is appropriate for your data type or whether the errors are concentrated in regions that the polisher cannot correct.

What should I do if my assembly has a high QV but many QUAST misassemblies?

A high QV with many QUAST misassemblies suggests that the assembly has structural errors that do not affect base-level accuracy. These errors may include misjoins, inversions, or relocations that would confound downstream analysis. Investigate the specific misassembly locations and consider whether a different assembly strategy might produce a structurally correct assembly. Polishing will not fix structural errors.

Is QUAST useful for non-model organisms without a reference genome?

QUAST requires a reference genome for comparison, so it is not directly useful for non-model organisms without a close reference. For such organisms, Merqury is the more appropriate tool because it provides reference-free quality estimates. If a reference from a related species is available, QUAST can still provide useful information about structural correctness, but the mismatch rate will be inflated by biological divergence.

How should I report assembly quality metrics in my publication?

Report the Merqury QV and completeness, the QUAST metrics including N50, misassembly count, and mismatch and indel rates, and the versions of all tools used. Describe the k-mer size and read set used for Merqury and the reference genome version used for QUAST. Many journals require that assembly quality metrics be reported in a standardized format, and the NCBI Data Resources provide guidance on assembly submission and quality reporting.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.