Evaluating Assembly Quality with k-mer Spectra: A Practical Guide to Merqury and Similar Tools

By Dr. Zubair Khalid, DVM, MS, PhD ·

Evaluating Assembly Quality with k-mer Spectra: A Practical Guide to Merqury and Similar Tools

Key Takeaways

  • K-mer spectrum analysis, as implemented by Merqury, provides a reference-free method for evaluating genome assembly quality by comparing k-mers from sequencing reads to those in the assembly. This approach quantifies base-level accuracy (QV), completeness, and detects issues like haplotype collapse or duplication, which contiguity statistics (e.g., N50) alone cannot reveal.
  • High-accuracy sequencing reads, such as PacBio HiFi or Illumina, are crucial for building a reliable k-mer database; lower-accuracy reads introduce noise that inflates the error shoulder and depresses completeness estimates. Basic read quality control, including adapter trimming and low-quality base removal, is a prerequisite for accurate k-mer database construction.
  • Merqury's QV (Quality Value) metric, a Phred-scaled probability of base correctness, is a primary indicator of base-level accuracy, with QV 40 signifying 1 error per 10,000 bases and QV 50 indicating 1 error per 100,000 bases; values in the mid-30s to mid-40s are typical for high-quality assemblies. Completeness, representing the fraction of read k-mers found in the assembly, should ideally be 95% or higher, with lower values suggesting unassembled regions or collapsed repeats.
  • Abnormalities in the k-mer copy number spectrum, such as an excess of single-copy k-mers in diploid genomes, strongly suggest haplotype collapse, while an excess at twice the expected coverage indicates regional duplication, both requiring reassembly rather than polishing. Contamination is identified by a secondary k-mer peak at lower coverage, distinct from the main genomic peak.
  • A structured decision framework is essential for assembly acceptance, defining pre-set thresholds for QV and completeness based on downstream application needs before evaluation begins. Assemblies failing these thresholds are categorized into tiers: tier one for acceptance, tier two for polishing or targeted fixes, and tier three for fundamental structural issues necessitating reassembly.

Genome assembly quality assessment requires metrics that go beyond contiguity statistics. Researchers who generate de novo assemblies need to determine whether the sequence is complete, accurate, and free from haplotype duplication. K-mer spectrum analysis, implemented in tools like Merqury, provides a reference-free approach by comparing the k-mers in unassembled sequencing reads against those found in the assembly. This guide explains how k-mer spectra work, how to run Merqury, how to interpret its output, and what to do when quality metrics indicate problems.

The Problem with Contiguity-Only Quality Assessment

Assembly projects often report N50 values as the primary quality metric. A high N50 suggests that half of the assembled sequence resides in contigs or scaffolds of at least that length. While contiguity matters for downstream analysis, it does not indicate whether the sequence itself is correct. A genome assembly can have excellent N50 values while containing base-level errors, collapsed repeats, or misassembled regions.

Base-level accuracy and completeness require different measurement approaches. For organisms with a high-quality reference genome, you can align your assembly to that reference and count mismatches. However, many research projects focus on non-model organisms where no close reference exists. Recent long-read assemblies frequently exceed the quality and completeness of available reference genomes, which makes reference-based validation inadequate or misleading. K-mer-based methods solve this problem by using the sequencing reads themselves as the ground truth.

The core insight is simple. If you sequence a genome to sufficient depth, the k-mers present in the reads reflect the true genomic sequence. A high-quality assembly should contain nearly all of those k-mers, and it should contain them at the expected copy number. K-mer spectrum analysis quantifies how well the assembly matches the read data without needing any external reference.

What K-mer Spectra Measure

A k-mer is a substring of length k. For a genome of size G, the number of distinct k-mers is approximately G minus k plus one, assuming no sequencing errors. When you sequence that genome to depth d, each k-mer appears roughly d times in the read set. Sequencing errors create additional k-mers that appear at low frequency, typically once or twice, because errors are random and rarely recur at the same position.

The k-mer spectrum of the reads therefore has a characteristic shape. A large peak of k-mers at the expected coverage depth represents the true genomic sequence. A smaller shoulder at low frequencies represents sequencing errors. The k-mer spectrum of an assembly should match the read spectrum for the true genomic k-mers, minus the error k-mers.

Merqury uses efficient k-mer set operations to compare the k-mers in a de novo assembly to those found in unassembled high-accuracy reads. This comparison yields three main quality estimates. First, base-level accuracy, reported as QV, which is the Phred-scaled probability of a base being correct. Second, completeness, which is the fraction of read k-mers present in the assembly. Third, for trio-based projects, haplotype-specific accuracy, completeness, phase block continuity, and switch errors. The tool also generates k-mer spectrum plots for visual inspection.

At a Glance

Quality MetricWhat It MeasuresInterpretation Guidance
QV (base accuracy)Phred-scaled probability of base correctnessQV 40 means 1 error per 10,000 bases, QV 50 means 1 per 100,000 bases
CompletenessFraction of read k-mers found in assemblyValues near 95% or higher indicate most genomic sequence is present
Copy number spectrumExpected versus observed k-mer multiplicityExcess single-copy k-mers suggest haplotype collapse, excess multi-copy suggests duplication
K-mer spectrum plotVisual distribution of k-mer frequenciesShould show a clean main peak with minimal low-frequency error shoulder

Preparing Input Data for Merqury

Merqury requires two inputs. The first is the assembly you want to evaluate, in FASTA format. The second is a k-mer database built from the sequencing reads. The reads should be high accuracy, because the k-mer database serves as the reference truth. PacBio HiFi reads and Illumina short reads both work well. Lower accuracy reads such as traditional PacBio CLR or older Oxford Nanopore data will introduce noise into the k-mer spectrum and reduce the reliability of the quality estimates.

Before building the k-mer database, you should perform basic read quality control. Remove adapter sequences and trim low-quality bases. For Illumina data, check the per-base quality scores and remove reads that fail quality thresholds. For HiFi data, verify that the read quality exceeds the manufacturer recommended threshold. Poor quality input reads will generate spurious k-mers that inflate the error shoulder and depress the completeness estimate.

The choice of k matters. Merqury defaults to k equals 21, which works well for most genomes. Smaller k values increase sensitivity but also increase the chance of k-mers occurring by chance in repetitive sequence. Larger k values reduce random matches but require higher sequencing depth to observe every k-mer. For genomes with very high repeat content, you may need to test multiple k values and compare the resulting spectra.

Building the K-mer Database

Merqury uses Meryl to build the k-mer database from reads. Meryl counts k-mers and stores them in a compact database format. The counting step is computationally intensive but scales well across multiple threads. For a typical mammalian genome sequenced to 30-fold coverage, the k-mer database can take several hours to build and require tens of gigabytes of disk space.

The command structure follows a standard pattern. You first create a database from the reads, then you can query that database against the assembly. Meryl supports both exact counts and presence-absence queries. Merqury handles the database construction internally when you provide it with the read files, or you can pre-build the database and pass it directly.

For large projects, consider building the k-mer database once and reusing it across multiple assembly versions. This approach saves time during iterative assembly and polishing cycles. Store the database on fast storage because repeated queries against a large database can become input-output bound.

Running Merqury

Merqury takes the assembly FASTA and the k-mer database as inputs. The basic command specifies the database, the assembly, and an output prefix. The tool generates several output files including the QV estimate, completeness statistics, and the k-mer spectrum plot.

The runtime depends on genome size, assembly contiguity, and the size of the k-mer database. For a 1 gigabase assembly, Merqury typically completes in under an hour on a standard server. Larger genomes such as the 4.74 gigabase land hermit crab genome require more time and memory. The tool is designed to be fast and robust for assembly validation, making it practical to run after every assembly iteration.

Merqury also supports trio analysis. When you have sequencing data from both parents and the offspring, you can build haplotype-specific k-mer databases and evaluate phasing accuracy. This mode reports phase block continuity and switch errors, which matter for projects that aim to produce haplotype-resolved assemblies.

Interpreting QV Values

The QV value is the most commonly cited Merqury output. It represents the Phred-scaled base accuracy of the assembly. A QV of 40 corresponds to one error per 10,000 bases, which means a 1 gigabase assembly would have approximately 100,000 errors. A QV of 50 corresponds to one error per 100,000 bases.

Published reference-quality assemblies report a range of QV values. The telomere-to-telomere assembly of the critically endangered tree Magnolia longipedunculata achieved a Merqury QV of 41.70. The yellow-throated marten genome reached a QV of 43.75. The land hermit crab genome, which is larger and more repetitive, reported a QV of 35.88. These examples show that QV values in the mid-30s to mid-40s are typical for high-quality assemblies, with higher values indicating better base-level accuracy.

When you see a QV below 30, the assembly likely has systematic errors that need correction. Common causes include incomplete polishing, misassembled regions, or contamination. When you see a QV above 50, the assembly is approaching the accuracy limit of the sequencing technology itself. Further polishing may not yield meaningful improvements.

Interpreting Completeness Statistics

Completeness in Merqury is the fraction of read k-mers that appear in the assembly. A completeness value of 95% means that 95% of the k-mers observed in the reads are present in the assembly. The missing 5% could represent unassembled genomic regions, collapsed repeats, or sequence that was incorrectly removed during assembly.

The yellow-throated marten genome was reported as 94.95% complete by Merqury evaluation. This value is typical for chromosome-level assemblies that do not achieve telomere-to-telomere status. The missing sequence often corresponds to highly repetitive regions such as centromeres, ribosomal DNA arrays, and segmental duplications that resist assembly with current technologies.

Completeness values below 90% warrant investigation. Check whether the missing k-mers cluster in specific genomic regions or are distributed uniformly. If they cluster, the assembly may have collapsed a large repeat array or missed an entire chromosome arm. If they are distributed, the problem may be insufficient sequencing depth or a systematic issue with the assembler.

The K-mer Spectrum Plot

The k-mer spectrum plot is a visual diagnostic that complements the numeric metrics. The plot shows the number of distinct k-mers at each coverage depth for both the reads and the assembly. A healthy assembly produces a spectrum that closely tracks the read spectrum at the main peak.

The main peak position indicates the average k-mer coverage. The width of the peak reflects coverage variation across the genome. A narrow peak indicates uniform coverage, while a broad peak suggests variable coverage that may complicate assembly. The low-frequency shoulder represents sequencing errors and should be minimal in the assembly spectrum.

Deviations from the expected spectrum shape indicate specific problems. An excess of k-mers at half the expected coverage suggests haplotype collapse in a diploid genome. An excess at twice the expected coverage suggests duplication of genomic regions. A secondary peak at lower coverage can indicate contamination from a second organism.

Haplotype Duplication and Collapse

Diploid genomes present a special challenge for assembly quality assessment. True haplotypes differ from each other by polymorphic sites. When an assembler correctly phases the two haplotypes, each haplotype-specific k-mer appears at half the coverage of the homozygous k-mers. When the assembler collapses the haplotypes into a single sequence, the haplotype-specific k-mers are missing entirely.

Merqury detects these patterns through the copy number spectrum. The expected copy number for each k-mer is calculated from the read depth. The observed copy number in the assembly is compared against this expectation. Excess single-copy k-mers in a diploid assembly indicate that the assembler collapsed divergent haplotypes. Excess multi-copy k-mers indicate that the assembler duplicated regions that should be single copy.

For haplotype-resolved projects, Merqury provides additional metrics. When parental data are available, the tool can assign k-mers to each haplotype and evaluate whether the assembly correctly separates them. Phase block continuity measures how long the assembly maintains a single haplotype without switching. Switch errors count the positions where the assembly incorrectly switches from one haplotype to the other.

Practical Workflow for Assembly Evaluation

Start with a clean assembly FASTA. Remove contigs that are shorter than a minimum length threshold, typically 1 kilobase, because short contigs often contain errors and contribute little to downstream analysis. Remove contigs that map to common contaminants such as bacterial or human sequences if you are working with a non-human sample.

Build the k-mer database from the highest accuracy reads available. For most projects, this means PacBio HiFi reads or Illumina short reads. If you have both, build separate databases and compare the results. Discrepancies between the two evaluations can reveal technology-specific artifacts.

Run Merqury and record the QV, completeness, and copy number statistics. Generate the k-mer spectrum plot and inspect it visually. Compare your metrics against published assemblies of similar genome size and complexity. The Magnolia, wasabi, hermit crab, and marten genomes provide useful reference points across a range of genome sizes and repeat contents.

If the QV is below 30, run additional polishing rounds. If completeness is below 90%, investigate whether specific regions are missing. If the copy number spectrum shows haplotype collapse, consider whether the assembly strategy needs to change. Document all metrics in your project records so you can track improvements across assembly iterations.

Records and Measurements to Keep

Maintain a quality assessment log for every assembly version. Record the assembly date, the assembler and version, the input read data and coverage, the k-mer database parameters, and all Merqury output metrics. This log allows you to compare iterations and identify which changes improved or degraded quality.

Record the exact commands used to generate each metric. Reproducibility requires that another researcher can repeat your evaluation and obtain the same results. The Galaxy Training Network provides accessible workflow training that emphasizes reproducible analysis practices. Following these practices ensures that your quality assessment is transparent and verifiable.

Store the k-mer spectrum plots alongside the numeric metrics. The plots capture information that summary statistics miss, such as the shape of the coverage distribution and the presence of secondary peaks. If a downstream analysis produces unexpected results, the spectrum plot can help you determine whether the assembly or the analysis is at fault.

Common Failure Patterns and Troubleshooting

Low QV with normal completeness usually indicates base-level errors that polishing should correct. Run additional rounds of polishing and re-evaluate. If QV does not improve, check whether the polishing tool is using the correct read data. Polishing with reads that have systematic errors will not improve accuracy.

High completeness with low QV can indicate that the assembly contains the right sequence but with many small errors. This pattern is common after initial assembly before polishing. The fix is straightforward: polish with high-accuracy reads and re-run Merqury.

Low completeness with normal QV suggests that the assembly is missing genomic regions. Check the assembly size against the expected genome size. If the assembly is substantially smaller, the assembler may have failed to capture repetitive regions. Consider whether additional sequencing depth or a different assembler would help.

Abnormal copy number spectra indicate structural problems. If the spectrum shows a large excess of single-copy k-mers in a diploid genome, the haplotypes were collapsed. If it shows excess multi-copy k-mers, regions were duplicated. Both problems require reassembly instead of polishing.

Contamination produces a distinctive pattern. The k-mer spectrum shows a secondary peak at lower coverage, corresponding to the contaminant genome. Check the assembly for contigs with atypical GC content or coverage. Remove contaminant contigs and re-evaluate.

Limitations of K-mer-Based Evaluation

K-mer methods cannot detect every assembly error. Errors that occur in repetitive regions may be invisible because the k-mers in those regions are identical to k-mers elsewhere in the genome. A misassembly that joins two copies of a repeat array may not change the k-mer spectrum if the repeat copies are identical.

K-mer methods require sufficient sequencing depth. If the coverage is too low, some true k-mers will be missing from the read set, and the completeness estimate will be artificially low. The minimum depth depends on the genome size and the k-mer length, but 30-fold coverage of high-accuracy reads is a reasonable target for most projects.

K-mer methods are reference-free, which is their main advantage, but this also means they cannot detect errors that are present in the reads themselves. If the sequencing technology introduces systematic errors, the k-mer spectrum will reflect those errors as if they were true sequence. Using high-accuracy reads minimizes this problem.

The QV estimate from Merqury is a global average. It does not identify which specific bases are wrong or where errors are concentrated. For error correction, you need additional tools that align reads to the assembly and identify discrepancies at specific positions.

Comparison with Other Quality Assessment Tools

Merqury is one of several tools for assembly quality assessment. BUSCO evaluates completeness by searching for a set of conserved single-copy orthologs. The Magnolia assembly reported a BUSCO completeness of 96.6% alongside a Merqury QV of 41.70. The wasabi genome was validated with three methods: BUSCO, Merqury, and Inspector. The hermit crab genome reported both BUSCO and Merqury metrics.

BUSCO and Merqury measure different aspects of quality. BUSCO measures whether the assembly contains a set of genes that are expected to be present and single copy. Merqury measures whether the assembly matches the sequencing reads. A genome can score well on one and poorly on the other. For example, an assembly that contains all BUSCO genes but has many base errors will score well on BUSCO and poorly on Merqury QV.

Inspector is another reference-free tool that evaluates assembly quality using read alignments. It provides complementary information to Merqury, including error positions and misassembly breakpoints. Using multiple tools gives a more complete picture than relying on any single metric.

The choice of tools depends on your project goals. If you need a single quality metric for a genome report, Merqury QV is widely accepted. If you need to identify specific errors for correction, Inspector or similar alignment-based tools are more useful. If you need to verify gene content, BUSCO is essential.

Quality Controls and Reproducibility

Reproducible quality assessment requires fixed parameters and documented procedures. The nf-core documentation describes community standards for reproducible workflow configuration. Following these standards ensures that your quality assessment can be repeated by others and compared across projects.

Version control your analysis scripts and record the versions of all software tools. Merqury, Meryl, and the associated plotting tools are under active development, and output formats can change between versions. A QV of 40 from one version may not be directly comparable to a QV of 40 from another version.

Use the same k-mer length for all assemblies you want to compare. Different k values produce different spectra and different QV estimates. If you change k, you must re-run the evaluation for all assemblies in the comparison.

Document the read data used for the k-mer database. The database should be built from the same reads used for assembly, or from a higher quality subset. If you use different reads for assembly and evaluation, the completeness estimate will be lower because the evaluation reads contain k-mers that the assembly reads did not cover.

Professional Escalation Criteria

Know when to seek help. If your Merqury QV remains below 30 after multiple polishing rounds, the problem may be in the assembly strategy instead of the polishing. Consult with colleagues who have experience with your sequencing platform and genome type.

If completeness is below 90% and the missing sequence appears to be concentrated in specific regions, you may need additional sequencing. Targeted approaches such as ultra-long reads or Hi-C can help resolve regions that standard sequencing cannot assemble.

If the copy number spectrum shows severe haplotype collapse, consider whether your assembly goal requires haplotype resolution. For many projects, a collapsed assembly is acceptable. For projects that need haplotype-resolved sequences, you may need to use trio data or a different assembler.

If you are preparing a genome report for publication, follow the reporting standards used in your field. The published genomes cited in this guide report Merqury QV and completeness alongside BUSCO and other metrics. Reviewers expect these values, and you should be prepared to explain any values that fall outside the typical range for your genome type.

Safety and Data Management Context

Genome assembly projects generate large data volumes. The raw reads, the k-mer databases, and the assembly files can consume terabytes of storage. Plan your storage needs before starting the project, and archive intermediate files according to your institution's data management policies.

The NCBI provides data resources for storing and sharing genome assemblies and sequencing reads. Depositing your data in public repositories ensures that your quality assessment can be verified by others and that your assembly contributes to the broader research community. The EMBL-EBI training materials describe data submission procedures and best practices for data management.

K-mer databases are derived from sequencing reads and may contain sensitive information. If your project involves human data, follow the relevant ethical and legal requirements for data handling. For non-human projects, check whether any export controls or material transfer agreements apply to your samples.

Building a Decision Framework for Assembly Acceptance and Iteration

K-mer quality metrics only become useful when you translate them into concrete decisions about whether to accept an assembly, continue polishing, or restart with a different strategy. Many research groups run Merqury, record the QV and completeness values, and then struggle to determine what those numbers mean for their specific project timeline and budget. A structured decision framework converts the raw metrics into actionable choices at each stage of the assembly process.

Defining Acceptance Thresholds Before You Start

The most common mistake in assembly quality assessment is deciding what constitutes acceptable quality after seeing the results. This approach invites bias, where you rationalize whatever values the assembly produces. Instead, define your acceptance thresholds before the first Merqury run, based on the downstream application requirements and the biological characteristics of your organism.

Start by listing the analyses that will use the assembly. A genome intended for gene annotation and comparative genomics has different quality requirements than one intended for identifying structural variants or for population genetics studies. For gene-focused projects, BUSCO completeness often matters more than QV, because gene presence and integrity drive the downstream results. For variant discovery, QV becomes critical because base errors masquerade as true polymorphisms. For repeat biology, the copy number spectrum matters most because collapsed or expanded repeats distort repeat family sizes.

Write your thresholds as explicit numeric criteria. For example, you might require a minimum QV of 40, a minimum completeness of 95%, and no evidence of large-scale haplotype collapse. These thresholds should be recorded in your project notebook or laboratory information management system before you run the evaluation. When the Merqury results arrive, you compare them against the pre-registered thresholds instead of deciding after the fact whether the values are acceptable.

The published genomes cited in this guide provide useful reference ranges. The Magnolia assembly achieved QV 41.70 with 96.6% BUSCO completeness. The yellow-throated marten reached QV 43.75 with 94.95% Merqury completeness. The hermit crab genome, which is substantially larger and more repetitive, reported QV 35.88. These values suggest that QV in the mid-30s to mid-40s represents the achievable range for chromosome-level assemblies, with genome size and repeat content influencing where your assembly will fall.

The Three-Tier Decision Matrix

Organize your decision process into three tiers based on the severity of the quality problems. Tier one includes assemblies that meet all acceptance thresholds and can proceed to downstream analysis. Tier two includes assemblies that fail one or more thresholds but show patterns that polishing or targeted fixes can address. Tier three includes assemblies with fundamental structural problems that require reassembly.

For tier one decisions, verify that the assembly meets every pre-registered threshold, then document the evidence and proceed. Keep the Merqury output files, the spectrum plots, and the exact commands used to generate them. This documentation becomes part of the assembly record and supports the claims you make in publications or data releases.

For tier two decisions, identify which specific metric failed and what the failure pattern indicates. A QV below threshold with normal completeness suggests base-level errors that additional polishing rounds can correct. Low completeness with normal QV suggests missing sequence that may require targeted gap filling or additional sequencing. Abnormal copy number spectra indicating haplotype collapse require more careful consideration, because polishing cannot fix collapsed haplotypes.

For tier three decisions, the assembly has structural problems that polishing cannot resolve. Severe haplotype collapse across large genomic regions, extensive misassembly indicated by fragmented spectra, or completeness below 85% with no clear explanation all fall into this category. Reassembly with different parameters, additional data, or a different assembler is the appropriate response. Continuing to polish a structurally flawed assembly wastes computational resources and produces a polished but still incorrect result.

Creating a Quality Scorecard for Each Assembly Version

A quality scorecard provides a standardized format for recording and comparing assembly versions. Create a table with rows for each quality metric and columns for each assembly version. Include the assembly date, the assembler and version, the input data and coverage, the k-mer database parameters, and the Merqury output values. This scorecard becomes the central record for tracking quality across the iterative assembly process.

The scorecard should include both the numeric metrics and the qualitative observations from the spectrum plots. Record whether the main peak is clean and narrow, whether the error shoulder is minimal, and whether any secondary peaks appear. These visual observations often reveal problems that the summary statistics miss. A secondary peak at half the main peak coverage in a diploid genome indicates haplotype structure that the numeric completeness value does not capture.

Update the scorecard after every assembly iteration, even when the changes are minor. The comparison across versions reveals which assembly parameters improved quality and which had no effect or made things worse. This information guides future assembly decisions and prevents you from repeating unsuccessful strategies. The Galaxy Training Network provides accessible workflow training that emphasizes reproducible analysis practices, which supports consistent scorecard documentation across projects and team members.

The Polishing Decision Loop

When the QV falls below your threshold but completeness is acceptable, enter the polishing decision loop. This loop has four steps: identify the polishing target, run the polishing tool, re-evaluate with Merqury, and compare the results against the previous version. Each loop iteration should produce a measurable improvement in QV, typically 1 to 5 points per round depending on the starting quality and the error type.

Before polishing, verify that the polishing tool is using the correct read data. Polishing with reads that have systematic errors will not improve accuracy and may introduce new errors. For PacBio HiFi assemblies, polish with the same HiFi reads used for assembly. For Oxford Nanopore assemblies, use the highest quality reads available, ideally after basecalling with the latest models. The wasabi genome project used PacBio CLR reads for assembly and validated with multiple methods including Merqury, demonstrating that even older long-read technologies can produce acceptable assemblies when combined with appropriate polishing.

After each polishing round, re-run Merqury with the same k-mer database and parameters. Record the new QV and completeness values in the scorecard. If the QV improves by less than 1 point after two consecutive polishing rounds, further polishing is unlikely to yield meaningful gains. The assembly has reached the accuracy limit of the current data and strategy. At this point, decide whether the achieved QV meets the project requirements or whether you need to generate additional sequencing data.

When to Stop Polishing and Accept the Assembly

Knowing when to stop is as important as knowing when to continue. The polishing decision loop should have a defined stopping rule based on your pre-registered thresholds and the observed improvement trajectory. A common stopping rule is to continue polishing until the QV meets the acceptance threshold or until two consecutive polishing rounds produce less than 1 point of improvement, whichever comes first.

The accuracy limit of the sequencing technology itself provides a practical ceiling. QV values above 50 approach the error rate of the sequencing platform, and further polishing cannot push the assembly beyond the information content of the reads. The published assemblies in this guide report QV values between 35.88 and 43.75, suggesting that this range represents realistic expectations for most projects. If your assembly reaches QV 45 or higher, you have likely extracted most of the available accuracy from your data.

Accepting an assembly that falls short of your ideal thresholds requires explicit justification. Document why the assembly is acceptable despite the shortfall, what the remaining errors are likely to be, and how those errors might affect downstream analyses. This justification becomes part of the assembly record and demonstrates that you made an informed decision instead of accepting a flawed assembly out of convenience.

Handling the Haplotype Collapse Decision

Haplotype collapse presents a distinct decision point because it cannot be fixed by polishing. When the copy number spectrum shows excess single-copy k-mers in a diploid genome, the assembler has merged divergent haplotypes into a single sequence. This problem requires either accepting the collapsed assembly or restarting with a haplotype-aware assembly strategy.

The decision depends on your project goals. For many applications, a collapsed assembly is acceptable. Gene annotation, repeat analysis, and comparative genomics often work well with a single haploid representation. The Magnolia and marten genomes, both of which report strong quality metrics, do not emphasize haplotype resolution. If your downstream analyses do not require haplotype-specific information, accept the collapsed assembly and document the limitation.

If your project requires haplotype resolution, the decision becomes whether to generate additional data or use a different assembly strategy. Trio-based approaches, where parental data inform haplotype assignment, provide the most reliable path to haplotype-resolved assemblies. Merqury supports trio analysis and can evaluate haplotype-specific accuracy, completeness, phase block continuity, and switch errors when parental data are available. The wasabi genome project achieved haplotype resolution and validated the result with Merqury, demonstrating the feasibility of this approach.

The Reassembly Decision Framework

Reassembly is the most expensive decision in the quality assessment process because it consumes additional computational resources and time. The decision to reassemble should follow a structured evaluation of the failure pattern, the available data, and the likelihood that different parameters will produce a better result.

First, characterize the failure pattern precisely. Low completeness with normal QV suggests missing sequence, which additional sequencing depth or different assembler parameters may address. Abnormal copy number spectra suggesting duplication or collapse indicate structural problems that parameter changes may not fix. Contamination patterns require removing the contaminant sequence and re-evaluating, not reassembling from scratch.

Second, assess whether you have the data needed for a successful reassembly. If the original assembly failed because of insufficient sequencing depth, generating additional reads before reassembling is necessary. If the failure resulted from assembler parameter choices, the same data may produce a better assembly with different settings. The nf-core documentation describes community standards for reproducible workflow configuration, which helps ensure that reassembly attempts use consistent and well-documented parameters.

Third, estimate the probability of improvement before committing resources. If the failure pattern suggests a fundamental limitation of the current data, reassembly will likely produce similar results. If the failure pattern suggests correctable parameter issues, reassembly has a reasonable chance of success. Document your reasoning in the scorecard so that the reassembly decision is transparent and reviewable.

Integrating Merqury Results with Other Validation Tools

The decision framework should incorporate multiple validation tools instead of relying on Merqury alone. Each tool measures different aspects of assembly quality, and the combination provides a more complete picture than any single metric. The wasabi genome project used three validation methods: BUSCO, Merqury, and Inspector. The hermit crab genome reported BUSCO, Merqury QV, and short-read mapping ratio.

BUSCO completeness and Merqury completeness measure different things. BUSCO searches for conserved single-copy orthologs, which represent gene content. Merqury completeness measures the fraction of read k-mers present in the assembly, which represents overall sequence representation. An assembly can score well on one and poorly on the other. A genome with complete gene content but missing repetitive regions will score high on BUSCO and lower on Merqury completeness. A genome with all sequence present but fragmented genes will score lower on BUSCO and higher on Merqury completeness.

Inspector provides error positions and misassembly breakpoints, which Merqury does not report. When Merqury indicates low QV, Inspector can identify where the errors are located and whether they cluster in specific regions. This information guides polishing efforts and helps determine whether the errors are random base substitutions or systematic misassemblies.

The decision framework should define how to reconcile conflicting results from different tools. If BUSCO reports high completeness but Merqury reports low completeness, investigate the discrepancy. The missing k-mers may be in repetitive regions that BUSCO does not survey. If Merqury reports high QV but Inspector identifies misassembly breakpoints, the misassemblies may be in regions where the k-mer spectrum cannot detect them, such as identical repeat copies.

Recording the Decision Rationale

Every acceptance, polishing, or reassembly decision should be recorded with its rationale. The record should include the quality metrics that triggered the decision, the alternatives considered, the reasoning behind the chosen action, and the expected outcome. This documentation serves multiple purposes: it supports publication claims, enables reproducibility, and provides a basis for future decisions when similar situations arise.

The EMBL-EBI training materials describe best practices for data management and documentation in bioinformatics projects. Following these practices ensures that your decision records are complete, organized, and accessible to collaborators and reviewers. The Carpentries lessons provide foundational training in data organization and project management that supports consistent documentation habits.

Store the decision records alongside the quality scorecard and the Merqury output files. Use a consistent naming convention that links each decision to the assembly version it concerns. This linkage allows anyone reviewing the project to trace the quality assessment history and understand why the final assembly was accepted.

Professional Escalation Criteria for Quality Decisions

Some quality problems exceed the scope of routine decision frameworks and require consultation with specialists. Define escalation criteria before problems arise so that you recognize when to seek help. Escalate when the QV remains below 30 after multiple polishing rounds, when completeness is below 85% with no clear explanation, or when the copy number spectrum shows patterns you cannot interpret.

Consult colleagues who have experience with your sequencing platform and genome type. Genome assembly challenges vary substantially across taxa. The hermit crab genome, at 4.74 gigabases with high repeat content, presented different challenges than the 2.16 gigabase Magnolia genome. Colleagues who have assembled similar genomes can provide practical advice that general documentation cannot.

For persistent quality problems, consider whether the issue lies in the sequencing data instead of the assembly. Low-quality reads produce low-quality assemblies regardless of assembler parameters. If the read data have systematic errors, generating new sequencing data may be more effective than continued assembly attempts. The NCBI provides resources for evaluating read quality and comparing your data against publicly available datasets from similar projects.

Escalate when the quality assessment reveals problems that affect the scientific conclusions you plan to draw. If the assembly will be used for conservation genetics, as with the Magnolia genome, quality problems that affect variant calling or population analyses require resolution before proceeding. The cost of resolving quality issues after downstream analyses have begun is substantially higher than addressing them during the assembly phase.

Frequently Asked Questions

What is the difference between Merqury QV and BUSCO completeness?

Merqury QV measures base-level accuracy by comparing k-mers in the assembly against k-mers in the sequencing reads. BUSCO completeness measures whether the assembly contains a set of conserved single-copy genes. The two metrics are complementary. An assembly can have high BUSCO completeness but low QV if it contains the right genes with many base errors. It can have high QV but low BUSCO completeness if it is accurate but missing gene-containing regions.

How much sequencing depth do I need for reliable Merqury evaluation?

The required depth depends on the genome size and the k-mer length. Higher depth gives more reliable k-mer counts and reduces the chance of missing true k-mers. For most projects, 30-fold coverage of high-accuracy reads is a reasonable target. If your completeness estimate is lower than expected, check whether the sequencing depth is sufficient before assuming the assembly is incomplete.

Why does my k-mer spectrum show a shoulder at low coverage?

The low-coverage shoulder in the read spectrum represents sequencing errors. Each error creates k-mers that appear only once or twice in the read set. The assembly spectrum should not show this shoulder because the assembly is a single consensus sequence. If the assembly spectrum shows a low-coverage shoulder, the assembly may contain errors that were not corrected during polishing.

What does a QV of 40 mean in practical terms?

A QV of 40 corresponds to one error per 10,000 bases. For a 1 gigabase assembly, this means approximately 100,000 errors total. Whether this is acceptable depends on your downstream analysis. For gene annotation and comparative genomics, QV 40 is generally sufficient. For applications that require single-base accuracy, such as identifying disease-causing variants, you may need higher QV.

Can I use Merqury with Oxford Nanopore reads?

Merqury requires high-accuracy reads for the k-mer database because the reads serve as the reference truth. Traditional Oxford Nanopore reads have lower accuracy and will introduce noise into the k-mer spectrum. If you have Oxford Nanopore reads, consider polishing them or using a different evaluation approach. The yellow-throated marten genome used Oxford Nanopore for assembly but the quality evaluation still benefited from high-accuracy read data.

How do I know if my assembly has haplotype collapse?

Haplotype collapse produces a characteristic pattern in the copy number spectrum. In a diploid genome, true haplotypes differ at polymorphic sites. If the assembler collapses the haplotypes, the haplotype-specific k-mers appear at half the expected coverage or are missing entirely. Merqury reports copy number statistics that reveal this pattern. If you have parental data, the trio mode can provide a more detailed assessment.

What should I do if my QV does not improve after polishing?

If QV remains below 30 after multiple polishing rounds, check whether the polishing tool is using the correct reads. Polishing with reads that have systematic errors will not improve accuracy. Also check whether the assembly has structural problems such as misjoins or collapsed repeats. Polishing corrects base errors but cannot fix structural errors. You may need to reassemble with different parameters or a different assembler.

How do I compare quality metrics across different assembly projects?

Compare metrics only when they were generated with the same tools and parameters. Use the same k-mer length, the same read data type, and the same software versions. The published genomes cited in this guide report Merqury QV values ranging from 35.88 to 43.75. These values provide a useful reference range, but direct comparisons are only valid when the evaluation methods match.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.