N50, L50, and Beyond: A Field Guide to Contiguity Metrics for Genome Assemblies
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- N50 and L50 are length-weighted summary statistics describing the contiguity of genome assemblies, not their correctness or completeness. N50 is the length of the shortest sequence among the largest sequences accounting for at least 50% of the total assembly length, while L50 is the number of such sequences.
- Contiguity metrics are crucial for downstream analyses such as gene annotation, comparative genomics, and variant calling, as fragmented assemblies hinder gene model resolution and accurate alignment.
- NG50, calculated against the estimated genome size rather than the assembly length, offers a more robust comparison between assemblies of the same species, mitigating inflation from contamination or incomplete haplotype separation.
- Scaffold N50 should be reported separately from contig N50 to reflect the impact of scaffolding on ordering and orienting contigs, with a large difference indicating successful pseudomolecule construction.
- BUSCO scores, assessing the presence of conserved single-copy orthologs, provide a complementary measure of gene-space completeness and should be reported alongside contiguity metrics for a comprehensive quality assessment.
- Misassemblies, chimeric joins, and base-level errors are not detected by N50 or L50; these require orthogonal validation methods like read mapping consistency checks or Hi-C data analysis.
Genome assembly contiguity metrics such as N50 and L50 are summary statistics that describe how much of an assembly is contained within its largest sequences. These metrics appear in genome papers, database submissions, and pipeline outputs, yet they are frequently misinterpreted or treated as standalone measures of assembly quality. This article explains the mathematical definitions of N50 and L50, their relationship to NG50 and complementary metrics such as auN, and how to interpret them when comparing assemblies or reporting quality in your own work. The practical goal is to help you avoid common reporting errors, choose appropriate metrics for your assembly type, and communicate contiguity accurately to reviewers, collaborators, and database curators.
What Contiguity Metrics Actually Measure
Contiguity metrics describe the length distribution of sequences in an assembly, not the correctness of those sequences. A contig is a contiguous stretch of sequence assembled without gaps, and a scaffold is a set of contigs ordered and oriented with gaps of known or estimated size between them. N50 and L50 are calculated from the sorted lengths of these sequences and provide a single number that summarizes how much of the assembly is represented by its largest pieces.
The N50 statistic is defined as the length of the shortest sequence among the set of largest sequences that together account for at least 50 percent of the total assembly length. To calculate N50, you sort all sequences by length from longest to shortest, then add lengths cumulatively until you reach or exceed half of the total assembly length. The length of the sequence that pushes the cumulative sum past this threshold is the N50. The L50 is the number of sequences required to reach that same 50 percent threshold. For example, if an assembly has an N50 of 28.53 Mb and an L50 of 11, this means that 11 of the largest sequences together account for at least half of the total assembly length, and the shortest of those 11 sequences is 28.53 Mb long.
These metrics are length-weighted summaries. They do not tell you whether the sequences are correct, whether they represent true biological chromosomes, or whether any sequence contains misjoins. A highly contiguous assembly can still contain base-level errors, collapsed repeats, or chimeric joins. Conversely, a fragmented assembly can be perfectly accurate at the base level but difficult to use for downstream analyses that require long-range context.
Why N50 and L50 Matter in Practice
Contiguity affects nearly every downstream use of an assembly. Gene annotation benefits from long sequences because genes can span multiple contigs, and fragmented assemblies make it harder to resolve complete gene models. Comparative genomics and chromosome evolution studies require chromosome-level scaffolds to identify syntenic blocks and rearrangements. Population genetics and variant calling depend on accurate alignment, and misassembled or overly fragmented sequences reduce mapping confidence. For species with large genomes or high repeat content, achieving high contiguity is often the primary technical challenge.
The practical value of N50 and L50 is that they allow quick comparison between assemblies produced by different methods or assemblers. When you are deciding whether a new assembly is an improvement over an existing one, N50 and L50 give you a first-pass answer. However, these metrics must be interpreted in the context of genome size, sequencing technology, and assembly strategy. An N50 of 2.2 Mb might be excellent for a highly repetitive 0.91 Gb genome assembled from short reads, while the same N50 would be considered poor for a bacterial genome assembled from long reads.
At a Glance: Contiguity Metrics Decision Table
| Metric | What It Measures | When to Report | Interpretation Caveat |
|---|---|---|---|
| N50 | Length of the sequence at the 50 percent cumulative length threshold | Standard reporting for any assembly | Does not account for genome size, a large N50 in a small assembly may be misleading |
| L50 | Number of sequences needed to reach 50 percent of total assembly length | Report alongside N50 for context | A low L50 is generally good, but it does not indicate correctness |
| NG50 | N50 calculated against the estimated genome size instead of assembly size | Use when genome size is known or estimated | More comparable across assemblies of the same species than N50 |
| auN | Area under the Nx curve, a single value summarizing the entire contiguity distribution | Use for ranking assemblies with different length distributions | Less intuitive than N50 but more informative for overall contiguity |
| Scaffold N50 | N50 calculated on scaffolds instead of contigs | Report separately from contig N50 | Scaffold N50 can be inflated by gap-filling or mis-scaffolding |
Core Principles of Contiguity Interpretation
N50 Is Not a Measure of Completeness
A common error is to interpret N50 as the median sequence length or as a measure of how complete an assembly is. Neither interpretation is correct. The median sequence length is the length of the middle sequence when all sequences are sorted, which is usually much smaller than N50 because most assemblies contain many small sequences. Completeness is better assessed with metrics such as BUSCO, which checks for the presence of conserved single-copy genes, or by comparing the total assembly length to the expected genome size.
The distinction matters when you are reporting assembly quality. An assembly with a high N50 can still be missing large portions of the genome, particularly if the assembler collapsed repetitive regions or failed to assemble heterochromatic sequences. Conversely, an assembly with a modest N50 can be nearly complete if the genome is highly repetitive and the assembler produced many small but accurate contigs.
N50 Depends on Total Assembly Length
Because N50 is calculated against the total assembly length, it is sensitive to the amount of sequence in the assembly. If an assembly contains a large amount of contamination or duplicated sequence, the total length increases, and the N50 can shift even if the underlying sequences are unchanged. This is why NG50, which uses the estimated genome size as the denominator, is often preferred for comparing assemblies of the same species.
For example, consider an assembly of a 1.5 Gbp genome. If the assembly contains 1.7 Gbp of sequence due to contamination or incomplete haplotype separation, the N50 will be calculated against 1.7 Gbp instead of 1.5 Gbp. The N50 value will be higher than it would be if calculated against the true genome size, giving a false impression of contiguity. NG50 corrects for this by using the estimated genome size, so you can compare assemblies even when their total lengths differ.
L50 Complements N50
N50 and L50 are paired statistics. N50 tells you the length threshold, and L50 tells you how many sequences meet that threshold. A low L50 indicates that a small number of sequences account for half the assembly, which is generally desirable. However, L50 alone is not informative without N50. An L50 of 11 could describe an assembly with an N50 of 28.53 Mb or an assembly with an N50 of 1 Mb, and these are very different situations.
When you report contiguity, always include both N50 and L50. Some journals and databases also request the number of sequences, the total assembly length, and the largest sequence length. These additional values help reviewers and database users understand the full length distribution instead of relying on a single summary statistic.
Practical Workflow for Calculating and Reporting Contiguity
Step 1: Obtain or Generate the Assembly
The first step is to have your assembly in a standard format, typically FASTA. Most assemblers produce FASTA files directly, and you can also retrieve assemblies from public databases such as NCBI for comparison purposes. The NCBI Data Resources provide access to assembled genomes, sequence records, and analysis tools that can help you understand how contiguity metrics are reported in official database records.
If you are working with a pipeline that generates assemblies, you should verify that the output FASTA contains only the sequences you intend to analyze. Some pipelines include decoy sequences, adapters, or mitochondrial contigs that should be removed before calculating contiguity metrics. The nf-core Documentation describes community standards for pipeline usage and configuration, and many nf-core pipelines include assembly quality assessment modules that calculate N50 and related metrics automatically.
Step 2: Calculate Contiguity Metrics
Several tools calculate N50, L50, NG50, and related metrics. QUAST is a widely used assembly evaluation tool that reports these statistics along with many other quality measures. The AquaaG pipeline integrates QUAST for assembly quality assessment and can automate the calculation of contiguity metrics as part of a reproducible annotation workflow, as described in the AquaaG pipeline publication. If you are working in R, the Bioconductor project provides packages for genomic analysis that can calculate these metrics from FASTA files or from assembly summary tables.
When you calculate N50, be clear about whether you are using contigs or scaffolds. Contig N50 is calculated on the unjoined contig sequences, while scaffold N50 is calculated on the scaffolded sequences that include gaps. These values can differ substantially. In the Diadema setosum reference genome, the contig N50 was 2.2 Mb while the scaffold N50 was 39.8 Mb, reflecting the successful scaffolding of contigs into chromosome-level pseudomolecules.
Step 3: Compare Against Genome Size
To calculate NG50, you need an estimate of the genome size. This can come from flow cytometry, k-mer analysis, or a closely related species with a known genome size. The NG50 is calculated the same way as N50, but the cumulative sum must reach 50 percent of the estimated genome size instead of 50 percent of the assembly length.
NG50 is particularly useful when you are comparing assemblies of the same species produced by different methods. Because the denominator is fixed, differences in NG50 reflect differences in the assembly length distribution instead of differences in total assembled sequence. If you do not have a reliable genome size estimate, you should report N50 and note that NG50 could not be calculated.
Step 4: Report Additional Context
Contiguity metrics are most useful when reported alongside other assembly statistics. At minimum, include the total assembly length, the number of sequences, the largest sequence length, and the GC content. If you are reporting a chromosome-level assembly, also report the number of chromosome-scale scaffolds and the proportion of the assembly contained within them.
The Pseudobagrus ussuriensis genome assembly provides a good example of thorough reporting. The authors reported a total assembly size of 741.97 Mb, with 26 chromosome-level contigs covering 97.34 percent of the assembly, an N50 of 28.53 Mb, and an L50 of 11. These values together give a clear picture of the assembly structure: most of the genome is contained in a small number of large contigs, and the remaining sequence is in smaller fragments.
Options and Tradeoffs in Assembly Strategies
Short-Read Assemblies
Short-read assemblies, typically from Illumina sequencing, produce highly accurate but fragmented contigs. The N50 for a short-read assembly of a eukaryotic genome is often in the range of tens to hundreds of kilobases, depending on the genome complexity and coverage. Repeats longer than the read length cannot be resolved, and the assembly is broken at repeat boundaries. Short-read assemblies are still useful for bacterial genomes, where the genome is small and repeats are limited, but they are generally insufficient for chromosome-level assembly of complex eukaryotic genomes.
Long-Read Assemblies
Long-read technologies such as PacBio HiFi and Oxford Nanopore produce reads that span repeats and provide long-range context. HiFi reads, which have high accuracy, enable haplotype-mixed assemblies that reconstruct a mosaic of both haplomes without requiring parental data or additional sequencing technologies. The MGA tool demonstrates that near-complete haplotype-mixed assemblies can be generated from HiFi reads alone, substantially outperforming existing haplotype-mixed assemblers. For many species and applications, such assemblies provide most of the benefits of fully phased diploid assemblies at lower cost.
Long-read assemblies typically have much higher N50 values than short-read assemblies. Contig N50 values in the megabase range are common, and scaffolding can produce chromosome-level pseudomolecules. However, long-read assemblies require more sequencing per genome and more computational resources for assembly and polishing.
Hybrid Approaches
Hybrid assembly strategies combine short reads and long reads to leverage the accuracy of short reads and the contiguity of long reads. Short reads are used to correct errors in long reads or to polish the final assembly, while long reads provide the backbone for contig assembly. Hybrid approaches are useful when long-read coverage is limited or when the genome is particularly complex.
The choice of assembly strategy depends on your research question, budget, and available infrastructure. If you need chromosome-level contiguity for comparative genomics, long-read or hybrid approaches are necessary. If you are assembling a small genome or a bacterial isolate, short reads may be sufficient. The Galaxy Training Network provides accessible workflow training for assembly and quality assessment, and The Carpentries Lessons offer foundational computing lessons that can help you manage the computational aspects of assembly workflows.
Observations and Measurements in Real Assemblies
Chromosome-Level Assemblies
Chromosome-level assemblies are the gold standard for eukaryotic genomes. These assemblies place most of the genome into sequences that correspond to individual chromosomes, often with telomere-to-telomere coverage of the euchromatic regions. The scaffold N50 for a chromosome-level assembly is typically close to the size of the largest chromosomes, and the number of chromosome-scale scaffolds approximates the haploid chromosome number.
The Latrodectus genome assemblies illustrate the range of outcomes in chromosome-level projects. The L. katipo genome consisted of 13 scaffolds likely corresponding to chromosomes, containing 90 percent of the total length, plus 1267 short scaffolds containing the remaining 10 percent. The L. hasselti genome consisted of 379 scaffolds with a total length of 1.7 Gbp. Both assemblies had BUSCO scores above 94 percent, indicating high gene-space completeness despite the differences in scaffold number and contiguity.
Contig and Scaffold N50 Differences
The difference between contig N50 and scaffold N50 reflects the success of the scaffolding process. In the Diadema setosum assembly, the contig N50 was 2.2 Mb and the scaffold N50 was 39.8 Mb, a nearly 18-fold improvement. This indicates that the scaffolding step successfully ordered and oriented many contigs into large pseudomolecules, likely using Hi-C or other long-range data.
When you report contiguity, you should report both contig and scaffold N50 values. A large gap between them indicates successful scaffolding, while a small gap suggests that scaffolding did not add much value. If you are comparing assemblies, be aware that some assemblers report only scaffold N50 while others report only contig N50, and these values are not directly comparable.
BUSCO as a Complement to Contiguity
BUSCO (Benchmarking Universal Single-Copy Orthologs) assesses gene-space completeness by searching for conserved single-copy genes that are expected to be present in the genome. A high BUSCO score indicates that most expected genes are present in the assembly, which is a measure of completeness instead of contiguity. The Latrodectus assemblies had BUSCO scores of 94.9 percent and 95.4 percent, and the AquaaG pipeline integrates BUSCO for gene-space completeness evaluation alongside QUAST for assembly quality assessment.
BUSCO and N50 measure different aspects of assembly quality, and both should be reported. An assembly can have a high N50 but a low BUSCO score if it is missing genes, or a high BUSCO score but a low N50 if it is fragmented. Reporting both gives reviewers and database users a more complete picture of assembly quality.
Records and Measurements for Assembly Reporting
What to Record
When you generate an assembly, record the following information for each assembly version:
- Assembly method and version, including the assembler name and parameters
- Sequencing technology and coverage
- Total assembly length
- Number of sequences
- Largest sequence length
- Contig N50 and L50
- Scaffold N50 and L50
- NG50 if genome size is known
- BUSCO completeness score
- Number of chromosome-scale scaffolds and proportion of assembly they represent
This information should be recorded in a structured format, such as a spreadsheet or a YAML configuration file, so that it can be reproduced and compared across assembly versions. The AquaaG pipeline uses YAML configuration files and produces assembly quality reports that include these metrics, providing a reproducible framework for assembly assessment.
How to Record
Use standard tools and formats so that your records are comparable to those in public databases. QUAST produces HTML and text reports that include N50, L50, NG50, and many other metrics. If you are using a pipeline such as AquaaG or an nf-core pipeline, the quality assessment modules will generate these reports automatically.
For long-term record keeping, store the assembly FASTA file, the quality assessment report, and the configuration files used to generate the assembly. This allows you to reproduce the assembly and verify the reported metrics if questions arise during review or publication.
Common Failure Patterns in Reporting
Several common errors appear in assembly reporting. The first is reporting N50 without specifying whether it refers to contigs or scaffolds. The second is reporting N50 without L50, which makes it impossible to understand the length distribution. The third is comparing N50 values across assemblies with different total lengths without accounting for genome size. The fourth is interpreting N50 as a measure of completeness or accuracy.
Another failure pattern is reporting only the best assembly version without documenting the assembly process. Reviewers increasingly expect to see the assembly parameters, the input data, and the quality metrics for the final assembly. Reproducibility requires that you document the output and the steps that produced it. The nf-core Documentation emphasizes reproducibility as a core principle of community pipelines, and the Galaxy Training Network provides tutorials that teach reproducible assembly workflows.
Limitations of Contiguity Metrics
N50 Does Not Detect Misassembly
The most important limitation of N50 and L50 is that they do not detect misassembly. A chimeric sequence that joins two unrelated genomic regions will increase N50 if it is long, but it will also introduce errors into downstream analyses. Misassemblies can be detected by comparing the assembly to a reference genome, by checking for inconsistent read coverage, or by using Hi-C data to verify that the assembly is consistent with the three-dimensional structure of the genome.
N50 Is Sensitive to the Length Distribution
N50 is a single point on the cumulative length distribution, and it does not capture the shape of that distribution. Two assemblies can have the same N50 but very different length distributions. One might have many sequences close to the N50 length, while the other might have a few very long sequences and many very short ones. The auN metric, which is the area under the Nx curve, provides a single value that summarizes the entire distribution and is more robust for ranking assemblies.
N50 Does Not Account for Ploidy or Heterozygosity
For diploid or polyploid genomes, the assembly may contain haplotypes that are not fully separated. A haplotype-mixed assembly represents a mosaic of both haplomes, and the contiguity metrics will reflect the structure of that mosaic. Fully phased assemblies that reconstruct both haplomes separately will have different contiguity metrics than haplotype-mixed assemblies, and comparing them directly can be misleading.
The MGA tool was developed to generate near-complete haplotype-mixed assemblies from HiFi reads alone, and the authors note that such assemblies provide most of the same benefits for downstream analyses as fully phased assemblies. When you report contiguity for a diploid assembly, you should specify whether the assembly is haplotype-mixed or fully phased, because this affects the interpretation of N50 and L50.
Quality Controls in Assembly Projects
Quality Checks Before Reporting
Before you report contiguity metrics, verify that the assembly passes basic quality checks. These include checking for contamination, verifying that the assembly length is consistent with the expected genome size, and confirming that the GC content is appropriate for the species. The NCBI Data Resources provide tools and resources for checking assembly quality, and many genome databases require assemblies to pass certain quality filters before they are accepted.
Escalation Criteria
If your assembly has a low N50 relative to expectations for the species or sequencing technology, you should investigate the cause before proceeding with downstream analyses. Possible causes include insufficient sequencing coverage, high heterozygosity, repetitive genome content, or assembler parameter issues. If the N50 is much lower than expected, consider whether you need additional sequencing, a different assembler, or a different assembly strategy.
If you are comparing your assembly to a reference and find that the N50 is substantially lower, check whether the reference was generated with different technology or methods. A chromosome-level reference generated with Hi-C scaffolding will have a much higher scaffold N50 than a contig-level assembly generated from short reads, and this difference does not necessarily indicate a problem with your assembly.
Professional Escalation
If you are unable to resolve assembly quality issues on your own, consult with colleagues who have experience with your sequencing platform or genome type. Bioinformatics training resources such as the EMBL-EBI Training program and the Galaxy Training Network offer courses and tutorials that can help you troubleshoot assembly problems. The Carpentries Lessons provide foundational computing skills that are useful for managing assembly workflows and data.
Safety and Regulatory Context for Genome Assembly Reporting
Database Submission Requirements
If you plan to submit your assembly to a public database such as NCBI, you must follow the submission guidelines, which include requirements for assembly quality metrics and metadata. The NCBI Data Resources provide detailed documentation on assembly submission, including the required fields and the quality checks that are applied to submitted assemblies. Submitting an assembly with inaccurate or incomplete contiguity metrics can result in delays or rejection.
Publication Standards
Many journals now require authors to report assembly quality metrics in a standardized format. This often includes N50, L50, the number of sequences, the total assembly length, and BUSCO scores. Some journals also require that the assembly be deposited in a public database before publication. Familiarize yourself with the specific requirements of the journal you are targeting, and ensure that your assembly records include all required metrics.
Reproducibility Requirements
Funding agencies and journals increasingly require that assembly workflows be reproducible. This means that you should document the exact commands, parameters, and software versions used to generate the assembly. Pipelines such as nf-core and AquaaG provide built-in reproducibility features, including version tracking and configuration files. The Bioconductor project also emphasizes reproducible genomic analysis through versioned packages and workflow documentation.
A Decision Framework for Choosing and Comparing Contiguity Metrics
Selecting the right contiguity metric for a given assembly project is not a one-size-fits-all decision. The choice depends on the assembly stage, the biological question, the availability of a genome size estimate, and the audience that will interpret the reported values. This section provides a practical decision framework that you can apply when you generate an assembly, compare multiple assemblies, or prepare a report for publication or database submission.
Step 1: Define the Assembly Stage and Sequence Type
Before you calculate any metric, determine whether you are working with contigs, scaffolds, or chromosome-level pseudomolecules. This distinction changes which metric is meaningful and how you should interpret it. Contig N50 reflects the raw output of the assembler before scaffolding, while scaffold N50 reflects the result of ordering and orienting contigs using long-range information such as Hi-C, optical mapping, or linked reads. Chromosome-level assemblies typically report scaffold N50 values that approach the size of the largest chromosomes.
The Diadema setosum reference genome illustrates why this distinction matters. The assembly had a contig N50 of 2.2 Mb and a scaffold N50 of 39.8 Mb, a nearly 18-fold difference. If you report only the scaffold N50, you obscure the underlying contig structure. If you report only the contig N50, you understate the value added by scaffolding. Report both values separately and label them clearly.
For the assembly stage, use this rule: contig N50 is the primary metric during the assembly and polishing phase, and scaffold N50 becomes relevant after scaffolding. If you are comparing your assembly to a published reference, check whether the published N50 refers to contigs or scaffolds before drawing any conclusion about relative quality.
Step 2: Determine Whether a Genome Size Estimate Is Available
The availability of a reliable genome size estimate determines whether you can use NG50 or must rely on N50. NG50 normalizes the cumulative length calculation against the estimated genome size instead of the assembly length, which makes it more comparable across assemblies of the same species. If you have a genome size estimate from flow cytometry, k-mer analysis, or a closely related species, calculate and report NG50 alongside N50.
If you do not have a genome size estimate, report N50 and state explicitly that NG50 could not be calculated. Do not substitute the assembly length for the genome size, because this defeats the purpose of NG50. The assembly length can be inflated by contamination, duplicated sequence, or incomplete haplotype separation, and using it as the denominator will produce an N50 that overstates contiguity.
For cross-species comparisons, NG50 is more appropriate than N50 because it controls for differences in genome size. However, even NG50 does not account for differences in repeat content, heterozygosity, or assembly strategy. Use NG50 as a first-pass comparison and then examine the full length distribution and completeness metrics before making a final judgment.
Step 3: Match the Metric to the Biological Question
Different downstream analyses have different contiguity requirements, and the metric you emphasize should reflect the question you are trying to answer.
For gene annotation and gene-space completeness, contiguity matters less than completeness. A fragmented assembly can still yield accurate gene models if the genes are fully covered by individual contigs. In this case, report BUSCO scores alongside N50, and interpret N50 as a secondary indicator. The Latrodectus genome assemblies demonstrate this point: both species had BUSCO scores above 94 percent despite very different scaffold counts. The L. katipo assembly had 13 chromosome-scale scaffolds containing 90 percent of the total length, while the L. hasselti assembly had 379 scaffolds. Both assemblies were suitable for gene annotation and comparative genomics, but they would be reported differently.
For comparative genomics and chromosome evolution studies, chromosome-level contiguity is essential. You need to know how many sequences correspond to chromosomes and what proportion of the assembly they represent. The Pseudobagrus ussuriensis genome provides a model for this type of reporting: 26 chromosome-level contigs covered 97.34 percent of the assembly, with an N50 of 28.53 Mb and an L50 of 11. These values together tell a clear story about the assembly structure.
For population genetics and variant calling, the key question is whether the assembly provides enough long-range context for accurate read mapping and variant detection. A high N50 is generally beneficial, but misjoins can introduce spurious variants. In this case, prioritize accuracy metrics such as read mapping rates and check for misassembly using Hi-C or linkage data.
Step 4: Apply the Comparison Rules for Multiple Assemblies
When you compare two or more assemblies, apply the following rules to avoid common errors.
First, compare like with like. Do not compare a contig N50 to a scaffold N50, and do not compare an N50 calculated on a haplotype-mixed assembly to an N50 calculated on a fully phased assembly. The MGA tool generates haplotype-mixed assemblies from HiFi reads alone, and these assemblies have different contiguity characteristics than fully phased diploid assemblies. If you are comparing a haplotype-mixed assembly to a phased assembly, note this difference in your report.
Second, compare NG50 instead of N50 when genome size estimates are available. This controls for differences in total assembly length that can arise from contamination or incomplete haplotype separation.
Third, examine the full length distribution, beyond the N50 point. Two assemblies can have identical N50 values but very different distributions. One might have a few very long sequences and many short ones, while the other might have many sequences close to the N50 length. The auN metric, which summarizes the entire Nx curve, is more informative for ranking assemblies with different distributions. Report auN when you need to distinguish between assemblies that have similar N50 values.
Fourth, always pair N50 with L50. The L50 tells you how many sequences contribute to the N50 threshold, which helps you understand whether the N50 reflects a few dominant sequences or a broader distribution. An L50 of 11 with an N50 of 28.53 Mb, as in the P. ussuriensis assembly, indicates that a small number of large contigs account for half the assembly. An L50 of 100 with the same N50 would indicate a very different structure.
Step 5: Document the Decision and the Metrics
Record the rationale for your metric choices in the assembly report or methods section. State whether you used contigs or scaffolds, whether you had a genome size estimate, and why you chose N50, NG50, or auN as the primary metric. This documentation helps reviewers and database users interpret your values correctly.
The AquaaG pipeline provides a reproducible framework for this documentation. It integrates QUAST for assembly quality assessment, BUSCO for gene-space completeness, and produces assembly quality reports that include contiguity metrics and metadata. The pipeline uses YAML configuration files, so you can record the assembly parameters and the metric choices in a structured format that can be versioned and shared.
Step 6: Escalate When Metrics Fall Outside Expected Ranges
Establish expected ranges for your species or assembly type before you generate the assembly. These expectations can come from published genomes of closely related species, from the performance of the assembler on test data, or from the sequencing technology you used. When your metrics fall outside these ranges, investigate before proceeding.
If the N50 is much lower than expected, check for insufficient sequencing coverage, high heterozygosity, or repetitive genome content. If the N50 is much higher than expected, check for chimeric joins or contamination that might have inflated the assembly length. If the scaffold N50 is dramatically higher than the contig N50, verify that the scaffolding was supported by evidence and not by spurious joins.
The EMBL-EBI Training program and the Galaxy Training Network offer practical tutorials on assembly quality assessment that can help you troubleshoot these issues. The Carpentries Lessons provide foundational computing skills for managing the data and workflows involved in assembly evaluation.
Common Failure Patterns in Metric Selection
Several recurring errors appear when researchers choose and report contiguity metrics. The first is reporting scaffold N50 without contig N50, which hides the underlying contig structure. The second is comparing N50 values across assemblies with different total lengths without using NG50. The third is interpreting N50 as a measure of completeness or accuracy instead of contiguity. The fourth is failing to specify whether the reported N50 refers to contigs or scaffolds.
Another failure pattern is selecting the metric that makes the assembly look best instead of the metric that is most appropriate for the question. For example, reporting only scaffold N50 when the assembly has poor contig contiguity, or reporting N50 instead of NG50 when the assembly contains substantial contamination. These choices mislead reviewers and database users and can lead to incorrect conclusions about assembly quality.
A Practical Comparison Workflow
When you need to compare multiple assemblies, use this workflow. First, collect the assembly FASTA files and any available genome size estimates. Second, run QUAST or a similar tool on each assembly to generate N50, L50, NG50, auN, and the number of sequences. Third, create a comparison table that includes the assembly method, the sequencing technology, the total assembly length, the contig and scaffold N50 values, the L50 values, and the BUSCO scores. Fourth, examine the table for patterns. If one assembly has a higher N50 but a lower BUSCO score, investigate whether the higher N50 came at the cost of completeness. If two assemblies have similar N50 values but different L50 values, examine the length distributions to understand the difference.
The nf-core Documentation describes community pipelines that include assembly quality assessment modules, and many of these pipelines generate comparison tables automatically. The Bioconductor project provides R packages for genomic analysis that can calculate contiguity metrics and generate publication-ready tables and figures.
When to Use auN Instead of N50
auN, the area under the Nx curve, provides a single value that summarizes the entire contiguity distribution. It is calculated by integrating the Nx values across all x from 0 to 100, where Nx is the length of the shortest sequence among the largest sequences that together account for x percent of the assembly. auN is more informative than N50 for ranking assemblies with different length distributions because it captures information from the entire curve instead of a single point.
Use auN when you need to distinguish between assemblies that have similar N50 values but different distributions. For example, one assembly might have a high N50 because it contains a few very long sequences, while another might have a similar N50 because it contains many moderately long sequences. These assemblies would have different auN values, and auN would rank them correctly for most downstream applications.
However, auN is less intuitive than N50, and many reviewers and database users are not familiar with it. Report auN alongside N50 and L50, and explain what it adds to the interpretation. The NCBI Data Resources provide examples of how contiguity metrics are reported in official database records, which can help you understand the conventions used in your field.
Integrating Contiguity Metrics into a Quality Report
A complete assembly quality report should include contiguity metrics, completeness metrics, and accuracy metrics. The contiguity section should report contig N50, scaffold N50, L50, NG50 if a genome size estimate is available, and auN. The completeness section should report BUSCO scores and the total assembly length relative to the expected genome size. The accuracy section should report read mapping rates, base-level accuracy estimates, and any evidence of misassembly.
The AquaaG pipeline produces this type of comprehensive report automatically, integrating QUAST for assembly quality assessment, BUSCO for gene-space completeness, and annotation tools for functional analysis. Using a pipeline like AquaaG ensures that your quality report is reproducible and that the metrics are calculated consistently across assembly versions.
Decision Rules for Reporting to Different Audiences
The audience for your contiguity metrics affects how you present them. For a genome paper aimed at a broad biological audience, report N50 and L50 with a brief explanation of what they mean. For a methods paper or a database submission, report the full set of metrics including NG50, auN, and the number of sequences. For a comparative genomics study, report NG50 and auN to enable fair comparisons across species.
The NCBI Data Resources provide submission guidelines that specify the required assembly quality metrics and metadata. The EMBL-EBI Training program offers courses on genome assembly and annotation that cover reporting standards and best practices. The Galaxy Training Network provides hands-on tutorials for assembly quality assessment that can help you prepare reports for different audiences.
Escalation Criteria for Metric Anomalies
Establish clear escalation criteria for when contiguity metrics fall outside expected ranges. If the N50 is more than an order of magnitude below the expected value for your species and sequencing technology, stop and investigate before proceeding with downstream analyses. Check the sequencing coverage, the assembler parameters, and the genome complexity. If the scaffold N50 is more than 10 times the contig N50, verify that the scaffolding was supported by evidence. If the BUSCO score is below 90 percent while the N50 is high, investigate whether the assembly is missing genes or whether the BUSCO analysis was run with the wrong lineage dataset.
If you cannot resolve these issues on your own, consult with colleagues who have experience with your sequencing platform or genome type. The EMBL-EBI Training program and the Galaxy Training Network offer advanced courses that cover troubleshooting and quality control for genome assemblies. The Carpentries Lessons provide foundational skills in shell scripting, data management, and version control that are essential for managing assembly workflows and documenting your decisions.
Frequently Asked Questions
What is the difference between N50 and NG50?
N50 is calculated against the total length of the assembly, while NG50 is calculated against the estimated genome size. If the assembly length is close to the genome size, N50 and NG50 will be similar. If the assembly contains contamination, duplicated sequence, or incomplete haplotype separation, the assembly length will exceed the genome size, and N50 will be higher than NG50. NG50 is preferred for comparing assemblies of the same species because it controls for differences in total assembly length.
Why is L50 reported alongside N50?
L50 is the number of sequences required to reach 50 percent of the total assembly length. It complements N50 by indicating how many sequences contribute to the N50 threshold. A low L50 means that a small number of large sequences account for half the assembly, which is generally desirable. Reporting both N50 and L50 gives a more complete picture of the length distribution than either metric alone.
Can N50 be used to compare assemblies of different species?
N50 can be compared across species only with caution. The N50 value depends on the genome size, the repeat content, and the assembly strategy, so a high N50 in one species does not necessarily indicate a better assembly than a lower N50 in another species. NG50, which is normalized by genome size, is more appropriate for cross-species comparisons, but even NG50 should be interpreted in the context of genome complexity and assembly method.
Does a higher N50 always mean a better assembly?
A higher N50 generally indicates better contiguity, but it does not guarantee a better assembly. An assembly with a high N50 can contain misjoins, collapsed repeats, or missing sequence. Conversely, an assembly with a lower N50 can be more accurate if the assembler chose to break contigs at ambiguous regions instead of joining them incorrectly. Always evaluate contiguity alongside completeness metrics such as BUSCO and accuracy metrics such as read mapping rates.
What is auN and when should I use it?
auN is the area under the Nx curve, which summarizes the entire contiguity distribution instead of a single point. It is calculated by integrating the Nx values across all x from 0 to 100. auN is more informative than N50 for ranking assemblies with different length distributions, and it is increasingly used in assembly evaluation. However, it is less intuitive than N50, so you should report both when possible.
How do I calculate N50 for a scaffolded assembly?
To calculate scaffold N50, use the scaffold sequences as they appear in the assembly FASTA file, including the gap regions. The scaffold N50 will be higher than the contig N50 if scaffolding successfully joined contigs. Report both contig N50 and scaffold N50 separately, and specify which one you are reporting when you communicate your results.
What should I do if my assembly has a low N50?
If your assembly has a low N50 relative to expectations, first check whether the assembly is complete and free of contamination. Then consider whether the sequencing coverage was sufficient, whether the assembler parameters were appropriate, and whether the genome has features such as high heterozygosity or repeat content that make assembly difficult. You may need additional sequencing, a different assembler, or a hybrid assembly strategy.
How do BUSCO scores relate to N50?
BUSCO scores measure gene-space completeness by checking for the presence of conserved single-copy genes, while N50 measures sequence contiguity. They are independent metrics that assess different aspects of assembly quality. A high-quality assembly should have both high BUSCO scores and high N50 values, but either metric can be high while the other is low. Report both to give a complete picture of assembly quality.
Related Bioinformatics Guides
- Evaluating Genome Assembly Quality: Metrics and Tools
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Metagenomic Assembly Overview: Challenges and Applications
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Insights into chromosomal evolution and sex determination of Pseudobagrus ussuriensis (Bagridae, Siluriformes) based on a chromosome-level genome.. DNA research : an international journal for rapid publication of reports on genes and genomes, 2022.
- MGA: a tool for haplotype-mixed assembly of long and accurate reads.. 2026.
- Closely related, yet phenotypically different - Genome assemblies of two sister species of widow spiders: <,i>,Latrodectus hasselti<,/i>, and <,i>,L. katipo<,/i>,, Theridiidae.. 2026.
- ERGA-BGE reference genome of <,i>,Diadema setosum:<,/i>, the Black Longspine Urchin invading the Mediterranean sea.. 2026.
- AquaaG: A comprehensive pipeline for quality assessment and annotation of genomes.. 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.