# From Reads to Complete Genomes: A Step-by-Step Pipeline for Long-Read Assembly and Polishing


## Key Takeaways

-   **Long-read sequencing is essential for resolving complex genomic structures** such as repetitive regions and structural variants, which are intractable for short-read technologies, enabling complete chromosome reconstruction and haplotype-resolved assemblies.
-   **The pipeline employs a modular workflow** starting with read quality assessment (NanoPlot, FastQC), followed by filtering (Filtlong, Chopper), initial assembly with Flye, iterative polishing using Racon (2-4 rounds), and final high-accuracy correction with Medaka, which utilizes neural network models specific to sequencing chemistry.
-   **Assembly accuracy is critically dependent on polishing**, with Racon reducing errors through read alignment and consensus generation, and Medaka achieving >99.9% base accuracy by leveraging platform-specific neural network models.
-   **Coverage depth is a key determinant of assembly quality**, with benchmarks suggesting ~36x average coverage for clinical applications (rare disease detection) and ~50x for plastid genomes to achieve accurate and unbiased assemblies.
-   **Targeted polishing with tools like GoldPolish-Target offers a resource-efficient alternative** for improving accuracy in specific problematic regions, achieving substantial reductions in indel and mismatch errors with significantly lower computational demands than whole-genome polishing.

---

Long-read sequencing enables researchers to reconstruct complete chromosomes, resolve complex repeat regions, and produce haplotype-resolved assemblies that short-read technologies cannot achieve. This pipeline converts raw long-read data into a polished genome through a modular workflow with concrete tool choices, parameter guidance, expected runtimes, and troubleshooting strategies for common failure modes. The workflow serves biology students, researchers, laboratory professionals, and life-science practitioners who need a practical path from sequencing output to a high-quality genome assembly suitable for downstream analysis.

## Scope and Reader Context

This pipeline addresses the complete assembly journey: raw read quality assessment, read filtering and preprocessing, initial assembly with long-read assemblers, iterative polishing with error correction tools, and final quality evaluation. The primary tools discussed are Flye for assembly, Racon for iterative polishing, and Medaka for final error correction, with alternatives such as GoldPolish-Target for targeted polishing and ptGAUL for plastid genomes. The workflow assumes access to a Linux environment with standard bioinformatics tools, sufficient computational resources, and a basic understanding of command-line operations. For readers needing foundational computing skills, [The Carpentries Lessons](https://carpentries.org/lessons) provide structured training in shell, Git, and data analysis that supports the technical requirements of this pipeline.

The evidence base for this workflow draws from recent advances in long-read assembly methodology. A 2025 study in the American Journal of Human Genetics demonstrated that nanopore long-read sequencing achieved approximately 36x average coverage and 32-kilobase read N50 from a single flow cell, with a pipeline that generated assemblies, phased variants, and methylation calls for rare disease detection. This study showed that long-read sequencing covered coding exons in approximately 280 genes and about 5 known Mendelian disease-associated genes that were not covered by short-read sequencing, and completely phased 87% of protein-coding genes. These findings establish the practical utility of long-read assembly for clinical genomics applications and provide realistic performance benchmarks for the pipeline described here.

## At a Glance: Pipeline Overview and Tool Selection

The following table summarizes the core pipeline stages, recommended tools, primary functions, and key considerations for each step. This decision framework helps newcomers select appropriate tools based on their data type, computational resources, and assembly goals.

| Pipeline Stage | Recommended Tool | Primary Function | Key Considerations |
| --- | --- | --- | --- |
| Read Quality Assessment | NanoPlot, FastQC | Evaluate read length distribution, quality scores, and coverage | Run before filtering to identify data issues early, document metrics for reproducibility |
| Read Filtering and Preprocessing | Filtlong, Chopper | Remove short reads, trim adapters, and filter by quality | Filtering thresholds depend on assembler requirements and genome complexity |
| Initial Assembly | Flye | Generate draft genome from long reads | Works with ONT and PacBio data, adjust parameters for genome size and coverage |
| Iterative Polishing | Racon | Correct errors in draft assembly using read alignments | Run 2 to 4 rounds, each round improves accuracy but with diminishing returns |
| Final Error Correction | Medaka | Produce high-accuracy consensus sequence | Uses neural network models, requires model matching to sequencing chemistry |
| Targeted Polishing | GoldPolish-Target | Polish specific regions with elevated error rates | Resource-efficient for large genomes, reduces indel and mismatch errors substantially |
| Quality Evaluation | QUAST, BUSCO | Assess assembly completeness and accuracy | Compare against reference genomes when available, check gene content completeness |

The pipeline structure follows a modular design that allows substitution of tools at each stage based on specific research needs. For example, the plastid Genome Assembly Using Long-read data (ptGAUL) pipeline provides a specialized alternative for organelle genomes, while MetaBooster and MetaBooster-HiFi offer strain-aware assembly pipelines for metagenomic applications. The modular approach ensures that researchers can adapt the workflow to their specific organism, sequencing platform, and computational constraints without redesigning the entire process.

## Core Principles of Long-Read Assembly

### Why Long Reads Matter for Genome Assembly

Long-read sequencing technologies from Oxford Nanopore Technologies and Pacific Biosciences produce reads that span repetitive regions, structural variants, and complex genomic arrangements that confound short-read assemblers. The fundamental advantage of long reads is their ability to provide contiguous coverage across regions that would otherwise fragment into ambiguous contigs. A 2023 study in Molecular Ecology Resources highlighted that plastomes containing long repeat sequences exceeding 300 base pairs present substantial challenges for short-read assembly, leading to misassemblies and consensus sequences with spurious rearrangements. Single-molecule long-read sequencing overcomes these challenges by spanning the repeats with individual reads.

The clinical relevance of long-read assembly extends to rare disease diagnostics. The 2025 American Journal of Human Genetics study demonstrated that long-read sequencing captured variants inaccessible to short-read sequencing, including structural variants and tandem repeats, and provided haplotype-resolved methylation profiling. This capability to phase variants across long distances and detect complex genomic alterations makes long-read assembly essential for comprehensive genomic analysis.

### Error Profiles and Their Implications

Long reads typically have higher error rates than short reads, with error profiles that differ between Oxford Nanopore and Pacific Biosciences platforms. A 2025 BMC Bioinformatics study noted that if these errors remain unaddressed, genome assemblies may exhibit high base error rates that compromise downstream analysis reliability. The error correction challenge is substantial: insertion and deletion errors are more common in long-read data than mismatches, and these errors accumulate during assembly if not corrected through polishing.

The practical implication is that assembly polishing is a required stage in the pipeline. The GoldPolish-Target study demonstrated that targeted polishing can reduce insertion and deletion errors by up to 49.2% and mismatch errors by up to 55.4%, achieving base accuracy values above 99.9% with Phred scores greater than 30. This level of accuracy is comparable to the state-of-the-art Medaka polisher while exhibiting up to 27-fold shorter run times and consuming 95% less memory on average.

### Coverage Requirements and Data Planning

Coverage depth directly influences assembly quality and contiguity. The rare disease study achieved approximately 36x average coverage from a single flow cell, which proved sufficient for generating assemblies that captured clinically relevant variants. For plastid genome assembly, the ptGAUL pipeline demonstrated accurate and unbiased assemblies using only about 50x coverage of plastome data. These coverage benchmarks provide practical planning targets for researchers designing sequencing experiments.

The relationship between coverage and assembly quality is not linear. Insufficient coverage leads to fragmented assemblies with gaps and misjoins, while excessive coverage increases computational costs without proportional quality gains. The optimal coverage depends on genome size, complexity, and the specific assembler used. For bacterial genomes, 50 to 100x coverage typically suffices, while larger and more complex eukaryotic genomes may require 30 to 60x coverage with ultralong reads for chromosome-scale assembly.

## Practical Workflow: Step-by-Step Implementation

### Step 1: Environment Setup and Data Organization

Before beginning assembly, establish a reproducible computing environment with the necessary tools installed. The [Bioconductor Project](https://bioconductor.org/) provides official documentation for package installation and reproducible genomic analysis workflows, while the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training and analysis tutorials that support reproducible research practices. For researchers new to command-line environments, [The Carpentries Lessons](https://carpentries.org/lessons) provide foundational training in shell operations, data management, and programming that are prerequisites for efficient bioinformatics work.

Organize your project directory with clear structure:

```
project/
├── data/
│   ├── raw_reads/
│   └── filtered_reads/
├── assembly/
│   ├── draft/
│   ├── polished/
│   └── final/
├── quality/
│   ├── pre_assembly/
│   └── post_assembly/
├── scripts/
└── logs/
```

Document all software versions and parameters in a README file or computational notebook. The [nf-core Documentation](https://nf-co.re/docs) emphasizes community pipeline standards for usage, configuration, and reproducible workflow context, providing a model for structured project documentation.

### Step 2: Read Quality Assessment

Run quality assessment on raw reads before any filtering or assembly steps. This initial assessment establishes baseline metrics that guide downstream decisions and provides a reference point for evaluating the effects of filtering.

Key metrics to record:

- Total number of reads and total bases
- Read N50 and mean read length
- Quality score distribution
- Coverage estimate based on expected genome size
- Adapter contamination levels

For nanopore data, tools such as NanoPlot provide comprehensive visualizations of read length distributions, quality scores, and cumulative yield. For PacBio data, similar metrics can be obtained through platform-specific tools or general quality assessment packages. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of sequence databases and analysis services that support data management and quality assessment workflows.

Document these metrics in your project records. They serve as the baseline for evaluating whether filtering improves assembly outcomes and for troubleshooting unexpected assembly results.

### Step 3: Read Filtering and Preprocessing

Filter raw reads to remove low-quality sequences, short fragments, and adapter contamination before assembly. The specific filtering strategy depends on your assembler and genome characteristics.

Standard filtering approach:

1. Remove reads shorter than a minimum length threshold (commonly 500 to 1000 base pairs for bacterial genomes, 1000 to 2000 base pairs for eukaryotic genomes)
2. Trim adapter sequences and low-quality ends
3. Remove reads with average quality scores below a threshold
4. Optionally subsample reads to achieve target coverage

Tools such as Filtlong and Chopper provide flexible filtering options for long-read data. For nanopore data, consider whether to retain ultralong reads even if they have lower quality, as these reads provide valuable scaffolding information for repetitive regions.

Record the number of reads and bases removed during filtering. This information helps evaluate whether aggressive filtering improves assembly quality or removes useful data.

### Step 4: Initial Assembly with Flye

Flye is a robust long-read assembler that handles both Oxford Nanopore and Pacific Biosciences data. It builds an initial assembly graph, resolves repeats, and produces a draft genome with contigs representing the assembled sequences.

Basic Flye command structure:

```
flye --nano-raw filtered_reads.fastq --genome-size 5m --out-dir assembly/draft --threads 16
```

Key parameters to adjust:

- `--genome-size`: Provide an estimate of the genome size to guide assembly parameters. This value can be approximate, Flye adjusts internally based on read coverage.
- `--nano-raw` or `--pacbio-raw`: Specify the sequencing platform and read type
- `--threads`: Set the number of CPU threads based on available resources
- `--iterations`: Control the number of assembly refinement iterations (default is 2 for raw reads)

Expected runtimes vary substantially based on genome size and coverage. Bacterial genomes typically assemble in minutes to hours, while mammalian genomes may require days. The rare disease study achieved 32-kilobase read N50 from a single flow cell, demonstrating that high-quality nanopore data can produce contiguous assemblies with modest computational requirements.

After assembly, examine the draft assembly statistics:

- Number of contigs and total assembly length
- Contig N50 and longest contig
- Coverage distribution across the assembly
- Presence of circular contigs (for bacterial genomes and plastids)

### Step 5: Iterative Polishing with Racon

Racon performs error correction by aligning reads to the draft assembly and generating a consensus sequence. It is typically run in multiple rounds, with each round improving accuracy.

Racon polishing workflow:

1. Align reads to the draft assembly using minimap2
2. Run Racon with the alignment to generate a polished assembly
3. Repeat the alignment and polishing steps for 2 to 4 rounds

Command sequence for one polishing round:

```
minimap2 -t 16 -x map-ont draft_assembly.fasta filtered_reads.fastq > reads_to_draft.paf
racon -t 16 filtered_reads.fastq reads_to_draft.paf draft_assembly.fasta > polished_round1.fasta
```

The `-x map-ont` parameter is appropriate for Oxford Nanopore data, use `-x map-pb` for Pacific Biosciences data. Each polishing round typically takes less time than the initial assembly, but the computational cost depends on genome size and read depth.

Monitor accuracy improvement across rounds by comparing assembly statistics and, if a reference genome is available, by calculating alignment-based accuracy metrics. The improvement between rounds diminishes, and most assemblies reach their practical accuracy limit after 2 to 4 rounds of Racon polishing.

### Step 6: Final Error Correction with Medaka

Medaka produces a high-accuracy consensus sequence using neural network models trained on specific sequencing chemistries. It typically outperforms Racon in final accuracy but requires more computational resources.

Medaka workflow:

1. Align reads to the Racon-polished assembly
2. Run Medaka with the appropriate model for your sequencing chemistry
3. Generate the final polished assembly

Command structure:

```
minimap2 -t 16 -x map-ont polished_round4.fasta filtered_reads.fastq > reads_to_polished.paf
medaka_consensus -i filtered_reads.fastq -d polished_round4.fasta -o assembly/final -t 16 -m r1041_e82_400bps_sup_v4.3.0
```

The model parameter `-m` must match your sequencing chemistry and basecalling model. Consult the Medaka documentation for the correct model name for your data. Using an incorrect model produces suboptimal results.

The GoldPolish-Target study demonstrated that Medaka achieves state-of-the-art polishing accuracy, with base accuracy values above 99.9% and Phred scores greater than 30. This level of accuracy is suitable for most downstream applications, including variant calling and comparative genomics.

### Step 7: Targeted Polishing for Problematic Regions

After standard polishing, some regions may retain elevated error rates, particularly in areas with complex repeats, homopolymers, or low coverage. GoldPolish-Target provides a modular targeted polishing pipeline that isolates and polishes user-specified assembly loci, offering a resource-efficient means for polishing targeted regions of draft genomes.

The GoldPolish-Target study demonstrated that this approach reduces insertion and deletion errors by up to 49.2% and mismatch errors by up to 55.4% in targeted regions, achieving accuracy comparable to Medaka while exhibiting up to 27-fold shorter run times and consuming 95% less memory on average. This makes targeted polishing particularly valuable for large genomes where polishing the entire assembly is computationally expensive.

Identify problematic regions through:

- Coverage analysis to find low-coverage areas
- Alignment-based error assessment against reference genomes when available
- BUSCO analysis to identify missing or fragmented conserved genes
- Visual inspection of assembly graphs for suspicious structures

### Step 8: Quality Evaluation and Validation

Comprehensive quality evaluation is essential before using the assembly for downstream analysis. Use multiple complementary metrics to assess assembly completeness and accuracy.

Assembly quality metrics:

- Contiguity statistics: contig N50, number of contigs, total assembly length
- Completeness: BUSCO scores for conserved gene content
- Accuracy: alignment-based error rates when a reference genome is available
- Coverage uniformity: distribution of read coverage across the assembly
- Structural correctness: absence of misjoins and chimeric contigs

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and sequence databases that support quality evaluation through comparative analysis. For organisms with available reference genomes, whole-genome alignment provides the most direct assessment of assembly accuracy.

BUSCO analysis evaluates assembly completeness by searching for conserved single-copy orthologs expected in the target lineage. High BUSCO scores indicate that the assembly captures the expected gene content, while low scores suggest missing or fragmented regions that may require additional sequencing or assembly improvement.

## Options and Tradeoffs in Tool Selection

### Assembler Selection

Flye is recommended as the primary assembler for this pipeline due to its robust performance across diverse genome types and sequencing platforms. Alternative assemblers include:

- Canu: Provides high-quality assemblies but requires more computational resources and longer runtimes
- Raven: Offers fast assembly with lower memory requirements, suitable for smaller genomes
- Shasta: Designed for rapid nanopore assembly, particularly useful for bacterial genomes
- HiCanu: Optimized for PacBio HiFi data, producing highly accurate assemblies

The choice of assembler depends on genome complexity, available computational resources, and the specific error profile of your sequencing data. For metagenomic applications, MetaBooster and MetaBooster-HiFi provide strain-aware assembly pipelines that outperform general-purpose assemblers in genome fraction, contig length, and error rates, as demonstrated in a 2022 Frontiers in Genetics study.

### Polishing Strategy

The Racon plus Medaka combination represents the standard approach for long-read assembly polishing. Alternative strategies include:

- GoldPolish-Target for targeted polishing of specific regions
- NextPolish for additional error correction rounds
- MarginPolish followed by Merfin for variant-aware polishing

The choice of polishing strategy depends on the required accuracy level, available computational resources, and the specific error patterns in your data. The GoldPolish-Target study demonstrated that targeted polishing can achieve accuracy comparable to Medaka with substantially lower computational costs, making it an attractive option for large genomes or when computational resources are limited.

### Platform-Specific Considerations

Oxford Nanopore and Pacific Biosciences platforms produce reads with different error profiles and characteristics. Nanopore data typically has higher error rates but can produce ultralong reads exceeding 100 kilobases, which are valuable for spanning complex repeats. PacBio HiFi data has lower error rates but shorter read lengths, typically 10 to 25 kilobases.

The rare disease study used nanopore sequencing and achieved 32-kilobase read N50 from a single flow cell, demonstrating the practical utility of nanopore data for clinical genomics applications. The ptGAUL pipeline supports both Oxford Nanopore and Pacific Biosciences platforms for plastid genome assembly, providing flexibility for researchers with different sequencing infrastructure.

## Observations and Measurements

### Expected Runtime and Resource Requirements

Runtime and resource requirements vary substantially based on genome size, coverage, and available computational infrastructure. The following observations provide practical planning guidance:

- Bacterial genomes (5 megabases) at 100x coverage: assembly completes in 10 to 30 minutes with 16 threads, polishing adds 20 to 60 minutes
- Fungal genomes (30 to 50 megabases) at 50x coverage: assembly completes in 1 to 3 hours, polishing adds 2 to 6 hours
- Plant or animal genomes (500 megabases to 3 gigabases) at 30 to 50x coverage: assembly completes in 1 to 3 days, polishing adds 2 to 5 days

Memory requirements scale with genome size and read depth. The GoldPolish-Target study demonstrated that targeted polishing consumes 95% less memory than whole-genome polishing approaches, making it suitable for large genomes on modest computational infrastructure.

### Quality Metrics Across Pipeline Stages

Document quality metrics at each pipeline stage to track improvement and identify potential issues:

| Pipeline Stage | Expected Contig N50 | Expected Base Accuracy | Key Quality Indicators |
| --- | --- | --- | --- |
| Raw reads | Not applicable | 85 to 95% | Read N50, coverage, quality scores |
| Draft assembly | Variable by genome | 95 to 98% | Contiguity, BUSCO completeness |
| After Racon polishing | Same as draft | 98 to 99.5% | Reduced indel errors |
| After Medaka polishing | Same as draft | 99.5 to 99.9% | High Phred scores, low error rates |
| After targeted polishing | Same as draft | Above 99.9% | Q30 or higher in targeted regions |

These ranges represent typical outcomes from published studies and practical experience. Individual results vary based on data quality, genome complexity, and parameter choices.

### Record Keeping for Reproducibility

Maintain detailed records of all pipeline parameters, software versions, and quality metrics. The [nf-core Documentation](https://nf-co.re/docs) emphasizes community pipeline standards for reproducible workflow context, providing a model for structured documentation. The [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that supports reproducible analysis practices.

Essential records include:

- Software versions for all tools used
- Parameter values for each pipeline stage
- Read filtering statistics (reads removed, bases retained)
- Assembly statistics at each stage
- Quality metrics and evaluation results
- Computational resource usage (runtime, memory, storage)

These records enable troubleshooting, method comparison, and manuscript preparation. They also support compliance with data sharing and reproducibility requirements from journals and funding agencies.

## Common Failure Patterns and Troubleshooting

### Fragmented Assembly with Many Small Contigs

Fragmented assemblies typically result from insufficient coverage, poor read quality, or complex repeat structures that cannot be resolved with available data.

Troubleshooting steps:

1. Check coverage statistics to confirm sufficient read depth
2. Examine read length distribution for excessive short reads
3. Review quality scores for systematic quality issues
4. Consider whether additional sequencing is needed to resolve complex regions
5. Try alternative assemblers that may handle repeats differently

The ptGAUL study demonstrated that long repeats exceeding 300 base pairs make plastome assembly challenging with short-read data, leading to misassemblies and spurious rearrangements. Long-read assembly overcomes these challenges, but coverage of approximately 50x is required for accurate results.

### High Error Rates After Polishing

Persistent high error rates after polishing suggest issues with read quality, polishing parameters, or model selection.

Troubleshooting steps:

1. Verify that the correct Medaka model is used for your sequencing chemistry
2. Confirm that reads are properly filtered before assembly and polishing
3. Check for systematic errors in homopolymer regions
4. Consider additional polishing rounds or alternative polishing tools
5. Evaluate whether targeted polishing of problematic regions is needed

The GoldPolish-Target study demonstrated that many genome assembly workflows still produce regions with elevated error rates, such as gaps filled with unpolished or ambiguous bases. Targeted polishing addresses these persistent error regions efficiently.

### Assembly Size Mismatch with Expected Genome Size

Assemblies that are substantially larger or smaller than the expected genome size indicate potential contamination, misassembly, or incomplete assembly.

Troubleshooting steps:

1. Check for contamination by examining coverage distribution and taxonomic classification of contigs
2. Verify genome size estimates using flow cytometry or other independent methods
3. Examine assembly graphs for misjoins or collapsed repeats
4. Assess whether repetitive regions are properly resolved
5. Consider whether organelle genomes or plasmids are included in the assembly

### Excessive Computational Resource Usage

Long-read assembly can require substantial computational resources, particularly for large genomes. If resource usage exceeds available infrastructure, consider:

1. Reducing coverage by subsampling reads
2. Using targeted polishing instead of whole-genome polishing
3. Selecting more memory-efficient tools
4. Running assembly on cloud or high-performance computing resources
5. Optimizing thread usage and parallelization settings

The GoldPolish-Target study demonstrated that targeted polishing achieves substantial resource savings compared to whole-genome polishing, making it suitable for large genomes on modest infrastructure.

## Limitations and Interpretation Boundaries

### Coverage and Data Quality Limitations

The pipeline assumes sufficient coverage and read quality for successful assembly. The rare disease study achieved approximately 36x average coverage from a single flow cell, which proved sufficient for clinical genomics applications. However, complex genomes with high repeat content or extreme base composition may require higher coverage for complete assembly.

Low-coverage regions may remain unresolved, leading to gaps or misassemblies that affect downstream analysis. The ptGAUL study demonstrated that approximately 50x coverage of plastome data produces accurate and unbiased assemblies, but this benchmark may not transfer directly to nuclear genomes with different complexity profiles.

### Platform-Specific Error Patterns

Oxford Nanopore and Pacific Biosciences platforms produce different error profiles that affect assembly and polishing strategies. Nanopore data typically has higher error rates, particularly in homopolymer regions, while PacBio HiFi data has lower error rates but shorter read lengths. The polishing tools and parameters described in this pipeline are optimized for specific platform characteristics and may require adjustment when switching between platforms.

### Genome Complexity Constraints

Genomes with extreme complexity, such as those with large segmental duplications, high heterozygosity, or extensive repeat content, may require specialized assembly strategies beyond the standard pipeline described here. The MetaBooster study demonstrated that strain-aware assembly in metagenomic contexts requires specialized pipelines that account for population heterogeneity and strain variation.

### Computational Infrastructure Requirements

The pipeline requires a Linux environment with sufficient computational resources for assembly and polishing. The GoldPolish-Target study demonstrated that targeted polishing reduces computational requirements substantially, but initial assembly of large genomes still requires significant memory and processing power. Researchers without access to appropriate infrastructure should consider cloud computing or institutional high-performance computing resources.

## Safety and Regulatory Context

### Data Management and Privacy Considerations

Genomic data, particularly human data, is subject to privacy and regulatory requirements. The rare disease study highlights the clinical utility of long-read assembly for diagnostic applications, but clinical implementation requires compliance with applicable regulations and ethical standards. Researchers working with human data must ensure appropriate consent, de-identification, and data security measures.

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official guidance on data submission, access controls, and database usage that support compliant data management practices. Researchers should consult institutional review boards and data protection officers before initiating projects involving human genomic data.

### Reproducibility and Documentation Standards

Journals and funding agencies increasingly require reproducible analysis workflows with documented parameters and software versions. The [nf-core Documentation](https://nf-co.re/docs) provides community pipeline standards that support reproducible workflow context, while the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that emphasizes reproducibility. The [Bioconductor Project](https://bioconductor.org/) provides official documentation for reproducible genomic analysis workflows.

Adopting these standards from the outset of a project reduces the burden of retrospective documentation and supports compliance with publication requirements.

### Professional Escalation Criteria

Seek expert consultation when:

1. Assembly quality metrics fall substantially below expected ranges after multiple troubleshooting attempts
2. The target genome has known complex features that require specialized assembly strategies
3. Clinical or regulatory decisions depend on assembly accuracy
4. Computational resource requirements exceed available infrastructure
5. Unexpected assembly structures suggest potential data contamination or technical issues

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) provides bioinformatics learning pathways and practical analysis education that support skill development for researchers addressing complex assembly challenges.

## Decision Framework: Matching Pipeline Strategy to Genome Complexity and Project Goals

Before running any assembly command, define the decision criteria that determine which pipeline variant, coverage target, and polishing depth your project requires. This framework prevents wasted compute time and clarifies when the standard Flye plus Racon plus Medaka workflow is appropriate versus when specialized alternatives such as ptGAUL for plastid genomes or MetaBooster for metagenomes are necessary. The framework uses three assessment axes: genome complexity, downstream application requirements, and available computational resources. Each axis produces a concrete decision that changes pipeline parameters.

### Axis 1: Genome Complexity Assessment

Genome complexity determines the minimum read length, coverage depth, and assembly strategy required for a complete result. The 2023 ptGAUL study in Molecular Ecology Resources demonstrated that plastomes containing long repeat sequences exceeding 300 base pairs cause misassemblies and spurious rearrangements with short-read data, while long-read assembly resolves these structures accurately. This finding establishes that repeat content, not genome size alone, drives assembly difficulty.

Score your target genome across these four complexity indicators:

| Complexity Indicator | Low Complexity (Score 1) | Moderate Complexity (Score 2) | High Complexity (Score 3) |
| --- | --- | --- | --- |
| Repeat content | Below 5 percent | 5 to 20 percent | Above 20 percent |
| Heterozygosity | Below 0.5 percent | 0.5 to 2 percent | Above 2 percent |
| Ploidy | Haploid or haploid-representative | Diploid with moderate divergence | Polyploid or highly divergent haplotypes |
| Genome size | Below 100 megabases | 100 megabases to 1 gigabase | Above 1 gigabase |

Sum the four scores to classify your genome. A total of 4 to 6 indicates low complexity, 7 to 9 indicates moderate complexity, and 10 to 12 indicates high complexity. This classification directly determines coverage targets and polishing requirements.

For low-complexity genomes, the standard pipeline with 50 to 100x coverage and 2 to 3 Racon rounds followed by Medaka produces satisfactory results. For moderate-complexity genomes, increase coverage to 60 to 100x, use 3 to 4 Racon rounds, and budget for targeted polishing of unresolved regions. For high-complexity genomes, plan for ultralong reads exceeding 50 kilobases where possible, target 80 to 120x coverage, and expect to use specialized tools or multiple assembly strategies.

The 2025 American Journal of Human Genetics study provides a clinical benchmark for this assessment. The nanopore sequencing approach achieved approximately 36x average coverage with a 32-kilobase read N50 from a single flow cell, which proved sufficient for detecting structural variants, tandem repeats, and phased variants in a rare disease cohort. This coverage level worked because the target was the human genome with established reference resources and the analysis focused on variant detection instead of complete de novo assembly of complex regions.

### Axis 2: Downstream Application Requirements

The intended use of your assembly determines the required base accuracy and contiguity. Different applications tolerate different error profiles, and matching polishing depth to application requirements avoids unnecessary compute expenditure.

Define your primary downstream application and its accuracy requirements:

| Application | Required Base Accuracy | Required Contiguity | Polishing Strategy |
| --- | --- | --- | --- |
| Gene presence and absence screening | 95 to 98 percent | Contig N50 above gene length | Racon only, 2 rounds |
| Structural variant detection | 98 to 99.5 percent | Contig N50 above variant span | Racon plus Medaka |
| SNP and small variant calling | Above 99.9 percent | Chromosome-scale or reference-guided | Racon plus Medaka plus targeted polishing |
| Clinical or diagnostic reporting | Above 99.9 percent with validation | Reference-grade or complete chromosome | Full pipeline with independent validation |
| Comparative genomics and phylogenetics | Above 99.5 percent | Contig N50 above conserved blocks | Racon plus Medaka |
| Gene annotation | 99 to 99.9 percent | Contig N50 above gene span | Racon plus Medaka |

The GoldPolish-Target study in BMC Bioinformatics demonstrated that targeted polishing achieves base accuracy above 99.9 percent with Phred scores greater than 30, comparable to Medaka, while running up to 27-fold faster and consuming 95 percent less memory. This finding supports a tiered polishing strategy: use Racon for initial correction, Medaka for whole-genome consensus, and GoldPolish-Target only for regions that remain problematic after standard polishing.

For clinical applications, the 2025 American Journal of Human Genetics study demonstrated that long-read assembly with approximately 36x coverage detected diagnostic variants in 11 probands from a 41-family rare disease cohort. The study identified diverse genetic causes including de novo variants, compound heterozygous variants, large-scale structural variants, and epigenetic modifications. This evidence supports the utility of long-read assembly for diagnostic workflows, but clinical implementation requires additional validation steps beyond the standard pipeline.

### Axis 3: Computational Resource Budget

Computational resources constrain the feasible pipeline strategy. Assess your available infrastructure before selecting tools and parameters, because some polishing approaches require substantially more memory and runtime than others.

Calculate your resource budget across these dimensions:

1. Available RAM in gigabytes
2. Available CPU threads
3. Wall-clock time limit for the assembly project
4. Storage space for intermediate files
5. Access to cloud or high-performance computing

The GoldPolish-Target study provides concrete resource benchmarks. The targeted polishing approach consumed 95 percent less memory than whole-genome polishing on average, making it suitable for large genomes on modest infrastructure. This resource efficiency enables a practical strategy: run the standard pipeline for initial assembly and whole-genome polishing, then use targeted polishing for specific regions instead of repeated whole-genome polishing rounds.

For bacterial genomes under 10 megabases, a standard workstation with 16 threads and 32 gigabytes of RAM handles the complete pipeline. For fungal genomes between 30 and 100 megabases, plan for 32 to 64 gigabytes of RAM and 16 to 32 threads. For plant or animal genomes above 500 megabases, budget for 128 gigabytes or more of RAM and access to high-performance computing or cloud resources.

When resources are constrained, apply these adjustments in order of preference:

1. Reduce coverage by subsampling reads to the minimum required for your complexity class
2. Use Racon only for polishing when base accuracy requirements are below 99.5 percent
3. Use GoldPolish-Target for targeted polishing instead of additional Medaka rounds
4. Use the ptGAUL pipeline for plastid genomes instead of whole-genome assembly tools
5. Use MetaBooster for metagenomic samples instead of single-genome assemblers

### Decision Matrix for Pipeline Selection

Combine the three axes into a single decision matrix that selects the pipeline variant and parameter set:

| Genome Complexity | Application Requirement | Resource Budget | Recommended Pipeline |
| --- | --- | --- | --- |
| Low | Gene screening | Limited | Flye plus 2 Racon rounds |
| Low | Variant calling | Standard | Flye plus Racon plus Medaka |
| Low | Clinical reporting | Standard | Flye plus Racon plus Medaka plus targeted polishing |
| Moderate | Gene screening | Limited | Flye plus 3 Racon rounds |
| Moderate | Variant calling | Standard | Flye plus Racon plus Medaka |
| Moderate | Clinical reporting | High | Flye plus Racon plus Medaka plus targeted polishing |
| High | Gene screening | High | Flye plus 4 Racon rounds plus Medaka |
| High | Variant calling | High | Flye plus Racon plus Medaka plus targeted polishing |
| High | Clinical reporting | High | Specialized assembly strategy with expert consultation |

For plastid genomes, substitute ptGAUL for the standard assembly pipeline. The 2023 Molecular Ecology Resources study demonstrated that ptGAUL produces accurate and unbiased assemblies using approximately 50x coverage of plastome data from either Oxford Nanopore or Pacific Biosciences platforms. This specialized pipeline handles long repeats and rearrangements that challenge general-purpose assemblers.

For metagenomic samples, substitute MetaBooster or MetaBooster-HiFi for the standard pipeline. The 2022 Frontiers in Genetics study demonstrated that these pipelines outperform state-of-the-art de novo metagenome assemblers in genome fraction, contig length, and error rates. Strain-aware assembly requires specialized handling of population heterogeneity that single-genome assemblers do not provide.

### Record System for Pipeline Decisions

Document the decision framework results before starting assembly. This record supports troubleshooting, method comparison, and manuscript preparation. The [nf-core Documentation](https://nf-co.re/docs) emphasizes community pipeline standards for reproducible workflow context, and the [Galaxy Training Network](https://training.galaxyproject.org/) offers accessible workflow training that supports reproducible analysis practices.

Create a pipeline decision record with these fields:

1. Genome name and taxonomic identifier
2. Estimated genome size and source of estimate
3. Complexity scores for repeats, heterozygosity, ploidy, and size
4. Total complexity classification
5. Primary downstream application and required accuracy
6. Available computational resources
7. Selected pipeline variant and rationale
8. Coverage target and read length requirements
9. Polishing strategy and expected number of rounds
10. Quality metrics that define successful completion

Store this record alongside the assembly outputs and quality metrics. The [Bioconductor Project](https://bioconductor.org/) provides official documentation for reproducible genomic analysis workflows that support structured record keeping. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) offers bioinformatics learning pathways that include practical analysis education for documenting and reproducing genomic analyses.

### Common Decision Errors and Their Consequences

Several recurring decision errors produce suboptimal assembly outcomes. Recognizing these patterns helps avoid wasted compute time and improves assembly quality.

**Error 1: Overestimating coverage requirements.** Running excessive coverage increases computational cost without proportional quality gains. The ptGAUL study demonstrated that approximately 50x coverage produces accurate plastid assemblies, and the rare disease study achieved clinically useful results with approximately 36x coverage. For low-complexity genomes, 50 to 100x coverage typically suffices. Additional coverage beyond 150x rarely improves assembly quality and substantially increases runtime and memory usage.

**Error 2: Underestimating repeat complexity.** Genomes with long repeats exceeding 300 base pairs require long-read assembly regardless of coverage depth. The ptGAUL study demonstrated that short-read data produces misassemblies and spurious rearrangements in these regions. If your genome has known repeat structures, plan for ultralong reads and verify that your read N50 exceeds the longest repeat length.

**Error 3: Applying whole-genome polishing when targeted polishing suffices.** The GoldPolish-Target study demonstrated that targeted polishing achieves accuracy comparable to Medaka with substantially lower computational cost. For large genomes where only specific regions have elevated error rates, targeted polishing provides a resource-efficient alternative to additional whole-genome polishing rounds.

**Error 4: Using the wrong Medaka model.** The Medaka model must match the sequencing chemistry and basecalling version. Using an incorrect model produces suboptimal consensus accuracy. Verify the model name against sequencing metadata before running Medaka, and consult the Medaka documentation for the complete model list.

**Error 5: Ignoring coverage distribution.** Average coverage can mask substantial regional variation. Low-coverage regions remain unresolved and produce assembly gaps or errors. Examine coverage distribution across the assembly and identify regions below the minimum threshold for your application. The rare disease study achieved approximately 36x average coverage, but regional coverage variation likely affected assembly quality in specific loci.

### Professional Escalation Criteria for Pipeline Decisions

Seek expert consultation when the decision framework produces ambiguous or conflicting guidance. Specific escalation triggers include:

1. Genome complexity scores above 10, indicating high complexity across multiple indicators
2. Clinical or diagnostic applications where assembly errors could affect patient care decisions
3. Repeated assembly failures despite parameter adjustment and troubleshooting
4. Computational resource requirements that exceed available infrastructure by more than 2-fold
5. Unexpected assembly structures that suggest data contamination or technical issues
6. Target genomes with known complex features such as large segmental duplications or extreme base composition

The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and sequence databases that support comparative analysis and troubleshooting. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) offers bioinformatics learning pathways that support skill development for addressing complex assembly challenges. For foundational computing skills that support efficient pipeline implementation, [The Carpentries Lessons](https://carpentries.org/lessons) provide structured training in shell, Git, and data analysis.

### Implementing the Decision Framework in Practice

Apply the decision framework before any assembly command. This implementation sequence takes approximately 30 minutes and prevents costly errors:

1. Record the genome name, estimated size, and complexity scores
2. Calculate the total complexity classification
3. Define the primary downstream application and required accuracy
4. Inventory available computational resources
5. Select the pipeline variant from the decision matrix
6. Set coverage targets and read length requirements
7. Document the decision record in the project directory
8. Proceed with read quality assessment and filtering

This structured approach ensures that assembly parameters match project requirements and resource constraints. The decision record provides a reference point for troubleshooting and method comparison, supporting reproducible research practices that journals and funding agencies increasingly require.

## Frequently Asked Questions

### What is the minimum coverage required for long-read genome assembly?

Coverage requirements depend on genome complexity and assembly goals. The rare disease study achieved approximately 36x average coverage from a single flow cell and generated assemblies suitable for clinical variant detection. The ptGAUL study demonstrated accurate plastid genome assembly with approximately 50x coverage. For bacterial genomes, 50 to 100x coverage is typically sufficient, while complex eukaryotic genomes may require 30 to 60x coverage with ultralong reads for chromosome-scale assembly. Higher coverage improves assembly contiguity and accuracy but increases computational costs.

### How many rounds of Racon polishing are recommended?

Most assemblies reach their practical accuracy limit after 2 to 4 rounds of Racon polishing. The improvement between rounds diminishes substantially, with the first round providing the largest accuracy gain. After 4 rounds, additional Racon polishing typically provides minimal improvement, and transitioning to Medaka for final error correction is more effective. Monitor accuracy metrics across rounds to determine when additional polishing is no longer beneficial.

### What is the difference between Racon and Medaka polishing?

Racon performs error correction through read-to-assembly alignment and consensus generation, providing fast and memory-efficient polishing. Medaka uses neural network models trained on specific sequencing chemistries to produce higher accuracy consensus sequences but requires more computational resources. The GoldPolish-Target study demonstrated that Medaka achieves state-of-the-art polishing accuracy with base accuracy values above 99.9% and Phred scores greater than 30. The standard approach uses Racon for initial polishing rounds followed by Medaka for final error correction.

### How do I choose the correct Medaka model for my data?

The Medaka model must match your sequencing chemistry and basecalling model. Consult the Medaka documentation for the complete list of available models and their corresponding sequencing platforms and basecalling versions. Using an incorrect model produces suboptimal results, so verify the model name against your sequencing metadata before running Medaka. If you are unsure which model applies to your data, consult your sequencing facility or platform documentation.

### What should I do if my assembly is fragmented with many small contigs?

Fragmented assemblies typically result from insufficient coverage, poor read quality, or complex repeat structures. Check coverage statistics to confirm sufficient read depth, examine read length distributions for excessive short reads, and review quality scores for systematic issues. Consider whether additional sequencing is needed to resolve complex regions, and try alternative assemblers that may handle repeats differently. The ptGAUL study demonstrated that long repeats exceeding 300 base pairs present substantial challenges for short-read assembly but can be resolved with long-read data.

### How do I evaluate the quality of my final assembly?

Use multiple complementary metrics to assess assembly quality. Contiguity statistics such as contig N50 and number of contigs indicate assembly completeness. BUSCO scores evaluate conserved gene content and provide a reference-free completeness assessment. When a reference genome is available, whole-genome alignment provides the most direct accuracy measurement. Coverage uniformity analysis identifies potential misassemblies or collapsed regions. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference genomes and sequence databases that support comparative quality evaluation.

### Can I use this pipeline for metagenomic or plastid genome assembly?

The standard pipeline is designed for single-genome assembly and may require modification for metagenomic or organelle genome applications. The MetaBooster and MetaBooster-HiFi pipelines provide strain-aware assembly for metagenomic data, outperforming general-purpose assemblers in genome fraction, contig length, and error rates. The ptGAUL pipeline provides a specialized approach for plastid genome assembly using long-read data, producing accurate and unbiased assemblies with approximately 50x coverage. For these specialized applications, use the dedicated pipelines instead of adapting the standard workflow.

### What computational resources do I need for long-read assembly?

Resource requirements depend on genome size, coverage, and the specific tools used. Bacterial genomes can be assembled on a standard workstation with 16 threads and 32 gigabytes of memory. Larger genomes require substantially more resources, with mammalian genomes potentially requiring hundreds of gigabytes of memory and multiple days of compute time. The GoldPolish-Target study demonstrated that targeted polishing consumes 95% less memory than whole-genome polishing, making it suitable for large genomes on modest infrastructure. Consider cloud computing or institutional high-performance computing resources for large or complex genomes.

## Related Bioinformatics Guides

- [Long-Read Genome Assembly and Polishing Strategies](/knowledge/bioinformatics/long-read-genome-assembly-and-polishing-strategies)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Metagenomics Pipeline: From Raw Reads to Taxonomic and Functional Profiles](/knowledge/bioinformatics/metagenomics-pipeline-from-raw-reads-to-taxonomic-and-functional-profiles)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection.](https://pubmed.ncbi.nlm.nih.gov/39862869). American journal of human genetics, 2025.
- [Plastid Genome Assembly Using Long-read data.](https://pubmed.ncbi.nlm.nih.gov/36939021). Molecular ecology resources, 2023.
- [GoldPolish-target: targeted long-read genome assembly polishing.](https://pubmed.ncbi.nlm.nih.gov/40055584). BMC bioinformatics, 2025.
- [Enhancing Long-Read-Based Strain-Aware Metagenome Assembly.](https://pubmed.ncbi.nlm.nih.gov/35646097). Frontiers in genetics, 2022.
- [Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection.](https://pubmed.ncbi.nlm.nih.gov/39228712). medRxiv : the preprint server for health sciences, 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.