PacBio CLR Sequencing: Mechanism, Workflow, and Applications
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to PacBio CLR Sequencing
Continuous long read (CLR) sequencing is a mode of single-molecule real-time (SMRT) sequencing developed by Pacific Biosciences (PacBio) that produces reads spanning tens to hundreds of kilobases. In CLR mode, the sequencing instrument captures fluorescence signals continuously as a DNA polymerase incorporates nucleotides into a growing complementary strand, generating a single long read per DNA molecule without the circular consensus step that defines PacBio's higher-accuracy HiFi sequencing. CLR reads are characterized by their exceptional length—routinely exceeding 20 kb and frequently surpassing 100 kb—but carry a per-base error rate of approximately 10–15%, distributed randomly across the read.
What is CLR Sequencing?
CLR sequencing represents the original data collection mode of PacBio's Sequel and RS II platforms. During a CLR run, the polymerase processively synthesizes DNA for as long as the enzyme remains active and the template remains intact. The instrument records the entire polymerase activity as a single continuous read, which may span the full length of the template molecule. Because the template is a linear double-stranded DNA molecule with hairpin adapters, the polymerase can theoretically circle the template multiple times; however, in CLR mode, the sequencing software does not collapse these passes into a consensus sequence. Instead, each pass is treated as part of the same long read, and the raw signal is reported directly.
The key distinction from HiFi sequencing lies in the data processing. In HiFi mode, the polymerase's multiple passes around a circular template are computationally identified and collapsed into a single high-accuracy consensus read. In CLR mode, the read is reported as the continuous polymerase trace, including all passes, without consensus correction. This yields reads that are longer than HiFi reads (which are typically capped at 15–25 kb by the circular consensus approach) but with substantially lower per-base accuracy.
CLR vs. HiFi Reads
The choice between CLR and HiFi sequencing involves a trade-off between read length and accuracy. CLR reads offer the longest possible single-molecule reads, which are invaluable for resolving complex genomic structures, spanning repetitive elements, and assembling genomes with high contiguity. However, the 10–15% error rate necessitates either deep coverage for error correction or specialized algorithms that can tolerate noisy reads.
HiFi reads, by contrast, achieve >99.9% accuracy (Q30 or better) by generating a circular consensus sequence from multiple passes of the polymerase around a shorter template. The trade-off is read length: HiFi reads are typically 10–25 kb, limited by the need to keep the template short enough for the polymerase to complete multiple passes before the enzyme loses activity. For applications requiring both long-range information and high per-base accuracy—such as variant calling in repetitive regions—HiFi is generally preferred. For applications where maximum read length is paramount, such as closing gaps in complex assemblies or characterizing ultra-long structural variants, CLR remains a viable option.
Underlying Technology: SMRT Cells and Zero-Mode Waveguides
The physical foundation of PacBio sequencing is the SMRT cell, a consumable chip containing hundreds of thousands of zero-mode waveguides (ZMWs). Each ZMW is a nanophotonic confinement structure that enables the detection of single fluorescent molecules in real time.
Zero-Mode Waveguides
A ZMW is a nanoscale well, approximately 100 nanometers in diameter and 100 nanometers deep, fabricated in a metal film deposited on a glass substrate. The diameter of the well is smaller than the wavelength of visible light, which creates an optical phenomenon known as the zero-mode waveguide effect. When light is directed at the bottom of the well, the evanescent field decays exponentially with distance from the glass surface, creating an illumination volume of only approximately 20 zeptoliters (20 × 10⁻²¹ liters). This extraordinarily small detection volume is critical because it confines the observation window to a single polymerase molecule.
The polymerase enzyme is immobilized at the bottom of each ZMW via a biotin–streptavidin linkage. The template DNA, prepared with a biotinylated adapter, is bound to the polymerase before loading. Because the illumination volume is so small, background fluorescence from unincorporated nucleotides in the bulk solution is effectively excluded; only nucleotides within the active site of the polymerase—where incorporation occurs—remain in the detection volume long enough to generate a measurable signal.
Real-Time Nucleotide Incorporation
The four deoxyribonucleotide triphosphates (dNTPs) used in SMRT sequencing are each labeled with a distinct fluorophore attached to the terminal phosphate group. When a labeled nucleotide enters the polymerase active site and is incorporated into the growing DNA strand, the fluorophore is held in the ZMW detection volume for the duration of the incorporation event, typically tens of milliseconds. During this time, the instrument excites the fluorophore with a laser and records the emitted fluorescence at the characteristic wavelength for that nucleotide.
After incorporation, the polymerase cleaves the terminal phosphate linkage, releasing the fluorophore along with the pyrophosphate group. The fluorophore diffuses out of the detection volume, and the fluorescence signal returns to baseline. The instrument records the sequence of fluorescence pulses, with each pulse corresponding to a single nucleotide incorporation event. The time between pulses—the interpulse duration—reflects the kinetics of the polymerase, which can be influenced by DNA modifications such as methylation or damage.
The key advantage of this real-time detection scheme is that it observes the polymerase in action, providing both sequence information and kinetic information. The latter can be used to detect base modifications, although this capability is more commonly exploited in specialized applications than in standard CLR sequencing.
Library Preparation for CLR Sequencing
Library preparation for CLR sequencing is designed to produce long, intact DNA molecules with the necessary adapters for polymerase binding and SMRT cell loading. The workflow differs from short-read library preparation in its emphasis on minimizing DNA damage and avoiding size-selection steps that would discard long fragments.
DNA Shearing and Size Selection
High-molecular-weight (HMW) genomic DNA is the starting material for CLR library preparation. The DNA must be extracted using methods that minimize mechanical shearing, such as agarose plug lysis or gentle phenol-chloroform extraction followed by dialysis. Once extracted, the DNA is sheared to a target size range using either hydrodynamic shearing (e.g., g-TUBE, Covaris) or needle shearing. For CLR sequencing, the target fragment size is typically 20–50 kb, although fragments up to 100 kb or more can be sequenced if the library preparation is optimized.
Size selection is performed using either pulsed-field gel electrophoresis or a BluePippin system, which uses electrophoresis to select DNA fragments within a defined size range. The size-selection step is critical because it removes short fragments that would otherwise dominate the sequencing output, wasting SMRT cell capacity on reads that are too short to provide useful long-range information. For CLR sequencing, a typical size-selection window might be 20–40 kb, with the lower cutoff chosen to exclude fragments below 10 kb.
Damage Repair and End Repair
DNA damage, particularly nicks and oxidized bases, poses a significant challenge for long-read sequencing because the polymerase will stall or dissociate when it encounters a damaged site. To address this, the library preparation includes a damage repair step using a cocktail of enzymes: DNA repair enzymes such as Escherichia coli endonuclease VIII and T4 pyrimidine dimer glycosylase remove oxidized bases, while T4 DNA polymerase and T4 polynucleotide kinase repair nicks and dephosphorylate the 5' ends.
Following damage repair, the DNA undergoes end repair to create blunt ends with 5' phosphate groups. This is achieved using T4 DNA polymerase, which has both 3'→5' exonuclease and 5'→3' polymerase activity, in the presence of dNTPs. The reaction is typically performed at 20°C for 30 minutes to balance exonuclease and polymerase activities, yielding blunt-ended, phosphorylated fragments suitable for adapter ligation.
Adapter Ligation and Annealing
The SMRTbell adapter is a hairpin structure—a single-stranded oligonucleotide that folds back on itself to form a double-stranded stem with a loop at one end. The stem contains the sequences required for polymerase binding and primer annealing, while the loop connects the two ends of the insert DNA. Ligation of SMRTbell adapters to both ends of the insert creates a circular template: the insert is flanked by adapter sequences on both sides, and the hairpin loops connect the top and bottom strands.
Adapter ligation is performed using T4 DNA ligase in a reaction buffer containing ATP and polyethylene glycol (PEG) to promote intermolecular ligation. The molar ratio of adapter to insert is typically 10:1 to 20:1 to ensure that both ends of each insert receive an adapter and to minimize concatemer formation. After ligation, the reaction is treated with exonuclease III and exonuclease VII to digest any DNA molecules that lack adapters on both ends, enriching for complete SMRTbell templates.
The final library is annealed with a sequencing primer that is complementary to the adapter sequence, and the primer-annealed template is bound to the polymerase. The polymerase–template complex is then purified to remove excess polymerase and primer, and the complex is ready for loading onto SMRT cells. For a detailed comparison of library preparation approaches across sequencing platforms, see Library Prep in Sequencing.
Sequencing Run: From Primer Annealing to Base Calling
The sequencing run converts the prepared library into raw sequence data through a series of coordinated steps involving polymerase binding, SMRT cell loading, optical imaging, and computational base calling.
Primer Annealing and Polymerase Binding
The sequencing primer is a 20–30 nucleotide oligonucleotide complementary to the adapter sequence adjacent to the insert. Annealing is performed by heating the library to 80°C for 2 minutes, then slowly cooling to 37°C over 30 minutes in the presence of excess primer. This thermal ramp ensures that the primer anneals specifically to the adapter sequence without disrupting the hairpin structure.
The annealed template is then incubated with the DNA polymerase at a molar ratio of approximately 1:10 (template to polymerase) to ensure that each template molecule is bound by a polymerase. The polymerase used in SMRT sequencing is a modified phi29 DNA polymerase, engineered for high processivity and strand displacement activity. The polymerase binds to the primer–template junction and forms a stable complex that can be stored at 4°C for several hours before loading.
Loading and Imaging
The polymerase–template complexes are loaded onto SMRT cells by diffusion. A SMRT cell contains approximately 1 million ZMWs, but only a fraction of these will contain a functional polymerase–template complex after loading. The loading density is optimized to maximize the number of ZMWs with a single active complex while minimizing the number of ZMWs with multiple complexes, which would produce mixed signals.
Once loaded, the SMRT cell is placed in the sequencing instrument, where it is maintained at a controlled temperature (typically 30°C for CLR runs) and supplied with a continuous flow of sequencing reagents. The reagent mix contains the four labeled dNTPs, each at a concentration of approximately 100 nM, along with cofactors such as magnesium ions and a phosphodiesterase to cleave the fluorophore–pyrophosphate product.
The instrument illuminates the ZMWs with a laser and captures fluorescence emission using a CCD or CMOS camera. The imaging system records a movie of the SMRT cell, capturing frames at a rate of approximately 100 frames per second. Each ZMW produces a time trace of fluorescence intensity at four wavelengths, corresponding to the four nucleotides. The sequencing run continues for a user-defined duration, typically 10–20 hours for CLR mode, during which the polymerase processively synthesizes DNA until it either dissociates from the template or encounters a blocking lesion.
Base Calling and Quality Scores
Base calling is performed in real time by the instrument's software, which analyzes the fluorescence traces from each ZMW. The algorithm identifies pulses—transient increases in fluorescence at a specific wavelength—and assigns each pulse to a nucleotide. The key challenge is distinguishing true incorporation events from noise and from transient binding of nucleotides that do not lead to incorporation.
The base calling algorithm uses a hidden Markov model that accounts for the expected kinetics of the polymerase: a pulse must exceed a minimum intensity threshold and persist for a minimum duration to be called as an incorporation event. The algorithm also assigns a quality score to each base, reflecting the confidence in the call. For CLR reads, the quality scores are typically low (Q10–Q20, corresponding to 90–99% accuracy per base) because each base is called from a single observation, without the consensus correction used in HiFi mode.
The output of the sequencing run is a set of raw reads in BAM format, each with a read name, sequence, and per-base quality scores. The read length distribution reflects the polymerase processivity and the template length; for a well-prepared CLR library, the median read length is typically 15–30 kb, with a substantial fraction of reads exceeding 50 kb.
Error Profile and Accuracy of CLR Reads
The error profile of CLR reads is fundamentally different from that of short-read sequencing technologies. Understanding this error profile is essential for selecting appropriate analysis tools and interpreting results.
Error Rate and Types
CLR reads have an overall error rate of approximately 10–15% per base, but the distribution of errors is not uniform across the read. The dominant error type is insertion–deletion (indel) errors, which account for roughly 80% of all errors. These errors arise from the polymerase incorporating a nucleotide that is not immediately detected, or from the detection algorithm misidentifying a pulse. Substitution errors are less common, accounting for the remaining 20% of errors.
The key feature of CLR errors is their randomness. Unlike sequencing-by-synthesis platforms where errors are often biased by sequence context (e.g., homopolymers or GC-rich regions), CLR errors are largely stochastic. This means that at any given position, the probability of an error is approximately constant, and errors are not systematically biased toward specific sequence motifs. This property is exploited by error correction algorithms, which align multiple reads to the same genomic region and use the consensus to eliminate errors.
The error rate is not constant across the read. The first 50–100 bases of a CLR read often have a higher error rate due to the polymerase establishing processive synthesis and the detection algorithm calibrating to the signal. Similarly, the last few hundred bases of a read may have elevated error rates as the polymerase begins to lose processivity and the signal-to-noise ratio degrades.
Impact on Downstream Analysis
The high error rate of CLR reads has profound implications for downstream analysis. For genome assembly, the errors must be corrected either by using the reads themselves (self-correction) or by using high-accuracy short reads as a scaffold. The random nature of the errors means that with sufficient coverage—typically 50–100× for CLR-only assembly—the consensus sequence can be determined with high accuracy.
For read mapping, the high error rate requires aligners that can tolerate indels and mismatches. Traditional short-read aligners such as BWA-MEM and Bowtie2 are not designed for reads with 10–15% error rates and will fail to align a substantial fraction of CLR reads. Instead, long-read aligners such as minimap2, NGMLR, and BLASR are required. These aligners use seed-and-extend strategies with longer seeds and more permissive scoring matrices to accommodate the error profile.
For variant calling, the high error rate means that individual CLR reads cannot be used to call variants with confidence. Instead, variants are called from the consensus of multiple aligned reads, or from the assembly graph. Structural variants, which involve large-scale genomic rearrangements, are more amenable to CLR detection because the breakpoints can be identified from the alignment of long reads even in the presence of high base-level error.
Bioinformatics Analysis of CLR Data
The computational analysis of CLR data requires specialized tools designed to handle long, error-prone reads. The analysis pipeline typically involves error correction, assembly or mapping, and downstream variant detection.
Error Correction Strategies
Error correction of CLR reads can be performed using either a self-correction approach or a hybrid approach. Self-correction, also known as pre-assembly correction, aligns the CLR reads to each other and uses the multiple alignments to generate a consensus sequence for each read. Tools such as Canu and Falcon implement this approach, using the overlap information to identify and correct errors. The corrected reads are then used as input to the assembly algorithm.
Hybrid error correction uses high-accuracy short reads (e.g., Illumina reads) to correct the CLR reads. The short reads are mapped to the CLR reads, and the consensus of the short-read alignments is used to correct the long-read sequence. Tools such as Pilon and LoRDEC implement this approach. Hybrid correction is computationally efficient but requires the additional cost of generating short-read data.
For CLR-only assembly, the error correction step is integrated into the assembly algorithm. Canu, for example, performs error correction as the first stage of its pipeline, using the overlap graph to identify and correct errors before the assembly stage. The corrected reads are then assembled using a string graph or overlap-layout-consensus approach.
Genome Assembly with CLR Reads
The assembly of CLR reads into a genome sequence is performed using long-read assemblers that can handle the high error rate. The two main approaches are the overlap-layout-consensus (OLC) approach and the string graph approach.
The OLC approach, implemented in Canu and Falcon, begins by computing all pairwise overlaps between reads. The overlaps are used to build an overlap graph, in which nodes represent reads and edges represent overlaps. The graph is then simplified by removing redundant edges and resolving repeats, and the remaining paths through the graph represent contigs. The final step is consensus generation, in which the multiple sequence alignment of reads along each path is used to determine the consensus sequence.
The string graph approach, implemented in Flye and Miniasm, uses a similar principle but represents the assembly as a graph of sequences rather than reads. Flye, in particular, uses a repeat graph construction that is designed to handle the high error rate of CLR reads and produces highly contiguous assemblies.
For bacterial genomes, CLR sequencing can produce a complete genome in a single contig with sufficient coverage. For larger genomes, the assembly typically produces hundreds to thousands of contigs, which can be further scaffolded using optical mapping or Hi-C data. The Pacbio Only Bacterial Sequencing approach is a common application of this technology.
Variant Calling and Structural Variant Detection
Variant calling from CLR reads is challenging due to the high error rate, but it is feasible for structural variants and for single-nucleotide variants (SNVs) when sufficient coverage is available. For SNV calling, the reads are aligned to the reference genome, and the allele frequency at each position is estimated from the aligned reads. Because the error rate is random, a variant allele present at a frequency of 50% (heterozygous) can be distinguished from sequencing error with sufficient coverage—typically 30–50× for CLR reads.
Structural variant detection is a major strength of CLR sequencing. The long reads can span entire structural variants, allowing the breakpoints to be identified directly from the alignment. Tools such as Sniffles and pbsv use the alignment information to identify deletions, insertions, duplications, inversions, and translocations. The long read length also enables the detection of structural variants in repetitive regions that are inaccessible to short-read sequencing.
Applications and Use Cases
CLR sequencing is used in a range of applications where read length is the primary consideration. The technology is particularly well suited to de novo genome assembly, full-length transcript sequencing, and metagenomics.
De Novo Genome Assembly
The primary application of CLR sequencing is de novo genome assembly. The long reads can span repetitive elements, resolve complex genomic regions, and produce highly contiguous assemblies. For bacterial genomes, a single SMRT cell can provide sufficient coverage to assemble the genome into a single contig. For eukaryotic genomes, CLR sequencing produces assemblies with contig N50 values in the megabase range, far exceeding what is achievable with short-read sequencing alone.
The assembly of plant and animal genomes has been transformed by CLR sequencing. The ability to span long repeats and segmental duplications has enabled the assembly of complex genomes that were previously intractable. The Sequencing Coverage required for CLR assembly is typically 50–100×, depending on the genome size and complexity.
Full-Length Transcript Sequencing
CLR sequencing can be used to sequence full-length cDNA molecules, providing information about transcript isoforms that is lost in short-read RNA-seq. The workflow involves reverse transcription of mRNA to cDNA, followed by amplification and SMRTbell library preparation. The resulting reads span the full length of the transcript, allowing the identification of alternative splicing isoforms, alternative transcription start sites, and polyadenylation sites.
The Pacbio cDNA Sequencing Genewiz service is an example of a commercial offering that uses PacBio technology for transcript sequencing. The long reads provide isoform-level resolution that is not achievable with short-read RNA-seq, which must infer isoforms from exon junction reads.
Metagenomics and Microbial Genomics
CLR sequencing is well suited to metagenomics, where the goal is to assemble genomes from complex microbial communities. The long reads can span entire genes and operons, enabling the assembly of complete microbial genomes from metagenomic samples. This is particularly valuable for identifying novel species and characterizing the functional potential of microbial communities.
For microbial genomics, CLR sequencing can produce complete genomes for bacteria and archaea, including the resolution of repetitive elements such as rRNA operons and insertion sequences. The Pacbio Only Bacterial Sequencing approach is a cost-effective strategy for generating complete bacterial genomes.
Common Pitfalls and Practical Considerations
CLR sequencing is a powerful technology, but it has specific requirements and failure modes that must be understood to obtain high-quality data.
DNA Quality and Quantity
The most common cause of CLR sequencing failure is poor DNA quality. The polymerase is highly sensitive to DNA damage, and nicks, oxidized bases, and abasic sites will cause the polymerase to stall or dissociate. DNA should be extracted using methods that minimize shearing and oxidation, such as agarose plug lysis or gentle phenol-chloroform extraction. The DNA should be assessed for integrity using pulsed-field gel electrophoresis or a fragment analyzer before library preparation.
The quantity of DNA required for CLR sequencing depends on the library preparation method and the number of SMRT cells to be sequenced. A typical library preparation requires 5–10 μg of HMW DNA, and each SMRT cell requires approximately 500 ng to 1 μg of library. For a standard CLR run with 8 SMRT cells, a total of 10–20 μg of DNA is recommended. The Prepare Sample for Sequencing guide provides general guidance on sample preparation.
Coverage and Cost
The high error rate of CLR reads necessitates deep coverage for most applications. For de novo assembly, 50–100× coverage is recommended, which for a bacterial genome (5 Mb) corresponds to 250–500 Mb of sequence data. For a human genome (3 Gb), 50× coverage corresponds to 150 Gb of sequence data, which requires approximately 15–30 SMRT cells at current throughput.
The cost of CLR sequencing is higher than short-read sequencing on a per-base basis, but the long reads can reduce the overall cost of a project by eliminating the need for additional scaffolding or gap-closing experiments. For bacterial genome assembly, a single SMRT cell can produce a complete genome, making CLR sequencing cost-competitive with short-read sequencing for this application.
Data Analysis Pitfalls
A common pitfall in CLR data analysis is using inappropriate tools. Short-read aligners and assemblers will fail with CLR reads, producing poor results or crashing entirely. It is essential to use tools designed for long-read data, such as minimap2 for alignment and Canu or Flye for assembly.
Another pitfall is insufficient coverage for error correction. If the coverage is too low, the error correction step will not be able to distinguish true sequence from errors, resulting in a fragmented assembly or incorrect consensus sequence. It is important to calculate the expected coverage before starting a CLR sequencing project and to adjust the number of SMRT cells accordingly.
Finally, the interpretation of CLR data requires an understanding of the error profile. Variants called from individual CLR reads should be treated with caution, and structural variant calls should be validated by visual inspection of the alignments or by orthogonal methods such as PCR or optical mapping.
Frequently Asked Questions
What is PacBio CLR sequencing?
PacBio CLR (continuous long read) sequencing is a single-molecule real-time sequencing mode that produces reads of tens to hundreds of kilobases. It uses a DNA polymerase immobilized in a zero-mode waveguide to incorporate fluorescently labeled nucleotides, with the fluorescence signal recorded in real time. CLR reads have a high error rate (10–15%) but provide the longest reads available from any sequencing platform.
How does CLR sequencing differ from HiFi sequencing?
CLR sequencing reports the continuous polymerase trace without consensus correction, producing long reads with high error rates. HiFi sequencing uses the same SMRT technology but generates a circular consensus sequence from multiple passes of the polymerase around a shorter template, producing reads of 10–25 kb with >99.9% accuracy. The choice between CLR and HiFi depends on whether read length or accuracy is the priority.
What is the error rate of PacBio CLR reads?
The error rate of CLR reads is approximately 10–15% per base, with insertion–deletion errors being the most common type. The errors are randomly distributed across the read, which allows them to be corrected by consensus when sufficient coverage is available.
What are the main applications of CLR sequencing?
The main applications are de novo genome assembly, full-length transcript sequencing, and metagenomics. CLR sequencing is also used for structural variant detection and for closing gaps in reference genomes.
How much DNA is needed for PacBio CLR sequencing?
A typical CLR library preparation requires 5–10 μg of high-molecular-weight DNA. Each SMRT cell requires approximately 500 ng to 1 μg of library, so a standard run with 8 SMRT cells requires 10–20 μg of DNA.
Can CLR reads be used for variant calling?
Yes, CLR reads can be used for variant calling, particularly for structural variants. For single-nucleotide variants, sufficient coverage (30–50×) is required to distinguish true variants from sequencing errors. The random error profile of CLR reads makes them suitable for consensus-based variant calling.
What are the common pitfalls in CLR data analysis?
Common pitfalls include using short-read analysis tools that cannot handle the error rate, insufficient coverage for error correction, and misinterpreting errors as true variants. It is essential to use long-read-specific tools and to understand the error profile when interpreting results.
Key Takeaways
- CLR sequencing produces the longest reads of any sequencing platform, routinely exceeding 20 kb and often surpassing 100 kb.
- The technology relies on zero-mode waveguides to observe single polymerase molecules in real time, with fluorescently labeled nucleotides detected as they are incorporated.
- CLR reads have a stochastic error rate of 10–15%, dominated by insertion–deletion errors, which can be corrected by consensus with sufficient coverage.
- Library preparation requires high-molecular-weight DNA, with damage repair, end repair, and SMRTbell adapter ligation as critical steps.
- CLR sequencing is the method of choice for de novo genome assembly, particularly for complex genomes with long repeats, and for structural variant detection.
- Analysis requires long-read-specific tools such as minimap2, Canu, and Flye; short-read tools are not suitable for CLR data.
- The main trade-off is between CLR (longest reads, high error) and HiFi (shorter reads, high accuracy); the choice depends on the application.
Further Reading
- Liu Y et al. Comparison of structural variants detected by PacBio-CLR and ONT sequencing in pear. BMC genomics. 2022. PubMed 36517766
- Wei ZG, Zhang SW. NPBSS: a new PacBio sequencing simulator for generating the continuous long reads with an empirical model. BMC bioinformatics. 2018. PubMed 29788930
- Yang H, Wang Y. From fragmentation to resolution: high-fidelity genome assembly of zancudomyces culisetae through comparative insights from PacBio, Nanopore, and Illumina sequencing. G3 (Bethesda, Md.). 2025. PubMed 40888030