# PacBio Only Bacterial Sequencing: Methods and Pitfalls

## Introduction to PacBio-Only Bacterial Sequencing

### What is PacBio-Only Sequencing?

PacBio-only bacterial sequencing refers to the complete workflow of generating, assembling, and analyzing a bacterial genome using exclusively single-molecule real-time (SMRT) sequencing data from Pacific Biosciences platforms, without supplementary Illumina short-read data. Unlike hybrid approaches that merge short and long reads, PacBio-only sequencing relies entirely on the long reads produced by SMRT technology to achieve both assembly contiguity and base-level accuracy.

The defining feature of this approach is that the same sequencing run provides three simultaneous data types: the primary nucleotide sequence, base modification information from polymerase kinetics, and coverage depth sufficient for de novo assembly. For bacterial genomes, which typically range from 1.5 to 12 Mb, a single SMRT cell on modern platforms can generate enough data to close a genome completely, producing a single circular contig with no gaps.

PacBio-only sequencing is distinct from [Pacbio CLR Sequencing](/knowledge/molecular-biology/pacbio-clr-sequencing) in that the modern workflow predominantly uses circular consensus sequencing (CCS) to generate high-fidelity (HiFi) reads, whereas older continuous long read (CLR) approaches produced lower-accuracy single-pass reads. The choice between these modes has profound implications for downstream assembly and analysis.

### Advantages Over Short-Read or Hybrid Approaches

The primary advantage of PacBio-only sequencing is operational simplicity combined with superior assembly outcomes. A complete bacterial genome assembled from PacBio-only data typically consists of one contig per replicon (chromosome and plasmids), with no gaps or ambiguous bases. Short-read assemblies, by contrast, invariably produce fragmented draft genomes with multiple contigs whose order and orientation are uncertain, particularly in regions containing repetitive elements such as rRNA operons, insertion sequences, and transposases.

Hybrid approaches that combine Illumina short reads with Oxford Nanopore or PacBio CLR reads can achieve complete assemblies, but they require two sequencing platforms, two library preparations, and more complex bioinformatic pipelines. PacBio-only HiFi sequencing eliminates this complexity: a single library, a single sequencing run, and a single assembly algorithm are sufficient.

A second major advantage is the simultaneous detection of DNA methylation. SMRT sequencing records the kinetics of nucleotide incorporation, and the presence of modified bases such as N6-methyladenine (m6A), N4-methylcytosine (m4C), and 5-methylcytosine (m5C) alters these kinetic signatures. This means that a single PacBio run provides both the genome sequence and its epigenome, information that is biologically significant for understanding restriction-modification systems, gene regulation, and virulence in pathogenic bacteria. Hybrid approaches require a separate [Bisulfite Sequencing](/knowledge/molecular-biology/bisulfite-sequencing) experiment to obtain equivalent methylation information, and bisulfite conversion cannot detect m6A at all.

## PacBio Sequencing Technology and Data Characteristics

### SMRT Sequencing Chemistry

SMRT sequencing occurs in zero-mode waveguides (ZMWs), which are nanophotonic wells approximately 70 nm in diameter and 100 nm deep, fabricated in a metal film on a glass substrate. Each ZMW confines light to a detection volume of approximately 20 zeptoliters (20 × 10⁻²¹ L), allowing the observation of single fluorescent molecules against a low background.

The sequencing reaction uses a DNA polymerase (typically a modified phi29 DNA polymerase) immobilized at the bottom of each ZMW. A single SMRTbell template—a closed, hairpin-adapter-flanked double-stranded DNA molecule—diffuses into the ZMW, and the polymerase begins rolling-circle replication. As each nucleotide is incorporated, a fluorescently labeled phosphate-linked nucleotide is held in the active site for tens of milliseconds, and the emitted fluorescence is detected in real time. The label is cleaved upon incorporation, and the polymerase continues processively around the circular template.

The key kinetic parameters recorded for each incorporation event are the inter-pulse duration (IPD), which is the time between successive nucleotide incorporations, and the pulse width, which is the duration of the fluorescence signal. These parameters are influenced by the local sequence context and, critically, by the presence of modified bases in the template. A methylated adenine, for example, causes a characteristic increase in IPD at the modified position and at the adjacent positions, providing the basis for methylation detection.

### Continuous Long Reads vs HiFi Reads

PacBio platforms can generate two fundamentally different types of reads, and understanding the distinction is essential for experimental design.

**Continuous long reads (CLR)** are the product of a single pass around the SMRTbell template. The polymerase reads the template once, producing a read whose length is limited by polymerase processivity and sequencing run time. CLR reads are typically 10–25 kb in length but have per-base accuracy of only 85–92%. This low accuracy is due to the stochastic nature of single-molecule fluorescence detection and the intrinsic error rate of the polymerase. CLR reads are useful for genome scaffolding and for spanning long repetitive regions, but they require substantial computational correction, typically through self-correction during assembly or polishing with additional data.

**HiFi reads** (also called circular consensus sequences or CCS reads) are generated by sequencing the same SMRTbell template multiple times. The polymerase reads around the circular template, and the subreads from each pass are computationally collapsed into a single consensus sequence. Because errors in SMRT sequencing are largely random rather than systematic, the consensus of multiple passes achieves high accuracy. A typical HiFi read of 10–15 kb length is generated from 10–15 passes and has an accuracy of ≥99.9% (Q30 or better). The trade-off is read length: because the polymerase must complete multiple passes, the maximum HiFi read length is shorter than the maximum CLR read length, and the sequencing time per ZMW is longer.

For bacterial genomes, HiFi reads are almost always the preferred choice. The Q30 accuracy means that assembly errors are rare, and the 10–15 kb read length is sufficient to span the vast majority of bacterial repetitive elements. CLR reads are primarily used when ultra-long reads (>25 kb) are needed to resolve very large repeat structures, or when the highest possible throughput per SMRT cell is required at the expense of accuracy.

### Error Profiles and Their Impact

The error profile of SMRT sequencing is distinctly different from that of [Illumina sequencing](/knowledge/diagnostics/molecular/illumina-sequencing-principle-chemistry-and-workflow). Illumina errors are predominantly substitution errors that occur at specific sequence contexts, such as homopolymers and GC-rich regions. SMRT sequencing errors, in contrast, are predominantly insertion and deletion (indel) errors that occur randomly along the read. Substitution errors are less common.

This error profile has important consequences. Because errors are random, increasing coverage depth reduces the error rate of the consensus sequence in a predictable manner. A single HiFi read at Q30 has an expected error rate of 1 in 1000 bases; the consensus of multiple independent HiFi reads can achieve Q50 or better. In contrast, systematic errors—those that occur at the same position in every read—cannot be corrected by depth alone. Fortunately, systematic errors in SMRT sequencing are rare, which is why HiFi consensus accuracy is so high.

The random indel error profile also means that alignment-based polishing tools must be indel-aware. Tools that assume substitution-dominated errors, such as those designed for Illumina data, will perform poorly on PacBio data. This is a critical consideration when choosing polishing software.

## Library Preparation for Bacterial Genomes

### DNA Quality and Quantity Requirements

The quality of the input DNA is the single most important determinant of PacBio sequencing success. The SMRTbell library preparation requires high-molecular-weight (HMW) DNA, and the maximum read length achievable is directly proportional to the integrity of the input DNA. For bacterial genomes, where the target read length is 10–15 kb, the input DNA should have a median size of at least 30–50 kb.

[DNA extraction methods](/knowledge/diagnostics/molecular/dna-extraction-methods-a-comparative-review-of-techniques-and-their-applications) must avoid mechanical shearing. Gentle lysis protocols using lysozyme (20 mg/mL in 10 mM Tris-HCl, pH 8.0, 1 mM EDTA) followed by proteinase K digestion (20 mg/mL at 56°C for 1–2 hours) and phenol-chloroform extraction or commercial kits designed for HMW DNA (e.g., Qiagen Genomic-tip 100/G) are recommended. Bead-beating and vigorous vortexing should be avoided, as these introduce random double-strand breaks.

The quantity requirement depends on the library preparation method. Standard SMRTbell library preparation requires 1–5 μg of DNA. Low-input protocols, such as the SMRTbell Express Template Prep Kit 2.0, can work with as little as 100 ng, but the yield of SMRTbell templates and the final sequencing output will be correspondingly lower. For bacterial samples where DNA is limiting, such as clinical isolates or unculturable organisms, the low-input protocol is the only viable option. The [Prepare Sample for Sequencing](/knowledge/molecular-biology/prepare-sample-for-sequencing) page provides general guidance on sample handling that applies here.

### SMRTbell Template Preparation

The SMRTbell template is a closed circular DNA molecule consisting of the double-stranded insert flanked by two hairpin adapters. The preparation involves several steps:

1. **Shearing**: Input DNA is sheared to the desired insert size using a Covaris g-TUBE or Megaruptor. For bacterial genomes, a target insert size of 15–20 kb is typical. Shearing is performed by centrifugation at 4,000–6,000 rpm for 60 seconds, repeated until the desired size distribution is achieved.

2. **Damage repair and end repair**: The sheared DNA contains damaged bases and non-blunt ends. The damage repair step uses a cocktail of enzymes including DNA polymerase I, T4 polynucleotide kinase, and T4 DNA polymerase to repair nicks, remove damaged bases, and generate blunt ends. The reaction is incubated at 37°C for 30 minutes.

3. **Adapter ligation**: The hairpin adapters, which contain a recognition sequence for the sequencing primer and a nicking site, are ligated to the blunt-ended DNA using T4 DNA ligase. The ligation reaction is incubated at 25°C for 30 minutes.

4. **Exonuclease digestion**: The ligation product is treated with a combination of exonuclease III and exonuclease VII to digest any linear DNA molecules that failed to ligate adapters on both ends. Only closed circular SMRTbell templates survive this step.

5. **Size selection**: The library is size-selected using the BluePippin system to remove small fragments (<5 kb) that would otherwise consume sequencing capacity without contributing useful read length. The size selection cutoff is typically set to 5–10 kb.

6. **Cleanup and quantification**: The final library is purified using AMPure PB beads and quantified using a Qubit fluorometer and a Bioanalyzer or Fragment Analyzer to confirm the size distribution.

### Multiplexing and Barcoding

Barcoding allows multiple bacterial samples to be sequenced on a single SMRT cell, reducing per-sample cost. The barcoded adapters contain a unique 16-nucleotide sequence that is ligated to the SMRTbell template. After sequencing, reads are assigned to samples based on the barcode sequence.

The number of samples that can be multiplexed depends on the required coverage per sample and the throughput of the platform. For a Sequel II system producing 30–40 Gb per SMRT cell, a typical bacterial genome of 5 Mb requires 500 Mb of HiFi data for 100× coverage. This means that 60–80 bacterial samples could theoretically be multiplexed on a single SMRT cell. In practice, however, the number is usually limited to 8–24 samples to account for variable sequencing output and to ensure that each sample receives sufficient coverage.

Barcode design must account for the error profile of SMRT sequencing. Barcodes should differ by at least 4 nucleotides to allow error-tolerant assignment, and they should not contain homopolymers or self-complementary regions that could form secondary structures. The standard PacBio barcode sets have been designed with these criteria in mind.

## Sequencing Run Design and Coverage Considerations

### Choosing Sequencing Platform

The choice of PacBio platform determines throughput, read length, and cost. The Sequel system (now largely obsolete) produced approximately 10 Gb per SMRT cell with read lengths up to 20 kb. The Sequel II system increased throughput to 30–40 Gb per SMRT cell with similar read lengths. The Revio system, introduced in 2022, represents the current generation, with a throughput of 60–90 Gb per SMRT cell and improved read length.

For bacterial sequencing, the Revio system is the platform of choice when available, as its higher throughput allows more samples to be multiplexed per run. However, the Sequel II remains a viable option, particularly for laboratories that already have the instrument. The choice between platforms also affects the cost per genome, with the Revio offering a lower cost per base.

An important consideration is that the Revio system uses a different SMRT cell format and requires different consumables. Laboratories transitioning from Sequel II to Revio must validate their library preparation protocols on the new platform, as the loading conditions and sequencing chemistry differ.

### Coverage Depth for Assembly and Methylation

The required coverage depth depends on the intended analysis. For de novo assembly of a bacterial genome from HiFi reads, a minimum of 30× coverage is recommended, and 50–100× is typical. Higher coverage provides several benefits: it improves the accuracy of the consensus sequence, it increases the likelihood of spanning repetitive regions, and it provides the depth needed for methylation detection.

For methylation analysis, the coverage requirement is more stringent. Methylation detection relies on the statistical comparison of kinetic signatures between modified and unmodified positions. At 30× coverage, the kinetic signal is noisy, and many methylation events will be missed. A coverage of 100× or higher is recommended for reliable methylation calling, particularly for m5C, which produces a weaker kinetic signal than m6A or m4C.

The relationship between coverage and assembly quality is not linear. Below 30×, assembly quality degrades rapidly, with an increasing number of misassemblies and unresolved repeats. Above 100×, the marginal benefit of additional coverage diminishes, and the cost becomes prohibitive. The optimal coverage for a combined assembly and methylation project is therefore 100–150×.

### Pooling and Throughput

Pooling multiple bacterial samples on a single SMRT cell requires careful calculation of expected yield. The sequencing output of a SMRT cell is variable, and the actual yield can range from 50% to 120% of the expected value. To ensure that each sample receives sufficient coverage, the pooling strategy should be conservative.

A typical pooling calculation proceeds as follows: for a 5 Mb genome at 100× coverage, the required data per sample is 500 Mb. For a Sequel II SMRT cell with an expected yield of 30 Gb, the maximum number of samples that can be pooled is 30 Gb / 0.5 Gb = 60 samples. However, to account for yield variability, the actual number should be reduced to 40–48 samples. For the Revio system with 60 Gb expected yield, the corresponding numbers are 120 and 80–96 samples.

The pooling strategy also affects the sequencing run time. HiFi sequencing requires a longer run time than CLR sequencing because the polymerase must complete multiple passes around each template. A typical HiFi run on the Sequel II takes 30 hours, while a CLR run takes 10–20 hours. The Revio system has a similar run time for HiFi sequencing.

## Base Calling and Demultiplexing

### Generating HiFi Reads with CCS

The generation of HiFi reads from raw SMRT sequencing data is performed by the CCS algorithm, which is implemented in the SMRT Link software suite. The CCS algorithm processes the raw sequencing data from each ZMW, identifies the subreads corresponding to each pass around the SMRTbell template, and generates a consensus sequence.

The CCS algorithm operates as follows:

1. **Subread identification**: The raw polymerase read is segmented into subreads based on the adapter sequences that flank the insert. Each subread corresponds to one pass around the template.

2. **Subread alignment**: The subreads are aligned to each other using a partial order alignment algorithm. This step identifies the consensus sequence and the error profile.

3. **Consensus generation**: The aligned subreads are collapsed into a single consensus sequence using a hidden Markov model that accounts for the error profile of SMRT sequencing.

4. **Quality score assignment**: Each base in the consensus sequence is assigned a quality score (Q score) based on the number of subreads supporting it and the agreement among subreads.

The CCS algorithm has several parameters that can be adjusted. The most important are the minimum number of passes (default 3) and the minimum predicted accuracy (default 0.99, corresponding to Q20). Increasing the minimum number of passes produces longer, more accurate reads but reduces the yield, as ZMWs with fewer passes are discarded. For bacterial sequencing, a minimum of 5 passes and a minimum predicted accuracy of 0.99 (Q20) or 0.999 (Q30) is recommended.

The output of CCS is a BAM file containing HiFi reads with quality scores. The __MASK_4__ page describes a related workflow for cDNA, which shares the CCS step.

### Demultiplexing and Quality Filtering

Demultiplexing assigns HiFi reads to their sample of origin based on the barcode sequence. The demultiplexing algorithm in SMRT Link uses the barcode sequences to identify the sample, and it can tolerate a small number of mismatches (typically 1–2) to account for sequencing errors.

The demultiplexing process also removes adapter sequences and identifies chimeric reads. A chimeric read is a read that contains sequences from two different templates, typically arising from template switching during polymerase replication. Chimeric reads are identified by the presence of two different barcodes or by internal adapter sequences, and they are removed from the dataset.

Quality filtering is performed after demultiplexing. Reads with a predicted accuracy below the threshold (typically Q20 or Q30) are removed. Additionally, reads that are shorter than a minimum length (typically 1 kb) are removed, as they are unlikely to contribute useful information to the assembly.

The final output of the demultiplexing and quality filtering step is a set of FASTA or FASTQ files, one per sample, containing high-quality HiFi reads. These files are the input to the assembly step.

## Assembly of PacBio-Only Bacterial Genomes

### Choosing an Assembler

Several assemblers are available for PacBio-only bacterial genome assembly, and the choice depends on the read type (HiFi vs CLR) and the desired output.

**Canu** is a long-read assembler that was originally designed for CLR reads but has been adapted for HiFi reads. It uses an overlap-layout-consensus approach, where reads are first overlapped to identify their relationships, then laid out to form contigs, and finally consensus-called. Canu is robust and well-tested, but it is computationally intensive and can be slow for large datasets.

**Flye** is a de Bruijn graph-based assembler that is particularly well-suited for bacterial genomes. It constructs a repeat graph from the reads, resolves repeats using the read connectivity, and outputs contigs. Flye is fast and memory-efficient, and it handles HiFi reads well. For bacterial genomes, Flye typically produces a single circular contig per replicon.

**Hifiasm** is a haplotype-resolved assembler that was designed specifically for HiFi reads. It uses a string graph approach and is capable of resolving haplotypes in diploid organisms. For haploid bacterial genomes, Hifiasm produces a single contig per replicon with high accuracy. Hifiasm is the fastest of the three assemblers for HiFi data.

**Unicycler** is a hybrid assembler that was designed for short-read data but has been adapted to use long reads. It uses a short-read assembly as a scaffold and then uses long reads to resolve repeats and close gaps. For PacBio-only data, Unicycler is not the optimal choice, as it does not fully exploit the information in long reads. However, it can be used as a cross-check on the assembly produced by other tools.

For most bacterial PacBio-only projects, **Flye** or **Hifiasm** is the recommended assembler. Flye is more forgiving of parameter misconfiguration, while Hifiasm produces slightly more accurate assemblies at the cost of being more sensitive to read quality.

### Circularization and Polishing

Bacterial chromosomes and most plasmids are circular molecules. The assembly of a circular genome from long reads should produce a single contig whose ends overlap. Circularization is the process of identifying this overlap and joining the ends to form a circular sequence.

Most assemblers automatically circularize contigs when the overlap is detected. However, the circularization is not always perfect, and manual curation may be required. The circlator tool can be used to circularize contigs and to rotate the sequence to a defined starting position.

Polishing is the process of correcting residual errors in the assembly. For HiFi reads, the assembly is already highly accurate, but polishing can improve the consensus quality further. The two most commonly used polishers are:

**Racon** is a fast consensus tool that uses partial order alignment to correct errors. It can be run with the original reads as input, and it is typically run for 2–3 rounds. Racon is fast and memory-efficient, making it suitable for bacterial genomes.

**Medaka** is a neural network-based polisher that was originally developed for Oxford Nanopore data but has been adapted for PacBio data. It is more accurate than Racon but slower. For bacterial genomes, the accuracy difference is small, and Racon is usually sufficient.

The polishing step is essential for CLR assemblies, where the raw read accuracy is only 85–92%. For HiFi assemblies, polishing is optional but recommended, as it can correct the few remaining errors and improve the quality of the final genome.

### Assessing Assembly Quality

The quality of a bacterial genome assembly should be assessed using multiple metrics:

**Contiguity**: The number of contigs and their sizes. A complete bacterial genome should have one contig per replicon. The N50 statistic, which is the length of the contig at which 50% of the assembly is contained, should be at least the size of the largest replicon.

**Completeness**: The proportion of the genome that is assembled. This can be assessed using BUSCO (Benchmarking Universal Single-Copy Orthologs), which checks for the presence of a set of conserved genes. A complete assembly should have >99% of BUSCO genes present and complete.

**Accuracy**: The base-level accuracy of the assembly. This can be assessed by mapping the reads back to the assembly and counting mismatches and indels. A high-quality assembly should have >99.99% identity between the reads and the assembly.

**Circularity**: The assembly should be circular for bacterial chromosomes and plasmids. This can be confirmed by checking that the ends of the contig overlap and that the sequence is circularly closed.

The QUAST tool provides a comprehensive summary of assembly quality metrics and is recommended for routine assessment.

## Methylation Analysis from PacBio Data

### Kinetic Information and Methylation Detection

The SMRT sequencing process records the kinetics of nucleotide incorporation for every position in the template. The two key kinetic parameters are the inter-pulse duration (IPD) and the pulse width. Modified bases alter these parameters in a characteristic manner, providing a direct readout of DNA methylation.

The detection of methylation from SMRT data is performed by comparing the observed IPD at each position to the expected IPD for an unmodified template with the same sequence context. The ratio of observed to expected IPD is called the IPD ratio, and a significant deviation from 1.0 indicates the presence of a modified base.

The three most common bacterial methylation types produce distinct kinetic signatures:

**N6-methyladenine (m6A)**: This modification produces a strong increase in IPD at the modified adenine and at the adjacent positions. The kinetic signal is robust and can be detected at relatively low coverage (30×).

**N4-methylcytosine (m4C)**: This modification produces a moderate increase in IPD at the modified cytosine and adjacent positions. The signal is weaker than for m6A but still detectable at moderate coverage (50×).

**5-methylcytosine (m5C)**: This modification produces a weak kinetic signal that is difficult to distinguish from background noise. Detection of m5C requires high coverage (100×) and careful [statistical analysis](/blog/guides/statistical-analysis). For this reason, [Bisulfite Sequencing](/knowledge/molecular-biology/bisulfite-sequencing) remains the gold standard for m5C detection, and PacBio-based m5C calling should be validated by an orthogonal method.

The kinetic information is extracted from the raw sequencing data during base calling. The SMRT Link software produces a modified base annotation file that lists the positions of putative modifications and their kinetic scores.

### Motif Analysis and Visualization

Once modified positions are identified, the next step is to determine the sequence motif recognized by the methyltransferase. This is performed using the MotifMaker tool, which is part of the SMRT Analysis suite.

MotifMaker operates as follows:

1. **Motif discovery**: The modified positions are analyzed to identify common sequence contexts. The tool searches for motifs of 4–8 bases that are enriched at modified positions.

2. **Motif refinement**: The identified motifs are refined to determine the exact recognition sequence and the position of the modified base within the motif.

3. **Motif assignment**: The refined motifs are compared to known methyltransferase recognition sequences to identify the responsible enzymes.

The output of MotifMaker is a list of motifs, the number of occurrences in the genome, the fraction of occurrences that are modified, and the predicted methyltransferase.

The methylation data can be visualized using the SMRT Link Genome Browser or exported to a BED file for viewing in other genome browsers. The visualization should show the position of modified bases relative to the motif and the fraction of sites that are modified.

A common cause of methylation calling artifacts is the presence of sequence variants. A single nucleotide polymorphism (SNP) at a motif position can abolish methylation, and the kinetic signature at that position will be different from the expected pattern. It is important to check that the methylation calls are consistent across the genome and that the motifs are not located in regions of sequence variation.

## Common Pitfalls and Troubleshooting

### Insufficient Coverage

The most common cause of failed PacBio-only bacterial sequencing projects is insufficient coverage. This can arise from several sources: over-pooling of samples, lower-than-expected SMRT cell yield, or inefficient library preparation.

The symptoms of insufficient coverage are fragmented assemblies, incomplete circularization, and poor methylation detection. The assembly may produce multiple contigs that cannot be joined, or the circularization step may fail because the ends of the chromosome are not covered by reads.

To troubleshoot insufficient coverage, first check the yield of the sequencing run. If the yield is lower than expected, the problem may be in the library preparation (e.g., low SMRTbell yield, excessive short fragments) or in the loading of the SMRT cell. If the yield is adequate but the coverage per sample is low, the problem is likely over-pooling. The pooling calculation should be repeated with a more conservative estimate of yield.

### Chimeric Reads and Contamination

Chimeric reads are reads that contain sequences from two different templates. They arise from template switching during polymerase replication, where the polymerase dissociates from one template and re-associates with another. Chimeric reads can cause misassemblies, particularly if the two templates share sequence similarity.

Contamination is the presence of DNA from another organism in the sample. This can arise from the DNA extraction step (e.g., reagents contaminated with bacterial DNA), from the sequencing platform (e.g., carryover from previous runs), or from the sample itself (e.g., a mixed culture).

The symptoms of contamination are an assembly with an unexpectedly large genome size, the presence of contigs with unusual GC content or coverage, or the detection of methylation motifs that are not expected for the target organism.

To troubleshoot contamination, first check the assembly for contigs with atypical coverage. Contaminating DNA is usually present at a different concentration than the target DNA, so the coverage of contaminating contigs will be different. The Kraken2 or Centrifuge tools can be used to classify the taxonomic origin of contigs. If contamination is confirmed, the source should be identified and eliminated.

### Parameter Selection Errors

The choice of parameters for assembly and polishing can have a significant impact on the final assembly quality. Common errors include:

**Incorrect genome size estimate**: Many assemblers require an estimate of the genome size. If this estimate is too large, the assembler may produce fragmented assemblies; if too small, it may produce collapsed assemblies.

**Incorrect read type specification**: Some assemblers have different modes for HiFi and CLR reads. Using the wrong mode can result in poor assemblies.

**Insufficient polishing rounds**: For CLR assemblies, a single round of polishing is rarely sufficient. At least 2–3 rounds of Racon or Medaka polishing are recommended.

**Incorrect coverage cutoff**: Some assemblers have a coverage cutoff parameter that removes reads with unusually high or low coverage. If this parameter is set incorrectly, it can remove legitimate reads.

To troubleshoot parameter errors, consult the documentation for the specific tool and check that the parameters are appropriate for the read type and genome size. The default parameters are usually a good starting point, but they may need adjustment for unusual datasets.

### Methylation Calling Artifacts

Methylation calling from SMRT data can produce false positives and false negatives. Common artifacts include:

**Low coverage**: At low coverage, the kinetic signal is noisy, and the statistical test for methylation may produce false positives. Increasing the coverage to 100× or higher reduces this problem.

**Sequence context effects**: The IPD is influenced by the local sequence context, and some contexts produce kinetic signatures that resemble methylation. The MotifMaker tool accounts for this by comparing the observed IPD to the expected IPD for the same sequence context, but the correction is not perfect.

**Strand bias**: Methylation is often strand-specific, and the kinetic signal may be stronger on one strand than the other. The methylation calling should be performed on both strands and the results combined.

**Mixed methylation states**: In a population of cells, not all copies of a motif may be methylated. The fraction of methylated sites is called the methylation fraction, and it can vary from 0 to 1. Low methylation fractions can be difficult to detect, particularly at low coverage.

To troubleshoot methylation artifacts, examine the kinetic scores at known methylated and unmethylated sites. The distribution of scores should be bimodal, with a clear separation between modified and unmodified positions. If the separation is poor, the coverage may be insufficient, or the methylation may be partial.

## Practical Summary and Best Practices

### Typical Workflow

A typical PacBio-only bacterial sequencing project follows these steps:

1. **DNA extraction**: Extract high-molecular-weight DNA using a gentle lysis protocol. Assess DNA quality by spectrophotometry (A260/A280 ratio of 1.8–2.0) and by [pulsed-field gel electrophoresis](/knowledge/diagnostics/molecular/pulsed-field-gel-electrophoresis) or a Fragment Analyzer.

2. **Library preparation**: Shear the DNA to 15–20 kb, repair damage, ligate SMRTbell adapters, and size-select to remove fragments below 5 kb. Quantify the library and assess the size distribution.

3. **Sequencing**: Load the library onto the SMRT cell and sequence using the HiFi mode. The run time is typically 30 hours on the Sequel II or Revio.

4. **Base calling and demultiplexing**: Process the raw data with SMRT Link to generate HiFi reads. Demultiplex barcoded samples and filter reads by quality and length.

5. **Assembly**: Assemble the reads using Flye or Hifiasm. Assess the assembly quality using QUAST and BUSCO.

6. **Circularization and polishing**: Circularize the contigs and polish with Racon or Medaka. Confirm the circularity and accuracy of the final genome.

7. **Methylation analysis**: Identify modified bases from the kinetic data, call motifs using MotifMaker, and visualize the results.

8. **Annotation and submission**: Annotate the genome using Prokka or [NCBI Prokaryotic Genome Annotation Pipeline](/blog/guides/ncbi-prokaryotic-genome-annotation-pipeline), and submit the assembly to GenBank.

### Quality Control Checkpoints

Quality control should be performed at every step of the workflow. The key checkpoints are:

**After DNA extraction**: Check the DNA concentration, purity (A260/A280), and integrity. The DNA should be >30 kb in size with no visible degradation.

**After library preparation**: Check the SMRTbell yield and size distribution. The library should have a median size of 15–20 kb with no significant peak below 5 kb.

**After sequencing**: Check the sequencing yield, read length distribution, and read quality. The yield should be within the expected range for the platform, and the median read length should be close to the library insert size.

**After assembly**: Check the assembly contiguity, completeness, and accuracy. The assembly should have one contig per replicon, >99% BUSCO completeness, and >99.99% read-to-assembly identity.

**After methylation analysis**: Check the number of modified positions, the fraction of methylated motifs, and the consistency of the methylation calls across the genome.

## Frequently Asked Questions

### What is PacBio-only bacterial sequencing?

PacBio-only bacterial sequencing is a complete workflow for generating a finished bacterial genome using exclusively Pacific Biosciences SMRT sequencing data. It involves library preparation, sequencing on a PacBio platform (Sequel, Sequel II, or Revio), and computational analysis including base calling, assembly, polishing, and methylation detection. No Illumina or other short-read data is used. The approach produces complete, closed bacterial genomes with simultaneous methylation information.

### How much coverage is needed for PacBio-only bacterial assembly?

For de novo assembly of a bacterial genome from HiFi reads, a minimum of 30× coverage is required, and 50–100× is typical. For combined assembly and methylation detection, 100–150× coverage is recommended, as methylation calling requires higher depth for reliable kinetic signal detection. Coverage below 30× results in fragmented assemblies and incomplete circularization.

### Can PacBio-only sequencing detect DNA methylation?

Yes. SMRT sequencing records the kinetics of nucleotide incorporation, and modified bases such as m6A, m4C, and m5C produce characteristic kinetic signatures. The detection of m6A and m4C is robust at moderate coverage, while m5C detection requires high coverage and careful analysis. The MotifMaker tool identifies the sequence motifs recognized by methyltransferases. For m5C, orthogonal validation with [bisulfite sequencing](/knowledge/molecular-biology/bisulfite-sequencing) is recommended.

### What assemblers work best for PacBio-only bacterial genomes?

Flye and Hifiasm are the recommended assemblers for PacBio-only bacterial genomes. Flye uses a repeat graph approach and is robust to parameter misconfiguration. Hifiasm uses a string graph approach and produces slightly more accurate assemblies. Canu is a viable alternative but is slower and more computationally intensive. Unicycler is not recommended for PacBio-only data, as it is designed for hybrid or short-read assembly.

### Do I need to polish PacBio assemblies?

For HiFi assemblies, polishing is optional but recommended. HiFi reads have ≥99.9% accuracy, and the assembly is already highly accurate. However, polishing with Racon or Medaka can correct the few remaining errors and improve the consensus quality. For CLR assemblies, polishing is essential, as the raw read accuracy is only 85–92%, and multiple rounds of polishing are required.

### What is the difference between HiFi and continuous long reads?

HiFi reads (circular consensus sequences) are generated by sequencing the same SMRTbell template multiple times and collapsing the subreads into a consensus. They are 10–15 kb in length with ≥99.9% accuracy. Continuous long reads (CLR) are single-pass reads that are 10–25 kb in length with 85–92% accuracy. HiFi reads are preferred for bacterial assembly due to their high accuracy, while CLR reads are used when ultra-long reads are needed.

### How do I avoid contamination in PacBio-only sequencing?

Contamination can arise from the DNA extraction step, the sequencing platform, or the sample itself. To avoid contamination, use sterile reagents and consumables, include a negative control in the DNA extraction, and use a fresh SMRT cell for each run. After assembly, check for contigs with atypical coverage or GC content, and classify the taxonomic origin of contigs using Kraken2 or Centrifuge.

### What are common mistakes in PacBio-only bacterial sequencing?

Common mistakes include insufficient coverage due to over-pooling, using degraded DNA that limits read length, incorrect assembler parameters, skipping polishing for CLR assemblies, and misinterpreting methylation data at low coverage. Other mistakes include failing to assess assembly quality with BUSCO and QUAST, and not validating m5C methylation calls with an orthogonal method.

## Key Takeaways

- PacBio-only bacterial sequencing uses exclusively SMRT data to produce complete, closed bacterial genomes with simultaneous methylation information, eliminating the need for hybrid approaches.
- HiFi reads (≥99.9% accuracy, 10–15 kb) are the preferred read type for bacterial assembly; CLR reads (85–92% accuracy, 10–25 kb) are used only when ultra-long reads are required.
- High-molecular-weight DNA (>30 kb) is essential for successful library preparation, and gentle lysis protocols that avoid mechanical shearing are required.
- A coverage of 50–100× is recommended for assembly, while 100–150× is needed for reliable methylation detection.
- Flye and Hifiasm are the recommended assemblers for PacBio-only bacterial genomes, producing single circular contigs per replicon.
- Polishing with Racon or Medaka is recommended for HiFi assemblies and essential for CLR assemblies.
- Methylation analysis from SMRT data detects m6A and m4C robustly, but m5C detection requires high coverage and orthogonal validation.
- Quality control checkpoints at every step—DNA extraction, library preparation, sequencing, assembly, and methylation analysis—are essential for a successful project.

## Further Reading

- Jin H et al. *Using PacBio sequencing to investigate the bacterial microbiota of traditional Buryatian cottage cheese and comparison with Italian and Kazakhstan artisanal cheeses*. Journal of dairy science. 2018. [PubMed 29753477](https://doi.org/10.3168/jds.2018-14403)
- Derakhshani H et al. *Completion of draft bacterial genomes by [long-read sequencing](/knowledge/bioinformatics/long-read-sequencing-technologies-pacbio-and-oxford-nanopore) of synthetic genomic pools*. BMC genomics. 2020. [PubMed 32727443](https://doi.org/10.1186/s12864-020-06910-6)
- Cook R et al. *The long and short of it: benchmarking viromics using Illumina, Nanopore and PacBio sequencing technologies*. Microbial genomics. 2024. [PubMed 38376377](https://doi.org/10.1099/mgen.0.001198)
- Teikari J, Baunach M, Dittmann E. *Cyanobacterial Genome Sequencing, Annotation, and Bioinformatics*. Methods in molecular biology (Clifton, N.J.). 2022. [PubMed 35524055](https://doi.org/10.1007/978-1-0716-2273-5_14)
- Yu T et al. *Effects of Waterlogging on Soybean Rhizosphere Bacterial Community Using V4, LoopSeq, and PacBio 16S rRNA Sequence*. Microbiology spectrum. 2022. [PubMed 35171049](https://doi.org/10.1128/spectrum.02011-21)
- De Maio N et al. *Comparison of [long-read sequencing technologies](/knowledge/bioinformatics/long-read-sequencing-technologies-pacbio-and-oxford-nanopore) in the hybrid assembly of complex bacterial genomes*. Microbial genomics. 2019. [PubMed 31483244](https://doi.org/10.1099/mgen.0.000294)

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)