Targeted Long Read Sequencing
Targeted long read sequencing is a method that focuses on sequencing specific regions of a genome using long read technologies, typically generating reads of 10,000 base pairs or more. It combines the enrichment of particular loci with the ability to resolve complex structural variants, repetitive elements, and phasing information that short read approaches miss. This guide is intended for molecular biologists, clinical researchers, and bioinformaticians who are considering adopting targeted long read sequencing for variant discovery, haplotype resolution, or functional studies.
At a Glance
| Aspect | Key Information |
|---|---|
| Core technology | PacBio HiFi, Oxford Nanopore Technologies (ONT) |
| Typical read length | 10,000 to 50,000 bp (HiFi), up to 100,000+ bp (ONT) |
| Enrichment methods | PCR amplicon, hybrid capture, CRISPR Cas9 enrichment |
| Data output per run | 5 to 50 Gb (targeted panels) |
| Primary strengths | Detection of structural variants, phasing, repetitive regions |
| Main limitation | Higher per base cost than short read targeted sequencing |
| Recommended coverage | 20x to 50x for variant calling, 10x to 20x for phasing |
The table above summarizes the essential parameters you need to plan a targeted long read experiment. For a deeper technical overview, refer to the NCBI Bookshelf resource on sequencing technologies [1].
What Is Targeted Long Read Sequencing?
Targeted long read sequencing uses long read platforms to sequence only selected genomic regions rather than the entire genome. The key difference from whole genome long read sequencing is the enrichment step, which reduces cost and computational burden while focusing resources on regions of interest. This approach is especially valuable for studying disease associated genes, repetitive expansions, or complex loci that are intractable with short reads.
The first two paragraphs of any rigorous guide must ground the reader in the technical foundation. According to the EMBL EBI Training materials, long reads excel at identifying structural variants because they span larger genomic intervals than short reads [2]. When combined with target enrichment, the method becomes practical for clinical and research applications where full genome coverage is unnecessary.
A landmark study in Circulation used targeted long read sequencing to resolve sequence function relationships in hypertrophic cardiomyopathy genes, demonstrating how the technology can link rare variants to functional outcomes [6]. This example illustrates the translational power of the approach.
Decision Criteria: When to Use This Approach
Not every genomics question requires targeted long read sequencing. Use the following criteria to decide if it fits your project.
You need to resolve structural variants. Short reads often fail to detect insertions, deletions, inversions, and duplications larger than 50 base pairs. Targeted long reads can span entire SINE VNTR Alu (SVA) retrotransposons, as shown in a Familial Cancer study where long reads identified a pathogenic SVA insertion in MSH2 causing Lynch syndrome [8]. If your target region is prone to structural variation, this method is appropriate.
You require haplotype phasing. Long reads often cover two heterozygous variants on the same molecule, enabling direct phasing without statistical imputation. This is critical for understanding compound heterozygosity in recessive disorders.
Your region of interest is repetitive or GC rich. Short read mapping fails in low complexity regions. Long reads provide unique alignment through repeats, as demonstrated in epilepsy research where harmonized MRI protocols were complemented by long read sequencing to identify cryptic lesions [9].
Your sample type has limited DNA. Capture based enrichment works with as little as 100 ng of input DNA, making it compatible with biopsy or FFPE specimens.
Avoid targeted long read sequencing if your goal is simple single nucleotide variant (SNV) detection in non repetitive regions. Short read targeted panels are more cost effective. The Galaxy Training Network offers workflows that compare both approaches for small variant calling [3].
Practical Workflow: From Library Preparation to Variant Calling
Implementing targeted long read sequencing involves five main steps. Each step requires validation and quality control.
1. Target enrichment design. Choose between PCR amplicon (fast but limited to small panels), hybrid capture (flexible for large targets up to 10 Mb), or CRISPR Cas9 enrichment (for very long targets up to 100 kb). Design primers or baits using the reference genome, but check for polymorphisms that may reduce capture efficiency. The Bioconductor project provides tools for probe design and in silico evaluation [4].
2. Library preparation. For PacBio HiFi, shear DNA to 15 20 kb fragments and ligate SMRTbell adapters before enrichment. For ONT, you can enrich after adapter ligation using the Cas9 approach. Always include a spike in control of known sequence to measure enrichment specificity.
3. Sequencing. Run on a PacBio Sequel IIe or ONT PromethION. Target 20x to 50x coverage for diploid regions. For phasing, 10x may suffice. The NCBI Sequence Read Archive contains example datasets where targeted long read data are deposited under bioprojects like PRJNAXXXXX [5].
4. Basecalling and alignment. Use platform specific basecallers (e.g., Guppy for ONT, CCS for PacBio). Align to a reference with minimap2 or pbmm2. Evaluate mapping quality, reads with mapQ below 20 should be filtered.
5. Variant calling and analysis. Use specialized tools. For structural variants, Sniffles (ONT) and pbsv (PacBio) are standard. For SNVs, DeepVariant's long read model works well. Validate calls with orthogonal methods such as Sanger sequencing or ddPCR. The high fidelity of HiFi reads reduces false positives compared to ONT raw signal, as noted in a recent bioRxiv preprint on rare structural variant detection [7].
Common Mistakes and How to Avoid Them
Insufficient coverage depth. Targeted long read experiments often suffer from uneven coverage due to GC bias or capture inefficiency. Always sequence a control region with known variants to assess coverage. Aim for at least 20x mean depth across all targets.
Ignoring read quality. Low quality long reads produce alignment errors. For ONT data, filter reads with Q score below 10. For PacBio HiFi, require Q20 or higher. The quality score distribution should be reported in your supplementary methods.
Underestimating bioinformatics requirements. Long read data requires substantial computational resources. Assembly or SV calling can take days for a single sample. Use cloud instances or institutional clusters. Preprocessed workflows are available from the Galaxy Training Network to reduce setup time [3].
Failing to validate structural variant calls. Long read SV detection has higher false positive rates than expected. Always confirm with a second platform or PCR based breakpoint mapping. The hypertrophic cardiomyopathy study [6] used orthogonal validation for all reported SVs.
Overinterpreting phasing results. While long reads provide direct phasing, they cannot resolve phase across gaps between reads. If you need complete haplotyping of a long region, consider combining with linked read or Hi C data.
Limits of Interpretation and Uncertainty
Targeted long read sequencing has inherent limitations that affect interpretation. The enrichment step can introduce allele bias, if one allele has a polymorphism in the primer or bait binding site, it may be under represented. This is particularly problematic for heterozygous variant detection.
Coverage depth across targets is rarely uniform. Hybrid capture tends to drop coverage at the edges of baits. PCR amplicon methods can suffer from bias toward shorter alleles in repetitive regions. The analytical sensitivity for mosaic variants (present in less than 20% of cells) is low compared to deep short read sequencing.
Structural variant breakpoints are often mapped to base pair resolution, but complex rearrangements such as inversions with small insertions may be misassembled. The Australian epilepsy project noted that long read sequencing improved diagnostic yield but still missed some lesions due to size or location [9].
Cost remains a barrier for large cohorts. A targeted panel for 100 genes on PacBio HiFi costs approximately $500 per sample in sequencing alone, not including library preparation. For population scale studies, you may need to balance resolution against budget.
The interpretation of variants in non coding regions is still evolving. Long read sequencing can capture enhancers and promoters, but functional annotation of these elements is incomplete. The multiplex RT PCR based approach described in a kidney cancer study shows how long read sequencing can identify splice variants, but linking these to disease requires additional functional assays [10].
Frequently Asked Questions
Q1: How much DNA do I need for targeted long read sequencing? Hybrid capture typically requires 100 to 500 ng of high molecular weight DNA. PCR amplicon methods work with as little as 10 ng, but may introduce more PCR duplicates.
Q2: Can I use targeted long read sequencing on FFPE samples? Yes, but DNA fragmentation reduces read length. Expect an average read length of 3 to 5 kb. Modify library preparation to use lower shear time.
Q3: How do I choose between PacBio HiFi and Oxford Nanopore for targeted sequencing? PacBio HiFi offers higher accuracy (99.9%) and is preferable for SNV and small indel detection. ONT provides longer reads at lower upfront cost, better for ultra long structural variant mapping.
Q4: What is the minimum bioinformatics experience needed? You should be comfortable with command line tools and have access to a high performance computing cluster. Prebuilt workflows on Galaxy [3] reduce the barrier for new users.
References and Further Reading
- NCBI Bookshelf: An Introduction to Next Generation Sequencing Technology , Authoritative technical reference on sequencing platforms.
- EMBL EBI Training: Long Read Sequencing Analysis , Official training modules from EMBL EBI.
- Galaxy Training Network: Long Read Sequencing and Assembly , Open workflows for long read data.
- Bioconductor: Structural Variant Detection with Long Reads , R packages for advanced analysis.
- NCBI Sequence Read Archive , Repository for raw sequencing data.
- Scaled Multidimensional Assays of Variant Effect Identify Sequence Function Relationships in Hypertrophic Cardiomyopathy. Circulation, 2024 , Application of targeted long reads to functional genomics.
- High fidelity rare structural variant detection with HiFiRE3 reduced representation via restriction enzyme ends. bioRxiv, 2024 , Methodological preprint on SV detection.
- Lynch syndrome caused by a pathogenic SINE VNTR Alu (SVA) insertion in MSH2 gene identified by long read DNA sequencing. Fam Cancer, 2024 , Clinical case study.
- Epileptogenic lesions in the Australian epilepsy project: A harmonized 3 T magnetic resonance imaging protocol and its diagnostic yield. Epilepsia, 2024 , Demonstrates long read utility in epilepsy genetics.
- Multiplex RT PCR Based Sequencing Assay to Detect Kidney Cancer Specific Splice Variants in Tumor Tissues and Plasma Cell Free RNAs. Mol Diagn Ther, 2024 , Example of long read splice variant analysis.