Amplicon Sequencing
Amplicon sequencing is a targeted next generation sequencing method that amplifies specific genomic regions using PCR, then sequences the resulting amplicons to high depth. It provides cost effective, high coverage data for focused genomic regions and is widely used in microbial community profiling, variant detection, forensic identification, and clinical diagnostics. This guide is for researchers, graduate students, and laboratory technicians planning amplicon sequencing experiments who need a source bounded practical framework. NCBI Bookshelf offers foundational reference material on sequencing technologies. Understanding the key decision points from target selection to sequencing platform choice can save time and money. [EMBL EBI Training](https://www.e bi.ac.uk/training/) provides official resources on bioinformatics workflows for amplicon data.
At a Glance
| Aspect | Key Points |
|---|---|
| Definition | Targeted PCR amplification of selected genomic regions followed by high throughput sequencing |
| Primary Applications | 16S rRNA and ITS microbial profiling, targeted variant screening, forensic short tandem repeat analysis, pathogen detection |
| Typical Workflow | DNA extraction, PCR amplification, library preparation, sequencing, quality filtering, taxonomic or variant assignment, diversity analysis |
| Quality Control | Primer specificity checks, inclusion of positive and negative controls, spike in standards for quantification, read quality trimming |
| Common Pitfalls | Primer bias, insufficient sequencing depth, index hopping, overinterpretation of rare sequences, batch effects |
Core Concepts and Decision Points
Amplicon sequencing begins with selecting a genomic target. In microbial ecology, the 16S rRNA gene (bacteria) or internal transcribed spacer (fungi) are standard choices because they contain both conserved and variable regions. Forensic applications might target human short tandem repeats or mitochondrial DNA segments. Custom panels for cancer mutations or pathogen resistance markers are also common. Primer design is critical: you must ensure specificity to the target group while minimizing amplification of nontarget DNA. The Galaxy Training Network offers workflows that include in silico primer evaluation. Amplicon length matters for sequencing platform compatibility. Illumina platforms work well with amplicons up to 600 base pairs, longer reads from PacBio or Oxford Nanopore can span entire 16S genes. Multiplexing many samples in a single run reduces cost but introduces risks of index misassignment, so dual indexing is recommended. Sequencing depth depends on the question. For microbial community beta diversity, 10,000 reads per sample may suffice, but rare species detection may require deeper coverage. The NCBI Sequence Read Archive hosts thousands of amplicon datasets that can help you benchmark depths for similar studies.
A practical decision framework includes the following steps. First, define the biological question: do you need species level resolution or broader taxonomic patterns? Second, choose a target region and verify that published primers cover your expected diversity. Third, decide on a sequencing platform considering read length, throughput, and cost. Fourth, plan controls: negative extraction controls, no template PCR controls, and a mock community of known composition. Fifth, allocate sufficient sequencing depth per sample. For example, in a study of gut microbiome signatures for postmortem interval estimation Construction of a Segmented PMI Estimation Model Integrating Intestinal Microbial Signatures and Machine Learning in Nude Mice, researchers used 16S amplicon sequencing to track microbial succession, which required high depth and careful time series controls.
Practical Workflow or Implementation Sequence
A robust amplicon sequencing workflow can be divided into six steps.
Step 1: Sample collection and DNA extraction. Use extraction methods that minimize bias, such as bead beating for microbial cells. Include negative controls throughout.
Step 2: Primer selection and PCR optimization. Test primer annealing temperature gradients and cycle numbers to reduce nonspecific products. Run agarose gels to confirm single bands. Bioconductor packages like DECIPHER can simulate primer binding across target taxa.
Step 3: Library preparation. Attach sequencing adapters and indexes via a second PCR or ligation. Purify amplicons to remove primers and dimers. Quantify libraries accurately with fluorometric methods.
Step 4: Sequencing. Pool libraries equimolarly. Use a sequencing platform appropriate for your amplicon length. For Illumina MiSeq, a 2x300 cycle kit can cover the V3-V4 16S region. Include a PhiX control spike in for low diversity libraries.
Step 5: Bioinformatics processing. Demultiplex reads, trim primers, filter low quality bases, merge paired ends, and denoise to amplicon sequence variants (ASVs) or cluster into operational taxonomic units (OTUs). The Dada2 package in R (part of Bioconductor) is widely used for this step. Assign taxonomy using reference databases like SILVA or Greengenes. Compute alpha and beta diversity metrics.
Step 6: Interpretation and validation. Compare your results with positive and negative controls. Validate key findings with quantitative PCR or independent methods. For environmental samples, consider that amplicon data reflect relative abundance, not absolute cell counts. The study on microbial succession on plastics in freshwater streams Microbial succession and early transformation signatures of low density polyethylene and polylactic acid in an urban freshwater stream used a combination of amplicon sequencing and chemical analyses to link community shifts to polymer degradation.
Common Mistakes
Primer bias against certain taxa. Universal primers often underrepresent some groups. Test your primers against known communities or use multiple primer sets. The oral gut joint axis study The oral gut joint axis in osteoarthritis: a multiomics case control study used both 16S and shotgun metagenomics to cross validate findings.
Insufficient sequencing depth for rare taxa. If you target low abundance community members, increase depth and include positive controls with known rare sequences.
Index hopping between samples. This occurs when free indexes in a pooled library recombine during sequencing. Use unique dual indexes and avoid excessive PCR cycles. NCBI Bookshelf describes index hopping mechanisms and mitigation strategies.
Overinterpreting ASVs or OTUs as species. Amplicon sequences often cannot discriminate closely related species. Report at genus level if database coverage is low.
Ignoring batch effects. Process samples in randomized order and include technical replicates. A low lipid diet study in salmon Effects of a low lipid diet on the gut microbiome and head kidney transcriptome of juvenile Chinook Salmon highlighted how dietary changes can obscure true microbial signals if batch effects are not controlled.
Neglecting to verify PCR contamination. Always run no template controls and check for unexpected bands. In chicken anemia virus detection PCR based detection of Chicken Anemia Virus and Escherichia coli co infection in commercial poultry flocks in Lahore, Punjab, Pakistan, careful controls were essential to distinguish co infection from contamination.
Limits of Interpretation and Uncertainty
Amplicon sequencing provides relative abundance data, not absolute cell counts. Differences in rRNA gene copy numbers across taxa can bias perceived proportions. Additionally, amplicon data cannot distinguish live from dead cells without pretreatment. The method also suffers from PCR bias due to GC content and amplicon length. Taxonomic resolution depends on the hypervariable region chosen, the V1-V2 region may resolve some genera better than V3-V4. These uncertainties mean that conclusions about community composition should be validated with complementary techniques. For instance, a study on selenium hyperaccumulation effects on root microbiomes Effect of selenium hyperaccumulation on the root endophytic and rhizosphere microbiome of two Astragalus species combined amplicon sequencing with cultivation to confirm functional roles. Another limitation is that rare sequences may represent artifacts from sequencing errors or contaminant DNA, so use strict denoising and consider removing singletons. The practical limit of detection for a given taxon is typically around 0.01% relative abundance, but confidence intervals widen at low frequencies.
Frequently Asked Questions
What is the difference between amplicon sequencing and whole genome sequencing? Amplicon sequencing targets specific regions, while whole genome sequencing (WGS) sequences all DNA in a sample. Amplicon sequencing is cheaper and provides high depth per target, but WGS offers broader functional and taxonomic information. EMBL EBI Training explains the trade offs between targeted and shotgun approaches.
How many reads per sample do I need for 16S amplicon sequencing? For community composition analysis, 10,000 to 50,000 reads per sample is typical. Rarefaction curves can help determine if your depth captures diversity. For rare biosphere detection, 100,000 reads or more may be required.
Can I use amplicon sequencing for metagenomics? No. Amplicon sequencing targets only one or a few marker genes, whereas metagenomics sequences all genetic material. However, amplicon data is often used as a rapid screen before deeper metagenomic sequencing.
How do I control for contamination in amplicon sequencing? Include extraction blanks, PCR negatives, and mock communities. Filter out ASVs present in negative controls at similar abundance. Use a bioinformatics decontamination tool like the microDecon R package. Galaxy Training Network provides tutorials on contamination detection workflows.
References and Further Reading
- NCBI Bookshelf: Overview of DNA sequencing technologies and amplicon applications. NCBI Bookshelf
- EMBL EBI Training: Introduction to amplicon sequencing and data analysis. EMBL EBI Training
- Galaxy Training Network: Amplicon analysis workflows including 16S and ITS pipelines. Galaxy Training Network
- Bioconductor: Dada2 package documentation for amplicon denoising and taxonomic assignment. Bioconductor
- NCBI Sequence Read Archive: Repository for raw amplicon sequencing datasets used in published studies. NCBI Sequence Read Archive
- Construction of a Segmented PMI Estimation Model Integrating Intestinal Microbial Signatures and Machine Learning in Nude Mice. PubMed
- Microbial succession and early transformation signatures of low density polyethylene and polylactic acid in an urban freshwater stream. PubMed
- The oral gut joint axis in osteoarthritis: a multiomics case control study. PubMed
- Effects of a low lipid diet on the gut microbiome and head kidney transcriptome of juvenile Chinook Salmon. PubMed
- PCR based detection of Chicken Anemia Virus and Escherichia coli co infection in commercial poultry flocks. PubMed
- Effect of selenium hyperaccumulation on the root endophytic and rhizosphere microbiome of two Astragalus species. PubMed