# Gene Sequencing: Methods, Mechanisms, and Applications

## Introduction to Gene Sequencing

### What Is Gene Sequencing?

Gene sequencing is the process of determining the precise order of nucleotides—adenine (A), cytosine (C), guanine (G), and thymine (T)—within a DNA molecule. This order constitutes the genetic code that dictates cellular function, development, and phenotype. Sequencing can be applied to individual genes, exons, entire genomes, or transcriptomes, and the resulting data serve as the foundational substrate for virtually all modern molecular biology, from evolutionary genetics to precision medicine.

The purpose of gene sequencing extends beyond simply reading nucleotide order. It enables the identification of disease-causing mutations, the characterization of microbial communities, the quantification of gene expression, and the assembly of complete genomes from organisms previously uncharacterized at the molecular level. The scale of sequencing has expanded dramatically: a single modern instrument can generate more data in one run than the entire Human Genome Project produced over a decade.

### A Brief History of Sequencing Technologies

The history of sequencing begins with the development of two complementary methods in the 1970s. Frederick Sanger's chain-termination method, published in 1977, used dideoxynucleotide triphosphates to generate truncated fragments whose lengths revealed the underlying sequence. Walter Gilbert and Allan Maxam developed a chemical cleavage method that, while historically important, was technically demanding and hazardous due to its reliance on toxic chemicals. Sanger sequencing became the dominant approach and remained so for nearly three decades.

The next major leap came with the advent of next-generation sequencing (NGS) in the mid-2000s. Platforms such as Roche's 454 (pyrosequencing), Illumina's Solexa (sequencing-by-synthesis), and Applied Biosystems' SOLiD (sequencing-by-ligation) introduced massively parallel processing, enabling millions of fragments to be sequenced simultaneously. These technologies reduced the cost of sequencing a human genome from roughly $100 million in 2001 to under $1,000 by 2015.

The most recent wave of innovation has focused on [long-read sequencing](/knowledge/bioinformatics/long-read-sequencing-technologies-pacbio-and-oxford-nanopore). Pacific Biosciences introduced single-molecule real-time (SMRT) sequencing in 2011, and Oxford Nanopore Technologies launched its nanopore sequencer in 2014. These platforms produce reads of 10–100 kilobases or longer, overcoming the short-read limitations of NGS for resolving repetitive regions, structural variants, and full-length transcripts.

## The Molecular Basis of DNA Sequencing

### DNA Polymerases and Nucleotide Incorporation

All sequencing technologies rely on the fundamental enzymology of DNA polymerases—enzymes that synthesize DNA in a template-directed manner. These polymerases catalyze the nucleophilic attack of the 3′-hydroxyl group of the growing DNA strand on the α-phosphate of an incoming deoxynucleoside triphosphate (dNTP), releasing pyrophosphate and extending the strand by one nucleotide. The reaction requires a primer with a free 3′-OH, a single-stranded template, and magnesium ions (typically 2–8 mM MgCl₂) as a cofactor.

The fidelity of this process is remarkable. High-fidelity polymerases such as *Pfu* or Q5® incorporate the correct nucleotide with error rates of approximately 1 in 10⁶–10⁷ bases, owing to their 3′→5′ exonuclease proofreading activity. In sequencing applications, however, the polymerase must be modified to accept unnatural nucleotide analogs—either fluorescently labeled terminators (as in Sanger and Illumina sequencing) or phosphate-labeled nucleotides (as in Pacific Biosciences SMRT sequencing). These modifications require polymerases engineered to accommodate bulky fluorophores without losing processivity or speed.

### Chain Termination and Fluorescent Labeling

The conceptual breakthrough of Sanger sequencing was the use of dideoxynucleoside triphosphates (ddNTPs). These analogs lack the 3′-hydroxyl group essential for phosphodiester bond formation. When a ddNTP is incorporated, DNA synthesis halts irreversibly because no subsequent nucleotide can be attached. By performing four separate reactions—each containing a different ddNTP—or by using four differentially labeled ddNTPs in a single reaction, a population of fragments is generated, each terminating at a specific nucleotide position.

Modern Sanger sequencing uses fluorescently labeled ddNTPs, where each of the four terminators carries a distinct fluorophore. This allows all four reactions to be multiplexed in a single tube. The resulting fragments, differing in length by single nucleotides, are then separated by size using capillary electrophoresis. A laser excites the fluorophores as fragments migrate past a detection window, and the emission spectrum identifies the terminal nucleotide at each position.

For NGS platforms, the chemistry differs in a critical way: reversible terminators. Illumina sequencing-by-synthesis uses nucleotides that are both fluorescently labeled and chemically blocked at the 3′-OH position. After incorporation, the fluorophore is cleaved and the block removed, allowing the next nucleotide to be added. This cyclic process—incorporate, image, cleave, repeat—enables sequencing to proceed one base at a time across millions of spatially separated clusters.

## Sanger Sequencing: The Gold Standard for Targeted Reads

### The Chain-Termination Reaction

Sanger sequencing remains the method of choice for validating variants identified by NGS, sequencing individual genes, and analyzing small amplicons. The reaction begins with a PCR-amplified template, typically 200–1,000 base pairs in length. The template is denatured to single-stranded DNA, and a sequencing primer is annealed upstream of the region of interest. The reaction mixture contains a thermostable polymerase, a mixture of unlabeled dNTPs, and fluorescently labeled ddNTPs at a ratio of approximately 1:100 (ddNTP:dNTP).

The ratio of ddNTPs to dNTPs is critical. Too high a ddNTP concentration produces very short fragments; too low produces fragments too long to be resolved. A typical reaction uses 0.5–2 μM ddNTPs and 50–200 μM dNTPs. The polymerase extends the primer until a ddNTP is incorporated, at which point synthesis terminates. Because termination occurs stochastically at every position across the template population, the reaction produces a ladder of fragments, each differing in length by one nucleotide and each labeled with the fluorophore corresponding to its terminal base.

Thermal cycling (typically 25–35 cycles of 96°C denaturation, 50°C annealing, and 60°C extension) linearly amplifies the number of terminated fragments. After cycling, unincorporated ddNTPs are removed by ethanol precipitation or column purification to prevent background fluorescence during detection.

### Capillary Electrophoresis and Detection

The sequencing products are loaded onto a capillary array filled with a sieving polymer, typically polyacrylamide or a proprietary matrix. Under an applied electric field (approximately 300 V/cm), the negatively charged DNA fragments migrate toward the anode. The sieving polymer separates fragments by size, with smaller fragments migrating faster. The resolution is sufficient to distinguish fragments differing by a single nucleotide up to approximately 800–1,000 bases.

As each fragment passes through the detection window, a laser (typically 488 nm for argon-ion excitation) excites the fluorophore. Four emission wavelengths are detected simultaneously, and the sequence is reconstructed from the order of fluorescence peaks in the electropherogram. Base calling software assigns a quality score (Phred score, Q) to each base, where Q20 corresponds to 99% accuracy and Q30 to 99.9% accuracy. For routine applications, reads of 600–900 bases with Q30 or better are expected.

Despite its throughput limitations—a single capillary produces one read per run—Sanger sequencing remains indispensable. It is the standard for confirming clinically actionable variants, resolving ambiguous NGS results, and sequencing regions with high GC content or homopolymer tracts that challenge short-read platforms. For detailed procedural guidance, see the [Sanger Sequencing Protocol](/knowledge/molecular-biology/sanger-sequencing-protocol).

## Next-Generation Sequencing: High-Throughput Parallelization

### Library Preparation and Adapter Ligation

NGS begins with the construction of a sequencing library—a collection of DNA fragments flanked by known adapter sequences. The process starts with DNA fragmentation, typically by sonication (yielding fragments of 200–500 bp), enzymatic digestion, or tagmentation (a transposase-mediated reaction that simultaneously fragments and ligates adapters). Fragment size is critical: too short and the fragments lack unique sequence for alignment; too long and cluster generation becomes inefficient.

After fragmentation, end repair converts the heterogeneous ends into blunt ends using a combination of T4 DNA polymerase and T4 polynucleotide kinase. A single 3′-adenine overhang is then added using Taq polymerase or Klenow fragment (3′→5′ exo⁻) in the presence of dATP. This A-tailing step enables efficient ligation of adapters that carry a complementary 3′-thymine overhang. Adapters are double-stranded oligonucleotides containing several functional elements: the sequencing primer binding sites, the flow-cell annealing sequences (for Illumina), and a unique index or barcode sequence (6–10 bases) that allows multiplexing of multiple samples in a single run.

The ligated fragments undergo size selection—either by gel extraction or bead-based purification using SPRI (solid-phase reversible immobilization) beads—to remove adapter dimers and oversized fragments. Finally, the library is amplified by PCR (typically 8–12 cycles) to generate sufficient material for cluster generation. The number of PCR cycles must be minimized to reduce duplication artifacts and bias. For a comprehensive treatment of this step, see [Library Prep in Sequencing](/knowledge/molecular-biology/library-prep-in-sequencing).

### Bridge Amplification and Emulsion PCR

Before sequencing, individual library fragments must be clonally amplified to produce a signal strong enough for detection. Illumina platforms use bridge amplification on a flow cell surface. The flow cell is coated with two types of oligonucleotides complementary to the adapter sequences. When a single-stranded library fragment hybridizes to one of these surface-bound oligos, a polymerase extends it to create a double-stranded molecule. The free end then arches over and hybridizes to the second oligo, forming a bridge. Repeated cycles of extension and denaturation amplify this single molecule into a cluster of approximately 1,000 identical copies, all spatially localized within a few micrometers.

Ion Torrent and 454 platforms use emulsion PCR instead. Here, individual library fragments are captured on beads within water-in-oil droplets, with each droplet containing a single bead, a single DNA fragment, and PCR reagents. Amplification occurs within the droplet, producing a bead covered with clonally amplified copies of the original fragment. After breaking the emulsion, beads are deposited into wells on a semiconductor chip (Ion Torrent) or a PicoTiterPlate (454).

### Sequencing-by-Synthesis vs. Sequencing-by-Ligation

Illumina sequencing-by-synthesis uses reversible terminator chemistry. A polymerase incorporates a single fluorescently labeled, 3′-blocked nucleotide complementary to the template. After incorporation, the flow cell is imaged, and the fluorophore is cleaved while the 3′-block is removed, allowing the next cycle to proceed. Each cycle adds one base to every cluster, and the sequence is read from the order of fluorescence images. Read lengths are typically 150–300 base pairs (paired-end), and a single run can generate 1–3 billion reads.

Sequencing-by-ligation, as implemented in the SOLiD platform, uses a different approach. A universal primer is annealed to the adapter, and a ligase joins octamer probes that are labeled with one of four fluorophores. Each octamer contains a specific two-base interrogation position (e.g., bases 4–5 of the octamer). After ligation, the fluorophore is imaged, and the last three bases of the octamer are cleaved, leaving a 5-base overhang for the next ligation. This process is repeated, and the sequence is decoded from the two-base encoding scheme. While SOLiD is no longer commercially supported, its principle illustrates an alternative to polymerase-based sequencing.

The choice between synthesis and ligation methods affects error profiles. Sequencing-by-synthesis is prone to errors in homopolymer regions (runs of identical bases) due to incomplete incorporation or signal decay, whereas sequencing-by-ligation is more robust to homopolymers but has shorter read lengths and different systematic biases.

## [Long-Read Sequencing Technologies](/knowledge/bioinformatics/long-read-sequencing-technologies-pacbio-and-oxford-nanopore)

### Single-Molecule Real-Time (SMRT) Sequencing

Pacific Biosciences' SMRT sequencing operates on a fundamentally different principle: real-time observation of DNA synthesis without amplification. The platform uses zero-mode waveguides (ZMWs)—nanophotonic structures with a diameter of approximately 70 nm that confine light to a tiny observation volume (about 20 zeptoliters). A single DNA polymerase molecule is immobilized at the bottom of each ZMW, and a circular single-stranded template (a SMRTbell) is loaded.

During sequencing, the polymerase incorporates fluorescently labeled nucleotides, each labeled at the terminal phosphate rather than the base. When a nucleotide enters the active site, the fluorophore is held in the ZMW's detection volume for tens of milliseconds—long enough to record its emission. After incorporation, the fluorophore is cleaved with the pyrophosphate and diffuses away, returning the signal to background. This design eliminates the need for reversible terminators because the label is removed naturally during polymerization.

SMRT sequencing produces reads averaging 10–25 kilobases, with some exceeding 100 kilobases. The platform's key advantage is the ability to sequence through repetitive regions and GC-rich sequences that confound short-read platforms. The circular nature of the SMRTbell template enables repeated sequencing of the same molecule, generating high-accuracy consensus reads (HiFi reads) with accuracy exceeding 99.9% when the molecule is read multiple times.

### Nanopore Sequencing and Electrical Detection

Oxford Nanopore Technologies takes an entirely different approach: sequencing by measuring changes in ionic current as DNA passes through a protein nanopore. The pore, typically a modified version of the *Mycobacterium smegmatis* porin A (MspA) or the *E. coli* α-hemolysin, is embedded in an electrically resistant polymer membrane. A voltage of approximately 180 mV is applied across the membrane, driving an ionic current of about 100–200 pA through the pore.

A processive enzyme (a helicase or polymerase) controls the translocation of DNA through the pore at a rate of approximately 450 bases per second. As each nucleotide (or more precisely, each 5-nucleotide k-mer) passes through the pore's constriction zone, it partially blocks the current. The magnitude of the current change is characteristic of the specific k-mer, allowing the sequence to be decoded from the current trace.

Nanopore sequencing offers several unique advantages. It produces ultra-long reads (100 kilobases to megabases), does not require PCR amplification (avoiding associated biases), and can sequence RNA directly. The devices are also portable—the MinION is smaller than a smartphone—and provide real-time data streaming, enabling immediate analysis. The primary limitation is per-base accuracy, which is lower than Illumina's (approximately 95–99% for raw reads), though newer chemistry and base-calling algorithms have narrowed this gap. For a comparison of nanopore and Sanger approaches, see [Nanopore vs Sanger Sequencing](/knowledge/molecular-biology/nanopore-vs-sanger-sequencing).

## Bioinformatics and Data Analysis in Gene Sequencing

### From Raw Signals to Base Calls

The first computational step in any sequencing workflow is base calling—converting raw instrument signals into nucleotide sequences with associated quality scores. For Illumina platforms, base calling uses the intensity of four fluorescence channels across successive cycles. The current standard, implemented in Illumina's Real-Time Analysis software, uses a model that accounts for phasing (the loss of synchrony among molecules in a cluster) and pre-phasing (the incorporation of two bases in one cycle). The output is a BCL file, which is converted to FASTQ format containing the sequence and a Phred quality score for each base.

For nanopore sequencing, base calling is a machine-learning problem. The raw ionic current signal is segmented into events corresponding to individual k-mers, and a recurrent neural network (such as Guppy or Bonito) translates these events into base sequences. The accuracy of nanopore base calling has improved substantially with the introduction of newer pore versions (R10.4) and improved neural network architectures.

### Alignment and Variant Calling

Once base calls are generated, reads must be aligned to a reference genome. The standard tool is BWA-MEM (Burrows-Wheeler Aligner with Maximal Exact Matches), which uses a seed-and-extend strategy to find the optimal alignment for each read. The output is a SAM (Sequence Alignment/Map) file, typically compressed to BAM format. For RNA-seq data, splice-aware aligners such as STAR or HISAT2 are required because they must account for reads spanning exon-exon junctions.

Variant calling identifies positions where the sample differs from the reference. The Genome Analysis Toolkit (GATK) HaplotypeCaller is the most widely used tool for this purpose. It reassembles the active region around potential variants, aligns reads to the local haplotypes, and applies a Bayesian model to assign genotype likelihoods. Variants are output in VCF (Variant Call Format). Key quality metrics include read depth, mapping quality, and the balance of alleles supporting each genotype.

### De Novo Genome Assembly

When no reference genome exists, reads must be assembled de novo. Short-read assembly uses de Bruijn graphs, where reads are decomposed into k-mers (substrings of length k) that form nodes in a graph, with edges representing overlaps of k−1 bases. The assembler (e.g., SPAdes for bacteria, ABySS for large genomes) traverses the graph to reconstruct contigs. However, repetitive regions longer than the read length create ambiguities that fragment the assembly.

Long-read assembly uses a different strategy: overlap-layout-consensus. Reads are compared to find overlaps, a layout is constructed, and a consensus sequence is derived. Tools such as Flye and Canu handle the higher error rates of long reads by using error correction and repeat resolution. Hybrid approaches combine short reads (for accuracy) with long reads (for contiguity), producing chromosome-level assemblies when combined with optical mapping or Hi-C data.

## Applications of Gene Sequencing

### Whole-Genome and Whole-Exome Sequencing

Whole-genome sequencing (WGS) aims to determine the complete DNA sequence of an organism's genome. For humans, this means approximately 3.2 billion base pairs. WGS captures all coding and non-coding regions, including introns, regulatory elements, and structural variants. The primary challenge is coverage: to reliably detect variants, each position must be sequenced at a depth of at least 30× (30 independent reads), requiring approximately 90–100 gigabases of raw sequence per genome.

Whole-exome sequencing (WES) targets only the protein-coding regions, which constitute approximately 1–2% of the genome (about 30 megabases). WES uses hybridization capture with biotinylated probes complementary to exonic sequences, followed by sequencing of the captured fragments. The advantage is cost efficiency—WES can be performed at higher depth for the same cost as WGS—but it misses non-coding variants, deep intronic mutations, and structural variants. The choice between WGS and WES depends on the clinical or research question, with WGS preferred for undiagnosed genetic conditions and WES for known gene panels.

### RNA Sequencing and Transcriptomics

RNA sequencing (RNA-seq) quantifies gene expression by sequencing cDNA derived from RNA. The workflow involves RNA extraction, poly-A selection (or ribosomal RNA depletion), fragmentation, reverse transcription, and library preparation. The resulting reads are aligned to the genome or transcriptome, and expression levels are quantified as read counts per gene. Differential expression analysis (using tools such as DESeq2 or edgeR) identifies genes whose expression changes between conditions.

RNA-seq also enables the detection of alternative splicing, allele-specific expression, and fusion genes. Long-read RNA-seq (Iso-Seq on PacBio or direct RNA sequencing on Nanopore) captures full-length transcripts, resolving isoform structures that short reads cannot. For specialized applications, see [Bisulfite Sequencing](/knowledge/molecular-biology/bisulfite-sequencing) for DNA methylation analysis and [CHIP Sequencing](/knowledge/molecular-biology/chip-sequencing) for protein-DNA interactions.

### Clinical and Diagnostic Applications

Clinical sequencing has transformed the diagnosis of genetic disorders. Targeted gene panels sequence a defined set of genes known to be associated with a condition, offering high depth and rapid turnaround. Exome sequencing is used when the causative gene is unknown, and genome sequencing is increasingly used for critically ill neonates and undiagnosed diseases.

In oncology, sequencing of tumor biopsies identifies driver mutations that guide targeted therapy. Liquid biopsy approaches sequence [circulating tumor DNA](/knowledge/molecular-biology/circulating-tumor-dna) (ctDNA) from blood, enabling non-invasive monitoring of tumor burden and the detection of resistance mutations. Pharmacogenomic sequencing identifies variants affecting drug metabolism, guiding dose selection to minimize adverse reactions and maximize efficacy.

Infectious disease diagnostics use sequencing to identify pathogens, determine antimicrobial resistance profiles, and track outbreaks through [genomic epidemiology](/knowledge/bioinformatics/genomic-epidemiology-integrating-pathogen-genomics-into-outbreak-investigations). During the COVID-19 pandemic, [viral genome sequencing](/blog/guides/viral-genome-sequencing-a-project-planning-guide) enabled real-time tracking of SARS-CoV-2 variants and informed vaccine development.

## Common Pitfalls and Practical Considerations

### Avoiding Contamination and Bias

Contamination is a pervasive risk in sequencing. Even trace amounts of foreign DNA—from the operator, reagents, or previous samples—can be amplified and sequenced, producing misleading results. Prevention requires dedicated pre- and post-PCR areas, the use of filtered pipette tips, and the inclusion of negative controls (water or buffer processed through the entire workflow). For low-input samples, the risk is amplified because contaminating DNA may outnumber the sample DNA.

PCR bias is another concern. GC-rich regions amplify poorly, leading to underrepresentation in the final data. This bias can be mitigated by using polymerases engineered for GC-rich templates, adding DMSO or betaine to the reaction, or using amplification-free library preparation methods. PCR duplicates—identical reads arising from the same original fragment—inflate apparent coverage and must be removed computationally before variant calling.

### Choosing the Right Sequencing Platform

The choice of platform depends on the biological question, sample type, and budget. The table below summarizes key parameters:

| Platform | Read Length | Throughput per Run | Error Rate | Best For |
|----------|-------------|-------------------|------------|----------|
| Sanger | 600–900 bp | 1 read | ~0.1% | Variant validation, small amplicons |
| Illumina (NovaSeq) | 150–300 bp | 1–3 billion reads | ~0.1% | WGS, WES, RNA-seq, ChIP-seq |
| Ion Torrent | 200–400 bp | 5–80 million reads | ~1% (homopolymers) | Targeted panels, microbial sequencing |
| PacBio SMRT | 10–25 kb (HiFi) | 1–4 million reads | ~0.1% (HiFi) | De novo assembly, full-length transcripts |
| Oxford Nanopore | 10–100+ kb | 10–50 million reads | ~1–5% | Ultra-long reads, real-time diagnostics |

For targeted resequencing of a few genes, Sanger or Ion Torrent may be sufficient. For whole-genome studies, Illumina offers the best cost-per-base. For complex genomes with extensive repeats or structural variation, long-read platforms are essential. For a deeper discussion of depth requirements, see [Sequencing Coverage](/knowledge/molecular-biology/sequencing-coverage).

### Interpreting Variants with Caution

Variant interpretation is the most error-prone step in the sequencing workflow. A variant may be a true biological finding, a sequencing artifact, or a misalignment. Several safeguards are essential:

1. **Confirm with an orthogonal method.** Sanger sequencing remains the gold standard for validating clinically significant variants.
2. **Check read depth and quality.** Variants with fewer than 10 supporting reads or low mapping quality should be treated with suspicion.
3. **Consider population frequency.** A variant present at high frequency in population databases (e.g., gnomAD) is unlikely to be a rare disease-causing mutation.
4. **Assess segregation.** In family studies, the variant should co-segregate with the phenotype.
5. **Evaluate functional impact.** In silico predictors (PolyPhen, SIFT, CADD) provide supporting evidence but are not definitive.

The interpretation of variants of unknown significance (VUS) requires caution. A VUS should not be used for clinical decision-making without additional evidence, and patients should be counseled about the uncertainty.

## Frequently Asked Questions

### What are the different types of gene sequencing?

The main types are Sanger sequencing (targeted, single reads), next-generation sequencing (massively parallel short reads, including whole-genome, whole-exome, and targeted panels), and long-read sequencing (PacBio SMRT and Oxford Nanopore). Specialized variants include RNA-seq for transcriptomes, [bisulfite sequencing](/knowledge/molecular-biology/bisulfite-sequencing) for methylation analysis, and ChIP-seq for protein-DNA interactions.

### How does gene sequencing work?

All sequencing methods rely on DNA polymerase extending a primer along a template. Sanger sequencing uses chain-terminating dideoxynucleotides to generate length-laddered fragments. NGS uses reversible terminators or ligation to add one base at a time across millions of clusters. Long-read platforms observe polymerase activity in real time (PacBio) or measure ionic current through a nanopore (Oxford Nanopore).

### What is the difference between Sanger and NGS?

Sanger sequencing produces a single read of 600–900 bases per reaction, with very high accuracy, and is best for validating variants or sequencing small numbers of samples. NGS produces millions to billions of short reads (150–300 bases) in parallel, enabling whole-genome or whole-transcriptome analysis at scale, but requires substantial bioinformatics infrastructure.

### What are the applications of gene sequencing?

Applications include whole-genome and exome sequencing for rare disease diagnosis, RNA-seq for gene expression profiling, metagenomics for microbial community analysis, oncology for tumor profiling and liquid biopsy, pharmacogenomics for drug response prediction, and infectious disease surveillance.

### How long does gene sequencing take?

Sanger sequencing takes 2–4 hours from PCR to result. Illumina NGS takes 1–3 days depending on read length and throughput. Nanopore sequencing produces data in real time, with initial results available within minutes to hours. The total turnaround time, including library preparation and analysis, ranges from 1 day (targeted nanopore) to 2 weeks (high-throughput WGS with deep coverage).

### What is the cost of gene sequencing?

Costs have fallen dramatically. Sanger sequencing costs $5–20 per reaction. Targeted NGS panels cost $100–500 per sample. Whole-exome sequencing costs $300–800. Whole-genome sequencing costs $500–1,500 at production scale, though clinical-grade WGS with interpretation can cost $2,000–5,000. Long-read sequencing remains more expensive per base but is cost-effective for specific applications.

### What is the difference between whole-genome and whole-exome sequencing?

Whole-genome sequencing reads the entire genome, including non-coding regions, introns, and regulatory elements, at a typical depth of 30×. Whole-exome sequencing targets only protein-coding exons (about 1–2% of the genome) using hybridization capture, allowing higher depth at lower cost. WGS detects structural variants and non-coding mutations but is more expensive and requires more data storage; WES is more cost-effective for identifying coding variants but misses non-coding changes.

## Key Takeaways

- Gene sequencing determines the order of nucleotides in DNA and has evolved from Sanger's chain-termination method to massively parallel NGS and real-time long-read platforms.
- The molecular basis of sequencing relies on DNA polymerase extension, with chain-terminating dideoxynucleotides (Sanger) or reversible terminators (Illumina) enabling base-by-base reading.
- NGS achieves scale through library preparation, clonal amplification (bridge PCR or emulsion PCR), and sequencing-by-synthesis or by-ligation, generating millions to billions of reads per run.
- Long-read technologies (PacBio SMRT and Oxford Nanopore) overcome short-read limitations, enabling assembly of complex genomes, detection of structural variants, and [full-length transcript sequencing](/knowledge/bioinformatics/full-length-transcript-sequencing-unraveling-isoform-diversity-with-long-reads).
- Bioinformatics is integral to sequencing: base calling, alignment, variant calling, and de novo assembly require specialized tools and careful quality control.
- Sequencing applications span clinical diagnostics, oncology, infectious disease surveillance, transcriptomics, and metagenomics, with platform choice driven by read length, throughput, accuracy, and cost.
- Common pitfalls include contamination, PCR bias, insufficient coverage, and overinterpretation of variants; rigorous experimental design and orthogonal validation are essential for reliable results.

## Further Reading

- Wensel CR et al. *Next-generation sequencing: insights to advance clinical investigations of the microbiome*. The Journal of clinical investigation. 2022. [PubMed 35362479](https://doi.org/10.1172/JCI154944)
- Janda JM, Abbott SL. *16S rRNA gene sequencing for bacterial identification in the diagnostic laboratory: pluses, perils, and pitfalls*. Journal of [clinical microbiology](/knowledge/diagnostics/microbiology/clinical-microbiology-from-specimen-collection-to-pathogen-identification). 2007. [PubMed 17626177](https://doi.org/10.1128/JCM.01228-07)
- Regueira-Iglesias A et al. *Critical review of 16S rRNA gene sequencing workflow in microbiome studies: From primer selection to advanced data analysis*. Molecular oral microbiology. 2023. [PubMed 37804481](https://doi.org/10.1111/omi.12434)
- Li MN et al. *16S rRNA gene sequencing for bacterial identification and infectious disease diagnosis*. Biochemical and biophysical research communications. 2024. [PubMed 39550863](https://doi.org/10.1016/j.bbrc.2024.150974)
- Lv X et al. *Application of high-throughput gene sequencing in lymphoma*. Experimental and molecular pathology. 2021. [PubMed 33493455](https://doi.org/10.1016/j.yexmp.2021.104606)
- Osman MA et al. *16S rRNA Gene Sequencing for Deciphering the Colorectal Cancer Gut Microbiome: Current Protocols and Workflows*. Frontiers in microbiology. 2018. [PubMed 29755427](https://doi.org/10.3389/fmicb.2018.00767)

## Related Clinical & Scientific Guides

* [MAPK Pathway: Mechanism, Function, and Clinical Relevance](/knowledge/molecular-biology/mapk-pathway)
* [Mammalian Cell Culture Bioreactors: A Practical Guide](/knowledge/molecular-biology/mammalian-cell-culture-bioreactor)
* [Nucleotide Formation: Biosynthesis and Assembly of DNA/RNA Building Blocks](/knowledge/molecular-biology/nucleotide-formation)