# Exome Sequencing


## Key Takeaways

- Exome sequencing targets protein-coding regions (exons), representing 1-2% of the genome but harboring approximately 85% of known disease-causing variants, making it a cost-effective approach for identifying pathogenic mutations in rare genetic disorders and neurodevelopmental conditions.
- The workflow involves DNA extraction, library preparation with exonic capture probes, high-throughput sequencing (typically 80x-100x depth), and rigorous bioinformatics analysis including alignment, variant calling, and annotation against reference genomes and population databases.
- Exome sequencing is indicated when a broad scan of coding regions is necessary without a strong prior hypothesis, outperforming targeted gene panels for heterogeneous conditions, but it does not reliably detect structural variants, deep intronic mutations, or regulatory element changes.
- Critical quality control metrics include coverage uniformity (≥90% of target bases at ≥20x), mapping quality (≥90% of reads), appropriate variant quality filtering (e.g., GATK recommendations), and monitoring the transition to transversion ratio.
- Common pitfalls include overlooking splice site variants, using overly lenient coverage thresholds, failing to filter common population variants, and neglecting the detection of mosaic variants, which require deep coverage and careful interpretation.
- Limitations include lower sensitivity for copy number variations compared to dedicated platforms, inability to detect non-coding regulatory variants, potential for poor coverage in GC-rich or repetitive exonic regions, and ambiguity in variant interpretation requiring segregation studies or functional validation.

---

Exome sequencing is a targeted genomic approach that reads the protein coding regions of the genome, known as the exome, which comprises approximately 1% to 2% of total DNA but harbors about 85% of known disease causing variants. Researchers and clinicians investigating rare genetic disorders, neurodevelopmental conditions, and certain cancers should use this guide to understand when exome sequencing is appropriate, how to run the workflow, and where its limitations lie. A foundational overview of this technology is available through the [NCBI Bookshelf](https://www.ncbi.nlm.nih.gov/books/), which catalogs authoritative medical and genomic reference materials.

Unlike whole genome sequencing, exome sequencing focuses exclusively on exons and their flanking splice sites, making it more cost effective and computationally manageable while still capturing the majority of actionable variants. However, it misses structural variants, deep intronic mutations, and regulatory element changes. For a structured introduction to the underlying principles and data interpretation, the [EMBL-EBI Training](https://www.ebi.ac.uk/training/) platform offers modules on next generation sequencing analysis that apply directly to exome data.

## At a Glance

| Feature | Description | Key Considerations |
| --- | --- | --- |
| Target | Protein coding exons (~30 Mb) | Covers 1-2% of genome, ~85% of known disease variants |
| Typical Coverage | 80x-100x depth | Higher depth improves detection of somatic and mosaic variants |
| Cost | Moderate (lower than whole genome, higher than gene panels) | Balance between breadth and budget |
| Turnaround | 2-6 weeks (laboratory + bioinformatics) | Depends on throughput and analysis pipeline complexity |
| Common Applications | Rare disease diagnosis, cancer somatic mutation profiling, neurodevelopmental disorder studies | Requires careful variant interpretation and family segregation |

## Decision Criteria

Choose exome sequencing when the clinical or research question requires scanning all protein coding genes without a strong prior hypothesis. This approach outperforms targeted gene panels for heterogeneous conditions such as neurodevelopmental disorders or undiagnosed genetic syndromes. For example, a feasibility pilot study in children with nonsyndromic neurodevelopmental disorders demonstrated that exome sequencing provided a diagnosis in a significant proportion of cases that would have been missed by smaller panels ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42442334/)). Conversely, if a specific gene or pathway is already suspected, a targeted panel delivers faster and cheaper results. Whole genome sequencing becomes necessary when non coding variants, large structural rearrangements, or mitochondrial genome mutations are suspected, as exome sequencing does not reliably capture these. For somatic applications like hepatocellular carcinoma, exome sequencing can identify protein coding mutations, but mitochondrial genome mutations require dedicated mitochondrial sequencing ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42443345/)).

## Practical Workflow or Implementation Steps

A standard exome sequencing project follows a defined workflow. The [Galaxy Training Network](https://training.galaxyproject.org/) provides step by step tutorials for each phase.

1. **Sample Collection and DNA Extraction**  
   Obtain high molecular weight genomic DNA from blood, saliva, or tissue. Check DNA integrity and quantity using spectrophotometry and gel electrophoresis.

2. **Library Preparation**  
   Fragment DNA to 150-200 base pair fragments. Attach adapters and then hybridize to biotinylated probes that capture exonic regions. Wash away non target DNA and amplify the enriched library.

3. **Sequencing**  
   Sequence the library on a high throughput platform (e.g., Illumina) to achieve a mean target depth of 80x to 100x. Store raw sequencing reads in FASTQ format and deposit them in public repositories such as the [NCBI Sequence Read Archive](https://www.ncbi.nlm.nih.gov/sra) to comply with data sharing policies.

4. **Bioinformatics Analysis**  
   This critical step aligns reads to a reference genome (GRCh38), calls variants, annotates their functional impact, and filters against population databases. The [Bioconductor](https://bioconductor.org/) project offers R packages like `Rsubread`, `VariantAnnotation`, and `TVTB` that can handle alignment, variant calling, and annotation in a reproducible manner.

5. **Variant Interpretation**  
   Prioritize rare, protein altering variants (missense, nonsense, splice site, frameshift) that segregate with the phenotype. Use software tools and manual review of read alignments to confirm calls. Cross reference with clinical databases and peer reviewed literature to assess pathogenicity.

6. **Validation and Reporting**  
   Confirm candidate variants using an orthogonal method (e.g., Sanger sequencing). Generate a clinical report that lists pathogenic, likely pathogenic, and variants of uncertain significance.

## Quality Checks

When running or reviewing an exome sequencing analysis, monitor these quality metrics. The [EMBL-EBI Training](https://www.ebi.ac.uk/training/) materials emphasize the following:

- **Coverage Uniformity**: At least 90% of target bases should be covered at 20x or higher. Low coverage regions, especially GC rich exons, can lead to false negatives.
- **Mapping Quality**: The percentage of reads align correctly to the reference. Values below 90% indicate contamination or technical error.
- **Variant Quality Filters**: Apply GATK recommended filters (QD < 2.0, FS > 60.0, MQ < 40.0). Manually inspect suspicious calls.
- **Transition to Transversion Ratio**: Expected ratio is approximately 2.1 to 2.5. Deviations suggest systematic errors.

## Common Mistakes

Even experienced teams can stumble on recurrent pitfalls. One frequent error is attributing a variant solely to exonic sequences while ignoring critical splice site regions that extend into introns. Many pathogenic variants lie within the first two base pairs of intronic splice donor or acceptor sites, which are captured by standard exome probe designs. However, deeper intronic mutations that create cryptic splice sites are missed. Another mistake is selecting too lenient a coverage threshold. A study analyzing craniofacial features in children with neurodevelopmental disorders found that low coverage in certain exons led to undercalling of pathogenic variants ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42443842/)). Also, failing to filter out common population variants (e.g., >1% allele frequency in gnomAD) inflates the VUS list. Finally, researchers sometimes overlook mosaicism. Somatic mutations present at low allele fractions require deep coverage and cautious interpretation, especially in cancer and some neurodevelopmental cases.

## Limits and Uncertainty

Exome sequencing does not capture the complete genetic picture. It cannot reliably detect large deletions, duplications, inversions, or translocations unless dedicated copy number variant calling algorithms are applied, and even then sensitivity is lower than chromosomal microarray. Non coding regulatory variants that affect gene expression are invisible. The exome itself is a targeted capture, and probe design varies by manufacturer. Some exonic regions, particularly those with high GC content or repetitive sequences, may be poorly covered. Furthermore, variant interpretation can be ambiguous. Many missense changes are classified as variants of uncertain significance and cannot guide clinical decisions without further segregation or functional studies. Pathogenicity predictors are not definitive. Even when a clear diagnosis is obtained, genetic heterogeneity means that a negative result does not rule out a genetic cause, it may require reanalysis or a different technology. A clinical example from pediatric Wilson disease shows that loss of function variants in ATP7B have clear prognostic implications, but other less severe missense variants show variable penetrance ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42442737/)). For cases with non obstructive azoospermia, exome sequencing identified novel variants in meiotic genes, but not all affected individuals had detectable mutations, indicating that additional causes remain hidden ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42443317/)). These limitations underscore the need for careful pretest counseling and multidisciplinary interpretation.

## Frequently Asked Questions

**How many variants does exome sequencing typically find?**  
On average, an exome yields about 20,000 to 30,000 variants relative to the reference genome. After applying quality filters and excluding common polymorphisms, about 200 to 500 rare or novel coding variants remain. Further phenotype driven filtering usually narrows the list to a handful of candidates.

**Can exome sequencing detect copy number variations?**  
Yes, but with lower sensitivity than dedicated platforms. Algorithms such as ExomeDepth or CNVkit can infer deletions and duplications from read depth differences. However, small CNVs and those in repetitive regions may be missed. Confirmation by an orthogonal method is recommended.

**How is exome sequencing used in cancer?**  
In oncology, exome sequencing is applied to tumor/normal paired samples to identify somatic driver mutations, mutational signatures, and potential therapeutic targets. It can also reveal germline predisposing variants. One study found that somatic mitochondrial DNA mutations in hepatocellular carcinoma were prognostically significant, though such mutations would be missed without specific mitochondrial enrichment ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42443345/)).

**What is the clinical utility for rare diseases?**  
Exome sequencing achieves a diagnostic yield of 25% to 40% in undiagnosed rare disease cohorts. It is particularly useful for conditions with high genetic heterogeneity, such as Dandy-Walker malformation, where several genes contribute and exome analysis revealed multiple associated new variants ([PubMed](https://pubmed.ncbi.nlm.nih.gov/42443849/)). The diagnostic rate improves when proband and parents are studied together.

## Related Clinical & Scientific Guides

* [Observational vs. Experimental Studies: How to Tell Them Apart](/blog/guides/observational-vs-experimental-studies-how-to-tell-them-apart)
* [Astrocyte Single Cell Rna Seq](/blog/guides/astrocyte-single-cell-rna-seq)
* [Structural Genes](/blog/guides/structural-genes)


## References and Further Reading

- **NCBI Bookshelf** , Comprehensive textbook chapters on genetic testing technologies, variant interpretation, and clinical genetics. Access at [https://www.ncbi.nlm.nih.gov/books/](https://www.ncbi.nlm.nih.gov/books/).
- **EMBL-EBI Training** , Interactive courses on sequence analysis, including exome sequencing pipelines and quality control. Visit [https://www.ebi.ac.uk/training/](https://www.ebi.ac.uk/training/).
- **Galaxy Training Network** , Freely available tutorials for constructing exome analysis workflows, from raw data to annotated variants. Explore at [https://training.galaxyproject.org/](https://training.galaxyproject.org/).
- **Bioconductor** , R software repository with packages for variant calling, annotation, and visualization. Documentation at [https://bioconductor.org/](https://bioconductor.org/).
- **NCBI Sequence Read Archive** , Public database for depositing and accessing raw sequencing data. See [https://www.ncbi.nlm.nih.gov/sra](https://www.ncbi.nlm.nih.gov/sra).
- **Feasibility Pilot to Expand Exome Sequencing Access for Children With Nonsyndromic Neurodevelopmental Disorders** , Clinical evidence on the utility of exome sequencing in a pediatric cohort. [PubMed](https://pubmed.ncbi.nlm.nih.gov/42442334/).
- **Craniofacial Features and Pathogenic Variants in 1,252 Children With Neurodevelopmental Disorders** , Demonstrates the importance of coverage and phenotype driven filtering. [PubMed](https://pubmed.ncbi.nlm.nih.gov/42443842/).
- **Clinical Characteristics and Genetic Analysis of Different Dandy-Walker Malformation** , Illustrates exome sequencing in a heterogeneous condition. [PubMed](https://pubmed.ncbi.nlm.nih.gov/42443849/).
- **Prognosis of Pediatric Hepatic Wilson Disease With ATP7B Loss of Function Variants** , Showcases how exome findings translate into prognostic stratification. [PubMed](https://pubmed.ncbi.nlm.nih.gov/42442737/).
- **Novel Variants in LINC and TTM Complexes of Meiotic Chromosome Dynamics** , An example of exome sequencing applied to reproductive genetics. [PubMed](https://pubmed.ncbi.nlm.nih.gov/42443317/).

## Related Articles

- [Protein Synthesis](/blog/guides/protein-synthesis)
- [Incomplete Dominance Gene](/blog/guides/incomplete-dominance-gene)
- [Cell Membrane Function Biology](/blog/guides/cell-membrane-function-biology)
- [Dna Structure](/blog/guides/dna-structure)
- [Protein Structure](/blog/guides/protein-structure)