Exome Sequencing Workflow: From Capture to Variant Calling
Exome sequencing targets the protein-coding regions of the genome, approximately 1 to 2 percent of the total sequence, to identify disease-relevant variants at a lower cost and computational burden than whole-genome sequencing. The workflow spans nucleic acid extraction, library preparation, target enrichment, sequencing, and bioinformatics analysis, with each stage carrying distinct quality control requirements. This article walks laboratory students, technicians, researchers, and diagnostic professionals through each step, with emphasis on capture methods, sequencing platforms, and variant calling pipelines, along with practical quality checks and common failure patterns.
At a Glance
The table below summarizes the main stages of the exome sequencing workflow, the key decisions at each stage, and the primary quality metrics to monitor.
| Workflow Stage | Core Decision | Primary Quality Metric |
|---|---|---|
| Sample preparation | DNA extraction method and input quantity | DNA integrity and concentration |
| Library preparation | Fragmentation method and adapter ligation | Library yield and insert size |
| Target enrichment | Hybridization capture versus amplicon-based capture | Enrichment efficiency and on-target rate |
| Sequencing | Read length and depth selection | Mean depth and coverage uniformity |
| Bioinformatics | Alignment and variant calling tools | Transition-transversion ratio and variant quality score |
Scope and Context for Exome Sequencing
Whole-exome sequencing has become a standard approach for causal gene detection in disease diagnosis and treatment management. The technology enables analysis of all protein-coding sequences in the human genome, which is where the majority of disease-related genetic aberrations are located. For rare diseases, most of which affect children, exome sequencing has increased the discovery rate of causative genes and improved diagnosis. A better understanding of the genetic basis of rare disease translates to more accurate prognosis, management, surveillance, and genetic advice for affected families.
The clinical utility of exome sequencing is well documented across multiple specialties. In a study of 400 patients with rare Mendelian disorders, a phenotype-driven proband-only exome sequencing strategy achieved an overall molecular diagnostic yield of 53 percent, identifying 243 pathogenic variants in 210 cases, 85 of which were novel. In primary immunodeficiency, a longitudinal study of 878 patients found disease-causing variants in 498 probands, encompassing 152 distinct monogenic disorders, with a diagnostic yield of 56 percent for targeted panel sequencing and 45 percent for a whole-exome-only approach. For monogenic kidney diseases, clinical exome sequencing identified genetic variants potentially explaining the phenotype in 56.5 percent of 138 recruited patients.
The choice between targeted panels and whole-exome sequencing depends on clinical context, cost, and the likelihood of identifying a causative variant. Whole-exome sequencing offers a simplified workflow, reduced overall cost in certain scenarios, and the potential for identification of novel diseases. Cost analysis based on current commercial prices demonstrated savings ranging from 300 to 950 dollars with a whole-exome-only approach, depending on diagnostic yield. However, exome sequencing has limitations, including restricted ability to detect copy number variants, lower coverage compared to targeted sequencing, and a lack of consensus regarding references and minimal application requirements.
Sample Preparation and Quality Assessment
DNA Extraction and Quantification
The first step in the exome sequencing workflow is the extraction of genomic DNA from the appropriate sample type. Blood is the most common source for germline testing, while formalin-fixed paraffin-embedded tissue is frequently used in oncology applications. The extraction method must yield DNA of sufficient quantity and purity for downstream library preparation.
DNA quantification should be performed using a method that distinguishes double-stranded DNA from RNA and single-stranded DNA. Fluorometric methods are preferred over spectrophotometric methods because they are more specific and less affected by contaminants. The quality of the DNA should be assessed by electrophoresis or microfluidic analysis to determine the degree of fragmentation. High molecular weight DNA with minimal degradation is required for optimal library preparation, particularly for hybridization-based capture methods.
For formalin-fixed paraffin-embedded derived DNA, the quality is often compromised by crosslinking and degradation. A hybridization-based enrichment protocol optimized for such samples demonstrated that two rounds of short hybridization provided the highest enrichment efficiency while maintaining complete target coverage. The optimized protocol achieved approximately sixfold higher enrichment than the hybridization-based analogue, ensuring adequate coverage despite increased duplicate rates.
Input Quantity and Its Effect on Results
The amount of input DNA directly affects library complexity and the ability to detect variants at low allele frequencies. Insufficient input DNA leads to duplicate reads and reduced sensitivity for variant detection. The input requirement varies by capture method, with hybridization-based methods generally requiring more input DNA than amplicon-based methods.
For clinical samples with limited DNA availability, such as small biopsies or degraded formalin-fixed paraffin-embedded tissue, the laboratory must document the input quantity and assess whether it meets the minimum requirement for the chosen workflow. Samples below the minimum input should be flagged for potential reduced sensitivity, and this limitation should be communicated to the requesting clinician.
Library Preparation
Fragmentation and Size Selection
Genomic DNA must be fragmented into pieces of appropriate size for sequencing. Fragmentation can be achieved by physical methods such as sonication or enzymatic methods. The fragment size distribution affects the efficiency of target enrichment and the uniformity of coverage across the exome.
After fragmentation, size selection removes fragments that are too short or too long. The optimal insert size depends on the sequencing platform and the read length. For short-read sequencing, insert sizes typically range from 150 to 300 base pairs. The size selection step is critical because fragments that are too short may not hybridize efficiently to capture probes, while fragments that are too long may reduce cluster density on the flow cell.
End Repair, A-Tailing, and Adapter Ligation
The fragmented DNA undergoes end repair to create blunt ends, followed by A-tailing to add a single adenine to the 3-prime ends. Adapters containing sequencing primer binding sites and sample-specific barcodes are then ligated to the fragments. The barcodes enable multiplexing of multiple samples in a single sequencing run.
The efficiency of adapter ligation affects the library yield and the proportion of reads that contain the expected adapter sequences. Adapter dimers, which are adapter molecules ligated to each other without an insert, are a common artifact that wastes sequencing capacity. Quality control after adapter ligation should assess the library concentration and check for the presence of adapter dimers.
Target Enrichment Methods
Hybridization-Based Capture
Hybridization-based capture is the most widely used method for exome enrichment. In this approach, the library is denatured and hybridized to biotinylated probes that are complementary to the target regions. The hybridized fragments are then pulled down using streptavidin-coated beads, while non-target fragments are washed away.
The performance of hybridization-based capture depends on several factors, including hybridization temperature, duration, buffer composition, and the number of capture rounds. A study that systematically varied these conditions found that two rounds of short hybridization provided the highest enrichment efficiency while maintaining complete target coverage. The optimized protocol demonstrated approximately sixfold higher enrichment than the hybridization-based analogue, ensuring adequate coverage despite increased duplicate rates.
Hybridization-based capture offers several advantages for exome sequencing. It provides more uniform coverage across the target regions compared to amplicon-based methods, and it can be scaled from small panels to whole-exome capture without loss of enrichment efficiency or target depth. This scalability makes it suitable for precision oncology applications, including challenging formalin-fixed paraffin-embedded samples.
Amplicon-Based Capture
Amplicon-based methods use polymerase chain reaction to amplify target regions directly from genomic DNA. This approach requires less input DNA and has a faster turnaround time than hybridization-based capture. However, amplicon-based methods have limitations, including less uniform coverage and difficulty in detecting larger insertions or deletions.
The choice between hybridization-based and amplicon-based capture depends on the application. For targeted panels with a small number of genes, amplicon-based methods may be sufficient and more cost-effective. For whole-exome sequencing, hybridization-based capture is generally preferred because of its scalability and coverage uniformity.
Comparison of Capture Methods
Different exome enrichment kits exhibit variable efficiency across genomic regions, leading to systematic, non-biological batch effects that are much stronger than other technical factors. A workflow to minimize the effect of capture inconsistencies in single-nucleotide variant data includes quality control, mapping to the genome, variant calling, joint genotyping, and imputation of genotypes using reference haplotypes. Variants are then aggregated into gene-level features measuring the burden of deleterious mutations.
When samples are enriched with different capture kits across multiple centers, the detection rate of a gene may be low in samples enriched with one kit but high in samples enriched with other kits. Such differences are unlikely to reflect true biology and should be corrected through gene-level imputation. A study conducted on over a thousand breast cancer cases across 11 cohorts using 8 exome capture kits demonstrated that this pipeline leads to a considerable decrease in the batch effect signal, potentially increasing the likelihood of finding true biological signals.
Sequencing
Platform Selection and Read Length
The choice of sequencing platform affects read length, throughput, cost, and error profiles. Short-read platforms are the most commonly used for exome sequencing, with read lengths typically ranging from 75 to 150 base pairs. Longer reads can improve the ability to detect structural variants and resolve repetitive regions, but they are not yet standard for exome sequencing.
The sequencing platform must be selected based on the number of samples to be processed, the required depth of coverage, and the turnaround time. Scalability testing of an optimized hybridization-based enrichment protocol showed performance comparable to a commercial whole-exome kit, with maintained enrichment efficiency and target depth across large panels.
Depth of Coverage
Depth of coverage refers to the average number of times each base in the target region is sequenced. Higher depth increases the sensitivity for detecting variants, particularly those present at low allele frequencies. For germline variant detection, a mean depth of 80 to 100 reads is generally considered adequate. For somatic variant detection in oncology, higher depth may be required to detect variants present in a small fraction of cells.
Coverage uniformity is as important as mean depth. Regions with low coverage may harbor variants that are missed, while regions with very high coverage consume sequencing capacity without additional benefit. The percentage of target bases covered at a minimum depth, such as 20 reads, is a useful metric for assessing coverage uniformity.
Multiplexing and Batching
Multiplexing allows multiple samples to be sequenced in a single run by using sample-specific barcodes. The number of samples that can be multiplexed depends on the required depth of coverage and the throughput of the sequencing platform. Higher multiplexing reduces the cost per sample but also reduces the depth of coverage per sample.
Batching decisions affect the ability to compare variants across samples. Samples processed in the same batch are subject to the same technical conditions, which reduces batch effects. However, batch effects can still arise from differences in capture efficiency across genomic regions, particularly when different capture kits are used.
Bioinformatics Analysis Pipeline
Quality Control of Raw Sequencing Data
The first step in the bioinformatics analysis is quality control of the raw sequencing data. This involves assessing the quality scores of each base, the GC content distribution, the presence of adapter sequences, and the duplication rate. Tools such as FastQC are commonly used for this purpose.
Quality control metrics should be reviewed before proceeding with alignment. Samples with poor quality scores, high adapter contamination, or excessive duplication should be flagged for potential re-sequencing or re-analysis. The quality control step is critical because errors introduced at this stage propagate through the entire analysis pipeline.
Alignment to the Reference Genome
The quality-filtered reads are aligned to a reference genome, such as the human genome build hg38. The alignment step assigns each read to its genomic location, allowing for mismatches that may represent true variants or sequencing errors. The choice of alignment tool affects the accuracy and speed of the alignment.
After alignment, the results are typically stored in a binary alignment format. The alignment files are then sorted by genomic position, and duplicate reads are removed. Duplicate removal is important because duplicate reads can arise from polymerase chain reaction amplification during library preparation, and they can bias variant allele frequency estimates.
Variant Calling
Variant calling identifies positions where the sequenced sample differs from the reference genome. The output is a variant call format file that contains information about each variant, including its genomic position, reference and alternate alleles, quality scores, and genotype.
Variant calling algorithms for single-nucleotide variants range from standalone tools to machine learning-based combined pipelines. The choice of variant caller affects the sensitivity and specificity of variant detection. For copy number variants, tools compare the number of reads aligned to a dedicated segment, with the read depth serving as a proxy for copy number.
A targeted spatial masking strategy has been developed to suppress deterministic artifacts in short-read sequencing data while preserving clinically actionable variants outside low-complexity regions. This protocol removed thousands of sequencing and alignment artifacts while maintaining the retained biological callset, with negligible disease-associated diagnostic variants detected in the excluded artifact fraction. Low-complexity region masking preserved physiological transition-transversion and insertion-deletion profiles in retained calls, resolved pseudo-multiallelic noise, and distinguished excluded artifact calls by distorted mutational and variant allele frequency signatures.
Variant Annotation and Prioritization
The raw variant call format file contains many variants, most of which are benign or of unknown significance. Variant annotation adds information about each variant, including its location in genes, its predicted effect on protein function, and its frequency in population databases. Variant prioritization then ranks the variants based on their likelihood of being disease-causing.
Phenotype-driven variant filtration strategies have been shown to increase diagnostic yield. In a study of 400 patients with rare Mendelian disorders, a phenotype-driven proband-only exome sequencing strategy achieved an overall molecular diagnostic yield of 53 percent. The strategy involved filtering variants based on the patient's phenotype and prioritizing variants in genes known to be associated with the presenting features.
A semiautomated and phenotype-driven whole-exome sequencing diagnostic workflow, incorporating both the DRAGEN pipeline and the Exomiser variant prioritization tool, achieved a 41 percent molecular diagnostic rate for 66 duo-, quad-, or trio-whole-exome sequencing cases, and 28 percent for 40 singleton cases. Preliminary results were returned to ordering physicians within one week for 12 of 38 probands with positive findings, which were instrumental in guiding appropriate clinical management, especially in critical care settings.
Quality Control and Documentation
Quality Metrics at Each Step
Quality control should be performed at each step of the exome sequencing workflow. The table below summarizes the key quality metrics and the actions to take when metrics fall outside acceptable ranges.
| Quality Metric | Acceptable Range | Action When Out of Range |
|---|---|---|
| DNA integrity number | 7 or higher for fresh samples | Re-extract or request new sample |
| Library yield | Sufficient for capture and sequencing | Re-amplify or re-prepare library |
| On-target rate | 60 percent or higher for hybridization capture | Optimize hybridization conditions |
| Mean depth of coverage | 80 to 100 reads for germline | Increase sequencing output |
| Percentage of target bases at 20 reads | 95 percent or higher | Increase sequencing output |
| Transition-transversion ratio | 2.0 to 2.2 for whole-exome | Review variant calling parameters |
Laboratory Quality Management
The World Health Organization Laboratory Quality Management System Handbook provides guidance for establishing and maintaining quality in laboratory testing. Key elements include documentation of procedures, training of personnel, internal quality control, and external quality assessment. For exome sequencing, the laboratory should maintain records of all steps in the workflow, including sample receipt, DNA extraction, library preparation, capture, sequencing, and bioinformatics analysis.
Standard operating procedures should be written for each step and reviewed periodically. Personnel should be trained and assessed for competency. Internal quality control samples, such as a well-characterized reference sample, should be included in each batch to monitor assay performance over time. External quality assessment programs provide an independent assessment of the laboratory's performance.
Records and Traceability
Accurate records are essential for the interpretation and reporting of exome sequencing results. Each sample should be tracked through the workflow using a unique identifier. The records should include the sample type, extraction method, DNA quantity and quality, library preparation method, capture method, sequencing platform, and bioinformatics pipeline version.
The bioinformatics analysis should be reproducible, with the software versions and parameters documented. The variant call format file should be archived along with the alignment files and the raw sequencing data. The storage requirements for these files are substantial, and the laboratory should have a data management plan that addresses storage capacity, backup, and retention periods.
Common Failure Patterns and Troubleshooting
Low Enrichment Efficiency
Low enrichment efficiency results in a low proportion of reads mapping to the target regions. This can be caused by suboptimal hybridization conditions, insufficient probe concentration, or degradation of the probes. The optimization of hybridization conditions, including temperature, duration, buffer composition, and the number of capture rounds, can improve enrichment efficiency.
If the on-target rate is low, the laboratory should review the hybridization conditions and consider increasing the number of capture rounds. Two rounds of short hybridization provided the highest enrichment efficiency while maintaining complete target coverage in an optimized protocol.
High Duplicate Rate
A high duplicate rate indicates that many reads are identical, which reduces the effective depth of coverage. This can be caused by insufficient input DNA, excessive polymerase chain reaction amplification, or over-sequencing of the library. The duplicate rate should be monitored, and samples with high duplicate rates should be flagged for potential re-preparation.
Coverage Gaps
Coverage gaps are regions of the exome that are not sequenced at sufficient depth. These gaps can be caused by GC bias, repetitive sequences, or poor capture efficiency in specific regions. The percentage of target bases covered at a minimum depth should be reported, and regions with consistently low coverage should be documented as limitations of the assay.
Batch Effects
Batch effects are systematic, non-biological differences between samples processed in different batches or with different capture kits. These effects are much stronger than other technical factors in whole-exome sequencing. A workflow that includes quality control, mapping, variant calling, joint genotyping, and imputation using reference haplotypes can minimize the effect of capture inconsistencies.
If the detection rate of a gene is low in samples enriched with a given capture kit but high in samples enriched with other kits, missing values in the former group should be imputed, as such differences are unlikely to reflect true biology.
Limitations of Exome Sequencing
Detection of Copy Number Variants
Exome sequencing has a restricted ability to detect copy number variants compared to whole-genome sequencing. The non-uniform coverage of the exome makes it difficult to accurately estimate copy number from read depth. However, recent technological advances have enabled copy number variant calling from exome sequencing data with accurate and highly sensitive bioinformatic tools.
In a study of 920 patients referred for whole-exome sequencing, 454 unresolved cases were further analyzed using the ExomeDepth algorithm. Causative copy number variants were identified in 40 patients, increasing the diagnostic yield from 50.7 percent to 55 percent. Twenty-two copy number variants were available for validation and were all confirmed, of these, five were novel.
Coverage Limitations
Exome sequencing provides lower coverage compared to targeted sequencing. The depth of coverage across the exome is not uniform, and some regions may have insufficient coverage for reliable variant calling. The missing consensus regarding references and minimal application requirements is a limitation for clinical applications.
Variant Interpretation Challenges
The interpretation of variants identified by exome sequencing requires careful consideration of the evidence for pathogenicity. Variants are classified according to guidelines from professional organizations, and unreported variants require expert review. In a study of 764 individuals with dystonia, all considered variants were reviewed in expert round-table sessions to validate their clinical significance.
The availability of regional genomic references affects the interpretation of variants. In a study of three Middle Eastern pediatric patients with genodermatoses, whole-exome sequencing identified pathogenic variants in all three cases, but the efficacy remained contingent on the availability of regional genomic references.
Safety and Regulatory Context
Laboratory Biosafety
The World Health Organization Laboratory Biosafety Manual provides guidance for the safe handling of biological materials. The risk assessment should consider the sample type, the potential for infectious agents, and the procedures being performed. Standard precautions, including the use of personal protective equipment and hand hygiene, should be followed when handling blood and tissue samples.
The extraction of DNA from blood or tissue samples should be performed in a designated area with appropriate containment. Samples should be transported and stored according to the laboratory's biosafety procedures. Waste materials, including used extraction columns and contaminated consumables, should be disposed of according to local regulations.
Regulatory Considerations
The World Health Organization Laboratory Quality Management System Handbook emphasizes the importance of meeting regulatory requirements for laboratory testing. For diagnostic exome sequencing, the laboratory should be accredited by an appropriate body and participate in external quality assessment programs.
The U.S. Food and Drug Administration Bioanalytical Method Validation Guidance provides a framework for validating analytical methods, although it is primarily focused on pharmacokinetic and toxicokinetic studies. The principles of method validation, including accuracy, precision, sensitivity, and specificity, are applicable to exome sequencing assays.
The National Center for Advancing Translational Sciences Assay Guidance Manual provides recommendations for the development and validation of assays used in drug discovery and development. The principles of assay validation, including the use of appropriate controls and the assessment of assay performance, are relevant to exome sequencing.
Professional Escalation Criteria
When to Escalate to a Senior Scientist
Laboratory personnel should escalate to a senior scientist when quality metrics fall outside acceptable ranges and troubleshooting does not resolve the issue. Examples include persistent low enrichment efficiency, high duplicate rates, or coverage gaps that cannot be explained by sample quality.
When to Escalate to the Ordering Clinician
The laboratory should escalate to the ordering clinician when the results have implications for patient management that require clinical interpretation. Examples include the identification of variants in genes associated with actionable conditions, the identification of variants of unknown significance in genes relevant to the indication, or the failure to identify a causative variant despite adequate coverage.
When to Recommend Additional Testing
The laboratory should recommend additional testing when exome sequencing does not identify a causative variant but the clinical suspicion for a genetic etiology remains high. Additional testing may include whole-genome sequencing, targeted testing for specific variant types, or testing of additional family members.
Frequently Asked Questions
What is the difference between whole-exome sequencing and targeted panel sequencing?
Whole-exome sequencing analyzes all protein-coding sequences in the genome, while targeted panel sequencing analyzes a predefined set of genes. Whole-exome sequencing offers a simplified workflow, reduced overall cost in certain scenarios, and the potential for identification of novel diseases. Targeted panels may provide higher depth of coverage for the included genes and are often faster to interpret.
How much DNA is needed for exome sequencing?
The amount of DNA needed depends on the capture method and the library preparation protocol. Hybridization-based methods generally require more input DNA than amplicon-based methods. The laboratory should document the input quantity and assess whether it meets the minimum requirement for the chosen workflow.
What is the recommended depth of coverage for exome sequencing?
For germline variant detection, a mean depth of 80 to 100 reads is generally considered adequate. For somatic variant detection in oncology, higher depth may be required to detect variants present in a small fraction of cells. The percentage of target bases covered at a minimum depth, such as 20 reads, is a useful metric for assessing coverage uniformity.
How long does the exome sequencing workflow take?
The turnaround time depends on the workflow design and the urgency of the clinical situation. A semiautomated workflow returned preliminary results to ordering physicians within one week for 32 percent of probands with positive findings. The turnaround time includes sample preparation, library preparation, capture, sequencing, and bioinformatics analysis.
What are the main limitations of exome sequencing?
Exome sequencing has a restricted ability to detect copy number variants, lower coverage compared to targeted sequencing, and a lack of consensus regarding references and minimal application requirements. Some regions of the exome may have insufficient coverage for reliable variant calling.
How are copy number variants detected from exome sequencing data?
Copy number variants are detected by comparing the number of reads aligned to a dedicated segment, with the read depth serving as a proxy for copy number. Tools such as ExomeDepth have been used to identify causative copy number variants in patients who remained undiagnosed after initial variant analysis.
What is the diagnostic yield of exome sequencing?
The diagnostic yield varies by indication and patient population. Studies have reported yields of 53 percent in patients with rare Mendelian disorders, 56 percent in patients with primary immunodeficiency, and 56.5 percent in patients with monogenic kidney diseases. The yield depends on the clinical characteristics of the patient population and the bioinformatics analysis approach.
How are variants classified for clinical reporting?
Variants are classified according to professional guidelines, such as those from the American College of Medical Genetics and Genomics. Unreported variants are classified based on the evidence for pathogenicity, and all considered variants are reviewed in expert round-table sessions to validate their clinical significance.
Related Diagnostic Guides
- Quality Control Analysis: Methods for Monitoring Lab Performance
- Procedure for Quality Control: Step-by-Step Implementation in a Molecular Lab
- DNA Shearing for NGS Library Preparation: Methods and Quality Control
- How to Perform a Gram Stain: Protocol and Quality Control
- How to Interpret DNA Sequencing Chromatograms: Peaks, Quality, and Heterozygotes
References and Further Reading
- Laboratory Quality Management System Handbook. World Health Organization.
- Laboratory Biosafety Manual. World Health Organization.
- Assay Guidance Manual. National Center for Advancing Translational Sciences.
- Bioanalytical Method Validation Guidance. U.S. Food and Drug Administration.
- NCBI Literature Resources. National Center for Biotechnology Information.
- A semiautomated whole-exome sequencing workflow leads to increased diagnostic yield and identification of novel candidate variants.. Cold Spring Harbor molecular case studies, 2019.
- Monogenic variants in dystonia: an exome-wide sequencing study.. The Lancet. Neurology, 2020.
- Efficacy and economics of targeted panel versus whole-exome sequencing in 878 patients with suspected primary immunodeficiency.. The Journal of allergy and clinical immunology, 2021.
- Phenotype-driven variant filtration strategy in exome sequencing toward a high diagnostic yield and identification of 85 novel variants in 400 patients with rare Mendelian disorders.. American journal of medical genetics. Part A, 2021.
- Clinical exome sequencing is a powerful tool in the diagnostic flow of monogenic kidney diseases: an Italian experience.. Journal of nephrology, 2021.
- Paediatric genomics: diagnosing rare disease in children.. Nature reviews. Genetics, 2018.
- Unexplained Intellectual Disability: Diagnostic Workflow Moving Towards "Exome Sequencing First Approach"?. Indian journal of pediatrics, 2024.
- Meta-analysis of tumor- and T cell-intrinsic mechanisms of sensitization to checkpoint inhibition.. Cell, 2021.
- Optimization of a hybridization-based target enrichment protocol for precision oncology.. 2026.
- Whole-Exome Sequencing Identifies Recurrent Germline-Associated and Somatic Variants in Oral Squamous Cell Carcinoma from Southwest India. 2026.
- Targeted Genomic Region Masking Supports Accurate Variant Calling While Suppressing Low-Complexity Sequencing Artifacts.. 2026.
- A Practical Workflow for Correcting Kit-Specific Effects in Whole-Exome Sequencing Data. 2026.
- Whole Exome Sequencing Reveals Promising Genes Associated with Congenital Renal Parenchymal Anomalies in Greek Children. 2026.
- Fibroblast growth factor receptor inhibition for succinate dehydrogenase-deficient gastrointestinal stromal tumors: a phase 2 trial.. 2026.
- Leveraging Whole-Exome Sequencing to Decipher the Genetic Landscape of Three Genodermatoses' Cases in Middle Eastern Pediatric Patients.. 2026.
- Exome Sequencing Data Analysis and a Case-Control Study in Mexican Population Reveals Lipid Trait Associations of New and Known Genetic Variants in Dyslipidemia-Associated Loci. Frontiers in Genetics, 2022.
- Comprehensive Outline of Whole Exome Sequencing Data Analysis Tools Available in Clinical Oncology. Cancers, 2019.
- Whole Exome Sequencing Data Analysis for Detection of Breast Cancer Gene Variants and Pathway Study. International Journal of Current Research and Review, 2022.
- Scalable mixed model methods for set-based association studies on large-scale categorical data analysis and its application to exome-sequencing data in UK Biobank.. American Journal of Human Genetics, 2023.
- Clinical pharmacogenetic analysis in 5,001 individuals with diagnostic Exome Sequencing data. npj Genomic Medicine, 2022.
- Exome Sequencing Data Analysis. Encyclopedia of Bioinformatics and Computational Biology, 2019.
- Germline CNV Detection through Whole-Exome Sequencing (WES) Data Analysis Enhances Resolution of Rare Genetic Diseases. Genes, 2023.
- Whole Exome Sequencing Data Analysis Algorithms in Cancer Diagnostics. Prime Archives in Cancer Research, 2020.
- A highly sensitive and specific workflow for detecting rare copy-number variants from exome sequencing data. Genome Medicine, 2020.
- Leveraging the power of high performance computing for next generation sequencing data analysis: Tricks and twists from a high throughput exome workflow. Plos One, 2015.
- HPexome: An automated tool for processing whole-exome sequencing data. Softwarex, 2020.
- GermVarX: A Robust Workflow for Joint Germline Variant Exploration in whole-exome sequencing cohorts. Plos One, 2026.
This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.