Detecting Copy Number Variants from Whole-Exome Sequencing: Read-Depth Methods and Best Practices
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Read-depth methods for detecting copy number variants (CNVs) from whole-exome sequencing (WES) exploit the principle that duplicated genomic regions yield higher sequencing depth and deleted regions yield lower depth. This approach is crucial as standard variant callers often miss CNVs, which are clinically relevant in approximately 11% of unsolved rare monogenic disease cases.
- Normalization is paramount to account for systematic biases inherent in WES, such as probe hybridization efficiency and GC content, which significantly influence read depth. Reference-based normalization, comparing test samples to a set of 20+ similarly processed control samples, is a critical strategy to mitigate these technical variations.
- Exome-based CNV calling differs from whole-genome sequencing due to the fragmented, noncontiguous nature of exon capture, limiting resolution to target boundaries and necessitating more conservative statistical thresholds. Tools like ExomeDepth, which builds an optimized reference set per sample, and CNVkit, which incorporates off-target reads for smoother profiles, offer distinct but complementary approaches.
- Rigorous quality control and validation are indispensable for reporting defensible CNV findings. This includes assessing input data quality (e.g., mean depth >30x), defining accurate target regions, and critically, confirming candidate calls with independent methods such as qPCR, MLPA, or microarray before clinical reporting.
- Common failure patterns include single-exon artifacts, which are frequent false positives and require validation, and biases introduced by GC-rich regions or batch effects, necessitating careful reference set selection and management. Understanding these limitations is vital for accurate interpretation in a clinical context.
Whole-exome sequencing (WES) generates targeted sequence data that can be used to detect copy number variants (CNVs), but the noncontiguous nature of exon capture introduces systematic biases that require specialized read-depth analysis methods. This article explains how read-depth-based CNV detection works in exome data, which normalization strategies address capture biases, how tools such as ExomeDepth and CNVkit fit into a practical workflow, and what quality controls and interpretation limits researchers must apply before reporting findings.
The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has access to WES data and needs a reliable path from raw alignments to defensible CNV calls. The focus is on germline CNV detection from targeted exome capture, with attention to the decisions that determine whether a call is trustworthy enough for downstream interpretation.
The Problem That Read-Depth Methods Solve
Standard variant calling pipelines for WES data typically emphasize single-nucleotide variants and small insertions or deletions. Copy number variants, which involve larger duplications or deletions of genomic sequence, are often missed by these approaches because they do not manifest as a single altered base position. The clinical relevance of this gap is well documented. In a large study of families with suspected rare monogenic disease, reanalysis of exome data using additional analytic methods such as CNV calling identified pathogenic variants in roughly 11% of families who received a molecular diagnosis after prior nondiagnostic exome sequencing [<a href="#ref-1">1</a>]. This finding demonstrates that CNV detection from existing exome data can resolve cases that standard analysis leaves unsolved.
The challenge is that WES data are not naturally suited to copy number analysis. Unlike whole-genome sequencing, where reads are distributed across the genome in a relatively continuous manner, exome capture enriches only the protein-coding regions. The resulting coverage is fragmented, uneven, and influenced by the efficiency of hybridization probes, the GC content of captured sequences, and batch effects introduced during library preparation and sequencing [<a href="#ref-2">2</a>]. Read-depth methods exploit the simple principle that a duplicated region produces roughly twice the sequencing depth of a normal region, while a deleted region produces roughly half the depth. The difficulty lies in distinguishing these true copy number changes from the substantial technical noise that is inherent to targeted capture.
Core Principles of Read-Depth CNV Detection
Read-depth methods for CNV detection operate on a straightforward biological signal. The number of sequencing reads that align to a genomic interval is proportional to the number of copies of that interval in the input DNA. If a sample carries a heterozygous deletion, the depth in that region drops to approximately half the expected value. If a sample carries a duplication, the depth rises to approximately one and a half times the expected value for a heterozygous event or twice the expected value for a homozygous duplication.
The practical difficulty is that raw read depth is influenced by many factors unrelated to copy number. These include the efficiency of the capture probes for each exon, the GC content of the target sequence, the local sequence complexity, and the overall sequencing depth of the sample. A successful read-depth workflow must therefore estimate the expected depth for each target region under normal copy number conditions and then compare the observed depth in the test sample against that expectation.
The Role of Normalization
Normalization is the process of removing systematic technical variation so that the remaining depth signal reflects biological copy number. The most common approach is to compare the test sample against a set of reference samples that were processed with the same capture kit and similar sequencing protocols. The reference set provides an empirical estimate of the typical depth profile across all target regions. The test sample depth is then expressed as a ratio relative to this reference profile.
Several factors must be controlled during normalization. GC content has a well-known effect on sequencing depth, and most tools apply a GC correction step. Capture kit design also matters because different kits have different probe densities and target definitions. Batch effects, which arise when samples are processed at different times or in different laboratories, can introduce systematic differences that mimic copy number changes. The choice of reference samples is therefore one of the most consequential decisions in the entire workflow.
Why Exome CNV Calling Differs from Genome CNV Calling
Whole-genome sequencing data allow CNV callers to use sliding windows across the genome, which provides many data points and relatively stable statistical properties. Exome data do not offer this luxury. The targets are discrete exons separated by large intronic gaps, so the analysis is restricted to the captured regions. Each exon provides a limited number of reads, and the depth signal is noisier than in genome data. The noncontiguous nature of the capture is the primary bottleneck for accurate CNV detection in exome data [<a href="#ref-2">2</a>].
This difference has practical consequences. Exome-based CNV callers must be more conservative in their statistical thresholds, and the resolution of breakpoints is limited to the boundaries of the captured targets. A deletion that removes part of an exon may be difficult to distinguish from a deletion that removes the entire exon. Despite these limitations, exome-based CNV detection can achieve sensitivity comparable to exon-resolution chromosomal microarrays when multiple algorithms are integrated and the analysis is performed carefully [<a href="#ref-3">3</a>].
At a Glance: Read-Depth CNV Detection from WES
| Decision Point | Recommended Approach | Key Consideration |
|---|---|---|
| Input data | Aligned BAM files from WES with consistent capture kit | Mixing capture kits within one analysis introduces systematic bias |
| Reference set | 20 or more samples processed with the same kit and protocol | Smaller reference sets increase false positive rates |
| Normalization | GC correction plus reference-based depth ratio | Raw depth is not comparable across samples without normalization |
| Primary tool | ExomeDepth for germline exome CNV calling | Designed specifically for targeted capture data |
| Secondary tool | CNVkit for segmentation and visualization | Useful for orthogonal confirmation of candidate calls |
| Validation | Independent method such as qPCR, MLPA, or microarray | Read-depth calls require confirmation before clinical reporting |
| Reporting | Annotate with gene names, exon counts, and population frequencies | Interpret in the context of the phenotype and inheritance pattern |
Practical Workflow for Read-Depth CNV Detection
A reliable read-depth CNV workflow from WES data follows a sequence of steps that move from raw alignments to interpreted calls. Each step has specific inputs, outputs, and quality checks.
Step 1: Assess Input Data Quality
Before any CNV analysis begins, the aligned BAM files must be checked for basic quality metrics. The mean depth across target regions should be sufficient for confident calling, typically at least 30x for germline analysis. Samples with very low depth will produce noisy ratios that are difficult to interpret. The capture kit used for each sample must be recorded because the analysis depends on knowing which regions were targeted.
The alignment files should also be checked for consistency. If some samples were aligned to an older reference genome build, the coordinates will not match the target definitions used by the CNV caller. All samples in a single analysis should be aligned to the same reference build and processed with the same alignment and preprocessing pipeline.
Step 2: Define Target Regions
The target regions for the capture kit must be available in a format that the CNV caller can use. These regions are typically provided as BED files by the kit manufacturer. The target file defines the exons or capture probes that will be analyzed. Some tools require the targets to be flattened or merged to avoid overlapping intervals.
The quality of the target file directly affects the analysis. If the target file does not match the kit that was actually used for capture, the depth calculations will be wrong. Laboratory professionals should verify that the target file version matches the kit version recorded in the sample metadata.
Step 3: Build or Select a Reference Set
The reference set is the foundation of read-depth normalization. The ideal reference set consists of samples that were processed with the same capture kit, the same sequencing platform, and similar library preparation protocols. The samples should be free of known CNVs in the regions of interest, although in practice this is difficult to guarantee for every locus.
The size of the reference set matters. A larger reference set provides a more stable estimate of the expected depth profile and reduces the impact of any single outlier sample. Many tools recommend a minimum of 20 reference samples, and larger sets of 50 or more improve performance. The reference samples should be selected without knowledge of their CNV status at the loci being tested, because selecting samples based on normal results at one locus can bias the analysis at other loci.
Step 4: Run the CNV Caller
ExomeDepth is a widely used tool for germline CNV detection from exome data. It operates by constructing an optimized reference set for each test sample, selecting a subset of reference samples that best match the test sample's overall depth profile. This approach reduces the impact of batch effects because the reference is tailored to the test sample instead of being a fixed set.
CNVkit is another option that uses a different strategy. It models the read depth across the genome, including off-target reads that align outside the captured regions, to create a smoother depth profile. This approach can improve the resolution of CNV boundaries and is particularly useful for visualizing large-scale copy number changes.
The choice of tool depends on the specific question. ExomeDepth is designed for exome data and performs well for detecting exonic deletions and duplications. CNVkit provides additional flexibility for visualization and can be used as a secondary method to confirm calls made by ExomeDepth.
Step 5: Filter and Annotate Candidate Calls
The raw output of a CNV caller contains many candidate events, and most of them will be false positives. Filtering is essential to reduce the list to a manageable set of high-confidence candidates. Common filtering criteria include the number of exons affected, the magnitude of the depth ratio change, the statistical significance of the call, and the presence of the same call in multiple samples.
Calls that affect a single exon are more likely to be artifacts than calls that affect multiple consecutive exons. The depth ratio should be consistent with the expected value for a heterozygous event, approximately 0.5 for a deletion and 1.5 for a duplication. Calls with extreme ratios that are not biologically plausible should be treated with suspicion.
Annotation adds biological context to each call. The affected genes, the number of exons involved, and the predicted functional consequence should be recorded. Population frequency databases can help identify common polymorphisms that are unlikely to be pathogenic.
Step 6: Validate Candidate Calls
Validation is the step that separates research findings from clinically actionable results. Read-depth calls from exome data are probabilistic estimates, and they can be wrong. The standard approach is to confirm candidate CNVs with an independent method such as quantitative PCR, multiplex ligation-dependent probe amplification, or chromosomal microarray.
The need for validation is supported by the observation that even well-optimized exome CNV calling pipelines benefit from integration of multiple algorithms. In the Deciphering Developmental Disorders study, combining calls from multiple exome-based CNV algorithms using machine learning produced a higher quality data set than any individual algorithm alone [<a href="#ref-3">3</a>]. This finding suggests that no single method is sufficient for high-confidence calling.
Normalization Strategies and Their Tradeoffs
The normalization strategy is the most important technical decision in read-depth CNV detection. Different tools implement different normalization approaches, and the choice affects sensitivity and specificity.
Reference-Based Normalization
Reference-based normalization compares the test sample to a set of control samples. The depth at each target region in the test sample is divided by the median depth across the reference samples at that same region. This ratio is then adjusted for GC content and other technical factors.
The main advantage of reference-based normalization is that it removes systematic biases that affect all samples equally. The main disadvantage is that it requires a well-matched reference set. If the reference samples were processed under different conditions than the test sample, the normalization will be imperfect and may introduce artifacts.
Within-Sample Normalization
Some methods normalize the depth within a single sample without using external references. These methods assume that the overall depth distribution across all targets is representative of normal copy number, and they identify regions that deviate from this distribution.
Within-sample normalization is attractive because it does not require a reference set, but it is less powerful than reference-based normalization. A sample with a large duplication or deletion will have a skewed depth distribution, and the method may fail to detect the event or may produce false calls elsewhere.
GC Correction
GC content is a major source of bias in sequencing depth. Regions with extreme GC content, either very high or very low, tend to be underrepresented in sequencing libraries. Most CNV callers apply a GC correction that models the relationship between GC content and depth and adjusts the observed depth accordingly.
The GC correction is applied differently by different tools. Some tools fit a loess curve to the relationship between GC content and depth, while others use a simpler binning approach. The choice of correction method can affect the detection of CNVs in GC-rich or GC-poor regions.
Batch Effect Management
Batch effects are systematic differences between groups of samples that were processed at different times, in different laboratories, or with different reagent lots. These effects can mimic CNVs because they alter the depth profile in a consistent way across multiple samples.
The most effective way to manage batch effects is to include samples from the same batch in the reference set. ExomeDepth addresses this issue by selecting an optimized reference set for each test sample, choosing the reference samples that best match the test sample's depth profile. This approach reduces the impact of batch effects without requiring explicit batch metadata.
Tool Selection: ExomeDepth and CNVkit in Practice
The choice of CNV calling tool depends on the data type, the research question, and the available computational resources. ExomeDepth and CNVkit represent two different approaches that are both useful for exome data.
ExomeDepth
ExomeDepth is designed specifically for detecting germline CNVs from targeted capture data. It uses a beta-binomial model to compare the read counts in the test sample against an optimized reference set. The tool selects a reference set for each test sample by maximizing the correlation between the test sample and the reference samples, which reduces the impact of batch effects.
The output of ExomeDepth is a list of candidate CNVs with associated statistical scores. The tool also generates plots that show the read depth ratios across the affected region, which are useful for visual inspection of candidate calls.
ExomeDepth is well suited for detecting exonic deletions and duplications in germline samples. It is less well suited for detecting large-scale copy number changes that span many megabases, because the exonic targets provide limited resolution for breakpoint mapping.
CNVkit
CNVkit uses a different strategy that incorporates both on-target and off-target reads. The off-target reads, which align outside the captured regions, provide information about the copy number state of the intervening genomic regions. This information allows CNVkit to create a smoother depth profile and to estimate copy number across larger genomic intervals.
CNVkit is particularly useful for visualization. It generates copy number profiles that show the depth ratios across chromosomes, which can reveal large-scale events that might be missed by exon-focused methods. The tool also provides segmentation algorithms that identify regions of constant copy number.
The main limitation of CNVkit for exome data is that the off-target reads are sparse and noisy. The off-target signal is not as reliable as the on-target signal, and the copy number estimates in off-target regions should be interpreted with caution.
Integration of Multiple Tools
The evidence from large clinical cohorts supports the use of multiple algorithms for exome CNV detection. In the DDD study, integrating calls from multiple exome-based CNV algorithms using random forest machine learning generated a higher quality data set than using individual algorithms [<a href="#ref-3">3</a>]. This finding suggests that no single tool captures all true CNVs without also producing false positives.
A practical approach is to run two tools with different methodologies and to focus on calls that are supported by both. Calls that are detected by only one tool should be treated as lower confidence and may require additional validation. The integration of multiple tools increases the computational burden but improves the reliability of the final call set.
Quality Controls and Records
The reliability of read-depth CNV detection depends on the quality of the input data and the rigor of the analysis. Laboratory professionals should implement quality controls at each stage of the workflow and maintain records that allow the analysis to be reproduced.
Sample-Level Quality Metrics
Each sample should have a record of the capture kit, the sequencing platform, the mean depth across targets, and the percentage of targets with adequate coverage. Samples with low mean depth or poor coverage uniformity should be flagged because they will produce noisy CNV calls.
The alignment quality should also be recorded. The percentage of reads that map to the target regions, the duplication rate, and the fraction of reads with mapping quality above a threshold are all relevant metrics. Samples with unusual alignment statistics may indicate library preparation problems that affect CNV calling.
Analysis-Level Quality Metrics
The CNV calling analysis should produce metrics that describe the quality of the reference set and the fit of the model. The correlation between the test sample and the selected reference samples is a useful diagnostic. Low correlation indicates that the reference set is poorly matched to the test sample, and the resulting calls should be interpreted with caution.
The number of candidate CNVs per sample is another useful metric. Samples with an unusually high number of calls may have poor data quality or may be outliers in the reference set. The distribution of calls across samples should be examined to identify systematic artifacts.
Records for Reproducibility
The analysis should be recorded in a way that allows it to be reproduced. The version of the reference genome, the version of the capture kit target file, the version of the CNV calling tool, and the parameters used for the analysis should all be documented. The list of reference samples used for each test sample should also be recorded, because the choice of reference set affects the results.
Reproducible analysis practices are emphasized across bioinformatics training resources. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility in genomic analysis [<a href="#ref-4">4</a>], and the nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-5">5</a>]. These resources are useful for laboratories that are establishing their own CNV analysis pipelines.
Common Failure Patterns in Read-Depth CNV Calling
Understanding the common ways that read-depth CNV calling fails helps researchers interpret their results and avoid reporting false findings.
Single-Exon Artifacts
The most common false positive in exome CNV calling is a single-exon call that does not represent a real copy number change. Single exons have limited read counts, and the depth ratio is noisy. A single-exon deletion call should be treated with suspicion unless it is supported by an independent method.
The biological relevance of single-exon events varies by gene. Some genes have exons that are known to be recurrently deleted or duplicated in disease, and single-exon events in these genes may be real. In the absence of such evidence, single-exon calls should be considered low confidence.
GC-Rich Region Failures
Regions with very high GC content are difficult to capture and sequence. The depth in these regions is often low, and the depth ratio is noisy. CNV calls in GC-rich regions are prone to false positives, particularly for deletions.
The GC correction applied by the CNV caller may not fully compensate for the bias in these regions. Researchers should examine the raw depth in GC-rich regions before accepting a CNV call.
Batch Effect Artifacts
Batch effects can produce systematic false positives that appear in multiple samples from the same batch. These artifacts are particularly dangerous because they look like real biological events. A CNV that is called in many samples from the same sequencing run should be examined for batch effects.
The best defense against batch effects is a well-matched reference set. If the reference set includes samples from the same batch as the test samples, the batch effect will be absorbed by the normalization. If the reference set is from a different batch, the batch effect will appear as a CNV.
Reference Set Contamination
The reference set can contain samples with real CNVs at the loci being tested. If a reference sample carries a deletion at a particular locus, the median depth at that locus will be lower than expected, and the test sample will appear to have a duplication relative to the reference.
This problem is difficult to detect because the reference samples are assumed to be normal. The use of a large reference set reduces the impact of any single contaminated sample, but it does not eliminate the problem. Researchers should be cautious about calls that are made relative to a small reference set.
Interpretation Limits and Clinical Context
Read-depth CNV detection from exome data has inherent limitations that affect the interpretation of results. These limitations should be communicated clearly in any report or publication.
Resolution Limits
The resolution of exome CNV detection is limited by the spacing of the capture targets. Breakpoints can only be mapped to the boundaries of the captured exons, and the exact genomic extent of a CNV may be uncertain. A deletion that removes three exons may actually extend into the intronic regions between the exons, but the read-depth data cannot determine the precise breakpoints.
This limitation is important for clinical interpretation because the functional consequence of a CNV depends on which genes and regulatory elements are affected. The uncertainty in breakpoint location should be acknowledged in the interpretation.
Sensitivity Limits
Exome CNV detection is less sensitive than whole-genome CNV detection. The noncontiguous nature of the capture reduces the number of data points available for each region, and the statistical power to detect small CNVs is limited [<a href="#ref-2">2</a>]. Small deletions or duplications that affect a single exon may be missed, particularly if the exon has low coverage.
The sensitivity of exome CNV detection is comparable to exon-resolution chromosomal microarray when the analysis is performed carefully [<a href="#ref-3">3</a>]. This finding provides a benchmark for what can be expected from a well-optimized exome CNV pipeline.
The Role of Genome Sequencing
Genome sequencing can detect types of variation that are invisible to exome sequencing, including intronic variants, small structural variants, copy-neutral inversions, complex rearrangements, and tandem repeat expansions [<a href="#ref-1">1</a>]. When exome sequencing has not provided a diagnosis, genome sequencing may identify pathogenic variants that require genome-wide coverage.
The decision to pursue genome sequencing after a negative exome analysis depends on the clinical context. The diagnostic yield of genome sequencing after prior genetic testing is approximately 8%, and a substantial fraction of these diagnoses could be made by reanalysis of exome data, including CNV calling [<a href="#ref-1">1</a>]. This finding suggests that exome CNV analysis should be performed before considering genome sequencing.
Safety and Regulatory Context for Clinical Reporting
The use of read-depth CNV detection in clinical settings carries specific responsibilities. The accuracy of the analysis directly affects patient care, and errors can have serious consequences.
Validation Requirements
Clinical laboratories must validate their CNV calling pipelines before using them for patient testing. Validation typically involves testing samples with known CNVs and comparing the pipeline results against an independent method. The validation should establish the sensitivity and specificity of the pipeline for the types of CNVs that will be reported.
The validation data should be documented and maintained. The performance characteristics of the pipeline, including its limitations, should be communicated to the clinicians who order the testing.
Confirmation Before Reporting
Read-depth CNV calls from exome data should be confirmed by an independent method before they are reported as clinical findings. The confirmation method should be appropriate for the type and size of the CNV. Quantitative PCR and multiplex ligation-dependent probe amplification are commonly used for small exonic events, while chromosomal microarray is used for larger events.
The confirmation step is essential because read-depth calls are probabilistic estimates. The false positive rate of exome CNV calling is not negligible, and reporting an unconfirmed call could lead to incorrect clinical decisions.
Interpretation in Context
The clinical significance of a CNV depends on the gene involved, the size and location of the event, and the patient's phenotype. A CNV that is benign in one context may be pathogenic in another. The interpretation should be performed by a qualified professional who has access to the patient's clinical information.
Automated interpretation tools can assist with variant prioritization, but they have limitations. An evaluation of an automated genome interpretation model found that the model accuracy was reduced in some cases because of incomplete variant calling, including copy number variants [<a href="#ref-6">6</a>]. This finding underscores the need for human review of CNV calls in clinical settings.
Professional Escalation Criteria
Researchers and laboratory professionals should know when to escalate a CNV finding to a more senior colleague or to a clinical genetics service. The following situations warrant escalation.
Unexpected or Incidental Findings
A CNV that affects a gene associated with a serious disease, when the patient was not being tested for that disease, should be escalated to a clinical genetics professional. The finding may have implications for the patient's health that require genetic counseling and follow-up.
Uncertain Clinical Significance
A CNV that affects a gene of unknown function, or a CNV that has not been reported in the literature, should be escalated for expert review. The interpretation of such findings requires knowledge of the gene's function, the inheritance pattern of the disease, and the population frequency of the variant.
Discrepant Results
A CNV call that conflicts with other test results, such as a chromosomal microarray or a targeted gene panel, should be escalated. The discrepancy may indicate a technical problem with one of the assays, or it may indicate a complex rearrangement that requires additional investigation.
Technical Failures
A sample that produces an unusually high number of CNV calls, or a sample that fails quality control metrics, should be escalated to the laboratory director. The sample may need to be re-sequenced or excluded from the analysis.
Records and Measurements for Quality Assurance
The quality of a read-depth CNV analysis depends on the records that are kept and the measurements that are taken. The following records should be maintained for each analysis.
Sample Metadata
Each sample should have a record of the capture kit, the sequencing platform, the library preparation protocol, and the date of sequencing. This metadata is essential for building a well-matched reference set and for detecting batch effects.
Analysis Parameters
The parameters used for the CNV calling analysis should be recorded, including the version of the reference genome, the version of the target file, the version of the CNV calling tool, and any custom parameters. This record allows the analysis to be reproduced and audited.
Quality Metrics
The quality metrics for each sample and each analysis should be recorded. These metrics include the mean depth, the coverage uniformity, the correlation between the test sample and the reference set, and the number of candidate CNVs. Trends in these metrics over time can reveal changes in the sequencing process that affect CNV calling.
Call-Level Records
Each CNV call should have a record that includes the genomic coordinates, the number of exons affected, the depth ratio, the statistical score, and the annotation. The record should also indicate whether the call was confirmed by an independent method and whether it was reported to the clinician.
Common Failure Patterns and How to Avoid Them
The following failure patterns are commonly observed in read-depth CNV analysis from exome data. Awareness of these patterns helps researchers avoid them.
Using a Poorly Matched Reference Set
The most common cause of poor CNV calling performance is a reference set that does not match the test samples. Reference samples should be processed with the same capture kit, the same sequencing platform, and similar library preparation protocols. Mixing samples from different kits or different sequencing centers within one analysis introduces systematic bias.
Ignoring Batch Effects
Batch effects are a major source of false positives in exome CNV calling. Samples that are processed together tend to have similar depth profiles, and differences between batches can be mistaken for CNVs. The reference set should include samples from the same batch as the test samples whenever possible.
Accepting Single-Exon Calls Without Validation
Single-exon CNV calls are the most common false positive in exome CNV analysis. These calls should be treated as low confidence unless they are supported by an independent method. The validation of single-exon calls is particularly important when the call affects a gene of clinical significance.
Failing to Filter Low-Quality Samples
Samples with low mean depth or poor coverage uniformity produce noisy CNV calls. These samples should be identified and either excluded from the analysis or flagged for cautious interpretation. The quality thresholds should be established before the analysis and applied consistently.
Overinterpreting Breakpoint Boundaries
The breakpoints of exome CNV calls are imprecise because the analysis is limited to captured exons. The exact genomic extent of a CNV may extend beyond the called region. Researchers should not overinterpret the breakpoint boundaries in their interpretation of the clinical significance.
Practical Implementation Steps for a Laboratory
The following steps provide a practical path for implementing read-depth CNV detection from WES data in a laboratory setting.
Step 1: Inventory Available Data
Review the available WES data and record the capture kit, sequencing platform, and alignment pipeline for each sample. Identify samples that were processed with the same kit and platform, as these will form the basis for the reference set.
Step 2: Install and Configure Tools
Install ExomeDepth and CNVkit from the appropriate repositories. The Bioconductor project provides official package documentation for R-based tools such as ExomeDepth [<a href="#ref-7">7</a>]. Verify that the tools are compatible with the reference genome build used for the alignments.
Step 3: Build the Reference Set
Select reference samples that were processed with the same capture kit and platform as the test samples. The reference set should include at least 20 samples, and larger sets are preferable. Record the sample identifiers and the criteria used for selection.
Step 4: Run the Analysis
Run the CNV calling analysis using the selected tools. Record the parameters used for the analysis and the version of the target file. Generate the output files and the diagnostic plots for each sample.
Step 5: Filter and Annotate
Apply the filtering criteria to the raw calls. Remove calls that affect a single exon unless they are supported by additional evidence. Annotate the remaining calls with gene names, exon counts, and population frequencies.
Step 6: Validate and Report
Select candidate calls for validation by an independent method. Confirm the calls that will be reported. Prepare a report that includes the genomic coordinates, the affected genes, the depth ratio, and the validation status of each call.
Frequently Asked Questions
What is the difference between read-depth and allele-frequency methods for CNV detection?
Read-depth methods use the number of sequencing reads that align to a genomic region to estimate copy number. A duplicated region produces more reads, and a deleted region produces fewer reads. Allele-frequency methods use the ratio of alternate to reference alleles at heterozygous single-nucleotide polymorphisms to infer copy number. Read-depth methods are more commonly used for exome data because they can be applied to any captured region, while allele-frequency methods require informative heterozygous sites.
How many reference samples are needed for reliable exome CNV calling?
The number of reference samples needed depends on the variability of the data and the size of the CNVs being detected. A minimum of 20 reference samples is commonly recommended, and larger reference sets of 50 or more samples improve the stability of the depth ratio estimates. The reference samples should be processed with the same capture kit and sequencing platform as the test samples.
Can read-depth methods detect mosaic CNVs from exome data?
Read-depth methods can detect mosaic CNVs if the mosaic fraction is large enough to produce a measurable change in depth. A mosaic deletion that affects 50% of cells produces a depth ratio of approximately 0.75, which may be detectable with good data quality. Lower mosaic fractions are difficult to detect because the depth change is small relative to the technical noise.
What is the minimum CNV size that can be detected from exome data?
The minimum detectable CNV size depends on the spacing of the capture targets and the depth of sequencing. A CNV that affects a single exon can be detected if the exon has adequate coverage and the depth ratio is sufficiently different from normal. In practice, single-exon calls are noisy and require validation. Multi-exon CNVs are detected with higher confidence.
How do I choose between ExomeDepth and CNVkit for my data?
The choice depends on the research question and the data characteristics. ExomeDepth is designed specifically for exome data and performs well for detecting exonic deletions and duplications. CNVkit provides additional visualization capabilities and can estimate copy number across larger genomic intervals using off-target reads. Running both tools and focusing on calls supported by both is a practical approach.
Why do I get different CNV calls from different tools?
Different tools use different normalization strategies, statistical models, and filtering criteria. These differences lead to different call sets. The integration of multiple tools can improve the quality of the final call set, as demonstrated in large clinical cohorts [<a href="#ref-3">3</a>]. Calls that are supported by multiple tools are more likely to be real.
Should I validate every CNV call from exome data?
The need for validation depends on the purpose of the analysis. For research purposes, validation of a subset of calls may be sufficient to estimate the false positive rate. For clinical reporting, every reported CNV should be confirmed by an independent method. The validation method should be appropriate for the type and size of the CNV.
What should I do if my exome CNV analysis does not identify a cause for a suspected genetic disorder?
If exome CNV analysis does not identify a pathogenic variant, the next step depends on the clinical context. Reanalysis of the exome data with updated tools and annotations may identify variants that were missed in the initial analysis. Genome sequencing can detect types of variation that are invisible to exome sequencing, including intronic variants and complex rearrangements [<a href="#ref-1">1</a>]. The decision to pursue genome sequencing should be made in consultation with a clinical genetics professional.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Variant Calling in Whole Exome Sequencing (WES): Principles, Algorithms, and Veterinary Applications
- RNA Sequencing Methods: A Guide to Library Prep, Strandedness, and Sequencing Depth
- Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices
- Radiomics Feature Selection: Methods and Best Practices
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Genome Sequencing for Diagnosing Rare Diseases.](https://pubmed.ncbi.nlm.nih.gov/38838312). The New England journal of medicine, 2024. [2] [Polishing copy number variant calls on exome sequencing data via deep learning.](https://pubmed.ncbi.nlm.nih.gov/35697522). Genome research, 2022. [3] [Detection and characterization of copy-number variants from exome sequencing in the DDD study.](https://pubmed.ncbi.nlm.nih.gov/39669630). Genetics in medicine open, 2024. [4] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [5] [nf-core Documentation](https://nf-co.re/docs). nf-core. [6] [Evaluation of an automated genome interpretation model for rare disease routinely used in a clinical genetic laboratory.](https://pubmed.ncbi.nlm.nih.gov/36939041). Genetics in medicine : official journal of the American College of Medical Genetics, 2023. [7] [Bioconductor](https://bioconductor.org/). Bioconductor Project.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.