# Genotyping Arrays for Variant Call Validation: How to Use SNP Chips to Confirm NGS Results


## Key Takeaways

- Genotyping arrays (SNP chips) offer an orthogonal validation method for Next-Generation Sequencing (NGS) variant calls by employing hybridization-based chemistry, distinct from sequencing-by-synthesis, thereby mitigating systematic errors inherent to sequencing pipelines.
- Concordance rates, typically exceeding 99% for high-quality germline variants, are calculated by comparing array genotypes with NGS variant calls, with stratification by genotype class (homozygous reference, heterozygous, homozygous alternate) crucial for identifying platform-specific error patterns.
- Array validation is limited to variants present on the array design and is generally not suitable for novel variants, insertions/deletions, or somatic variants due to allelic fraction variability, though it can serve as a control for germline samples in somatic studies.
- Practical implementation involves selecting an array with relevant variants (checking manifest files), using the same DNA aliquot for both platforms, performing array hybridization and genotype calling, extracting NGS variant calls, and systematically comparing and investigating discordant results.
- Sample identity and kinship checks are critical early steps, utilizing fingerprinting markers on arrays to detect mix-ups that would otherwise lead to uniformly low concordance across all variants.
- Discordant calls warrant investigation, starting with raw data review (array intensities, sequencing reads) and considering technical issues (contamination, probe failure) or biological factors (copy number variation, mosaicism).

---

Researchers who generate next-generation sequencing (NGS) variant calls often need an orthogonal method to confirm those calls before reporting or publishing. Genotyping arrays, also called SNP chips, provide that orthogonal confirmation by measuring known single nucleotide variants through a different biochemical mechanism than sequencing. This article explains how to use genotyping arrays to validate NGS variant calls, how to calculate concordance rates, how to interpret discordant calls, and what practical steps to take when array and sequencing results disagree. The content is written for biology students, researchers, laboratory professionals, and life-science practitioners who need concrete decision criteria for incorporating array-based validation into their variant calling workflows.

## The Role of Genotyping Arrays in Variant Validation

Genotyping arrays detect genetic variants using hybridization-based chemistry instead of sequencing-by-synthesis. A typical array contains hundreds of thousands to millions of probes designed to interrogate specific known single nucleotide polymorphisms. When a DNA sample is applied to the array, the probes bind to complementary sequences, and the fluorescence signal indicates which alleles are present at each locus. This fundamentally different detection mechanism makes arrays useful for confirming NGS results because systematic errors in sequencing, such as base-calling artifacts or alignment mistakes, are unlikely to recur in the array platform.

The validation workflow addresses a specific problem in variant calling. NGS produces variant calls with associated quality scores, but those scores reflect the sequencing and bioinformatics pipeline, not an independent measurement. When a variant has clinical, diagnostic, or breeding implications, researchers need confirmation from a second method. Genotyping arrays provide that confirmation for the subset of variants that are present on the array design.

The National Center for Biotechnology Information maintains databases and search systems that researchers use to design validation studies, access reference sequences, and deposit or retrieve variant information [<a href="#ref-1">1</a>]. The European Bioinformatics Institute offers training pathways for bioinformatics analysis that cover data resources and practical analysis education relevant to variant validation workflows [<a href="#ref-2">2</a>]. These resources support the computational aspects of comparing array genotypes with sequencing calls.

## At a Glance: Array Validation Decision Table

| Validation Scenario | Recommended Approach | Expected Concordance | Action on Discordance |
| --- | --- | --- | --- |
| Germline variant confirmation in diagnostic samples | Run array on same DNA aliquot used for sequencing | Above 99% for high-quality calls | Review raw data, check sample identity, consider Sanger sequencing |
| Population-scale variant screening | Use array as primary genotyping method with sequencing for novel variants | Platform-specific, typically 95-99% | Exclude discordant markers from downstream analysis |
| Somatic variant validation in tumor samples | Use array only for germline controls, not tumor variants | Limited utility for somatic calls | Do not use arrays to validate somatic variants due to allelic fraction issues |
| Cross-platform data integration | Use imputation to harmonize array and sequencing datasets | Depends on reference panel and marker density | Document imputation quality metrics per variant |

## Core Principles of Array-Based Validation

### How Genotyping Arrays Measure Variants

Genotyping arrays rely on allele-specific probe hybridization. Each probe on the array is designed to match one allele of a known variant. The DNA sample is fragmented, labeled with a fluorescent dye, and hybridized to the array. The fluorescence intensity at each probe position indicates whether the sample contains the reference allele, the alternate allele, both alleles, or neither allele. The genotype calling algorithm converts these intensity values into discrete genotype calls: homozygous reference, heterozygous, or homozygous alternate.

The array design determines which variants can be validated. Arrays contain a fixed set of probes, so only variants that are present on the array can be confirmed. Variants that are not on the array design cannot be validated with that particular chip. Researchers must check whether the specific variants of interest are represented on the array before designing a validation study.

The low-density Infinium QC Array-24 BeadChip contains 15,949 markers and supports linkage analysis, HLA haplotyping, fingerprinting, ethnicity determination, mitochondrial genome variations, blood groups, and pharmacogenomics [<a href="#ref-3">3</a>]. This array serves as an independent quality control option for NGS-based diagnostic laboratories and provides cost-efficient means for determining gender, ethnic ancestry, and sample kinships that are important for data interpretation of NGS-based genetic tests [<a href="#ref-3">3</a>]. The accuracy and reproducibility of Infinium QC genotyping calls were evaluated by comparing them with genotyping data from other platforms and with whole genome and exome sequencing, and concordance of genotype calls between the Infinium QC array and other platforms was above 99% [<a href="#ref-3">3</a>].

### Why Orthogonal Validation Matters

Sequencing errors and array errors arise from different sources. Sequencing errors can come from base-calling algorithms, polymerase errors during amplification, phasing issues, or alignment artifacts. Array errors can come from probe design flaws, cross-hybridization, or intensity calling thresholds. When two independent methods produce the same genotype call, the probability that both methods made the same error at the same locus is very low. This is the statistical basis for using arrays as an orthogonal validation method.

The validation value depends on the array probes being truly independent from the sequencing process. If the array design was derived from the same reference genome and the same variant database used for sequencing analysis, the validation confirms the genotype but does not confirm the variant's existence in the sample. The array confirms that the sample carries the allele that the probe was designed to detect. This distinction matters when researchers interpret concordance rates.

### Concordance Metrics and Their Interpretation

Concordance is the proportion of variants where the array genotype and the sequencing genotype agree. The standard calculation divides the number of concordant calls by the total number of calls compared. Researchers should report concordance separately for homozygous reference, heterozygous, and homozygous alternate calls because error rates differ by genotype class.

A customized single nucleotide variant microarray designed to detect disease-causing variants and copy number variation in patients with primary immunodeficiency disorders achieved a genotype calling accuracy rate of 99.7% when compared with whole-genome sequencing data from 56 non-PID controls [<a href="#ref-4">4</a>]. The sensitivity for detecting rare PID variants was high at 87%, and the single sample replication in two runs was high at 94.9% [<a href="#ref-4">4</a>]. These figures illustrate the performance levels that can be expected when array and sequencing platforms are compared.

Concordance rates above 99% are typical for high-quality germline samples when the array and sequencing platforms are both performing well [<a href="#ref-3">3</a>]. Lower concordance rates indicate problems that need investigation. The investigation should consider sample quality, array performance, sequencing depth, variant quality scores, and the genomic context of the discordant variants.

## Practical Workflow for Array-Based Validation

### Step 1: Select the Appropriate Array Platform

The array selection depends on the variants that need validation. Researchers should obtain the array manifest file, which lists all variants on the array, and check whether the variants of interest are included. The manifest also provides information about probe design, flanking sequence, and variant annotation.

For diagnostic laboratories, the low-density Infinium QC Array-24 BeadChip provides a cost-efficient option for determining gender, ethnic ancestry, and sample kinships that are important for data interpretation of NGS-based genetic tests [<a href="#ref-3">3</a>]. The array's ancestry informative markers are sufficient for ethnicity determination at continental and sometimes subcontinental levels, with assignment accuracy varying with the coverage for a particular region and ethnic groups [<a href="#ref-3">3</a>]. Mean accuracies of provenance prediction at a regional level varied from 81% for Asia, to 89% for Americas, 86% for Africa, 97% for Oceania, 98% for Europe, and 100% for India [<a href="#ref-3">3</a>].

For research applications, genome-wide arrays provide broader coverage but may not include rare or private variants. Custom arrays can be designed to include specific variants of interest, as demonstrated by the customized primary immunodeficiency array that added probes for 9,415 variants of 277 PID-related genes to the genome-wide Illumina Global Screening Array [<a href="#ref-4">4</a>].

### Step 2: Prepare DNA Samples

The same DNA aliquot used for sequencing should be used for array genotyping when possible. Using the same aliquot eliminates DNA extraction variability as a source of discordance. If the same aliquot is not available, a replicate extraction from the same tissue or blood sample is acceptable, but the possibility of sample mix-up or contamination should be considered when interpreting results.

DNA quality requirements differ between sequencing and array platforms. Arrays typically require higher molecular weight DNA than some sequencing library preparation methods. Degraded DNA can cause failed genotype calls on arrays even when the same DNA produced acceptable sequencing data. Researchers should check DNA concentration and integrity before array hybridization.

### Step 3: Perform Array Hybridization and Genotype Calling

Array hybridization follows the manufacturer's protocol. The key steps are DNA fragmentation, labeling, hybridization, washing, and scanning. The scanner produces intensity data that the genotype calling software converts into discrete genotype calls.

The genotype calling software applies clustering algorithms to assign genotypes based on fluorescence intensity. The quality of the clustering depends on the number of samples in the batch, the allele frequency in the population, and the separation between intensity clusters. Samples with unusual ancestry or copy number variation may fall outside expected clusters and receive low quality scores.

The customized primary immunodeficiency array used Illumina GenomeStudio 2.0 for data analysis of the Global Screening Array [<a href="#ref-4">4</a>]. The choice of analysis software affects genotype calling, and researchers should document the software version and settings used for calling.

### Step 4: Extract Variant Calls from Sequencing Data

The sequencing variant calls need to be in a format that allows direct comparison with array genotypes. The standard approach is to generate a variant call format file from the sequencing pipeline and extract the genotypes for the variants that are present on the array.

The sequencing variant caller should be run with appropriate settings for the analysis type. Germline variant calling uses different algorithms and parameters than somatic variant calling. The variant caller output should include genotype calls and quality scores for each variant.

The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover variant calling and related bioinformatics analyses [<a href="#ref-5">5</a>]. The nf-core documentation describes community pipeline standards, usage, configuration, and reproducible workflow context that can help researchers structure their sequencing analysis pipelines [<a href="#ref-6">6</a>].

### Step 5: Compare Array and Sequencing Genotypes

The comparison requires matching variants between the two platforms. The matching should use genomic coordinates and allele representation. Variants on the array are typically represented by their reference genome coordinates and the reference and alternate alleles. The sequencing variant calls should use the same reference genome build and allele representation.

The comparison script should handle strand issues, allele swaps, and multi-allelic variants. Some array probes are designed on the opposite strand from the reference genome, and the genotype calls need to be converted to the reference strand before comparison. Multi-allelic variants on the array may need special handling because the array can only represent two alleles per probe.

The comparison output should include a table with the variant identifier, genomic coordinates, array genotype, sequencing genotype, and concordance status. This table becomes the primary record for the validation study.

### Step 6: Calculate Concordance Rates

The concordance rate is calculated as the number of concordant calls divided by the total number of calls compared. The calculation should exclude failed calls from either platform. Failed calls are genotypes that the array or sequencing platform could not determine with sufficient confidence.

Concordance should be reported overall and stratified by genotype class. Stratification reveals whether discordance is concentrated in specific genotype classes. For example, discordance concentrated in heterozygous calls may indicate allelic dropout or reference bias in the sequencing data. Discordance concentrated in homozygous alternate calls may indicate mapping issues or copy number variation.

The customized primary immunodeficiency array achieved a genotype calling accuracy rate of 99.7% when compared with whole-genome sequencing data [<a href="#ref-4">4</a>]. This level of concordance represents the performance that can be expected when both platforms are working well and the variants are common enough to be well represented on the array.

### Step 7: Investigate Discordant Calls

Discordant calls require investigation before they can be resolved. The investigation should follow a systematic process that considers technical and biological explanations.

Technical explanations include sample mix-up, DNA contamination, array probe failure, sequencing alignment errors, and genotype calling errors. Biological explanations include copy number variation, mosaicism, and variants in the probe binding site that prevent hybridization.

The investigation should start with the raw data. The array intensity data can be reviewed to check whether the discordant variant had a clear genotype cluster. The sequencing read data can be reviewed to check whether the variant had sufficient depth and quality. If the raw data look clean on both platforms, the discordance may reflect a genuine biological difference between the DNA samples used for each platform.

## Options and Tradeoffs in Array-Based Validation

### Genome-Wide Arrays versus Custom Arrays

Genome-wide arrays provide broad coverage of common variants across the genome. These arrays are manufactured at scale, which keeps per-sample costs low. The variant content is fixed by the manufacturer and may not include rare or population-specific variants of interest.

Custom arrays allow researchers to include specific variants that are relevant to their study. The customized primary immunodeficiency array added probes for 9,415 variants of 277 PID-related genes to the genome-wide Illumina Global Screening Array [<a href="#ref-4">4</a>]. This approach combined the benefits of genome-wide coverage with targeted content for the genes of interest.

The tradeoff is cost and turnaround time. Custom arrays require design and manufacturing lead time, and the per-sample cost may be higher for small batches. Genome-wide arrays are available off the shelf and can be processed immediately.

### Array Validation versus Sanger Sequencing Validation

Sanger sequencing is the traditional method for validating NGS variant calls. Sanger sequencing provides direct sequence information for the variant and its surrounding context, which can reveal whether the variant is real or an artifact. However, Sanger sequencing is low-throughput and becomes impractical when many variants need validation.

A molecular autopsy study using Fluidigm Access Array PCR-enrichment with Illumina HiSeq 2000 NGS demonstrated 100% sensitivity for pathogenic variants and 87.20% sensitivity and 99.99% specificity for all substitutions in an optimization subset of 46 patients [<a href="#ref-7">7</a>]. The positive predictive value for NGS for rare substitutions was 16.0%, meaning 27 confirmed rare variants out of 169 positive NGS calls in 151 additional cases [<a href="#ref-7">7</a>]. This study required confirmatory Sanger sequencing of mutational variants because the NGS platform produced many false positive calls for rare variants [<a href="#ref-7">7</a>].

Arrays provide higher throughput than Sanger sequencing but only validate variants that are present on the array. Sanger sequencing can validate any variant regardless of whether it is on an array. The choice between arrays and Sanger sequencing depends on the number of variants to validate and whether the variants are represented on available arrays.

### Array Validation in Crop and Animal Breeding

Next-generation sequencing has revolutionized plant and animal research by providing powerful genotyping methods, including whole-genome re-sequencing, SNP arrays, and reduced representation sequencing [<a href="#ref-8">8</a>]. These methods offer a wide range of applications from genome-wide analysis to routine screening with a high level of accuracy and reproducibility, and they provide a straightforward workflow to identify, validate, and screen genetic variants in a short time with a low cost [<a href="#ref-8">8</a>].

The main challenges facing breeders and geneticists are how to choose an appropriate genotyping method and how to integrate genotyping data sets obtained from various sources [<a href="#ref-8">8</a>]. Imputation methods can be used to fill in missing data in genotypic data sets and to integrate data sets obtained using different genotyping tools [<a href="#ref-8">8</a>]. This integration is particularly relevant when combining array genotypes from one set of samples with sequencing genotypes from another set.

In breeding programs, arrays serve both as validation tools and as primary genotyping methods. The choice between arrays and sequencing for routine screening depends on the number of markers needed, the cost per sample, and the turnaround time. Arrays provide a low-cost option for screening many samples at a fixed set of markers, while sequencing provides flexibility to discover new variants.

## Observations and Measurements in Array Validation

### Sample Identity and Kinship Checks

Arrays provide a built-in sample identity check through fingerprinting markers. The low-density Infinium QC Array-24 BeadChip enables fingerprinting and kinship determination [<a href="#ref-3">3</a>]. Gender determination was correct in all tested cases in the evaluation of the Infinium QC array [<a href="#ref-3">3</a>].

Sample identity checks are important in validation studies because sample mix-up is a common source of discordance. If the array genotype and sequencing genotype come from different individuals, the concordance rate will be low across all variants. Running identity checks before calculating concordance can prevent wasted effort investigating discordant calls that result from sample mix-up.

Kinship analysis can also reveal unexpected sample relationships. The Infinium QC array evaluation found that pairwise concordances of African samples with samples from any other super populations were the lowest at 0.39 to 0.43, while the concordances within the same population were relatively high at 0.55 to 0.61 [<a href="#ref-3">3</a>]. These population-specific concordance patterns matter when interpreting kinship results.

### Ancestry and Population Structure

Ancestry determination is a useful quality check in validation studies because ancestry affects variant frequencies and genotype cluster positions. The Infinium QC array's ancestry informative markers were sufficient for ethnicity determination at continental and sometimes subcontinental levels, with assignment accuracy varying with the coverage for a particular region and ethnic groups [<a href="#ref-3">3</a>]. Mean accuracy of ethnicity assignment predictions was 63% [<a href="#ref-3">3</a>].

Ancestry information helps interpret discordant calls. If a sample has ancestry that is underrepresented in the reference panel used for array genotype calling, the genotype clusters may be poorly defined, leading to calling errors. The same ancestry may affect sequencing variant calling if the reference genome does not represent the sample's population well.

### Copy Number Variation Detection

Some arrays can detect copy number variation in addition to single nucleotide variants. The customized primary immunodeficiency array detected copy number variation in patients with primary immunodeficiency disorders [<a href="#ref-4">4</a>]. The array generated a genetic diagnosis in 37 out of 95 patients, including 3 patients diagnosed by copy number variation analysis [<a href="#ref-4">4</a>].

Copy number variation can cause discordance between array and sequencing genotypes. A deletion that removes one copy of a variant will affect the array intensity and the sequencing read depth. The genotype calls from both platforms may be incorrect or may be correct for different underlying biological states. Researchers should consider copy number variation when investigating discordant calls.

## Records and Documentation for Array Validation

### Required Records for Validation Studies

Validation studies require documentation of the samples, platforms, and analysis parameters used. The records should include the sample identifiers, the DNA extraction method, the DNA quality metrics, the array platform and version, the array processing date, the sequencing platform and version, the sequencing pipeline version, and the variant caller version and settings.

The records should also include the concordance calculation method and the results. The concordance table should list each variant, the array genotype, the sequencing genotype, and the concordance status. Failed calls should be recorded separately from discordant calls.

The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help researchers structure their analysis code and records [<a href="#ref-9">9</a>]. Reproducible analysis requires version control of code and documentation of software versions.

### Data Storage and Access

The raw data from array and sequencing platforms should be stored in accessible formats. Array intensity data files and genotype call files should be archived. Sequencing read files and variant call files should be archived. The analysis scripts used to compare array and sequencing genotypes should be version-controlled.

The National Center for Biotechnology Information provides databases and search systems for sequence resources and analysis services [<a href="#ref-1">1</a>]. Researchers can deposit validated variant data in appropriate databases to make the results available to the broader community.

### Reporting Concordance Results

Concordance results should be reported with sufficient detail for others to evaluate the validation. The report should include the number of variants compared, the number of concordant calls, the number of discordant calls, the number of failed calls, and the concordance rate. The report should also describe the criteria used to define concordance and the handling of failed calls.

The report should stratify concordance by variant type, genotype class, and genomic context. Stratification by variant type distinguishes single nucleotide variants from insertions and deletions. Stratification by genotype class distinguishes homozygous reference, heterozygous, and homozygous alternate calls. Stratification by genomic context distinguishes variants in coding regions, regulatory regions, and intergenic regions.

## Common Failure Patterns in Array Validation

### Low Concordance Due to Sample Mix-Up

Sample mix-up produces low concordance across all variants. The pattern is uniform discordance with no clear relationship to variant type or genomic context. Sample identity checks using fingerprinting markers can detect mix-up before concordance calculation.

The response to sample mix-up is to repeat the analysis with correctly identified samples. If the mix-up occurred during array processing, the array results should be discarded and the sample reprocessed. If the mix-up occurred during sequencing, the sequencing results should be discarded and the sample resequenced.

### Discordance Concentrated in Heterozygous Calls

Discordance concentrated in heterozygous calls suggests allelic dropout or reference bias. Allelic dropout occurs when one allele fails to amplify or hybridize, causing a heterozygous sample to be called homozygous. Reference bias occurs when the sequencing or analysis pipeline preferentially maps reads containing the reference allele.

The investigation should review the sequencing read data for the discordant variants. If the alternate allele is present in the reads but at a lower frequency than expected for a heterozygous call, the variant caller may have failed to call the heterozygote. The array intensity data should also be reviewed to check whether the heterozygous cluster was well separated from the homozygous clusters.

### Discordance Concentrated in Rare Variants

Rare variants are more likely to be discordant than common variants. The customized primary immunodeficiency array had high sensitivity for detecting rare PID variants at 87% [<a href="#ref-4">4</a>]. The remaining 13% of rare variants that were not detected may reflect array probe design limitations or sequencing errors.

Rare variants may not be well represented in the genotype clustering reference panel. If the array genotype calling algorithm has not seen many examples of the rare allele, the intensity cluster for the rare homozygote may be poorly defined. The sequencing variant caller may also have difficulty with rare variants because the variant allele fraction is low in the reads.

### Discordance Due to Variants in Probe Binding Sites

Variants in the probe binding site can prevent hybridization, causing the array to fail to call the genotype or to call the wrong genotype. The probe binding site includes the variant position and the flanking sequence. If the sample has a variant in the flanking sequence that is not on the array, the probe may not bind efficiently.

The investigation should check the array manifest for the probe sequence and compare it with the sample sequence. If the sample has a variant in the probe binding site, the array genotype call should be considered unreliable. The variant should be validated by an alternative method such as Sanger sequencing.

## Limitations of Array-Based Validation

### Arrays Only Validate Known Variants

Arrays can only validate variants that are present on the array design. Novel variants discovered by sequencing cannot be validated with an array unless they are added to a custom array design. This limitation means that arrays are not suitable for validating all NGS variant calls.

The customized primary immunodeficiency array could not detect variants in genes that were not covered by the custom probes. Twenty-eight patients could not be detected due to the limited coverage of the custom probes [<a href="#ref-4">4</a>]. This limitation reflects the fundamental constraint of array-based validation: the array content determines what can be validated.

### Arrays Do Not Validate Insertions and Deletions

Most genotyping arrays are designed to detect single nucleotide variants. Insertions and deletions are difficult to represent on arrays because the probe design requires a fixed sequence context. Some arrays include probes for known indels, but the coverage is typically limited compared with single nucleotide variants.

Researchers who need to validate indels should use alternative methods such as Sanger sequencing or targeted resequencing. The array can validate the single nucleotide variants in the same sample, but the indel calls require separate validation.

### Arrays Have Limited Utility for Somatic Variants

Somatic variants in tumor samples have variable allele fractions that depend on tumor purity and copy number status. Arrays are designed to detect germline variants with allele fractions of 50% for heterozygotes and 100% for homozygotes. Somatic variants with low allele fractions may not be detected by arrays.

The use of arrays for somatic variant validation is limited to germline controls. The array can confirm that the germline genotype is correct, which helps distinguish germline variants from somatic variants. The somatic variants themselves should be validated by alternative methods such as targeted resequencing or digital PCR.

### Cross-Platform Concordance Is Not Perfect

Even when both platforms are working correctly, concordance is not 100%. The Infinium QC array evaluation found concordance of genotype calls between the array and other platforms was above 99% [<a href="#ref-3">3</a>]. The remaining discordance reflects genuine differences between platforms in how variants are measured.

Researchers should establish an acceptable concordance threshold before starting the validation study. The threshold should be based on the intended use of the validated variants. Diagnostic applications may require higher concordance than research applications. The threshold should be documented in the validation protocol.

## Quality and Welfare Controls in Array Validation

### Batch Effects and Batch Design

Array genotyping is sensitive to batch effects. Samples processed in different batches may have different intensity distributions, which can affect genotype calling. The batch design should include appropriate controls to detect and correct batch effects.

The single sample replication in two runs was high at 94.9% in the customized primary immunodeficiency array evaluation [<a href="#ref-4">4</a>]. This replication rate indicates that the same sample processed in different runs can produce slightly different genotype calls. Researchers should include replicate samples in each batch to assess batch-to-batch variability.

### Positive and Negative Controls

Validation studies should include positive controls with known genotypes and negative controls with no DNA. Positive controls confirm that the array is detecting the expected variants. Negative controls confirm that the array is not producing false signals from contamination or nonspecific binding.

The controls should be processed in the same batch as the test samples. The control results should be reviewed before the test sample results are interpreted. If the controls fail, the batch should be repeated.

### Data Quality Metrics

Array platforms produce quality metrics that should be reviewed before genotype calls are used. The metrics include call rate, cluster separation, and intensity distributions. Samples with low call rates or poor cluster separation should be flagged for review.

Sequencing platforms also produce quality metrics that should be reviewed. The metrics include read depth, mapping quality, and variant quality scores. Variants with low quality scores should be flagged for review before they are compared with array genotypes.

## Safety and Regulatory Context for Array Validation

### Diagnostic Validation Requirements

Diagnostic laboratories that use arrays for validation must meet regulatory requirements for test validation. The requirements vary by jurisdiction and by the intended use of the test. The validation should include analytical sensitivity, analytical specificity, accuracy, precision, and reportable range.

The low-density Infinium QC Array-24 BeadChip represents an attractive independent QC option for NGS-based diagnostic laboratories [<a href="#ref-3">3</a>]. The array provides cost-efficient means for determining gender, ethnic ancestry, and sample kinships that are important for data interpretation of NGS-based genetic tests [<a href="#ref-3">3</a>]. These quality control functions support the diagnostic use of NGS.

### Data Privacy and Sample Consent

Array genotyping produces genetic data that may be subject to privacy protections. Researchers should ensure that sample consent covers the array genotyping and the intended use of the data. The consent should specify whether the data will be shared and with whom.

The National Center for Biotechnology Information provides databases and search systems that may be used for data deposition and access [<a href="#ref-1">1</a>]. Researchers should follow the data sharing policies of their institution and the relevant databases.

### Reporting of Validation Results

Validation results should be reported in a format that supports the intended use of the data. Diagnostic reports should include the validation method, the concordance rate, and any limitations of the validation. Research reports should include the same information to support reproducibility.

The report should distinguish validated variants from unvalidated variants. Variants that were not present on the array should be reported as unvalidated. Variants that were discordant between the array and sequencing should be reported as discordant with an explanation of the investigation.

## Professional Escalation Criteria

### When to Escalate Discordant Calls

Discordant calls should be escalated when they involve variants with clinical, diagnostic, or breeding significance. The escalation should include a review by a second analyst and consultation with a molecular geneticist or bioinformatician with relevant expertise.

The escalation should also occur when the discordance rate exceeds the acceptable threshold. A concordance rate below 99% for high-quality germline samples warrants investigation [<a href="#ref-3">3</a>]. The investigation should determine whether the discordance reflects a systematic problem with the array, the sequencing, or the samples.

### When to Repeat the Analysis

The analysis should be repeated when sample mix-up is suspected, when the array or sequencing quality metrics are poor, or when the discordance pattern suggests a technical problem. The repeat analysis should use fresh reagents and include appropriate controls.

The repeat analysis should also be performed when the discordant variants are concentrated in a specific genomic region. Regional discordance may indicate a structural variant or a complex genomic region that is difficult to analyze with both platforms.

### When to Use an Alternative Validation Method

An alternative validation method should be used when the array cannot validate the variants of interest. Sanger sequencing is the standard alternative for validating individual variants. The molecular autopsy study required confirmatory Sanger sequencing of mutational variants because the NGS platform produced many false positive calls for rare variants [<a href="#ref-7">7</a>].

An alternative method should also be used when the array and sequencing results are discordant and the investigation cannot resolve the discordance. The alternative method provides a third measurement that can break the tie between the array and sequencing results.

## Frequently Asked Questions

### What concordance rate should I expect when comparing array genotypes with NGS variant calls?

Concordance rates above 99% are typical when comparing array genotypes with sequencing data from high-quality germline samples [<a href="#ref-3">3</a>]. The customized primary immunodeficiency array achieved a genotype calling accuracy rate of 99.7% when compared with whole-genome sequencing data [<a href="#ref-4">4</a>]. Lower concordance rates indicate problems that need investigation, such as sample mix-up, DNA quality issues, or platform-specific errors.

### Can genotyping arrays validate somatic variants from tumor sequencing?

Genotyping arrays have limited utility for validating somatic variants because arrays are designed to detect germline variants with allele fractions of 50% for heterozygotes and 100% for homozygotes. Somatic variants in tumor samples have variable allele fractions that depend on tumor purity and copy number status. Arrays can be used to validate the germline genotype of the same sample, which helps distinguish germline from somatic variants, but the somatic variants themselves should be validated by alternative methods.

### How do I handle discordant calls between array and sequencing platforms?

Discordant calls should be investigated systematically. Start by reviewing the raw data from both platforms to check whether the genotype calls were made with sufficient confidence. Check sample identity using fingerprinting markers to rule out sample mix-up. Consider copy number variation and variants in probe binding sites as potential explanations. If the discordance cannot be resolved, use an alternative validation method such as Sanger sequencing.

### What variants can be validated with a genotyping array?

Only variants that are present on the array design can be validated. The array manifest lists all variants on the array, and researchers should check whether their variants of interest are included. Most arrays are designed to detect single nucleotide variants, and coverage of insertions and deletions is typically limited. Novel variants discovered by sequencing cannot be validated with an array unless they are added to a custom array design.

### How many samples should I include in an array validation study?

The number of samples depends on the purpose of the validation. For validating a specific set of variants, the number of samples should be sufficient to observe each variant in at least one sample. For assessing platform concordance, the sample size should be large enough to estimate the concordance rate with acceptable precision. The customized primary immunodeficiency array evaluation compared the array with whole-genome sequencing data from 56 non-PID controls [<a href="#ref-4">4</a>].

### What quality metrics should I review before using array genotype calls?

Review the call rate, which is the proportion of variants that received a genotype call. Review the cluster separation, which indicates how well the genotype clusters are distinguished. Review the intensity distributions for each variant to check for outliers. Samples with low call rates or poor cluster separation should be flagged for review before their genotype calls are used for validation.

### How do I integrate array and sequencing datasets for downstream analysis?

Imputation methods can be used to fill in missing data in genotypic data sets and to integrate data sets obtained using different genotyping tools [<a href="#ref-8">8</a>]. The integration requires harmonizing variant representation between the platforms, including genomic coordinates, allele representation, and strand orientation. The imputation quality should be assessed per variant, and variants with poor imputation quality should be excluded from downstream analysis.

### What are the limitations of using arrays for validation in plant and animal breeding?

Arrays provide a low-cost option for screening many samples at a fixed set of markers, but they only validate known variants. The main challenges in breeding applications are choosing an appropriate genotyping method and integrating genotyping data sets obtained from various sources [<a href="#ref-8">8</a>]. Imputation methods can help integrate array and sequencing datasets, but the success depends on the reference panel and the marker density [<a href="#ref-8">8</a>].

## Related Bioinformatics Guides

- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Longitudinal Microbiome Data Analysis: Methods and Best Practices](/knowledge/bioinformatics/longitudinal-microbiome-data-analysis-methods-and-best-practices)
- [Metagenomic Contamination Control: Best Practices for Clean Data](/knowledge/bioinformatics/metagenomic-contamination-control-best-practices-for-clean-data)
- [Persistent Identifiers for Research Data: A Guide to Selection and Use](/knowledge/bioinformatics/persistent-identifiers-for-research-data-a-guide-to-selection-and-use)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Clinical utility of the low-density Infinium QC genotyping Array in a genomics-based diagnostics laboratory.](https://pubmed.ncbi.nlm.nih.gov/28985730). BMC medical genomics, 2017.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Rapid Low-Cost Microarray-Based Genotyping for Genetic Screening in Primary Immunodeficiency.](https://pubmed.ncbi.nlm.nih.gov/32373116). Frontiers in immunology, 2020.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Next-generation sequencing using microfluidic PCR enrichment for molecular autopsy.](https://pubmed.ncbi.nlm.nih.gov/31337358). BMC cardiovascular disorders, 2019.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [Efficient genome-wide genotyping strategies and data integration in crop plants.](https://pubmed.ncbi.nlm.nih.gov/29352324). TAG. Theoretical and applied genetics. Theoretische und angewandte Genetik, 2018.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.