KIR Gene Variant Calling: Challenges in Copy Number Variation and Paralogy

By Dr. Zubair Khalid, DVM, MS, PhD ·

KIR Gene Variant Calling: Challenges in Copy Number Variation and Paralogy

Key Takeaways

  • KIR gene variant calling is fundamentally challenging due to high sequence similarity among paralogs, preventing unique read assignment, and significant copy number variation (CNV) between individuals, invalidating standard diploid assumptions.
  • Accurate KIR genotyping necessitates a combined approach: read-depth analysis to determine gene presence and copy number, and assembly-based methods or specialized tools (e.g., T1K, BAKIR, Locityper) for precise allele resolution.
  • Whole-genome sequencing (WGS), particularly at 30x coverage or higher, offers the most comprehensive data for KIR analysis, while whole-exome sequencing (WES) and RNA-sequencing (RNA-seq) present limitations in copy number estimation and gene representation, respectively.
  • Quality control is paramount, involving assessment of read depth at variant sites, mapping quality, allele balance, and consistency across multiple analytical methods to mitigate false positives arising from paralogous mapping or reference bias.
  • Standard short-read variant calling pipelines (e.g., GATK) are unsuitable for KIR loci; specialized tools that jointly consider all KIR genes and account for CNV are required to overcome paralogy and CNV challenges.

Killer cell immunoglobulin-like receptor (KIR) genes present one of the most difficult variant calling problems in human genomics because they combine high sequence similarity across gene family members with variable gene presence and copy number between individuals. Standard short-read variant calling pipelines that align to a single linear reference genome produce unreliable results for KIR loci because reads from highly similar paralogs map ambiguously and because the expected diploid copy number assumption fails when genes are absent or duplicated. This article provides a practical workflow for researchers and laboratory professionals who need accurate KIR variant calls, combining read-depth analysis for copy number estimation with assembly-based methods for allele resolution. The guidance covers data input requirements, tool selection, quality control measures, interpretation limits, and criteria for escalating results to specialized analysis.

The Structural Biology of KIR Genes That Breaks Standard Pipelines

KIR genes encode receptors on natural killer cells that regulate innate immune activity through recognition of HLA class I ligands. The KIR gene family resides in the leukocyte receptor complex on chromosome 19 and contains multiple genes with closely related sequences that arose through duplication events. This genomic architecture creates two distinct problems for variant calling that do not affect most other loci.

The first problem is paralogy. Several KIR genes share such high sequence identity that short sequencing reads cannot be uniquely assigned to a single gene. The KIR2DL5A and KIR2DL5B genes exemplify this challenge because their sequences are nearly identical across much of their length, making standard read mapping tools unable to distinguish which gene a read originated from. When reads map to multiple locations with equal or near-equal quality, alignment tools either discard them as multi-mappers or assign them arbitrarily, and downstream variant callers then produce false variant calls or miss true variants entirely.

The second problem is copy number variation. Unlike most human genes that exist in two copies per diploid genome, KIR genes can be present, absent, or duplicated on individual haplotypes. The KIR gene content of a haplotype is organized into haplotypes that differ in which genes are present. This means the assumption of diploid copy number that underlies standard germline variant calling is invalid for KIR loci. A gene that is hemizygous will show half the read depth of a homozygous gene, and a gene that is duplicated will show increased depth. Standard variant callers interpret these depth differences as copy number alterations or fail to call variants in regions with unusual depth.

The combination of paralogy and copy number variation means that KIR genes cannot be genotyped with standard variant calling pipelines. This limitation is well documented in the literature. A 2023 study in Genome Research describing the T1K method noted that KIR genes are highly polymorphic, similar to each other in sequence, and may be absent from chromosomes, and that none of the tools developed for HLA genotyping work for KIR genes [<a href="#ref-1">1</a>]. The same study reported that even specialized KIR genotypers could not resolve all KIR genes, with the KIR2DL5A and KIR2DL5B genes being particularly difficult to distinguish [<a href="#ref-1">1</a>].

Why Read Depth and Assembly Methods Are Required

The solution to KIR variant calling requires moving beyond the standard map-and-call paradigm. Two complementary approaches have emerged: read-depth analysis for copy number estimation and assembly-based methods for allele resolution.

Read-depth analysis works because the number of sequencing reads that map to a genomic region is proportional to the number of copies of that region in the sample. For KIR genes, read depth can distinguish between a gene that is absent, present in one copy, or present in two or more copies. This information is essential because variant calling in a gene that is actually hemizygous requires different assumptions than variant calling in a homozygous gene. If a researcher assumes diploid copy number for a hemizygous gene, the variant caller will interpret the heterozygous genotype incorrectly.

Assembly-based methods work by reconstructing the actual sequence of the KIR locus from the sequencing data instead of by mapping reads to a reference. De novo assembly of the KIR region can resolve the sequence of each gene copy without relying on ambiguous read mapping. This approach is particularly valuable for identifying novel alleles that differ substantially from the known allele database. A 2024 preprint describing the BAKIR tool emphasized that traditional genotyping methods struggle with the variable nature of KIR genes and that these challenges extend to high-quality phased assemblies [<a href="#ref-2">2</a>]. BAKIR uses a multi-stage mapping, alignment, and variant calling process to achieve high-precision gene and allele identification while maintaining high recall for sequences that are significantly mutated or truncated relative to the known allele database [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

The practical implication is that accurate KIR variant calling requires a workflow that combines both approaches. Read-depth analysis provides the copy number context that determines which genes are present and how many copies exist. Assembly-based methods provide the sequence-level resolution needed to identify specific alleles and novel variants. Neither approach alone is sufficient.

At a Glance: KIR Variant Calling Workflow Overview

The table below summarizes the key workflow components, their purpose, and the primary considerations for each stage.

Workflow ComponentPrimary PurposeKey Considerations
Data input assessmentDetermine whether sequencing data type and depth support KIR analysisWhole-genome sequencing, whole-exome sequencing, and RNA-seq each have different strengths and limitations for KIR genotyping
Copy number estimationDetermine which KIR genes are present and how many copies existRead-depth analysis across the KIR locus distinguishes absent, hemizygous, and duplicated genes
Allele resolutionIdentify specific KIR alleles and novel variantsAssembly-based methods or specialized genotyping tools that jointly consider all KIR genes are required
Quality controlVerify that variant calls are reliable and reproducibleCompare results across methods, check read depth at variant sites, and validate against known allele databases

Data Input Requirements for KIR Variant Calling

The choice of sequencing data type has a major impact on the accuracy and completeness of KIR variant calls. Each data type provides different information and has different limitations.

Whole-Genome Sequencing

Whole-genome sequencing provides the most complete coverage of the KIR locus because it sequences all genomic regions without an enrichment step. Both short-read and long-read whole-genome sequencing can be used for KIR genotyping. A 2025 Nature Genetics study describing Locityper reported that the tool can genotype challenging genes using both short-read and long-read whole-genome sequencing, and that it achieves a median quality value above 35 from both data types across 256 challenging medically relevant loci [<a href="#ref-4">4</a>]. The same study reported that Locityper outperformed state-of-the-art Illumina and PacBio HiFi variant calling pipelines by 10.9 and 1.7 quality points, respectively [<a href="#ref-4">4</a>].

The depth of whole-genome sequencing matters. Standard 30x coverage whole-genome sequencing provides sufficient depth for read-depth-based copy number estimation and for most allele calling approaches. Lower coverage may be insufficient for accurate copy number calls, particularly for genes that are present in only one copy. Higher coverage provides more confidence but at increased cost.

Whole-Exome Sequencing

Whole-exome sequencing enriches for protein-coding regions before sequencing. This approach reduces cost but introduces complications for KIR analysis. The enrichment process can introduce bias in read depth across the KIR locus, which complicates copy number estimation. Additionally, the KIR genes may not be uniformly captured by exome enrichment probes, leading to uneven coverage that affects variant calling accuracy.

Despite these limitations, whole-exome sequencing can still be used for KIR genotyping with appropriate methods. The T1K method described in a 2023 Genome Research study was designed to work with whole-exome sequencing data in addition to whole-genome sequencing and RNA-seq data [<a href="#ref-1">1</a>]. Researchers using whole-exome data should be aware that copy number calls may be less reliable than those from whole-genome data and should validate important findings with an orthogonal method.

RNA Sequencing

RNA sequencing provides information about KIR gene expression instead of genomic sequence directly. This distinction is important because RNA-seq data reflect which KIR genes are transcribed in the sampled cell population, and the expression level may not correlate with genomic copy number. However, RNA-seq can be valuable for identifying expressed KIR alleles and for studying KIR expression in specific cell types.

The T1K method was applied to tumor single-cell RNA-seq data in the 2023 Genome Research study, demonstrating that KIR genotyping from RNA-seq data is feasible [<a href="#ref-1">1</a>]. The study found that KIR2DL4 expression was enriched in tumor-specific CD8+ T cells, showing the potential of this approach for immunology research [<a href="#ref-1">1</a>]. Researchers using RNA-seq for KIR analysis should recognize that absence of expression does not necessarily indicate absence of the gene, and that expression levels can vary substantially between cell types and conditions.

Long-Read Sequencing

Long-read sequencing technologies such as PacBio HiFi and Oxford Nanopore provide reads that span much larger portions of the KIR locus than short reads. This longer read length can resolve some of the ambiguity that plagues short-read mapping in paralogous regions. Long-read data also enable de novo assembly approaches that can reconstruct complete KIR haplotypes.

The value of long-read sequencing for KIR analysis is supported by recent work on near-complete human genome assemblies. A 2025 Nature Genetics study of Middle Eastern family trios used long-read sequencing to generate highly accurate, near-complete and phased genomes and identified 75 new HLA and KIR alleles [<a href="#ref-5">5</a>]. The same study reported that assembly-based variant calling identified de novo and recessive variants that were strong candidates for causing previously unresolved symptoms, underscoring the value of de novo assembly for disease variant discovery [<a href="#ref-5">5</a>].

Core Principles of KIR Variant Calling

Several principles guide accurate KIR variant calling regardless of the specific tools used.

Joint Consideration of All KIR Genes

KIR genes cannot be analyzed in isolation because their high sequence similarity means that reads from one gene may map to another. A variant caller that analyzes each gene independently will misassign reads and produce incorrect genotypes. Methods that jointly consider alleles across all genotyped genes can reliably identify present genes and distinguish homologous genes. The T1K method described in the 2023 Genome Research study uses this joint approach, which the authors reported enables it to reliably identify present genes and distinguish homologous genes including the challenging KIR2DL5A and KIR2DL5B genes [<a href="#ref-1">1</a>].

Copy Number Awareness

Every KIR variant call must be interpreted in the context of the gene copy number. A variant that appears heterozygous in a gene with two copies may actually be homozygous in a gene with one copy. Conversely, a variant that appears homozygous in a gene with one copy may be a false call if the gene is actually present in two copies with different alleles. Read-depth analysis must be performed before or alongside variant calling to establish the copy number context.

Reference Bias Awareness

KIR variant calling is susceptible to reference bias, where reads that match the reference genome are more likely to be mapped and called than reads that differ from the reference. This bias is particularly problematic for KIR genes because the allele diversity is high and many alleles differ substantially from the reference sequence. Methods that use a pangenome or a set of known haplotypes as the reference can reduce this bias. The Locityper method described in the 2025 Nature Genetics study recruits and aligns reads to locus haplotypes extracted from a pangenome, which the authors reported enables accurate genotyping of challenging genes [<a href="#ref-4">4</a>].

Validation Through Multiple Methods

Given the difficulty of KIR variant calling, results from a single method should be treated with caution. Where possible, validate important findings using an orthogonal approach. For example, a variant identified by short-read mapping could be confirmed by assembly-based analysis or by long-read sequencing. Discrepancies between methods should be investigated instead of ignored.

Practical Workflow for KIR Variant Calling

The following workflow provides a practical approach to KIR variant calling that combines read-depth and assembly-based methods. The workflow assumes the researcher has sequencing data in a standard format such as FASTQ or BAM and access to a Linux computing environment.

Step 1: Assess Data Suitability

Before beginning analysis, assess whether the available data are suitable for KIR variant calling. Check the sequencing type, read length, coverage depth, and data quality. Whole-genome sequencing at 30x coverage or higher is preferred. Whole-exome and RNA-seq data can be used but with the limitations described above. Long-read data enable the most complete analysis.

Verify that the data files are complete and uncorrupted. Check the number of reads, the read length distribution, and the base quality scores. Low-quality data will produce unreliable results regardless of the analysis method used.

Step 2: Perform Read-Depth Analysis for Copy Number Estimation

Read-depth analysis provides the copy number context needed for accurate variant calling. The analysis involves mapping reads to the KIR locus and counting the number of reads that map to each gene. Genes that are absent will show near-zero depth, genes that are hemizygous will show approximately half the depth of homozygous genes, and duplicated genes will show increased depth.

Several approaches can be used for read-depth analysis. Standard alignment tools can map reads to the reference genome, and depth can be calculated using tools such as samtools. However, the paralogy problem means that reads from one KIR gene may map to another, confounding depth estimates. More sophisticated approaches use the joint genotyping methods described above, which account for the sequence similarity between genes.

The copy number calls from read-depth analysis should be recorded for each KIR gene. These calls provide the context for interpreting variant calls and should be included in the final analysis report.

Step 3: Run Specialized KIR Genotyping Tools

Specialized tools that are designed for KIR genotyping should be used instead of standard variant calling pipelines. The T1K method is designed for efficient and accurate inference of KIR or HLA alleles from RNA-seq, whole-genome sequencing, or whole-exome sequencing data [<a href="#ref-1">1</a>]. The BAKIR tool is designed for KIR genotyping and annotation on high-quality, phased genome assemblies [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. The Locityper tool can genotype challenging genes using short-read and long-read whole-genome sequencing [<a href="#ref-4">4</a>].

Each tool has different input requirements and output formats. Read the documentation carefully before running the analysis. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can help researchers learn how to use these tools effectively. The nf-core documentation provides community pipeline standards and usage guidance for reproducible workflow context.

Step 4: Perform Assembly-Based Analysis

For the most accurate results, perform assembly-based analysis of the KIR locus. This approach reconstructs the actual sequence of the KIR region from the sequencing data instead of relying on mapping to a reference. De novo assembly can resolve the sequence of each gene copy and identify novel alleles that differ substantially from the known allele database.

Assembly-based analysis is particularly valuable for samples with novel or rare KIR alleles. The 2025 Nature Genetics study of Middle Eastern genomes demonstrated the value of assembly-based variant calling for identifying novel variation, reporting 42.2 Mb of new sequence with 13.8% impacting known genes and 75 new HLA and KIR alleles [<a href="#ref-5">5</a>]. The same study reported that assembly-based variant calling identified de novo and recessive variants that were strong candidates for causing previously unresolved symptoms [<a href="#ref-5">5</a>].

Assembly-based analysis requires more computational resources than mapping-based analysis and may not be feasible for all laboratories. When assembly is not feasible, specialized genotyping tools that use a pangenome reference can provide some of the benefits of assembly-based analysis.

Step 5: Compare and Validate Results

Compare the results from different methods to identify discrepancies. A variant that is called by multiple independent methods is more reliable than a variant called by only one method. Discrepancies should be investigated to determine whether they result from method-specific artifacts or from genuine biological variation.

Validation can also involve comparing called alleles to known allele databases. The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services that can be used for this purpose. The EMBL-EBI Training provides bioinformatics learning pathways and data-resource training that can help researchers understand how to use these databases effectively.

Step 6: Document and Report Results

Document the analysis methods, parameters, and results in a reproducible format. The Carpentries Lessons provide foundational computing, data, shell, Git, and programming training context that can help researchers develop reproducible analysis workflows. The nf-core documentation provides community pipeline standards for reproducible workflow context.

The report should include the copy number calls for each KIR gene, the called alleles, the variant calls, and the quality metrics for each call. The report should also note any limitations of the analysis, such as regions with low coverage or ambiguous calls.

Options and Tradeoffs in KIR Variant Calling Tools

Several tools are available for KIR variant calling, each with different strengths and limitations. The choice of tool depends on the data type, the research question, and the available computational resources.

T1K

T1K is a computational method for the efficient and accurate inference of KIR or HLA alleles from RNA-seq, whole-genome sequencing, or whole-exome sequencing data [<a href="#ref-1">1</a>]. The method jointly considers alleles across all genotyped genes, which enables it to reliably identify present genes and distinguish homologous genes [<a href="#ref-1">1</a>]. The 2023 Genome Research study reported that T1K can call novel single-nucleotide variants and process single-cell data [<a href="#ref-1">1</a>].

T1K is appropriate for researchers who need to genotype KIR and HLA genes from standard sequencing data types. The joint consideration of all genes is a key advantage for resolving the paralogy problem. The ability to process single-cell data makes T1K valuable for immunology research.

BAKIR

BAKIR is a biologically informed annotator for the KIR locus that is designed for KIR genotyping and annotation on high-quality, phased genome assemblies [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. The tool structures its annotation pipeline around identifying key functional mutations, which the 2024 preprint and 2024 Bioinformatics article reported improves the identification and subsequent relevance of gene and allele calls [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. BAKIR uses a multi-stage mapping, alignment, and variant calling process to ensure high-precision gene and allele identification while maintaining high recall for sequences that are significantly mutated or truncated relative to the known allele database [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

BAKIR is appropriate for researchers who have high-quality, phased genome assemblies and need accurate KIR gene annotations. The tool is freely available on GitHub and can be installed through multiple methods including pip, conda, and singularity container [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>]. The user-friendly command-line interface promotes adoption in the scientific community [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

Locityper

Locityper is a tool capable of genotyping challenging genes using short-read and long-read whole-genome sequencing [<a href="#ref-4">4</a>]. For each target, Locityper recruits and aligns reads to locus haplotypes extracted from a pangenome and finds the likeliest haplotype pair by optimizing read alignment, insert size, and read depth profiles [<a href="#ref-4">4</a>]. The 2025 Nature Genetics study reported that Locityper achieves a median quality value above 35 from both long-read and short-read data across 256 challenging medically relevant loci, outperforming state-of-the-art Illumina and PacBio HiFi variant calling pipelines [<a href="#ref-4">4</a>].

Locityper is appropriate for researchers who need to genotype KIR genes at biobank scale. The study reported a low running time of 1 hour 35 minutes per sample at eight threads, making the tool scalable to large cohorts [<a href="#ref-4">4</a>]. Locityper also provides access to hyperpolymorphic HLA genes and other gene families including MUC and FCGR [<a href="#ref-4">4</a>].

Standard Variant Calling Pipelines

Standard variant calling pipelines that align reads to a linear reference genome and call variants with tools such as GATK are not appropriate for KIR genes. The 2023 Genome Research study noted that KIR genes cannot be genotyped with standard variant calling pipelines because of their high polymorphism and sequence similarity [<a href="#ref-1">1</a>]. Researchers should not attempt to use standard pipelines for KIR analysis.

Tool Comparison for KIR Genotyping

The table below compares the primary specialized tools available for KIR variant calling across key dimensions relevant to tool selection.

ToolSupported Data TypesKey StrengthPrimary Limitation
T1KRNA-seq, whole-genome, whole-exomeJoint consideration of all KIR and HLA genes, single-cell supportMay not resolve all KIR genes in complex cases
BAKIRHigh-quality phased assembliesHigh precision with high recall for divergent sequencesRequires assembly-grade data input
LocityperShort-read and long-read whole-genomeScalable to biobank cohorts, pangenome-basedDesigned for targeted loci instead of full annotation

Observations and Measurements for Quality Control

Quality control is essential for KIR variant calling because the technical challenges can produce false calls that are difficult to distinguish from genuine variation. The following observations and measurements should be recorded for each analysis.

Read Depth at Variant Sites

The read depth at each variant site provides information about the confidence of the call. Low depth at a variant site suggests that the call may be unreliable. High depth provides more confidence but can also indicate mapping artifacts if reads from multiple paralogs are mapping to the same location.

Record the read depth for each variant call. Compare the depth at variant sites to the depth at nearby non-variant sites. Substantial differences may indicate mapping artifacts.

Mapping Quality

The mapping quality of reads at a variant site indicates how confidently the reads were assigned to that location. Low mapping quality suggests that reads may have originated from a different paralog. Record the mapping quality distribution for reads supporting each variant call.

Allele Balance

The allele balance is the proportion of reads supporting each allele at a heterozygous site. For a true heterozygous site in a diploid gene, the allele balance should be approximately 50%. Deviations from this expectation may indicate copy number variation or mapping artifacts. For example, a site that appears heterozygous in a hemizygous gene will show an allele balance that reflects the true genotype instead of a 50% split.

Copy Number Consistency

The copy number calls from read-depth analysis should be consistent across the KIR locus. Adjacent genes should show similar copy numbers unless there is a known structural variant boundary. Inconsistent copy number calls may indicate technical artifacts.

Comparison Across Methods

The most powerful quality control measure is comparison across methods. If two independent methods produce the same variant calls, the calls are more likely to be correct. If methods disagree, the discrepancy should be investigated.

Common Failure Patterns in KIR Variant Calling

Understanding common failure patterns can help researchers identify problems in their analysis and avoid incorrect conclusions.

False Variant Calls from Paralog Mapping

The most common failure pattern is false variant calls resulting from reads that map to the wrong paralog. When a read from KIR2DL5A maps to KIR2DL5B, the variant caller may identify a variant that does not exist in the sample. This failure pattern is characterized by variant calls that are supported by reads with low mapping quality or by an allele balance that does not match the expected pattern.

Missed Variants in Low-Complexity Regions

Some KIR regions have low sequence complexity that makes them difficult to sequence and map. Variants in these regions may be missed entirely because reads do not map or because the variant caller filters them out as low quality. This failure pattern is characterized by gaps in coverage or by variant calls that are absent from regions with known polymorphism.

Incorrect Copy Number Calls

Copy number calls can be incorrect when read depth is affected by factors other than copy number, such as GC bias, capture bias in exome sequencing, or mapping artifacts. An incorrect copy number call can lead to incorrect interpretation of variant calls. For example, a variant that is actually homozygous in a hemizygous gene may be called as heterozygous if the copy number is incorrectly assumed to be two.

Reference Bias

Reference bias occurs when reads that match the reference genome are more likely to be mapped and called than reads that differ from the reference. This bias can cause variants to be missed, particularly in samples with alleles that differ substantially from the reference. The bias is more pronounced for KIR genes because of their high diversity.

Failure to Detect Novel Alleles

Standard variant calling approaches that rely on mapping to a reference genome may fail to detect novel alleles that differ substantially from the reference. Assembly-based methods are better suited for detecting novel alleles. The 2024 BAKIR preprint reported that the tool maintains high recall for sequences that are significantly mutated or truncated relative to the known allele database [<a href="#ref-2">2</a>].

Limitations of KIR Variant Calling

Even with the best available methods, KIR variant calling has limitations that researchers should understand.

Incomplete Allele Databases

KIR allele databases are incomplete, and new alleles continue to be discovered. The 2025 Nature Genetics study of Middle Eastern genomes identified 75 new HLA and KIR alleles, demonstrating that the known allele diversity is incomplete [<a href="#ref-5">5</a>]. Variant calls that do not match known alleles may represent novel alleles or may represent artifacts. Distinguishing between these possibilities requires careful validation.

Population-Specific Variation

KIR variation is population-specific, and methods that work well for one population may perform poorly for another. The 2025 Nature Genetics study reported that the Middle Eastern genomes revealed unique variation relative to existing references, showing enhanced mappability and variant calling when using population-specific references [<a href="#ref-5">5</a>]. Researchers should be aware that reference bias may be more pronounced for samples from populations that are underrepresented in reference databases.

Technical Limitations of Short Reads

Short reads cannot resolve all KIR variation because of the high sequence similarity between genes and the presence of structural variation. Even with specialized methods, some variants may be missed or incorrectly called. Long-read sequencing provides more complete information but is more expensive and less widely available.

Interpretation Challenges

The functional significance of KIR variants is not fully understood. A variant that is called accurately may have unknown functional consequences. Researchers should be cautious about interpreting the biological significance of KIR variants without functional validation.

Safety and Regulatory Context

KIR variant calling is a research application, and the results should not be used for clinical decision-making without appropriate validation and regulatory approval. Researchers should be aware of the following considerations.

Research Use Only

KIR variant calling methods described in this article are research tools. The results should not be used for clinical diagnosis or treatment decisions without validation in a clinical laboratory and approval by the appropriate regulatory authorities. Researchers should clearly label their results as research findings.

Data Privacy

KIR genotyping data are genetic data and should be handled in accordance with applicable privacy regulations. Researchers should ensure that their data handling procedures comply with relevant laws and institutional policies. The NCBI provides official descriptions of databases and data resources that can help researchers understand data sharing and privacy requirements.

Reproducibility

Reproducibility is essential for research integrity. Researchers should document their analysis methods and parameters so that others can reproduce their results. The nf-core documentation provides community pipeline standards for reproducible workflow context. The Carpentries Lessons provide foundational computing and data training that can help researchers develop reproducible workflows.

Professional Escalation Criteria

Researchers should escalate KIR variant calling results to a specialist or to a more comprehensive analysis when certain conditions are met.

Escalate When Results Are Inconsistent Across Methods

If different methods produce conflicting results for the same sample, the results should be escalated to a specialist who can investigate the discrepancy. Inconsistent results may indicate technical artifacts or genuine biological complexity that requires expert interpretation.

Escalate When Novel Alleles Are Suspected

If the analysis identifies a potential novel allele that does not match any known allele in the database, the result should be escalated for confirmation. Novel alleles require careful validation, potentially including long-read sequencing or functional studies. The 2024 BAKIR preprint reported that the tool maintains high recall for sequences that are significantly mutated or truncated relative to the known allele database, but novel alleles still require confirmation [<a href="#ref-2">2</a>].

Escalate When Results Will Be Used for Clinical Decisions

If KIR genotyping results will be used for clinical decisions, the analysis should be performed in a clinical laboratory with appropriate validation and quality control. Research-grade results should not be used for clinical decision-making without confirmation in a clinical setting.

Escalate When Samples Have Complex Structural Variation

Samples with complex structural variation in the KIR locus may require specialized analysis approaches. If the standard workflow produces ambiguous results, the sample should be escalated for assembly-based analysis or long-read sequencing.

Frequently Asked Questions

What makes KIR genes different from other genes for variant calling?

KIR genes have two properties that break standard variant calling pipelines. First, several KIR genes have very similar sequences, so short reads cannot be uniquely assigned to a single gene. Second, KIR genes vary in copy number between individuals, so the diploid assumption of standard variant callers is invalid. The 2023 Genome Research study noted that KIR genes are highly polymorphic, similar to each other in sequence, and may be absent from chromosomes [<a href="#ref-1">1</a>].

Can I use standard variant calling pipelines for KIR genes?

No. Standard variant calling pipelines that align reads to a linear reference genome and call variants with tools such as GATK are not appropriate for KIR genes. The 2023 Genome Research study reported that KIR genes cannot be genotyped with standard variant calling pipelines and that even specialized KIR genotypers could not resolve all KIR genes [<a href="#ref-1">1</a>].

What is the difference between read-depth analysis and assembly-based methods?

Read-depth analysis counts the number of sequencing reads that map to each KIR gene to estimate copy number. Assembly-based methods reconstruct the actual sequence of the KIR locus from the sequencing data without relying on mapping to a reference. Read-depth analysis provides copy number context, while assembly-based methods provide sequence-level resolution for allele identification.

Which sequencing data type is best for KIR variant calling?

Whole-genome sequencing provides the most complete coverage of the KIR locus. Both short-read and long-read whole-genome sequencing can be used. The 2025 Nature Genetics study reported that Locityper achieves a median quality value above 35 from both long-read and short-read data [<a href="#ref-4">4</a>]. Whole-exome and RNA-seq data can be used but have limitations for copy number estimation.

How do I know if my KIR variant calls are reliable?

Reliability can be assessed through multiple measures. Check read depth at variant sites, mapping quality, allele balance, and copy number consistency. Compare results across multiple methods. Variants called by independent methods are more reliable than variants called by only one method.

What should I do if I find a novel KIR allele?

Novel alleles require careful validation. Confirm the finding with an orthogonal method, such as assembly-based analysis or long-read sequencing. Compare the sequence to known alleles in public databases. The NCBI provides official descriptions of databases and sequence resources that can be used for this purpose.

Can KIR genotyping be performed from RNA-seq data?

Yes. The T1K method described in the 2023 Genome Research study was designed to work with RNA-seq data in addition to whole-genome and whole-exome sequencing data [<a href="#ref-1">1</a>]. The study applied T1K to tumor single-cell RNA-seq data and found that KIR2DL4 expression was enriched in tumor-specific CD8+ T cells [<a href="#ref-1">1</a>]. However, RNA-seq data reflect gene expression instead of genomic sequence directly.

What are the limitations of KIR variant calling?

KIR variant calling has several limitations. Allele databases are incomplete, and new alleles continue to be discovered. Population-specific variation can affect accuracy, particularly for populations underrepresented in reference databases. Short reads cannot resolve all KIR variation. The functional significance of many KIR variants is not fully understood.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Efficient and accurate KIR and HLA genotyping with massively parallel sequencing data.](https://pubmed.ncbi.nlm.nih.gov/37169596). Genome research, 2023. [2] [Biologically-informed Killer cell immunoglobulin-like receptor (KIR) gene annotation tool.](https://pubmed.ncbi.nlm.nih.gov/39372800). bioRxiv : the preprint server for biology, 2024. [3] [Biologically-informed killer cell immunoglobulin-like receptor gene annotation tool.](https://pubmed.ncbi.nlm.nih.gov/39432666). Bioinformatics (Oxford, England), 2024. [4] [Locityper enables targeted genotyping of complex polymorphic genes.](https://pubmed.ncbi.nlm.nih.gov/41107551). Nature genetics, 2025. [5] [Near-complete Middle Eastern genomes refine autozygosity and enhance disease-causing and population-specific variant discovery.](https://pubmed.ncbi.nlm.nih.gov/40325133). Nature genetics, 2025.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.