Duplication Mutation: Mechanisms, Consequences, and Detection

By Dr. Zubair Khalid, DVM, MS, PhD ·

Duplication Mutation: Mechanisms, Consequences, and Detection

Introduction to Duplication Mutations

A duplication mutation is a type of chromosomal or gene-level alteration in which a segment of DNA is copied one or more times, resulting in extra genetic material. The duplicated segment can range from a single nucleotide to entire chromosomes or even whole genomes. Unlike Point Mutation in DNA, which alters a single base pair, duplications increase the copy number of a genomic region and therefore change the dosage of all genes and regulatory elements contained within that region.

Duplications belong to the broader class of structural variants (SVs), alongside deletions, inversions, and translocations. They are distinguished from insertions by the origin of the extra sequence: a duplication is a copy of sequence already present elsewhere in the genome, whereas an insertion introduces foreign sequence (e.g., a transposable element or a segment from a different genomic location). This distinction matters mechanistically and diagnostically, as the two events arise through different molecular pathways and have different mutational signatures.

The significance of duplication mutations is dual. On one hand, they are a primary source of genetic novelty. Gene duplication provides raw material for the evolution of new functions, a concept formalized by Susumu Ohno in his 1970 monograph Evolution by Gene Duplication. On the other hand, duplications are a major cause of human genetic disease. Copy number gains at loci such as PMP22 cause Charcot-Marie-Tooth disease type 1A, and duplications of MECP2 cause a severe neurodevelopmental disorder in males. Understanding the mechanisms, consequences, and detection of duplication mutations is therefore essential for both evolutionary biology and clinical genetics.

Definition and Types

Duplication mutations can be classified along several axes: size, orientation, and genomic location.

By size:

  • Small-scale duplications (1 bp to ~1 kb) typically arise from replication errors and are often called microduplications.
  • Segmental duplications (1–200 kb) are blocks of near-identical sequence that constitute ~5% of the human genome.
  • Chromosomal duplications (>1 Mb) are visible by cytogenetic analysis and include partial trisomies.
  • Whole-genome duplications (WGDs) double the entire complement of genetic material and have occurred at least twice in the vertebrate lineage.

By orientation and location:

  • Tandem duplications are adjacent to the original copy, arranged head-to-tail. These are the most common type and often arise from non-allelic homologous recombination (NAHR) or replication slippage.
  • Dispersed duplications are copies located at non-contiguous positions, often on different chromosomes. These typically arise through retrotransposition or translocational events.
  • Inverted duplications place the copy in reverse orientation relative to the original, often at the same locus.

Historical Context

The study of duplications began with cytogenetic observations in Drosophila by Calvin Bridges in the 1930s, who documented the Bar phenotype caused by a tandem duplication of the Bar locus on the X chromosome. This was the first demonstration that a duplication could produce a visible mutant phenotype. Alfred Sturtevant later showed that the Bar duplication could undergo unequal crossing over, generating flies with either one or three copies of the region, establishing the principle that duplications are dynamic and can expand or contract.

The molecular era began with the discovery of globin gene families, where Southern blotting revealed multiple related genes arranged in clusters. This led to the realization that gene duplication followed by divergence is a fundamental mechanism of genome evolution. The completion of genome sequencing projects in the 2000s revealed the extent of segmental duplications and copy number variation (CNV) in normal human genomes, showing that duplications are not rare anomalies but ubiquitous features of genome architecture.

Molecular Mechanisms Generating Duplications

Duplication mutations arise through several distinct molecular mechanisms, each with characteristic sequence signatures and genomic contexts. Understanding these mechanisms is essential for interpreting duplication breakpoints and predicting their functional consequences.

Non-allelic Homologous Recombination

Non-allelic homologous recombination (NAHR) is the most common mechanism generating large duplications in humans. It occurs during meiosis when highly similar sequences at different genomic locations misalign, allowing crossing over between non-allelic homologous regions. This requires stretches of near-identical sequence, typically segmental duplications or repetitive elements such as Alu repeats.

The process proceeds as follows:

  1. During prophase I of meiosis, homologous chromosomes align. If the genome contains two paralogous sequences (e.g., two Alu elements in the same orientation on the same chromosome), they can mispair.
  2. The recombination machinery resolves the Holliday junction using the mispaired homologs as templates.
  3. Crossing over between misaligned chromatids produces one chromatid with a deletion and one with a duplication of the intervening sequence.

NAHR is the mechanism behind the common PMP22 duplication in Charcot-Marie-Tooth disease type 1A. The PMP22 gene lies within a 1.4 Mb region flanked by two ~24 kb low-copy repeats in direct orientation. NAHR between these repeats produces a 1.4 Mb tandem duplication with a frequency of ~1 in 10,000 births.

Key features of NAHR-mediated duplications:

  • Breakpoints cluster within the homologous sequences, not at random positions.
  • The duplicated segment is always flanked by the paralogous sequences that mediated the event.
  • The size of the duplication is determined by the distance between the mispaired sequences.
  • NAHR duplications are typically tandem and direct in orientation.

Replication Slippage

Replication slippage, also called polymerase slippage or slipped-strand mispairing, generates small duplications (1 bp to a few hundred bp) during DNA replication. It occurs most frequently at regions of tandem repeats, where the template and nascent strands can misalign.

The mechanism involves:

  1. During replication, the nascent strand dissociates from the template at a repetitive sequence.
  2. The nascent strand reanneals at a downstream repeat unit, creating a loop in the template strand.
  3. Continued replication copies the looped-out template region a second time, producing a duplication.

The efficiency of replication slippage depends on the repeat unit length and the number of repeats. In the human genome, microsatellites (1–6 bp repeats) are particularly prone to slippage, with mutation rates of 10⁻³ to 10⁻⁴ per locus per generation—orders of magnitude higher than point mutation rates. Longer repeats are more unstable because they provide more opportunities for misalignment.

Replication slippage can also occur during double-strand break repair by microhomology-mediated break-induced replication (MMBIR), where a stalled replication fork uses microhomology (2–5 bp) to switch templates, generating duplications with characteristic short direct repeats at the breakpoint junctions.

Retrotransposition

Retrotransposition is a mechanism that generates dispersed duplications through an RNA intermediate. Processed pseudogenes are the classic example: an mRNA transcribed from a parent gene is reverse-transcribed by the enzymatic machinery of a retrotransposon (typically LINE-1) and inserted into a new genomic location.

The process requires:

  1. Transcription of the parent gene to produce an mRNA.
  2. Binding of LINE-1 ORF2 protein to the mRNA (in trans).
  3. Reverse transcription of the mRNA at a new genomic site, often a staggered nick in AT-rich sequence.
  4. Integration of the cDNA, typically with a poly(A) tail and flanked by target site duplications.

Processed pseudogenes lack introns and promoter sequences, so they are usually non-functional. However, if the retroposed copy inserts downstream of an existing promoter, it can acquire expression and potentially new function. The human genome contains approximately 8,000 processed pseudogenes, and several have been co-opted into functional genes, such as GLUD2, a glutamate dehydrogenase gene that arose by retrotransposition in the hominid lineage.

Retrotransposition also generates duplications of partial gene sequences. For example, SVA (SINE-VNTR-Alu) elements can transduce flanking genomic sequence, carrying exons from adjacent genes to new locations. This mechanism has created chimeric genes in primates, including the PMS2CL gene, which is a partial duplication of PMS2 generated by SVA-mediated transduction.

Chromosomal Rearrangements

Large duplications can also arise as byproducts of other chromosomal rearrangements, particularly translocations and inversions. During the repair of DNA double-strand breaks by non-homologous end joining (NHEJ), broken chromosome ends can be joined incorrectly, generating duplications at the junction.

One well-characterized pathway is the breakage-fusion-bridge cycle, first described by Barbara McClintock in maize. This cycle operates as follows:

  1. A chromosome break occurs, generating a telomere-less end.
  2. The broken chromosome replicates, and the sister chromatids fuse at their broken ends.
  3. During anaphase, the fused dicentric chromosome is pulled to both poles, breaking at a random position.
  4. The new break can generate a duplication of the region between the original break and the new breakpoint.

This cycle can amplify large genomic regions and is observed in many cancers, where it drives oncogene amplification.

Another mechanism is fork stalling and template switching (FoSTeS), which occurs when a replication fork stalls at a difficult-to-replicate region. The lagging strand dissociates and reanneals at a different replication fork nearby, using microhomology to prime DNA synthesis. This can generate complex duplications with multiple breakpoints and interspersed deletions, a pattern frequently observed in MECP2 duplications.

Consequences of Duplication Mutations

The functional consequences of duplication mutations depend on the size of the duplicated region, the genes it contains, and the regulatory context. Outcomes range from no phenotypic effect to embryonic lethality.

Gene Dosage Effects

The most direct consequence of a duplication is increased copy number of the genes within the duplicated region, leading to increased gene expression. This can be tolerated, beneficial, or pathogenic depending on the gene and the magnitude of the increase.

For many genes, expression is tightly regulated, and even a 1.5-fold increase in copy number can cause disease. This is termed dosage sensitivity. The MECP2 gene provides a striking example: loss-of-function mutations cause Rett syndrome in females, while duplications of MECP2 cause a distinct but equally severe neurodevelopmental disorder in males. The male phenotype demonstrates that too much MeCP2 protein is as harmful as too little, reflecting the protein's role as a transcriptional repressor that must be maintained within a narrow concentration range.

Dosage effects can also be mediated by regulatory elements within the duplicated region. If a duplication includes an enhancer but not its target gene, it can alter the expression of genes at the new location—a phenomenon called enhancer adoption. Conversely, duplicating a gene without its enhancer can reduce expression if the gene is now competing for limiting transcription factors.

Gene Fusion and Divergence

When a duplication places two genes in close proximity, it can create a fusion gene with novel properties. This is particularly common in cancer, where tandem duplications of oncogenes can generate fusion transcripts. The TMPRSS2-ERG fusion in prostate cancer arises from a deletion, but the BCR-ABL1 fusion in chronic myeloid leukemia results from a translocation that effectively duplicates the fusion junction.

In evolutionary contexts, gene duplication followed by divergence is the primary mechanism for generating new gene functions. The duplicated copy is initially redundant, freeing it from purifying selection. Over time, the copy can accumulate mutations that would be deleterious in a single-copy gene. This process can lead to:

  • Neofunctionalization: One copy acquires a new function while the other retains the ancestral function.
  • Subfunctionalization: The ancestral functions are partitioned between the two copies, with each copy losing some functions.
  • Pseudogenization: One copy accumulates inactivating mutations and becomes a non-functional pseudogene.

The globin gene family illustrates all three outcomes. The ancestral globin gene duplicated to produce myoglobin and hemoglobin subunits. Subsequent duplications produced the α-globin and β-globin clusters on chromosomes 16 and 11, respectively. Within the β-globin cluster, the HBG1 and HBG2 genes (fetal γ-globins) arose by a relatively recent duplication and have undergone subfunctionalization, with HBG1 encoding the Gγ chain and HBG2 encoding the Aγ chain, differing by a single amino acid.

Pathogenic Duplications

Duplications cause disease through several mechanisms:

  1. Direct gene dosage: Increased expression of a dosage-sensitive gene (e.g., PMP22, MECP2, SNCA in Parkinson's disease).
  2. Disruption of gene structure: A duplication breakpoint can fall within a gene, disrupting its coding sequence.
  3. Position effects: A duplication can place a gene near new regulatory elements, altering its expression pattern.
  4. Genomic instability: Duplicated regions provide substrates for further NAHR, increasing the risk of additional rearrangements.

The clinical consequences of pathogenic duplications are highly variable. The 22q11.2 duplication syndrome, for example, has highly variable expressivity, with some carriers being asymptomatic and others having congenital heart defects, palatal anomalies, and intellectual disability. This variability complicates genetic counseling and underscores the importance of understanding the specific genes and regulatory elements within the duplicated region.

Duplication Mutations in Evolution

Duplication mutations are the most important source of new genetic material for evolution. Without duplication, the genome can only reshuffle existing genes; with duplication, entirely new genes can arise.

Gene Family Expansion

Gene families expand through successive rounds of duplication. The olfactory receptor (OR) gene family in mammals illustrates this process dramatically. Humans have approximately 400 functional OR genes and 600 pseudogenes, all derived from a single ancestral gene through repeated duplications. The family expanded through both tandem duplications (creating clusters on most chromosomes) and whole-genome duplications.

The expansion of gene families is not random. Genes involved in environmental sensing (olfaction, taste, immunity) tend to expand and contract rapidly, reflecting selection for diversity. In contrast, genes involved in core cellular processes (DNA replication, transcription, translation) are typically single-copy or present in small families, reflecting strong purifying selection against duplication.

Whole Genome Duplication

Whole-genome duplication (WGD) events have occurred repeatedly in eukaryotic evolution. Two rounds of WGD occurred in the ancestor of vertebrates (the "2R" hypothesis), and additional WGDs occurred in the teleost fish lineage, in yeast, and in many plant lineages, including the ancestors of wheat and cotton.

The consequences of WGD are profound:

  1. Immediate redundancy: Every gene is present in two copies, providing extensive raw material for evolutionary innovation.
  2. Rapid gene loss: Most duplicated genes are lost within a few million years, returning the genome to a diploid state.
  3. Retention of specific gene classes: Genes involved in development, transcriptional regulation, and signal transduction are preferentially retained after WGD, likely because their dosage is buffered by complex regulatory networks.
  4. Subfunctionalization and neofunctionalization: The retained duplicates undergo functional divergence, contributing to the evolution of developmental complexity.

The yeast Saccharomyces cerevisiae provides a well-studied example. It underwent a WGD approximately 100 million years ago, followed by massive gene loss. The ~5,800 genes in the yeast genome are descended from ~4,400 ancestral genes, with ~550 pairs of ohnologs (genes retained from the WGD) still present. These ohnologs are enriched for genes involved in transcription and signal transduction, and many have undergone subfunctionalization.

Neofunctionalization and Subfunctionalization

The fate of duplicated genes depends on the balance between mutation, selection, and genetic drift. The classical model proposed by Ohno suggested that one copy is freed from selection and can accumulate mutations, occasionally acquiring a new function (neofunctionalization). However, this model has been refined by the duplication-degeneration-complementation (DDC) model, which proposes that subfunctionalization is more common.

In the DDC model, the ancestral gene has multiple functions or regulatory elements. After duplication, mutations that inactivate one function in one copy and a different function in the other copy can be tolerated because the two copies together still provide all ancestral functions. This partitions the ancestral functions between the copies and preserves both genes.

A classic example is the engrailed gene in zebrafish, which has two copies (eng1a and eng1b) that arose from the teleost WGD. The ancestral gene is expressed in both the pectoral fin buds and a subset of neurons. In zebrafish, eng1a is expressed only in the fin buds and eng1b only in the neurons—a clear case of subfunctionalization.

Methods for Detecting and Characterizing Duplications

The detection of duplication mutations requires methods that can identify both the presence of extra genetic material and its genomic location. The choice of method depends on the size of the duplication, the resolution required, and the sample type.

Cytogenetic Methods

Classical cytogenetic techniques can detect duplications larger than ~5–10 Mb.

  • G-banded karyotyping: Chromosomes are stained with Giemsa, producing a characteristic banding pattern. Duplications appear as enlarged bands or extra chromosomal material. This method detects duplications >5 Mb but cannot identify the exact breakpoints.
  • Fluorescence in situ hybridization (FISH): Fluorescently labeled DNA probes are hybridized to metaphase chromosomes or interphase nuclei. A duplication is detected as an extra fluorescent signal. FISH can detect duplications as small as ~100 kb and can determine whether the duplication is tandem or dispersed.
  • Chromosomal microarray (CMA): Although not strictly cytogenetic, CMA bridges cytogenetics and molecular methods. Array-based comparative genomic hybridization (aCGH) and single-nucleotide polymorphism (SNP) arrays can detect duplications as small as 1–10 kb, depending on probe density.

Array-Based Methods

Array-based methods are the current clinical standard for detecting copy number variants (CNVs), including duplications.

  • aCGH: Test and reference DNA are labeled with different fluorophores and co-hybridized to an array of oligonucleotide probes. The ratio of test to reference fluorescence at each probe indicates copy number. A duplication produces a ratio of ~1.5 (3 copies vs. 2 copies).
  • SNP arrays: These arrays detect both copy number and genotype. A duplication is indicated by increased signal intensity and altered allele frequencies. SNP arrays have the advantage of detecting loss of heterozygosity and uniparental disomy, which aCGH cannot.

Array-based methods cannot determine the orientation or genomic arrangement of a duplication. A tandem duplication and a dispersed duplication produce the same array profile.

Next-Generation Sequencing Approaches

Next-generation sequencing (NGS) has revolutionized duplication detection by providing base-pair resolution and the ability to determine breakpoint structure.

  • Whole-genome sequencing (WGS): Paired-end reads that map with an unexpected insert size or orientation indicate structural variants. A tandem duplication produces reads where the insert size is larger than expected or where read pairs map in a "tail-to-tail" orientation. Breakpoint resolution is achieved by analyzing split reads that span the duplication junction.
  • Whole-exome sequencing (WES): Exome capture followed by sequencing can detect exonic duplications by analyzing read depth. However, WES has limited sensitivity for duplications because capture efficiency varies across exons, and breakpoints in intronic or intergenic regions are not captured.
  • Long-read sequencing: Platforms such as Oxford Nanopore and Pacific Biosciences produce reads of 10–100 kb, which can span entire duplication breakpoints. This allows direct detection of the duplication junction and determination of orientation and arrangement.

The computational analysis of NGS data for duplications uses several approaches:

  1. Read-depth analysis: The number of reads mapping to a genomic interval is proportional to copy number. Tools such as CNVnator and Control-FREEC use read depth to identify duplications.
  2. Paired-end mapping: Discordant read pairs indicate structural variants. Tools such as BreakDancer and DELLY identify duplications from read pairs with abnormal insert sizes.
  3. Split-read analysis: Reads that span a breakpoint are split into two parts, each mapping to a different genomic location. Tools such as Pindel and LUMPY use split reads to localize breakpoints at base-pair resolution.

Bioinformatics Tools

The following table summarizes commonly used tools for duplication detection:

ToolMethodInputOutputStrengthsLimitations
CNVnatorRead depthWGS BAMCNV callsSensitive for large CNVsPoor breakpoint resolution
DELLYPaired-end + split-readWGS BAMSV callsBreakpoint resolutionRequires high coverage
LUMPYProbabilistic SV callingWGS BAMSV callsIntegrates multiple signalsComplex to configure
PindelSplit-readWGS BAMSV callsDetects small indels and duplicationsHigh false-positive rate
Control-FREECRead depthWGS/WES BAMCNV callsHandles GC biasRequires matched normal
XHMMRead depth (exome)WES BAMCNV callsOptimized for exome dataLimited to exonic regions

Studying Duplication Mutations in the Laboratory

Experimental approaches to study duplication mutations fall into two categories: generating duplications de novo and detecting/quantifying existing duplications.

Mutagenesis Screens

Classical genetic screens in model organisms have identified duplication mutations by their phenotypes. In Drosophila, screens for dominant visible mutations recovered duplications of the Bar locus. In yeast, screens for increased gene expression identified duplications of genes whose product is rate-limiting for growth on specific media.

A powerful approach is the use of selectable markers. In yeast, a duplication of a gene that complements an auxotrophic mutation can be selected by growth on medium lacking the required nutrient. This approach has been used to isolate duplications of the CUP1 gene (copper resistance) and the HXT hexose transporter genes.

CRISPR-Mediated Duplication

The CRISPR-Cas9 system can generate targeted duplications by creating two double-strand breaks flanking the region to be duplicated. The repair of these breaks by NHEJ can produce a tandem duplication if the broken ends are joined in a specific orientation.

The protocol for CRISPR-mediated duplication is as follows:

  1. Design two single-guide RNAs (sgRNAs) targeting sites flanking the region to be duplicated. The sgRNAs should be separated by the desired duplication size.
  2. Clone the sgRNAs into expression vectors or use synthetic sgRNAs complexed with recombinant Cas9 protein.
  3. Deliver the Cas9-sgRNA complexes to cells by transfection or electroporation. For ribonucleoprotein (RNP) delivery, mix 100 pmol Cas9 protein with 120 pmol sgRNA in 20 µL of buffer and incubate at 37°C for 10 minutes before electroporation.
  4. After 48–72 hours, harvest cells and screen for duplications by PCR across the predicted junction.
  5. Single-cell clone positive populations to isolate cells with the desired duplication.

The efficiency of CRISPR-mediated duplication varies with cell type and locus. In human cell lines, efficiencies of 1–10% are typical. The main limitation is that NHEJ repair is imprecise, often producing small insertions or deletions at the junction.

Copy Number Variant Assays

Quantitative assays for detecting and measuring duplications include:

  • Quantitative PCR (qPCR): Primers are designed within the duplicated region, and the amplification is compared to a reference gene with known copy number. A duplication produces a 1.5-fold increase in amplification relative to the reference. The reaction uses SYBR Green or TaqMan chemistry, with 40 cycles of amplification (95°C for 15 s, 60°C for 60 s) on a real-time PCR instrument.
  • Digital PCR (dPCR): DNA is partitioned into thousands of nanoliter droplets or wells, and each partition is scored for the presence or absence of the target sequence. The absolute copy number is calculated from the proportion of positive partitions using Poisson statistics. dPCR is more precise than qPCR for detecting small copy number changes.
  • Multiplex ligation-dependent probe amplification (MLPA): This method uses pairs of probes that hybridize to adjacent sequences in the target region. After ligation, the probes are amplified by PCR using universal primers, and the products are separated by capillary electrophoresis. The peak height of each probe is proportional to copy number. MLPA can detect duplications of individual exons and is widely used in clinical diagnostics.

Common Pitfalls in Interpreting Duplication Mutations

Accurate interpretation of duplication mutations requires awareness of several recurring problems.

Misassembly Artifacts

Reference genome assemblies contain errors, particularly in regions of segmental duplication. These regions are difficult to assemble because their high sequence similarity causes assemblers to collapse the copies into a single contig. As a result, read mapping to these regions is unreliable, and duplication calls may be false positives or false negatives.

The human reference genome (GRCh38) still has gaps and misassemblies in segmental duplication regions. When analyzing WGS data, it is essential to filter out calls in known segmental duplications or to use a reference that has been specifically improved for these regions, such as the telomere-to-telomere (T2T) assembly.

Breakpoint Complexity

Duplication breakpoints are often more complex than a simple junction. Many duplications are accompanied by small insertions, deletions, or inversions at the breakpoint. This complexity arises from the repair mechanisms that generate duplications—NHEJ and FoSTeS frequently produce "complex" structural variants with multiple breakpoints.

When validating a duplication by PCR, primers designed to flank the predicted breakpoint may fail to amplify if the actual breakpoint is shifted by a few base pairs. Sequencing across the junction is essential to confirm the exact breakpoint structure.

Functional Redundancy

A duplication may be present but have no phenotypic effect because the duplicated genes are functionally redundant with other genes in the genome. This is particularly common for genes that belong to large families. For example, a duplication of one olfactory receptor gene is unlikely to produce a phenotype because hundreds of other OR genes provide redundant function.

Conversely, the absence of a phenotype does not mean the duplication is benign. The duplication may have subtle effects on fitness that are not apparent in standard laboratory assays. This is a particular concern in clinical genetics, where a duplication of unknown significance (VUS) may be reported without clear evidence of pathogenicity.

Practical Summary and Key Takeaways

Duplication mutations are a fundamental class of genetic variation with dual significance: they are both a major source of evolutionary innovation and a common cause of human disease. The mechanisms that generate duplications—NAHR, replication slippage, retrotransposition, and chromosomal rearrangements—produce characteristic sequence signatures that can be used to identify the causative mechanism. The functional consequences of duplications range from benign to lethal, depending on the genes involved and their dosage sensitivity.

Detection of duplications requires methods appropriate to the size and context of the event. Cytogenetic methods detect large duplications, array-based methods detect CNVs at moderate resolution, and NGS methods provide base-pair resolution and breakpoint structure. Each method has limitations, and accurate interpretation requires careful validation and awareness of common pitfalls.

Frequently Asked Questions

What is a duplication mutation?

A duplication mutation is a genetic alteration in which a segment of DNA is copied, producing extra genetic material. The duplicated segment can range from a single nucleotide to an entire chromosome or genome. Duplications increase the copy number of the genes and regulatory elements within the duplicated region.

What is an example of a duplication mutation?

The most well-known example is the duplication of the PMP22 gene on chromosome 17, which causes Charcot-Marie-Tooth disease type 1A. Other examples include duplications of MECP2 causing a severe neurodevelopmental disorder, duplications of SNCA causing familial Parkinson's disease, and the Bar duplication in Drosophila that was the first duplication mutation described.

How does a duplication mutation occur?

Duplications arise through several mechanisms: non-allelic homologous recombination (NAHR) during meiosis, replication slippage at repetitive sequences, retrotransposition through an RNA intermediate, and repair of DNA double-strand breaks by non-homologous end joining or fork stalling and template switching.

Are duplication mutations harmful?

Not necessarily. Many duplications are benign, particularly small duplications in non-coding regions or duplications of genes with functional redundancy. However, duplications of dosage-sensitive genes can cause disease, and duplications can disrupt gene structure or regulatory elements. The effect depends on the size, location, and gene content of the duplication.

Can duplication mutations be inherited?

Yes. Duplications that arise in germ cells can be transmitted to offspring. Inherited duplications follow Mendelian inheritance patterns. However, many duplications arise de novo during gametogenesis or early development and are not present in either parent.

How are duplication mutations detected?

Duplications are detected by cytogenetic methods (karyotyping, FISH), array-based methods (aCGH, SNP arrays), and next-generation sequencing (whole-genome or whole-exome sequencing). The choice of method depends on the size of the duplication and the resolution required. Validation is typically performed by qPCR, digital PCR, or MLPA.

What is the difference between a duplication and an insertion mutation?

A duplication is a copy of sequence that is already present elsewhere in the genome, whereas an insertion introduces foreign sequence—such as a transposable element or a segment from a different genomic location. Duplications increase copy number of existing genes; insertions introduce new sequence that may or may not be functional.

Key Takeaways

  • Duplication mutations increase the copy number of genomic segments, ranging from single nucleotides to whole genomes, and are distinct from insertions in that the extra sequence originates from elsewhere in the genome.
  • The primary mechanisms generating duplications are non-allelic homologous recombination, replication slippage, retrotransposition, and chromosomal rearrangement repair pathways, each leaving characteristic sequence signatures.
  • Duplications cause disease primarily through gene dosage effects, but they can also disrupt gene structure or alter regulatory landscapes.
  • In evolution, duplications are the raw material for gene family expansion, neofunctionalization, and subfunctionalization, with whole-genome duplications having occurred repeatedly in eukaryotic history.
  • Detection methods span cytogenetics, array-based approaches, and next-generation sequencing, with long-read sequencing providing the most complete breakpoint resolution.
  • Interpretation of duplication mutations requires caution regarding reference genome misassembly, breakpoint complexity, and functional redundancy.
  • The dual nature of duplications—as both a source of genetic novelty and a cause of disease—makes them a central topic in molecular biology, evolution, and clinical genetics.

Further Reading

  • Rinaldi I et al. Prognostic Significance of Fms-Like Tyrosine Kinase 3 Internal Tandem Duplication Mutation in Non-Transplant Adult Patients with Acute Myeloblastic Leukemia: A Systematic Review and Meta-Analysis. Asian Pacific journal of cancer prevention : APJCP. 2020. PubMed 33112537
  • Ono S. Gene duplication, mutation load, and mammalian genetic regulatory systems. Journal of medical genetics. 1972. PubMed 4562432
  • Levis MJ et al. Gilteritinib as Post-Transplant Maintenance for AML With Internal Tandem Duplication Mutation of FLT3. Journal of clinical oncology : official journal of the American Society of Clinical Oncology. 2024. PubMed 38471061
  • Burchert A et al. Sorafenib Maintenance After Allogeneic Hematopoietic Stem Cell Transplantation for Acute Myeloid Leukemia With FLT3-Internal Tandem Duplication Mutation (SORMAIN). Journal of clinical oncology : official journal of the American Society of Clinical Oncology. 2020. PubMed 32673171
  • Liang C et al. Gilteritinib maintenance after allogeneic hematopoietic stem cell transplantation for relapsed/refractory acute myeloid leukemia with FLT3-internal tandem duplication mutation. Hematology (Amsterdam, Netherlands). 2025. PubMed 40525981
  • Liu F et al. KIT A502_Y503 duplication mutation serves as a potential and universal target for neoantigen peptide in Chinese GIST patients. Journal of gastroenterology and hepatology. 2023. PubMed 36879550

Related Topics

Related Clinical & Scientific Guides