# C4A Gene: Structure, Function, and Clinical Significance


## Key Takeaways

- C4A is a crucial component of the classical and lectin complement pathways, functioning as a covalent opsonin that tags immune complexes and microbial surfaces for clearance via amide bond formation with amino groups.
- Extraordinary copy number variation (CNV) of C4A, ranging from 0 to 8 copies per diploid genome, is a significant genetic risk factor for systemic lupus erythematosus (SLE), juvenile dermatomyositis (JDM), and age-related macular degeneration (AMD).
- Beyond immunity, C4A is expressed in the brain and plays a critical role in synaptic pruning, with increased C4A expression linked to elevated schizophrenia risk due to excessive synapse elimination by microglia.
- C4A deficiency, often caused by gene deletion or conversion within the polymorphic MHC class III region, is a strong risk factor for SLE, associated with earlier onset and more severe disease, and is diagnosed through protein allotyping and genetic testing.
- Pathogens employ complement evasion strategies against C4A, including viral glycoprotein binding to C4b and bacterial secretion of inhibitors, while C4A itself can modulate the gut microbiome and influence host-pathogen interactions.
- Therapeutic targeting of the complement system, while not directly focused on C4A, includes FDA-approved inhibitors of C5 (eculizumab, ravulizumab), C3 (pegcetacoplan), and factor B (iptacopan), with investigational agents targeting C1s (sutimlimab) and MASP-2 also impacting C4 activation.

---

## Executive Summary & Key Metadata

The complement component 4A (C4A) gene encodes the C4A isotype of complement component 4 (C4), a central protein of the classical and lectin complement activation pathways. C4A is a constituent of the thioester-containing protein (TEP) family and functions as a covalent opsonin, tagging immune complexes and microbial surfaces for clearance. Unlike its near-identical paralog C4B, C4A preferentially forms amide bonds with amino group-bearing antigens, a biochemical distinction that underpins its unique roles in immune complex solubilization and autoimmunity. The gene resides within the highly polymorphic major histocompatibility complex (MHC) class III region on chromosome 6p21.33, embedded in the RP-C4-CYP21-TNX (RCCX) modular structure. Copy number variations (CNVs) of C4A are among the most significant genetic risk factors for systemic lupus erythematosus (SLE), and structural variation in C4A has been robustly associated with schizophrenia, age-related macular degeneration (AMD), and various autoimmune and infectious disease phenotypes. This manual provides a comprehensive, publication-grade reference on the genomic architecture, structural biology, signaling pathways, pathogenic mutations, host-pathogen interactions, pharmacogenomics, and bioinformatic resources pertaining to C4A.

| **Attribute** | **Detail** |
|---|---|
| **HGNC Symbol** | C4A |
| **UniProt Accession** | P0C0L4 |
| **Representative PDB ID** | true (structural models available via homology; see Section 2) |
| **Chromosomal Locus** | 6p21.33 (MHC class III region) |
| **Primary Molecular Function** | Complement component C4A; covalent opsonin; classical/lectin pathway mediator; immune complex clearance; synaptic pruning |
| **Disease & Pathology Associations** | Systemic lupus erythematosus (SLE), schizophrenia, age-related macular degeneration (AMD), juvenile dermatomyositis, autoimmune hepatitis, Graves' disease, recurrent respiratory infections, Alzheimer's disease, type 1 diabetes |
| **Gene Size** | ~20.6 kb (long form); ~14.6 kb (short form, due to HERV-K insertion polymorphism) |
| **Protein Length** | 1,744 amino acids (mature protein after signal peptide cleavage) |
| **Expression Pattern** | Liver (hepatocytes), macrophages, monocytes, brain (neurons, glia), adrenal cortex (cryptic transcripts) |
| **Copy Number Range** | 0–8 copies per diploid genome (population-dependent) |

---

## 1. Genomic Locus, Chromosomal Organization & Isoforms

### 1.1 Chromosomal Location and RCCX Modular Architecture

The C4A gene is located on the short arm of chromosome 6 at cytogenetic band 6p21.33, within the class III region of the human major histocompatibility complex (MHC). This genomic segment is one of the most gene-dense and polymorphic regions of the human genome. The C4A gene is embedded within a tandemly repeated modular structure known as the RCCX module, which consists of four genes in a fixed linear arrangement: **RP** (also called STK19, serine/threonine kinase 19), **C4** (complement component 4, either C4A or C4B), **CYP21** (steroid 21-hydroxylase, either CYP21A2 functional gene or CYP21A1P pseudogene), and **TNX** (tenascin-X, either TNXB functional gene or TNXA pseudogene). The RCCX module exists as a bimodular, trimodular, or quadrimodular structure on individual haplotypes, with the number of modules varying from one to four. Each module contains either a C4A or C4B gene, and the duplication/deletion of entire modules is the primary mechanism generating CNV at the C4A and C4B loci.

The complete exon-intron structure of the human C4A gene was first determined by Chack-Yung Yu in 1991, revealing a gene of approximately 20.6 kilobases (kb) in its long form, composed of 41 exons. The gene contains a large intron of 6–7 kb between exons 9 and 10, which is the site of a critical size dichotomy: the presence or absence of a human endogenous retrovirus (HERV-K) insertion in this intron determines whether the gene is "long" (C4L, ~20.6 kb) or "short" (C4S, ~14.6 kb). The HERV-K element is a 6.4 kb insertion that is present in approximately 60–70% of C4 genes in Caucasian populations. This size dichotomy has functional consequences, as the long form of C4A (C4AL) has been specifically associated with increased schizophrenia risk.

### 1.2 Promoter Architecture and Transcriptional Regulation

The C4A promoter region lacks a canonical TATA box but contains multiple GC-rich elements and binding sites for ubiquitous transcription factors including Sp1, AP-1, and NF-κB. The basal promoter activity is primarily driven by a cluster of Sp1 binding sites located within 200 base pairs upstream of the transcription start site. The 5' untranslated region (UTR) is encoded by exons 1 and 2, with the ATG initiation codon located in exon 2.

Tissue-specific regulation of C4A expression is complex. While the liver is the primary site of C4 synthesis, extrahepatic expression occurs in monocytes, macrophages, and microglia. In the brain, C4A is expressed by neurons and glial cells, with expression levels modulated by inflammatory stimuli. A particularly intriguing regulatory feature was identified by Tee et al. (1995), who discovered a cryptic promoter within intron 35 of the human C4A gene that initiates abundant adrenal-specific transcription of a 1 kb RNA. This adrenal-specific transcript is initiated from a promoter element that shares homology with the CYP21 promoter, suggesting a complex evolutionary relationship between the C4A and CYP21 genes within the RCCX module. The functional significance of this adrenal transcript remains incompletely understood, but it may represent a regulatory RNA or a truncated protein product.

### 1.3 Alternative Splicing and Isoforms

The C4A gene undergoes alternative splicing that generates multiple mRNA isoforms. The canonical full-length transcript encodes the 1,744-amino acid pre-pro-protein, which is proteolytically processed into three polypeptide chains (α, β, and γ) linked by disulfide bonds. However, several splice variants have been characterized:

1. **Full-length C4A transcript**: Encodes the complete C4A protein with all three chains.
2. **Intron 9 retention variant**: A splice variant that retains part of intron 9, leading to a truncated protein. This variant was found to be upregulated in mastitis-infected dairy cattle, suggesting a role in inflammatory responses.
3. **Adrenal-specific short transcript**: Initiated from the intron 35 promoter, producing a ~1 kb RNA of unknown function.

In dairy cattle, a novel splice variant of C4A was identified that is differentially expressed in mastitis-infected mammary tissue, indicating that alternative splicing of C4A may be regulated by inflammatory signals. The expression of C4A splice variants in the bovine mammary gland parenchyma during staphylococcal infection further supports the role of C4A splicing in host defense.

### 1.4 Copy Number Variation and Haplotype Diversity

The C4A gene exhibits extraordinary copy number variation (CNV) in human populations. The diploid copy number of C4A ranges from 0 to 8, with the most common copy number being 2 (one per chromosome 6). However, the distribution varies significantly across ethnic groups. East Asian populations have a higher frequency of C4A deficiency (zero copies) compared to Europeans, while the genetic mechanisms underlying C4A deficiency differ between populations. In Europeans, C4A deficiency is primarily caused by gene deletion through unequal crossover events in the RCCX module, whereas in East Asians, a substantial proportion of C4A deficiency results from a single-nucleotide polymorphism that converts C4A to C4B (the "C4A-to-C4B conversion").

The RCCX module structure and C4A CNV are tightly linked to HLA haplotypes. Specific HLA haplotypes, such as HLA-B44031;DRB1*1503, which are common in sub-Saharan African populations, carry C4A gene deletions and contribute to ethnicity-specific lupus susceptibility. The C4A*4, C4B*2 haplotype is associated with C2 deficiency, and the C4A*2A*3 duplication haplotype has been documented in the Old Order Amish.

---

## 2. 3D Protein Domain Architecture & Structural Biology

### 2.1 Primary Structure and Proteolytic Processing

The C4A protein is synthesized as a single-chain pre-pro-protein of 1,744 amino acids (molecular weight ~200 kDa) with an N-terminal signal peptide of 19 amino acids. During biosynthesis, the pro-protein undergoes sequential proteolytic cleavages:

1. **Signal peptide cleavage**: Removes the 19-amino acid signal peptide in the endoplasmic reticulum.
2. **Internal cleavage at Arg-654**: Generates the β-chain (residues 20–654) and the α-chain (residues 655–1,441).
3. **Furin-mediated cleavage at Arg-1,441**: Releases the γ-chain (residues 1,442–1,744).

The mature C4A protein is a heterotrimer composed of three disulfide-linked chains: β-chain (~70 kDa), α-chain (~95 kDa), and γ-chain (~33 kDa). The three chains remain covalently associated through inter-chain disulfide bonds, forming the functional C4A molecule.

### 2.2 Domain Architecture

The C4A protein can be divided into several functional domains from the N-terminus to the C-terminus:

**β-Chain (Residues 20–654)**:
- **N-terminal domain (residues 20–200)**: Contains the binding site for C1s and MASP-2, the proteases that cleave C4 to generate C4a and C4b.
- **Central β-sheet domain (residues 200–450)**: Forms the core of the β-chain and contributes to the overall structural stability of the molecule.
- **C-terminal β-chain domain (residues 450–654)**: Contains the C4b-binding protein (C4BP) interaction site.

**α-Chain (Residues 655–1,441)**:
- **C4a anaphylatoxin domain (residues 655–761)**: This N-terminal segment of the α-chain is released as C4a upon activation by C1s/MASP-2. C4a is a weak anaphylatoxin that signals through the C3a receptor (C3aR) and has recently been shown to bind oxytocin.
- **Thioester domain (residues 762–1,100)**: Contains the catalytic thioester bond formed between Cys-991 and Gln-994. This is the most functionally critical domain of C4A. The thioester bond is buried in a hydrophobic pocket in the native protein and becomes exposed upon activation, allowing covalent attachment to target surfaces.
- **Central α-chain domain (residues 1,100–1,300)**: Contains the binding site for complement receptor 1 (CR1/CD35) and the C4b-binding protein.
- **C-terminal α-chain domain (residues 1,300–1,441)**: Contains the factor I cleavage sites and the C4c/C4d junction.

**γ-Chain (Residues 1,442–1,744)**:
- **γ-chain domain**: Contains the C-terminal portion of the molecule and contributes to the overall stability of the C4b fragment. The γ-chain also contains the binding site for the C4b-binding protein.

### 2.3 The Thioester Bond and Isotype-Specific Reactivity

The defining biochemical feature of C4A is the intramolecular thioester bond formed between the side chains of Cys-991 and Gln-994 within the thioester domain. This bond is chemically reactive and mediates the covalent attachment of C4b to target surfaces. The critical difference between C4A and C4B lies in the amino acid residues at positions 1,101–1,106 (the "isotypic" residues). C4A has the sequence **PCPVLD** at these positions, while C4B has **LSPVIH**. This six-amino acid difference determines the chemical reactivity of the thioester:

- **C4A** preferentially reacts with amino groups (-NH₂) to form amide bonds. This makes C4A more efficient at binding to immune complexes and protein antigens.
- **C4B** preferentially reacts with hydroxyl groups (-OH) to form ester bonds. This makes C4B more efficient at binding to carbohydrate-rich surfaces, such as bacterial cell walls.

The structural basis for this differential reactivity was established by Law, Dodds, and Porter (1984), who demonstrated that the isotypic residues influence the orientation and charge distribution around the thioester bond, thereby modulating its nucleophilic selectivity. The C4A isotype's preference for amino groups is functionally significant for immune complex clearance, as C4A can covalently bind to the antibody component of immune complexes, facilitating their solubilization and clearance by erythrocyte CR1.

### 2.4 Three-Dimensional Structure

While a high-resolution crystal structure of full-length human C4A has not been determined, the structure of C4B has been solved, and the near-identical sequence homology (99% identity) between C4A and C4B allows for reliable homology modeling. The overall fold of C4A is predicted to be highly similar to that of C4B and other thioester-containing proteins (TEPs) such as C3 and α₂-macroglobulin. The protein adopts a multi-domain architecture with a central β-sheet core surrounded by α-helical domains. The thioester bond is located in a hydrophobic pocket that shields it from premature hydrolysis.

The structural model of C4A reveals the following key features:
- The β-chain forms a globular domain that is largely composed of β-sheets.
- The α-chain contains a series of α-helical bundles and β-sheets, with the thioester domain forming a distinct subdomain.
- The γ-chain is a small, predominantly α-helical domain that protrudes from the main body of the molecule.

> **Interactive 3D Protein Visualizer: Load C4A (PDB: true)**  
> [Interactive 3D Protein Visualizer: Load C4A (PDB: true)](/tools/protein-structure-viewer?source=alphafold&accession=P0C0L4)  
> *Explore the three-dimensional architecture of C4A, including the thioester domain, isotypic residues, and proteolytic cleavage sites. The visualizer allows rotation, zoom, and residue-level inspection.*

---

## 3. Cellular Signaling Pathways & Molecular Function

### 3.1 The Complement Cascade

C4A is a central component of the classical and lectin complement pathways. The complement system is a proteolytic cascade that serves as a major effector arm of innate immunity. The activation of C4A occurs through the following sequence of events:

1. **Initiation**: The classical pathway is initiated by the binding of C1q to immune complexes (antigen-antibody complexes) or to pathogen surfaces. The lectin pathway is initiated by the binding of mannose-binding lectin (MBL) or ficolins to carbohydrate patterns on microbial surfaces.

2. **C1s/MASP-2 activation**: The binding of C1q or MBL triggers the autoactivation of associated serine proteases (C1r/C1s for the classical pathway; MASP-1/MASP-2 for the lectin pathway). Activated C1s or MASP-2 then cleaves C4 at a specific arginine residue (Arg-761 in the α-chain), releasing the C4a anaphylatoxin fragment and exposing the thioester bond in the C4b fragment.

3. **Covalent attachment**: The activated C4b fragment undergoes a conformational change that exposes the highly reactive thioester bond. C4b can then form a covalent amide bond with amino groups on the target surface (for C4A) or an ester bond with hydroxyl groups (for C4B). This covalent attachment "tags" the target for subsequent complement activation.

4. **C3 convertase formation**: Surface-bound C4b binds C2, which is then cleaved by C1s or MASP-2 to generate the classical pathway C3 convertase (C4b2a). This enzyme cleaves C3 to generate C3b, which can then form the C5 convertase (C4b2a3b), leading to the formation of the membrane attack complex (MAC) and cell lysis.

### 3.2 Immune Complex Clearance and B Cell Regulation

Beyond its role in the complement cascade, C4A has specialized functions in immune complex handling and B cell tolerance. The C4A isotype's preference for forming amide bonds with amino groups makes it particularly efficient at opsonizing immune complexes. C4A-coated immune complexes are recognized by complement receptor 1 (CR1/CD35) on erythrocytes, which transport them to the liver and spleen for phagocytic clearance. This mechanism prevents the deposition of immune complexes in tissues, which is a key pathogenic event in SLE.

Recent work by Simoni et al. (2020) demonstrated that C4A regulates autoreactive B cells in murine lupus. C4A deficiency leads to impaired clearance of apoptotic cells and immune complexes, resulting in the accumulation of self-antigens and the activation of autoreactive B cells. This study provided a mechanistic link between C4A deficiency and the breakdown of B cell tolerance that characterizes SLE.

### 3.3 Synaptic Pruning and Neurodevelopmental Signaling

A paradigm-shifting discovery in 2016 revealed that C4A is expressed in the central nervous system and plays a critical role in synaptic pruning. In the brain, C4A is localized to synapses and functions as a "eat-me" signal that tags synapses for elimination by microglia. This process is essential for normal brain development, as it allows the refinement of neural circuits during adolescence. However, excessive C4A expression leads to over-pruning of synapses, which is a hallmark of schizophrenia pathology.

The mechanism of C4A-mediated synaptic pruning involves the following steps:
1. C4A is secreted by neurons and binds to synapses that are destined for elimination.
2. C4A deposition on synapses triggers the activation of the classical complement cascade locally.
3. C3b is deposited on the tagged synapses, which are then recognized by complement receptor 3 (CR3) on microglia.
4. Microglia phagocytose the C4A/C3b-tagged synapses, leading to their elimination.

Yilmaz et al. (2020) generated a transgenic mouse model overexpressing human C4A and demonstrated that C4A overexpression leads to excessive synaptic loss and behavioral changes reminiscent of schizophrenia. This study provided causal evidence for the role of C4A in synaptic pruning and established a direct link between C4A expression levels and schizophrenia-related phenotypes.

### 3.4 Oxytocin Binding and Social Behavior

A novel and unexpected function of C4A was recently identified: C4A binds to oxytocin (OT), a neuropeptide critical for social behavior. Yamamoto et al. (2025) demonstrated that complement component C4a (the anaphylatoxin fragment released upon C4 activation) binds to oxytocin and modulates plasma oxytocin concentrations in mice. Mice deficient in the C4 gene (Slp, the mouse homolog of C4A) exhibited increased allogrooming behavior during social interactions, suggesting that C4A modulates social behavior through its interaction with oxytocin. This finding expands the functional repertoire of C4A beyond immunity and implicates it in neuroendocrine signaling and social cognition.

### 3.5 Protein-Protein Interaction Networks

C4A participates in a complex network of protein-protein interactions that extend beyond the complement cascade. Key interaction partners include:

| **Interaction Partner** | **Function** | **Reference** |
|---|---|---|
| C1s/MASP-2 | Proteolytic activation of C4A | |
| C2 | Formation of C3 convertase (C4b2a) | |
| C3 | Downstream complement activation | |
| CR1 (CD35) | Immune complex clearance | |
| C4b-binding protein (C4BP) | Regulation of C4b activity | |
| Factor I | Proteolytic inactivation of C4b | |
| Oxytocin | Neuroendocrine signaling | |
| CSMD1 | Inhibitor of C4A-mediated synaptic pruning | |

The interaction between C4A and CSMD1 (CUB and sushi multiple domains 1) is particularly relevant to schizophrenia. CSMD1 is a complement control protein that inhibits C4A activity, and its expression is deregulated in first-episode psychosis. The balance between C4A and CSMD1 expression may determine the extent of synaptic pruning and thus influence schizophrenia risk.

### 3.6 Signaling Pathway Diagram

```mermaid
flowchart TD
    A["Immune Complex / Pathogen Surface"] -->|"C1q binding"| B["C1q-C1r-C1s Complex"]
    B -->|"Autoactivation"| C["Activated C1s"]
    C -->|"Cleavage"| D["C4A Pro-protein"]
    D -->|"Release of C4a"| E["C4a Anaphylatoxin"]
    D -->|"Exposure of thioester"| F["C4b Fragment"]
    F -->|"Covalent amide bond"| G["C4b-Antigen Complex"]
    G -->|"Binding of C2"| H["C4b2a C3 Convertase"]
    H -->|"Cleavage of C3"| I["C3b Deposition"]
    I -->|"Opsonization"| J["Phagocytosis by CR3+ Cells"]
    I -->|"C5 convertase"| K["Membrane Attack Complex"]
    
    L["Neuronal C4A Expression"] -->|"Synaptic tagging"| M["C4A-labeled Synapse"]
    M -->|"C3b deposition"| N["Microglial Recognition via CR3"]
    N -->|"Phagocytosis"| O["Synaptic Pruning"]
    
    P["C4a Fragment"] -->|"Binding"| Q["Oxytocin"]
    Q -->|"Modulation"| R["Plasma OT Levels"]
    R -->|"Regulation"| S["Social Behavior"]
```

---

## 4. Pathogenic Hotspot Mutations & Clinical Differentials

### 4.1 C4A Deficiency and Null Alleles

C4A deficiency (also referred to as C4A*Q0, for "quantitative zero") is the most clinically significant genetic alteration of the C4A gene. C4A deficiency can arise through multiple molecular mechanisms:

1. **Gene deletion**: Complete deletion of the C4A gene through unequal crossover events in the RCCX module. This is the most common mechanism in European populations.
2. **Gene conversion**: Conversion of C4A to C4B through recombination, resulting in the absence of C4A protein despite the presence of a C4 gene.
3. **Point mutations**: Nonsense, missense, or frameshift mutations that abolish protein expression.
4. **Splice site mutations**: Mutations that disrupt mRNA splicing and lead to nonfunctional transcripts.

The first molecular characterization of a C4A null allele was performed by Barba, Rittner, and Schneider (1993), who identified a point mutation leading to nonexpression of C4A. This mutation was a single nucleotide change that introduced a premature stop codon in the coding sequence. Subsequently, Lokki et al. (1999) identified identical frameshift mutations in both C4A and C4B genes in a Finnish SLE patient with complete C4 deficiency. The molecular basis of complete C4A and C4B deficiencies was further elucidated by Rupert et al. (2002), who characterized a patient with homozygous C4A and C4B mutant genes.

### 4.2 C4A Copy Number Variation and Disease Risk

The association between C4A CNV and disease susceptibility has been extensively studied, with the most robust findings in autoimmune diseases:

**Systemic Lupus Erythematosus (SLE)**:
- Low C4A copy number is a strong, independent risk factor for SLE.
- Homozygous C4A deficiency (zero copies) confers the highest risk, with odds ratios ranging from 3 to 10 depending on the population.
- C4A deficiency is associated with earlier disease onset and more severe disease manifestations, including lupus nephritis.
- The risk attributable to C4A deficiency is independent of and additive to HLA-DR3 and HLA-DR2 associations.
- In East Asian populations, C4A deficiency is more common than in Europeans, and the genetic basis differs (gene conversion vs. deletion).
- C4A deficiency is associated with specific clinical phenotypes, including anti-Ro/SSA and anti-La/SSB autoantibodies.

**Juvenile Dermatomyositis (JDM)**:
- Low C4A copy number is a significant risk factor for JDM.
- C4A deficiency is associated with the presence of myositis-specific autoantibodies.
- C4A CNV influences clinical manifestations and disease severity.

**Age-Related Macular Degeneration (AMD)**:
- Multiallelic CNV in C4A is associated with late-stage AMD.
- Rare protein-coding variants of C4A can confer either risk or protection for AMD.
- The complement system's role in AMD pathogenesis is well-established, and C4A variants contribute to the genetic architecture of this disease.

**Schizophrenia**:
- Increased C4A expression, driven by structural variation in the C4 gene, is associated with elevated schizophrenia risk.
- The C4AL (long form of C4A) variant is specifically associated with increased schizophrenia risk.
- C4A copy number is positively correlated with neuropil contraction in schizophrenia patients.
- C4A expression is deregulated in first-episode psychosis and is linked to cognitive deficits.
- Sex-dependent effects of C4A CNV have been observed in treatment-resistant schizophrenia.
- C4A CNV influences serum immune protein profiles in a sex-specific manner.

**Other Autoimmune Diseases**:
- C4A gene deletion is associated with Graves' disease.
- C4A gene deletion is associated with autoimmune chronic active hepatitis.
- Low C4A copy number is a risk factor for idiopathic inflammatory myopathies.
- C4A deficiency is associated with recurrent respiratory infections in children.
- C4A CNV influences susceptibility to type 1 diabetes and may serve as a biomarker for partial disease remission.
- C4A gene deletion is associated with IgA vasculitis and giant cell arteritis.

### 4.3 Specific Pathogenic Mutations

While CNV is the dominant form of C4A genetic variation, several specific point mutations have been characterized:

| **Mutation** | **Type** | **Consequence** | **Disease Association** | **Reference** |
|---|---|---|---|---|
| c.IVS9+1G>A | Splice site | Exon 9 skipping, frameshift | C4A deficiency | |
| c.1216C>T (p.Gln406Ter) | Nonsense | Premature termination | C4A deficiency | |
| c.1930_1931insA | Frameshift | Premature termination | Complete C4 deficiency | |
| c.1101C>A (p.Asp367Glu) | Missense | Altered thioester reactivity | C4A-to-C4B conversion | |
| c.1102T>C (p.Cys368Arg) | Missense | Disrupted thioester bond | C4A deficiency | |
| c.1103G>T (p.Pro369Leu) | Missense | Altered isotypic specificity | C4A-to-C4B conversion | |

The missense mutations at codons 367–369 are particularly significant, as these residues are part of the isotypic region that determines C4A versus C4B reactivity. A single nucleotide change at codon 367 (Asp→Glu) or codon 369 (Pro→Leu) can convert C4A to C4B-like reactivity, effectively eliminating C4A function.

### 4.4 Clinical Differentials and Diagnostic Considerations

The clinical presentation of C4A deficiency overlaps with several other complement deficiencies and autoimmune conditions. Key differential diagnoses include:

1. **C4B deficiency**: C4B deficiency presents with increased susceptibility to bacterial infections, particularly encapsulated organisms, but is not as strongly associated with SLE as C4A deficiency.
2. **C2 deficiency**: The most common complement deficiency in Caucasians, presenting with SLE-like symptoms and recurrent infections. C2 deficiency is linked to the C4A*4, C4B*2 haplotype.
3. **C1q deficiency**: A rare but highly penetrant risk factor for SLE, with nearly 90% of affected individuals developing the disease.
4. **C3 deficiency**: Presents with recurrent pyogenic infections and is associated with membranoproliferative glomerulonephritis.

Diagnostic evaluation of C4A deficiency requires both protein-level and genetic testing. Serum C4 protein levels can be measured by nephelometry or ELISA, but these assays cannot distinguish between C4A and C4B. Protein allotyping by immunofixation electrophoresis can differentiate C4A from C4B based on charge differences. Genetic testing using real-time PCR or multiplex ligation-dependent probe amplification (MLPA) can determine C4A and C4B copy numbers. Long-range PCR and Southern blotting can detect gene deletions and the long/short dichotomy.

---

## 5. Host-Pathogen & Viral Interactions

### 5.1 Complement Evasion by Pathogens

C4A, as a central component of the complement system, is a target for immune evasion strategies employed by various pathogens. Several mechanisms have been described:

**Viral Complement Evasion**:
- Herpesviruses, including herpes simplex virus (HSV) and cytomegalovirus (CMV), encode complement control proteins that inhibit C4 activation. The HSV glycoprotein C (gC) binds C4b and accelerates its decay, preventing C3 convertase formation.
- Poxviruses encode complement control proteins, such as vaccinia virus complement control protein (VCP), which bind C4b and C3b and inactivate them.
- Retroviruses, including human immunodeficiency virus (HIV), incorporate host complement regulatory proteins (CD46, CD55, CD59) into their envelopes, protecting them from complement-mediated lysis.

**Bacterial Complement Evasion**:
- *Staphylococcus aureus* produces staphylococcal complement inhibitor (SCIN), which stabilizes the C3 convertase in an inactive state, preventing downstream complement activation.
- *Streptococcus pyogenes* produces the M protein, which binds C4b-binding protein (C4BP), recruiting it to the bacterial surface and inactivating C4b.
- *Neisseria meningitidis* and *Neisseria gonorrhoeae* express porins that bind C4BP, providing resistance to complement-mediated killing.

**Parasitic Complement Evasion**:
- *Trypanosoma cruzi* expresses a complement regulatory protein (CRP) that binds C4b and C3b, inactivating them.
- *Schistosoma mansoni* acquires host C4BP on its surface, protecting it from complement attack.

### 5.2 C4A and Viral Infections

The interaction between C4A and viral pathogens has been studied in several contexts. Avian pathogenic *Escherichia coli* (APEC) alters complement gene expression in chicken erythrocytes, including C4A, suggesting that bacterial pathogens can modulate C4A expression as part of their pathogenic strategy. In dairy cattle, mastitis caused by staphylococci leads to altered expression of C4A splice variants in mammary tissue.

The RCCX module contains human endogenous retrovirus (HERV) elements, and the presence of HERV-K insertions in the C4 gene is associated with the long form of C4A (C4AL). Mariaselvam et al. (2024) demonstrated that higher HERV gene insertion contributes to increased risk of SLE, with the combination of low C4A copy number and higher HERV insertion conferring the greatest risk. This finding suggests a potential interaction between endogenous retroviral elements and C4A expression in autoimmune disease pathogenesis.

### 5.3 C4A and the Gut Microbiome

The complement system, including C4A, plays a role in shaping the gut microbiome. C4B gene copy number has been shown to influence intestinal microbiota through complement activation in patients with pediatric-onset inflammatory bowel disease. While this study focused on C4B, the high homology between C4A and C4B suggests that C4A may also contribute to microbiome composition. Furthermore, complement C4 associations with altered microbial biomarkers have been demonstrated in schizophrenia, exemplifying gene-by-environment interactions. The gut-brain axis and the role of complement in mediating the effects of the microbiome on brain function represent an emerging area of research.

### 5.4 C4A and Immune Evasion in Cancer

The complement system has dual roles in cancer, promoting both tumor elimination and tumor progression. C4A expression in the tumor microenvironment can either enhance anti-tumor immunity through opsonization of tumor cells or promote tumor growth through chronic inflammation. While direct studies of C4A in cancer are limited, the complement system's role in cancer immunoediting is well-established. The C4A-mediated clearance of immune complexes may be particularly relevant in cancers that generate high levels of circulating immune complexes.

---

## 6. Pharmacogenomics, Drug Targets & Small-Molecule Inhibitors

### 6.1 Therapeutic Targeting of the Complement System

The complement system has emerged as a major therapeutic target for a range of diseases, and C4A is an attractive target for modulation in specific contexts. Several therapeutic strategies have been developed or are in development:

**FDA-Approved Complement Inhibitors**:
- **Eculizumab (Soliris)**: A humanized monoclonal antibody against C5 that blocks the terminal complement pathway. While it does not directly target C4A, it is used to treat paroxysmal nocturnal hemoglobinuria (PNH), atypical hemolytic uremic syndrome (aHUS), and myasthenia gravis.
- **Ravulizumab (Ultomiris)**: A next-generation C5 inhibitor with extended half-life, used for the same indications as eculizumab.
- **Pegcetacoplan (Empaveli)**: A C3 inhibitor that targets the central complement component, used for PNH.
- **Iptacopan (Fabhalta)**: A factor B inhibitor that blocks the alternative pathway, recently approved for PNH.

**Investigational Complement Inhibitors Targeting C4**:
- **C1 esterase inhibitors (C1-INH)**: Plasma-derived or recombinant C1-INH (Berinert, Cinryze, Ruconest) inhibits C1s and MASP-2, thereby preventing C4 cleavage. These agents are approved for hereditary angioedema and are being investigated for other complement-mediated diseases.
- **Sutimlimab (Enjaymo)**: A monoclonal antibody against C1s that blocks the classical pathway, approved for cold agglutinin disease. By inhibiting C1s, sutimlimab prevents C4 cleavage and downstream complement activation.
- **Narsoplimab**: A monoclonal antibody against MASP-2 that blocks the lectin pathway, in clinical trials for hematopoietic stem cell transplant-associated thrombotic microangiopathy.
- **ANX005**: A monoclonal antibody against C1q that blocks the classical pathway, in clinical trials for Guillain-Barré

## Related Clinical & Scientific Guides

* [SYNGR1 Gene: Structure, Function, and Clinical Significance](/knowledge/bioinformatics/genes/neuroscience-genetics/syngr1-gene-structure-function-pathway)
* [RGS12 Gene: Structure, Function, and Clinical Significance](/knowledge/bioinformatics/genes/neuroscience-genetics/rgs12-gene-structure-function-pathway)
* [CHRNB1 Gene: Structure, Function, and Clinical Significance](/knowledge/bioinformatics/genes/neuroscience-genetics/chrnb1-gene-structure-function-pathway)