HITS-CLIP: Mapping Protein-RNA Interactions

By Dr. Zubair Khalid, DVM, MS, PhD ·

HITS-CLIP: Mapping Protein-RNA Interactions

HITS-CLIP stands for high-throughput sequencing of RNA isolated by crosslinking and immunoprecipitation. It is the founding member of the CLIP family of methods, and it answers a question that no computational prediction can settle on its own: where, in a living cell, does a specific RNA-binding protein (RBP) physically touch its target RNAs? By the end of this article you will be able to design a HITS-CLIP experiment from the crosslinking step through library amplification, set up the three controls that separate signal from artifact, and read the two main classes of output data (cluster peaks and crosslink-induced mutation sites) well enough to judge whether a published map is trustworthy.

HITS-CLIP is a bench method, so it helps to know what you need before you start. You need a cell line or tissue that expresses your RBP of interest at reasonable levels, a validated antibody that immunoprecipitates that RBP under native conditions, a 254 nm UV source (a Stratalinker-style crosslinker is standard), RNase A or a comparable single-strand-specific ribonuclease, a denaturing polyacrylamide gel system, nitrocellulose or PVDF membrane, and a small-RNA or CLIP-compatible library preparation kit. You also need access to an Illumina sequencer and a bioinformatician, or at least a working knowledge of a Unix shell and a read aligner.

The method was developed to convert a transient, noncovalent biochemical contact into a covalent mark that survives every harsh step that follows. That single idea, UV light welding an amino acid side chain to a pyrimidine base, is what makes HITS-CLIP more specific than older RNA immunoprecipitation (RIP) methods, which cannot distinguish direct binding from co-association in a large ribonucleoprotein particle [1]. The trade-off is that HITS-CLIP is technically demanding, takes roughly eight days from crosslinking to RNA ready for sequencing once immunoprecipitation conditions are established, and generates data that require careful controls to interpret [1].

What HITS-CLIP Measures and Why It Matters

An RNA-binding protein does not float free in the cytoplasm. It binds a subset of transcripts, often at specific sequence or structural elements, and that binding determines whether an mRNA is spliced, exported, stabilized, degraded, or translated. Standard expression profiling tells you how much RNA is present. It does not tell you which protein was sitting on it. HITS-CLIP closes that gap by capturing the RNA that was directly bound by your protein of interest at the moment of crosslinking.

The method has been applied across a wide range of RBPs. It has been used to map CstF-64, a component of the cleavage stimulation factor, on nascent RNAs in testis [2]. It has been used extensively with Argonaute proteins to map microRNA binding sites, because Argonaute is the effector protein that holds a microRNA against its target mRNA [1][3]. In human cardiac tissue, Ago2 HITS-CLIP detected roughly 4,000 binding sites across more than 2,200 target transcripts [3]. In mouse liver across developmental stages, Ago HITS-CLIP produced a genome-wide map of microRNA-mRNA interactions [4]. In white and brown adipose tissue, AGO HITS-CLIP revealed more than 20,000 unique AGO binding sites [5]. The same approach has been adapted to less conventional organisms, including the parasitic flatworm Schistosoma japonicum, where an antibody against native Argonaute was used to pull down crosslinked Argonaute-RNA complexes from adult worm extract [6].

Two features distinguish HITS-CLIP from a simple pulldown. First, the crosslink is covalent and occurs in intact cells, so the interaction you capture reflects the in vivo state rather than a reannealed artifact of lysate preparation [1]. Second, because the crosslinked peptide leaves a physical footprint on the RNA, the recovered sequence reads mark the actual contact point, not just the general region of a transcript. That footprint is what allows the data to be refined to single-nucleotide resolution.

The HITS-CLIP Workflow Step by Step

The workflow below follows the standard HITS-CLIP protocol as described for Argonaute and conventional RBPs [1], with practical notes drawn from published optimizations [2][7]. Each step has a purpose, and skipping or rushing a step usually shows up later as a specific artifact.

flowchart TD
    A[Live cells or tissue] --> B[UV 254 nm crosslink]
    B --> C[Lyse and treat with RNase]
    C --> D[Immunoprecipitate RBP RNA complex]
    D --> E[Ligate 3 prime linker]
    E --> F[SDS PAGE and membrane transfer]
    F --> G[Proteinase K release RNA]
    G --> H[Reverse transcribe and PCR]
    H --> I[High throughput sequencing]
    I --> J[Align and call peaks]
    J --> K[Refine to crosslink sites]

Step 1: UV 254 nm Crosslinking of Live Cells

Grow your cells to the density your protocol requires, then remove medium and place the dish on ice or in a cold room. Irradiate with 254 nm ultraviolet light. The energy dose matters. Too little and you recover almost no crosslinked product. Too much and you damage the RNA and drive nonspecific crosslinks. Most protocols deliver energy in the range of a few hundred millijoules per square centimeter, and the exact setting is calibrated for the crosslinker in use. For tissue, you typically irradiate a chilled suspension or a thinly sliced preparation rather than a whole organ, because UV light penetrates poorly.

UV 254 nm preferentially crosslinks uridine to aromatic and sulfur-containing amino acid side chains, particularly phenylalanine, tyrosine, and cysteine. This is why the crosslink site in the RNA tends to sit at or near a uridine, and it is also why the crosslink is not equally efficient for every RBP. A protein whose RNA-binding surface is rich in these residues will crosslink well. A protein that grips RNA mainly through electrostatic contacts with basic residues may crosslink poorly no matter how well it binds. Crosslink efficiency is therefore a property of the protein, not just the experiment, and a low yield does not automatically mean the protein does not bind RNA.

Step 2: Lysis and Partial RNase Digestion

After crosslinking, lyse the cells under conditions that keep the RBP-RNA complex intact. A typical lysis buffer contains a nonionic detergent, a salt such as sodium chloride or potassium chloride, a buffer at neutral pH, and a protease inhibitor cocktail. Add a ribonuclease inhibitor only if your downstream steps require intact long RNA. For HITS-CLIP you will deliberately digest the RNA, so the goal is controlled fragmentation, not protection.

Add RNase A (or a comparable single-strand-specific RNase) at a low concentration and incubate briefly. The purpose is to reduce the bound RNA to short fragments, typically in the range of a few tens of nucleotides, so that the crosslinked peptide-RNA complex migrates as a discrete band on a gel. This is the step most likely to ruin an experiment. Too much RNase and you digest the RNA down to nothing, leaving no sequence to recover. Too little and the complex stays large and heterogeneous, smears on the gel, and co-migrates with contaminating ribonucleoprotein complexes. Titrate the RNase on a pilot sample before committing your main prep. The correct amount is the amount that produces a sharp, well-resolved signal at the expected molecular weight of your RBP plus a short RNA.

Step 3: Immunoprecipitation of the RBP-RNA Complex

Add your validated antibody to the clarified lysate and incubate with gentle rotation. Then capture the immune complexes on protein A or protein G beads, or on antibody-conjugated magnetic beads. Wash the beads stringently. The washes remove noncovalently associated RNA and protein, which is exactly what you want, because only the covalently crosslinked RNA survives stringent washing. This is the step that gives HITS-CLIP its specificity advantage over RIP.

The antibody is the single most important reagent. It must recognize your RBP under native, nondenaturing conditions, and it must not cross-react with other proteins. A monoclonal antibody is preferable when one is available. The CstF-64 protocol, for example, relies on a well-characterized monoclonal antibody designated 3A7 [2]. If you are working in a nonmodel organism, you may need to raise and validate a custom antibody, as was done for Schistosoma japonicum Argonaute [6].

Step 4: 3' Linker Ligation

After washing, the RNA still tethered to the beads via the crosslinked protein receives a 3' adapter. This adapter is a short synthetic oligonucleotide that provides a known sequence for reverse transcription and PCR priming. Ligation is performed while the complex is still on the beads or after elution, depending on the protocol. The efficiency of this ligation depends on the availability of a free 3' hydroxyl on the RNA fragment. Fragments that are blocked, or that are still buried in the protein, ligate poorly.

This step is also where a specific and well-documented artifact can enter. If the 3' linker or the reverse transcription primer can anneal to internal sequences in the RNA, reverse transcription can initiate at the wrong place and produce reads that are complementary to the primer rather than to the bound RNA. In published HITS-CLIP libraries, up to 45 percent of peaks were attributable to this mispriming artifact, and most libraries showed detectable levels of it [7]. A modified protocol that eliminates this artifact improves both sensitivity and library complexity [7]. If your library shows an unexpected enrichment of sequences matching your primer, mispriming is the first thing to check.

Step 5: SDS-PAGE and Membrane Transfer

Run the immunoprecipitated material on a denaturing SDS-polyacrylamide gel. The crosslinked RBP-RNA complex migrates at a position corresponding to the protein plus the attached RNA fragment, which is slightly above the position of the naked protein. Transfer the resolved proteins to a nitrocellulose or PVDF membrane. Nitrocellulose is the traditional choice because it binds protein well and has low RNA-binding background, which reduces the recovery of free RNA that was never crosslinked [2].

Visualize the RNA by autoradiography if the RNA was radiolabeled during the 3' linker ligation, or by a fluorescent or chemiluminescent method if a nonradioactive label was used. You are looking for a signal at the expected size of your RBP. Excise that region of the membrane. This gel purification step is a second specificity filter. It removes free RNA, unbound antibody, and crosslinked complexes of the wrong size.

Step 6: Proteinase K Release of RNA

Treat the excised membrane slice with proteinase K. This protease digests the crosslinked protein, releasing the RNA fragment into solution. The RNA is then extracted, typically with phenol-chloroform followed by ethanol precipitation or a column cleanup. At this point you have a pool of short RNA fragments, each derived from a site where your RBP was crosslinked to an RNA in the living cell.

Step 7: Reverse Transcription and PCR

Reverse transcribe the released RNA into cDNA using a primer that anneals to the 3' linker. Then amplify the cDNA by PCR with primers that add the sequences needed for cluster generation on the sequencer. Many labs use a commercial small-RNA library preparation kit for this step rather than building the cloning scheme from scratch, because the ligation and amplification chemistry is finicky and the kits are well optimized [2]. The number of PCR cycles must be kept as low as possible. Each additional cycle increases the chance that abundant species dominate the library and that duplicate reads accumulate.

Step 8: Sequencing and Read Alignment

Sequence the library on a high-throughput platform. The reads are short and derive from the RNA fragments that were crosslinked to your protein. Align them to the reference genome or transcriptome. Reads that map to the same genomic position in the same orientation are collapsed into a single tag to avoid counting PCR duplicates. Clusters of overlapping tags, called peaks or clusters, mark regions where your RBP bound. The next section explains how those clusters are refined.

Controls You Cannot Skip

Three controls separate a real HITS-CLIP result from a plausible-looking artifact. Run all three.

No-antibody control. Process an identical lysate through the entire workflow but omit the primary antibody. This control reveals what binds nonspecifically to the beads, the linker, or the membrane. A library that looks similar to your experimental library in the no-antibody control means your signal is background.

No-UV control. Process an identical sample through the entire workflow but do not irradiate it. This control reveals RNA that co-immunoprecipitates with your RBP without a covalent crosslink. Because stringent washing removes noncovalent interactions, the no-UV library should be much weaker than the crosslinked library. If it is not, your washes are too gentle or your antibody is pulling down a stable ribonucleoprotein particle rather than a specific RBP-RNA contact.

Input RNA library. Sequence the total RNA from the same cells, ideally using the same linker and amplification chemistry so the data are directly comparable. This is the reference for steady-state RNA abundance. Without it, you cannot tell whether a peak reflects specific binding or simply reflects the fact that the underlying transcript is highly expressed. A directional RNA-seq method that uses the same oligonucleotides and PCR steps as the CLIP library lets you analyze both datasets with the same pipeline and reduces protocol-specific bias [8]. Size-matched input controls are also used in related CLIP variants to improve peak calling and reduce false positives [9].

Reading the Output: Clusters and Crosslink Sites

HITS-CLIP data come in two layers of resolution, and understanding the difference is essential to interpreting any published map.

The first layer is the cluster or peak. This is a region of the transcript, often tens of nucleotides long, where multiple overlapping sequence tags pile up. Clusters tell you where your RBP bound at the resolution of a gel fragment. They are useful for identifying target transcripts and for comparing binding between conditions, but they do not pinpoint the exact contact.

The second layer is the crosslink-induced mutation site, abbreviated CIMS. When reverse transcriptase encounters the residual crosslinked peptide or the modified base at the crosslink position, it often stalls or misincorporates a nucleotide. These events appear in the sequencing reads as deletions or substitutions at a specific position. Because the crosslink occurs at the physical contact between protein and RNA, CIMS mark the binding site at single-nucleotide resolution [1]. CIMS analysis is what turns a regional binding map into a precise footprint. The same principle underlies the truncated cDNA approach, in which reverse transcription stops at the crosslink and the position of the truncation marks the contact.

This distinction matters because a cluster can span a region much larger than the true binding element, and two different RBPs that bind nearby can produce overlapping clusters. CIMS resolves that ambiguity. If you are comparing your data to a motif or to a structural model, use CIMS positions rather than cluster boundaries.

Common Mistakes and Limitations

HITS-CLIP is unforgiving, and most failures trace back to one of a small number of steps. The table below summarizes the major pitfalls, how they present, and how to address them.

PitfallHow it presentsWhat to do
Low crosslink efficiencyVery few reads, weak or absent signal on the gelConfirm the RBP crosslinks well at 254 nm, optimize UV dose, increase starting material, consider a PAR-CLIP variant that uses photoactivatable ribonucleosides [10][11]
RNase over- or under-digestionSmear on the gel, or no recoverable RNATitrate RNase on a pilot sample, aim for a sharp band at the expected complex size
Poor antibody specificitySignal in the no-antibody control, wrong-size bandValidate the antibody by western blot and immunoprecipitation, use a monoclonal when possible [2]
PCR duplicationMany identical reads, low library complexityReduce PCR cycles, collapse duplicate reads during alignment
Mispriming during reverse transcriptionEnrichment of sequences complementary to the primer, artifactual peaksUse the modified protocol that removes the mispriming artifact [7]
Genome alignment ambiguityReads mapping to repetitive regions or multiple lociUse a splice-aware aligner, filter multimapping reads, inspect problematic loci manually
Confusing expression with bindingPeaks that track transcript abundanceInclude an input RNA library and normalize against it [8]

A few limitations are inherent to the method rather than fixable by technique. Crosslink efficiency is protein-dependent, so some RBPs will always be harder to map than others. UV 254 nm crosslinking favors uridine contacts, which can bias the recovered sites toward uridine-rich binding elements. The method captures a snapshot, so it reports where the protein was bound at the moment of irradiation, not the full dynamic range of its behavior. And because the library is amplified, abundant binding sites are easier to detect than rare ones, so absence of a peak is weaker evidence than presence of one.

For RBPs that resist 254 nm crosslinking, related methods offer alternatives. PAR-CLIP incorporates photoactivatable ribonucleosides such as 4-thiouridine into nascent RNA, which crosslinks more efficiently at 365 nm and produces a characteristic nucleotide transition that marks the crosslink site [10][11]. iCLIP and its variants improve single-nucleotide resolution by capturing truncated cDNAs directly [12]. Newer approaches use RNA deaminases fused to the RBP of interest to profile binding sites from low-input samples with simpler procedures [13]. These are complementary tools, not replacements, and the choice depends on your protein, your starting material, and the resolution you need.

One more practical limitation is starting material. Because the RNA yield after crosslinking, immunoprecipitation, gel electrophoresis, membrane transfer, and extraction is small, HITS-CLIP is usually performed from relatively large amounts of lysate or tissue homogenate [14]. If your biological question concerns a small subcellular compartment, such as neuronal synapses, you may need to isolate that compartment first and then scale the protocol accordingly [14].

Frequently Asked Questions

What does HITS-CLIP actually measure?

HITS-CLIP measures the sites where a specific RNA-binding protein was in direct physical contact with RNA in living cells. UV light covalently links the protein to the RNA at the contact point, and sequencing the recovered RNA fragments maps those contacts across the transcriptome [1].

How is HITS-CLIP different from RIP-seq?

RIP-seq captures RNA that co-immunoprecipitates with a protein, which can include indirect interactions in a larger ribonucleoprotein complex. HITS-CLIP adds a covalent UV crosslink and stringent washes, so only RNA directly bound by the protein survives, giving higher accuracy and resolution [1].

Why is UV crosslinking done on live cells?

Crosslinking in intact cells freezes the in vivo interaction before lysis, so the captured contacts reflect the protein's real binding state rather than reassociation that can occur in a lysate. It also makes the interaction covalent, which allows stringent washing without losing the signal [1].

What is the difference between a cluster and a CIMS?

A cluster is a region where multiple sequence tags pile up and marks where the protein bound at fragment resolution. A crosslink-induced mutation site (CIMS) is a specific nucleotide position where reverse transcription stalled or misincorporated, marking the contact at single-nucleotide resolution [1].

Do I need a no-UV control?

Yes. The no-UV control shows how much RNA co-immunoprecipitates without a covalent crosslink. It should be much weaker than the crosslinked sample. If it is not, your washes are too gentle or your antibody is pulling down a stable complex rather than a specific contact.

What causes mispriming artifacts in HITS-CLIP libraries?

Mispriming happens when the reverse transcription primer anneals to internal RNA sequences instead of the ligated linker, producing reads complementary to the primer. This artifact accounted for up to 45 percent of peaks in publicly available libraries, and a modified protocol eliminates it [7].

Can HITS-CLIP be used for microRNA targets?

Yes. Argonaute HITS-CLIP is one of the main ways to map microRNA binding sites, because Argonaute holds the microRNA against its target mRNA. It has been used to build transcriptome-wide microRNA target maps in tissues including heart, liver, and adipose tissue [3][4][5].

What if my protein crosslinks poorly at 254 nm?

Consider a variant that uses photoactivatable ribonucleosides, such as PAR-CLIP, which crosslinks at 365 nm and often gives higher yield for proteins that resist 254 nm crosslinking [10][11]. RNA deaminase-based methods are another option for low-input samples [13].

Related Articles

Sources

  1. Mapping Argonaute and conventional RNA-binding protein interactions with RNA at single-nucleotide resolution using HITS-CLIP and CIMS analysis.
  2. High-throughput sequencing of RNA isolated by cross-linking and immunoprecipitation (HITS-CLIP) to determine sites of binding of CstF-64 on nascent RNAs.
  3. Elucidation of transcriptome-wide microRNA binding sites in human cardiac tissues by Ago2 HITS-CLIP.
  4. Genome-wide identification of microRNA targets reveals positive regulation of the Hippo pathway by miR-122 during liver development.
  5. AGO HITS-CLIP reveals distinct miRNA regulation of white and brown adipose tissue identity.
  6. High-throughput sequencing of RNAs isolated by cross-linking immunoprecipitation (HITS-CLIP) reveals Argonaute-associated microRNAs and targets in Schistosoma japonicum.
  7. Improvements to the HITS-CLIP protocol eliminate widespread mispriming artifacts.
  8. Solid-Support Directional (SSD) RNA-Seq as a Companion Method to CLIP-Seq.
  9. Improved discovery of RNA-binding protein binding sites in eCLIP data using DEWSeq.
  10. Protocol for detecting RBM33-binding sites in HEK293T cells using PAR-CLIP-seq.
  11. Transcriptome-wide identification of in vivo interactions between RNAs and RNA-binding proteins by RIP and PAR-CLIP assays.
  12. Transcriptome-Wide Characterization of Protein:RNA Interactions with UV-Crosslinking Immunoprecipitation and Sequencing.
  13. Identification of RBP binding sites using RNA deaminases.
  14. Identification of RNA-RBP Interactions in Subcellular Compartments by CLIP-Seq.