Claude Discovers a CRISPR-Like Enzyme System in Phage DNA: The History and Future of Computational Discovery in Biology

By Dr. Zubair Khalid, DVM, MS, PhD ·

Claude Discovers a CRISPR-Like Enzyme System in Phage DNA: The History and Future of Computational Discovery in Biology

Key Takeaways

  • On 23 September 2026, Anthropic reported that autonomous Claude agents found a previously unknown family of reverse transcriptases in bacteriophage genomes, named ART for array-associated reverse transcriptase [1][2][3].
  • The ART enzyme system sits beside a long array of repeating DNA that looks somewhat like a CRISPR array, but its function is not yet known [1][2].
  • The agents surveyed reverse transcriptase loci across 1.9 billion protein clusters, ran 119 tasks and 949 agent sessions, and recovered about 200,000 reverse transcriptase clusters [2].
  • CRISPR itself was found in sequence data long before it was understood: an unexplained repeat in E. coli in 1987, a named locus in 2002, experimental proof of immunity in 2007, and an approved therapy in 2023 [4][6][11][14].
  • What is new here is not the mining of sequence data, which has a fifty-year history, but that the agents chose their own searches, read raw DNA, and noticed the anomaly themselves [2].
  • The discovery has not been shown to be programmable or useful, and the array was missed in ten reruns of the same campaign, so the result needs the same experimental scrutiny as any other claim [2].

On 23 September 2026, Anthropic announced that autonomous Claude agents had found a family of enzymes hiding in bacteriophage DNA. The enzymes are reverse transcriptases, the class of protein that copies RNA back into DNA. What made the find unusual was what sat next to the enzyme gene: a long array of evenly spaced repeating DNA, arranged in a pattern that looks somewhat like the CRISPR arrays that bacteria use to store memories of past infections. Anthropic named the family ART, for array-associated reverse transcriptase [1][2][3].

The function of ART is unknown. That is not a footnote. It is the central fact of the story. The preprint states plainly that whether the reverse transcriptase and its partner protein interact, and what the system does for the phage, are currently unknown [2]. No one has shown that the enzyme is active, that the short RNAs made from the array are its substrates, or that the system can be programmed to do anything at all. What the agents found was a pattern in sequence data, a pattern that resembles systems known to manipulate DNA, and a set of questions that now have to be answered at the bench.

Feng Zhang, the MIT and Broad Institute scientist who helped turn CRISPR into a genome editing tool, reviewed the work. "This is an exciting example of how AI agents can contribute to biological discovery," he said. "The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation." [3]

The reason this announcement landed with weight is that biology has been here before. The most powerful tool in modern genetics began as an odd repeated sequence in a bacterial genome that nobody could explain. It took 36 years to go from that unexplained repeat to an approved medicine [4][14]. If ART follows anything like that path, the interesting work has barely started.

What Claude Found in Phage DNA

The ART system has three parts. There is a reverse transcriptase gene. There is a dedicated partner gene next to it. And there is a long array of non-coding DNA repeats, evenly spaced, sitting alongside the enzyme [2].

The array is the part that catches the eye. The repeats run 15 to 49 nucleotides long, each with a palindromic core of about 15 nucleotides and less conserved edges. Repeats are shared within an ART clade but not across clades. The spacers between repeats run 120 to 220 nucleotides long [2]. That spacer length is notable: the spacers are much longer than typical CRISPR spacers, so ART is not simply a CRISPR variant with a different name.

The reverse transcriptase itself keeps the catalytic YxDD motif that defines the family, but it carries an unusually long N-terminus, with about 180 residues preceding the polymerase domain. Most reverse transcriptases have about 50 or fewer [2].

The partner genes are stranger still. Several unrelated partner protein types appear in different ART lineages, with no detectable similarity to each other. That pattern suggests the reverse transcriptase has been paired with different partners more than once over evolutionary time [2].

ART shows up mainly in jumbo phages, the very large bacteriophages, including cultured phages and phage contigs predicted from metagenomes [2]. These are viruses with big genomes, which means a lot of DNA to read, and the repeats sit in non-coding DNA that workflows built around predicted genes can easily pass over.

There is lab evidence that the array is not silent. In published RNA-sequencing data from Staphylococcus phage SA1 infection, at 15 minutes post infection, the array-derived RNAs account for up to 8 percent of phage RNAs, making them among the most abundant phage transcripts [2]. The Anthropic team also expressed the SA1 ART system on plasmids in Escherichia coli and performed small-RNA sequencing. The array was expressed as multiple distinct short RNAs with reproducible boundaries [1][2]. CRISPR arrays are also processed into short RNAs, which is part of why the resemblance is interesting.

That is the full extent of what has been shown. The system is transcribed. Its function is not known.

How the AI Agents Searched 1.9 Billion Protein Clusters

The scale of the search is the part that separates this campaign from a graduate student with a laptop and a hypothesis.

The agents surveyed reverse transcriptase loci across 1.9 billion protein clusters. They built their own profile hidden Markov models and reference sets, and recovered about 200,000 reverse transcriptase clusters [2]. They scored 3,564 candidate partner protein families and filed 19 campaign reports for human scientists [2].

Funnel from 1.9 billion protein clusters to about 200,000 reverse transcriptase clusters, 3,564 partner families, 19 reports and the ART system

The preprint describes the effort in precise terms: "The full campaign comprised 119 tasks and 949 agent sessions, amounting to 77 agent-hours and 215.6 million tokens over 21.5 hours of wall-clock time." [2]

The agents ran with the Claude Mythos 5 model inside Claude Code instances [2]. They used standard bioinformatics software, the same tools a human computational biologist would reach for: HMMER, MMseqs2, BLAST, MAFFT, FastTree, geNomad, Infernal, ViennaRNA, Foldseek, TM-align, AlphaFold2 and ESMFold for structure prediction, and Prodigal-gv for gene prediction. Their data sources included the Logan assemblies of public sequencing data, NCBI, UniProt, the AlphaFold database, the Protein Data Bank and Rfam [2].

None of that tooling is new. BLAST has been around since 1990 [16]. Profile hidden Markov models are decades old. What is new is who was driving. In earlier genome mining campaigns, a human designed the pipeline around a hypothesis, for example "look near cas genes" or "look in defense islands," and the pipeline produced a candidate list that experts then read. In the ART campaign, the agents chose the searches, built the profiles, read the loci and wrote the reports themselves [2].

The key observation came from reading raw sequence. The agents noticed repeats in non-coding DNA next to a reverse transcriptase gene. The preprint reports that the discovery depended on the model recognizing DNA repeats directly [2]. That is a different kind of task from ranking candidates by a score. It requires looking at the actual letters of the genome and noticing that something is off.

The reproducibility limit matters here. When the same campaign was run ten more times, agents sampled ART loci, and in two campaigns investigated the lineage. "However, none read the DNA upstream of the RTs, and the array was missed in every rerun." [2] The authors attribute this to the large search space and the non-deterministic behavior of the harness. A single successful run is not the same as a reliable method, and the preprint says so.

Why a Repeat Array Next to an Enzyme Matters

A repeat array beside an enzyme is not automatically interesting. Biology is full of repetitive DNA with no function worth writing home about. What makes this pattern worth a second look is the company it keeps.

CRISPR arrays are the obvious comparison. A CRISPR locus consists of short repeats separated by spacers, with cas genes nearby. The spacers are derived from phage and plasmid DNA, and they function as a molecular memory of past infections. The array is transcribed and processed into short guide RNAs that direct a nuclease to matching sequences [11][12]. That architecture, repeats plus spacers plus an adjacent enzyme, is the signature of CRISPR's programmable design.

Side-by-side comparison of a CRISPR-Cas locus, where guide RNAs direct a nuclease, and an ART locus with a reverse transcriptase, partner gene and repeat array whose function is unknown

ART has the repeats and the adjacent enzyme. It does not have the same spacer length, and it does not have cas genes. The preprint compares ART with known reverse-transcriptase systems, including retrons, diversity-generating retroelements, group II introns and CRISPR-associated reverse transcriptases, and concludes that ART does not fit any of them [2].

Reverse transcriptases are worth understanding on their own terms. They copy RNA into DNA, the reverse of the usual direction of transcription. Their discovery in retroviruses earned David Baltimore and Howard Temin, along with Renato Dulbecco, the 1975 Nobel Prize in Physiology or Medicine [37]. They are already central to genome editing: prime editing fuses a Cas9 nickase to an engineered reverse transcriptase so that a guide RNA can template a precise edit without double-strand breaks [33]. Bacterial retrons, which are reverse-transcriptase-based elements, function in anti-phage defense [23].

So the ingredients of ART are all familiar. A reverse transcriptase can write DNA from an RNA template. A repeat array can be processed into short RNAs. A partner gene can provide a second activity. What is missing is the wiring diagram. No one knows whether the ART reverse transcriptase uses the array-derived RNAs as templates, whether the partner protein binds the array, or whether the whole system does anything for the phage beyond being transcribed.

Anthropic noted that programmable molecular systems from microbes have repeatedly become core tools: restriction enzymes, Taq polymerase and CRISPR [1]. Restriction enzymes earned Werner Arber, Daniel Nathans and Hamilton Smith the 1978 Nobel Prize in Physiology or Medicine [35]. Taq polymerase was purified from the hot-spring bacterium Thermus aquaticus by Chien and colleagues in 1976 [34], and PCR earned Kary Mullis a share of the 1993 Nobel Prize in Chemistry [36]. The pattern is real. It is also retrospective. Most repetitive DNA in phage genomes has never become a tool, and there is no way to know in advance which oddity will matter.

A Short History of Finding Biology in Sequence Data

The idea that you can find biology by comparing sequences is older than most of the databases now used to do it.

In 1970, Needleman and Wunsch published a general dynamic-programming method for aligning two protein sequences [15]. That algorithm, and the faster variants that followed, made it possible to ask whether two sequences are related by descent. A pairwise alignment tool is the direct descendant of that work.

In 1990, BLAST made fast similarity search of large databases routine [16]. Before BLAST, comparing a new sequence against everything known was a project. After BLAST, it was a few minutes of compute. The change was technical, and it also changed what questions were worth asking.

In 1995, the first complete genome of a free-living organism, Haemophilus influenzae Rd, about 1.8 million base pairs, was sequenced and assembled by whole-genome shotgun sequencing [17]. A complete bacterial genome became a thing you could hold in a file and search.

In 2004, shotgun sequencing of Sargasso Sea water reported more than 1.2 million previously unknown genes [18]. That was an early view of how much biology sits in uncultured microbes, organisms that no one had grown in a lab and might never grow. Metagenomics turned the ocean, the soil and the gut into sequence libraries.

In 2024, the Logan project assembled the public Sequence Read Archive, 27.3 million datasets or about 5 x 10^16 base pairs, into contigs and compressed the data more than 100-fold for petabase-scale search [32]. Logan assemblies were among the data sources the ART agents drew on [2].

Each of these steps expanded the space of sequences that could be searched. Each one also created a new problem: more data than any human could read. The history of computational biology is partly a history of building tools to manage that surplus, and partly a history of the discoveries that fall out when someone finally looks.

How CRISPR Was Discovered: Computation First, Then the Bench

The CRISPR story is the clearest example of computation leading, with experiments following years later. It is also the story most often told badly, as if a single lab had a single insight. The record shows a slow accumulation of sequence observations, each one adding a piece that the next group could use.

Timeline of CRISPR from odd repeats in E. coli in 1987 to the first CRISPR therapy in 2023

In 1987, Ishino and colleagues reported an unusual set of repeated sequences next to the iap gene in E. coli while sequencing that gene. Its meaning was unknown [4]. That is the founding observation, and it was a side note in a paper about something else.

In 2000, Mojica and colleagues recognized a family of regularly spaced repeats across many archaea and bacteria [5]. The pattern was not a one-off. It was widespread.

In 2002, Jansen and colleagues named the loci CRISPR, for clustered regularly interspaced short palindromic repeats, and identified cas genes that sit beside them [6]. Naming matters. It gives a field something to search for.

In 2005, three groups, Mojica; Pourcel; and Bolotin, compared spacer sequences with databases and found that spacers match phage and plasmid DNA, implying an immune function [7][8][9]. This is the computational turn. The spacers looked like fragments of viruses, which suggested the array was a record of past infections.

In 2006, Makarova, Koonin and colleagues used computational analysis to predict that CRISPR-Cas is an RNA-interference-like immune system in prokaryotes [10]. The prediction was specific enough to test.

In 2007, Barrangou and colleagues proved experimentally that Streptococcus thermophilus gains phage resistance by adding phage-derived spacers [11]. The computational prediction held up at the bench.

In 2012, Jinek and colleagues showed that Cas9 is a programmable dual-RNA-guided DNA endonuclease [12]. That is the mechanistic insight that made editing possible.

In 2020, Emmanuelle Charpentier and Jennifer Doudna received the Nobel Prize in Chemistry for developing CRISPR-Cas9 genome editing [13]. In December 2023, the US FDA approved Casgevy, the first CRISPR-based therapy, for sickle cell disease in patients 12 and older [14].

Roughly 36 years separate an unexplained repeat in a sequence file from an approved medicine [4][14]. The computational work did not do the experiments. It told the experimentalists where to look.

YearMilestoneWhat computation contributedRef
1970Needleman-Wunsch alignmentDynamic programming for comparing two protein sequences[15]
1987Odd repeats beside iap in E. coliSequence reading surfaced an unexplained pattern[4]
1990BLASTFast similarity search of large databases[16]
1995First free-living genome sequencedWhole-genome shotgun assembly of H. influenzae Rd[17]
2000Repeats found across archaea and bacteriaComparative sequence survey revealed a family[5]
2002Locus named CRISPR, cas genes identifiedGene finding next to the array[6]
2004Sargasso Sea shotgun sequencingMetagenomics exposed uncultured microbial diversity[18]
2005Spacers match phage and plasmid DNADatabase comparison implied immune function[7][8][9]
2006CRISPR-Cas predicted as RNAi-like immunityComputational analysis generated a testable model[10]
2007Phage resistance proven in S. thermophilusPrediction confirmed at the bench[11]
2012Cas9 shown to be programmableMechanism work enabled editing[12]
2020Nobel Prize in ChemistryRecognition of CRISPR-Cas9 editing[13]
2023First CRISPR therapy approvedClinical translation of a sequence observation[14]
2024Logan assembles public SRAPetabase-scale search infrastructure[32]
2026Claude agents find ARTAutonomous survey and anomaly detection[1][2]

The Genome Mining Era: Cas12a, Cas13, Defense Islands, OMEGA and Bridge RNAs

Once CRISPR was understood, the search for more systems became systematic. The mining era is the direct ancestor of the ART campaign, and it shows how computational discovery matured from reading a single locus to surveying entire biospheres.

In 2015, Zetsche and colleagues characterized Cpf1, now Cas12a, as a single RNA-guided endonuclease [19]. In the same year, Shmakov and colleagues used computational searches to find and characterize new class 2 CRISPR-Cas systems [20]. The search was no longer opportunistic. It was a pipeline.

In 2016, C2c2, now Cas13, was shown to be a programmable RNA-targeting effector [21]. The toolkit expanded from DNA to RNA.

In 2018, Doron and colleagues used the clustering of immune genes in "defense islands" to predict and then experimentally validate previously unknown antiphage defense systems [22]. This is the key methodological move: instead of searching for a specific gene, search for neighborhoods where defense genes cluster, and see what else lives there.

In 2020, retrons were shown to function in anti-phage defense [23]. That work tied a reverse-transcriptase element directly to phage defense.

In 2021, Altae-Tran and colleagues showed that IscB, IsrB and TnpB from the IS200/IS605 transposon family are programmable RNA-guided DNA nucleases, the OMEGA systems, and IscB worked for editing in human cells [24]. The search had moved beyond CRISPR proper into the transposons that may have given rise to it.

In 2023, Fanzor was shown to be a eukaryotic programmable RNA-guided endonuclease [25]. In the same year, the FLSHclust algorithm clustered massive sequence sets and identified 188 previously unreported CRISPR-linked gene modules [26]. The scale of the search kept growing.

In 2024, Durrant and colleagues showed that IS110 elements make a bridge RNA that programs a recombinase for insertion, excision and inversion of DNA [27]. Matthew Durrant is also an author on the ART preprint [2]. The bridge RNA work is a reminder that the most interesting discoveries often come from systems that do not fit existing categories.

The pattern across the mining era is consistent. Computational searches generate candidates. Experimentalists characterize them. Some turn out to be tools. Most do not. The ART campaign is the latest entry in that sequence, with one difference: the searching was done by agents, not by a human-designed pipeline.

Protein Structure at Planetary Scale

Sequence alone does not tell you what a protein does. Structure helps, and structure prediction has become another form of computational discovery.

In 2021, AlphaFold2 predicted protein structures with accuracy often competitive with experiment [28]. The AlphaFold Protein Structure Database now holds predictions for more than 200 million proteins [29]. In 2023, ESMFold used a protein language model to predict structures up to about 60 times faster, enabling structure prediction for hundreds of millions of metagenomic proteins [30]. In 2024, the Nobel Prize in Chemistry went to David Baker for computational protein design and to Demis Hassabis and John Jumper for protein structure prediction [31].

In the ART work, agents used predicted structures to compare the reverse transcriptase with known reverse transcriptase families [2]. That is a standard use of structure prediction, and it helped place ART relative to known reverse-transcriptase families. The comparison showed that ART does not fit any known family cleanly.

Structure prediction does not answer the functional question. A predicted fold can tell you that a protein belongs to a family. It cannot tell you what the protein does in a phage, what its partner does, or whether the array-derived RNAs are its substrates. Those questions still require experiments.

What Changes When AI Agents Do the Searching

The ART campaign differs from earlier mining work in three ways, and it is worth being precise about each.

First, the agents chose their own searches. Previous pipelines were designed by humans around a hypothesis. The agents built their own profile hidden Markov models and reference sets, and decided what to look for next based on what they found [2]. That is a meaningful shift in who sets the agenda.

Second, the agents read primary data. The key observation, repeats in non-coding DNA next to a reverse transcriptase, came from reading raw sequence, not from a precomputed annotation [2]. The preprint reports that the discovery depended on the model recognizing DNA repeats directly [2]. Annotations are summaries, and summaries lose things. Repeats in non-coding DNA are easy to miss in workflows built around predicted genes.

Third, the agents wrote reports for human scientists. They filed 19 campaign reports [2]. The human role shifted from reading candidate lists to reading synthesized findings.

The authors' framing is that agents which read primary data directly can supply the judgment that limits discovery wherever data have outgrown the attention of human scientists, and that noticing an anomaly, which only exists against an expectation, is the start of discovery [2]. Human scientists did the lab work. The agents did the survey and triage.

That division of labor is the honest description. The agents did not run a gel, purify a protein, or test a phage. They searched, noticed and reported. The experiments that would establish what ART does have not been done.

What We Still Do Not Know About ART

The list of unknowns is longer than the list of knowns, and the preprint is direct about it.

"Whether the RT and its partner interact, and what the system does for the phage, are currently unknown." [2]

It is also not yet shown that the reverse transcriptase is active or that the short RNAs are its substrates [2]. The array is transcribed, and the transcripts have reproducible boundaries, but transcription is not function. Plenty of non-coding RNAs are made and do nothing useful.

ART does not fit any known reverse-transcriptase system, including retrons, diversity-generating retroelements, group II introns and CRISPR-associated reverse transcriptases [2]. That is a statement about classification, not about mechanism. Not fitting a category means the category does not describe it. It does not mean ART is novel in a way that makes it useful.

The reproducibility limit is a real caveat. In ten reruns of the same campaign, the array was missed every time [2]. The authors attribute this to the large search space and the non-deterministic behavior of the harness. A discovery that depends on a single successful run needs independent confirmation, and the authors report the reruns openly.

What can be said is that ART is a real pattern in real sequence data, that it is transcribed in at least one phage infection, and that it resembles systems known to manipulate DNA. Anthropic noted that only a handful of known systems share its features, and those are able to cut, copy and paste DNA [1]. That is a statement about the comparison set, not about ART. Whether ART does any of those things is unknown.

The Future of Computational Discovery in Biology

The bottleneck is shifting. Finding candidates is getting cheaper. Testing them is not.

Discovery loop from public sequence data to AI agents flagging anomalies, scientists designing tests, wet-lab experiments and new tools and medicines

An agent campaign can survey 1.9 billion protein clusters in 21.5 hours of wall-clock time [2]. A single wet-lab experiment to test whether a reverse transcriptase is active can take weeks. The asymmetry means that the value of computational discovery depends on how well it prioritizes. A list of 200,000 reverse transcriptase clusters is not useful. A short list of systems that look like they might do something specific is.

Negative results and reruns matter more in this world, not less. An agent campaign is not deterministic. ART was missed in ten reruns [2]. If a finding cannot be reproduced by the same method, it needs to be reproduced by a different one. That is normal science, and it applies to AI-led discovery as much as to any other kind.

Claims need the same experimental standard as any discovery. ART's function and programmability are not established [2]. The fact that an AI found it does not lower the bar. If anything, it raises the bar, because the finding is unusual and the method is new.

Public data resources make agent surveys possible. Logan assembled the public Sequence Read Archive into searchable contigs [32]. The AlphaFold database holds predictions for more than 200 million proteins [29]. Without these, an agent campaign would have nothing to search. The infrastructure that supports computational biology is as important as the models that use it.

The CRISPR history is the best guide to what comes next. Turning a sequence observation into a tool took years of mechanism work and experiments [4][11][12][14]. The observation in 1987 was not a tool. The tool came in 2012, after the mechanism was understood. If ART turns out to be something, the timeline will look similar: a pattern, then a function, then a mechanism, then maybe an application. The pattern is the easy part.

Common Misconceptions

  • Misconception: Claude discovered a new CRISPR system. Correction: ART is a reverse transcriptase with a repeat array that looks somewhat like a CRISPR array, but it does not fit any known reverse-transcriptase system, including CRISPR-associated reverse transcriptases, and its function is unknown [2].
  • Misconception: ART is programmable and can edit DNA. Correction: No one has shown that ART is programmable, active, or useful. Anthropic noted that a handful of known systems with similar features can cut, copy and paste DNA, but that is a statement about those systems, not about ART [1][2].
  • Misconception: The AI did the science. Correction: The agents did the survey and triage. Human scientists did the lab work, including the small-RNA sequencing of the SA1 ART system expressed in E. coli [1][2].
  • Misconception: The discovery is reproducible because it was published. Correction: The array was missed in ten reruns of the same campaign, which the authors attribute to the large search space and non-deterministic behavior of the harness [2].
  • Misconception: CRISPR was discovered by Doudna and Charpentier. Correction: Charpentier and Doudna received the 2020 Nobel Prize for developing CRISPR-Cas9 genome editing, but the locus was named in 2002 by Jansen and colleagues, and the immune function was proven in 2007 by Barrangou and colleagues [6][11][13].
  • Misconception: Reverse transcriptases only matter for retroviruses. Correction: Reverse transcriptases are central to prime editing, where a Cas9 nickase is fused to an engineered reverse transcriptase, and bacterial retrons function in anti-phage defense [23][33].

Limitations

The limitations of the ART discovery are specific and worth stating plainly. The function of the system is unknown [2]. The reverse transcriptase has not been shown to be active [2]. The short RNAs have not been shown to be its substrates [2]. The interaction between the reverse transcriptase and its partner protein has not been demonstrated [2]. The finding was not reproduced in ten reruns of the same campaign [2]. Apart from RNA sequencing of one phage infection and small-RNA sequencing in E. coli, the evidence so far is computational [2].

The limitations of AI-led discovery in general are different. Agents can survey data at a scale no human can match, but they cannot run experiments, and they cannot yet be trusted to be reliable across runs. The ART campaign succeeded once and failed ten times. That is a signal about the method as well as about the result. Agent campaigns also depend on the quality of public data, which is uneven. And they depend on the judgment of the humans who read the reports, because an anomaly is only interesting if someone recognizes it as one.

There is also a structural limitation. The preprint argues that agents can supply the judgment that limits discovery wherever data have outgrown the attention of human scientists [2]. That is a claim about attention, not about understanding. Agents can notice more. In this case they have not explained what they noticed. The explanation still comes from experiments and from people who know what questions to ask.

Frequently Asked Questions

What is the ART enzyme system?

ART stands for array-associated reverse transcriptase. It is a family of reverse transcriptases found in bacteriophage genomes, each sitting next to a dedicated partner gene and a long array of evenly spaced non-coding DNA repeats. The system was reported by Anthropic on 23 September 2026 after autonomous Claude agents surveyed reverse transcriptase loci across 1.9 billion protein clusters [1][2].

Did Claude discover a new CRISPR?

No. ART has a repeat array that looks somewhat like a CRISPR array, and the array is processed into short RNAs, which is also true of CRISPR. But ART does not fit any known reverse-transcriptase system, including CRISPR-associated reverse transcriptases, and its function is unknown [2]. The resemblance is architectural, not functional.

Who discovered CRISPR?

No single person did. Ishino and colleagues reported the odd repeats in E. coli in 1987. Mojica and colleagues recognized the family in 2000. Jansen and colleagues named the locus CRISPR in 2002. Barrangou and colleagues proved the immune function in 2007. Jinek and colleagues showed Cas9 was programmable in 2012. Charpentier and Doudna received the 2020 Nobel Prize in Chemistry for developing CRISPR-Cas9 genome editing [4][5][6][11][12][13].

What does a reverse transcriptase do?

A reverse transcriptase copies RNA into DNA, the reverse of the usual direction of transcription. The enzyme was discovered in retroviruses, which earned David Baltimore and Howard Temin, along with Renato Dulbecco, the 1975 Nobel Prize in Physiology or Medicine [37]. Reverse transcriptases are used in prime editing and are the core of bacterial retrons, which function in anti-phage defense [23][33].

Why was ART easy to miss?

The repeats sit in non-coding DNA next to the reverse transcriptase gene, and noticing them meant reading the raw sequence. The preprint reports that the discovery depended on the model recognizing DNA repeats directly [2]. In ten reruns of the same campaign, no agent read the DNA upstream of the reverse transcriptases, and the array was missed every time [2].

Is ART useful for medicine or biotechnology?

Unknown. No one has shown that ART is active, programmable, or capable of manipulating DNA. Anthropic noted that a handful of known systems with similar features can cut, copy and paste DNA, but that is a statement about those systems, not about ART [1][2]. Any application would require years of mechanism work and experiments, as the CRISPR history shows [4][11][12][14].

References

  1. Anthropic. Claude discovers a novel enzyme system (announcement, 23 September 2026)
  2. Yoon PH, Athukoralage JS, Ameisen E, et al. Autonomous AI agents discover reverse transcriptases with tandem repeat arrays. Anthropic preprint, 2026
  3. Tech Times. Claude finds hidden enzyme system in viral DNA; CRISPR pioneer calls it intriguing (24 September 2026)
  4. Ishino Y, Shinagawa H, Makino K, et al. Nucleotide sequence of the iap gene in Escherichia coli. Journal of Bacteriology, 1987
  5. Mojica FJM, Diez-Villasenor C, Soria E, Juez G. Biological significance of a family of regularly spaced repeats in the genomes of Archaea, Bacteria and mitochondria. Molecular Microbiology, 2000
  6. Jansen R, van Embden JDA, Gaastra W, Schouls LM. Identification of genes that are associated with DNA repeats in prokaryotes. Molecular Microbiology, 2002
  7. Mojica FJM, Diez-Villasenor C, Garcia-Martinez J, Soria E. Intervening sequences of regularly spaced prokaryotic repeats derive from foreign genetic elements. Journal of Molecular Evolution, 2005
  8. Pourcel C, Salvignol G, Vergnaud G. CRISPR elements in Yersinia pestis acquire new repeats by preferential uptake of bacteriophage DNA. Microbiology, 2005
  9. Bolotin A, Quinquis B, Sorokin A, Ehrlich SD. CRISPRs have spacers of extrachromosomal origin. Microbiology, 2005
  10. Makarova KS, Grishin NV, Shabalina SA, Wolf YI, Koonin EV. A putative RNA-interference-based immune system in prokaryotes. Biology Direct, 2006
  11. Barrangou R, Fremaux C, Deveau H, et al. CRISPR provides acquired resistance against viruses in prokaryotes. Science, 2007
  12. Jinek M, Chylinski K, Fonfara I, et al. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, 2012
  13. The Nobel Prize in Chemistry 2020, press release
  14. US FDA. FDA approves first gene therapies to treat patients with sickle cell disease (December 2023)
  15. Needleman SB, Wunsch CD. A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journal of Molecular Biology, 197090057-4)
  16. Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. Journal of Molecular Biology, 199080360-2)
  17. Fleischmann RD, Adams MD, White O, et al. Whole-genome random sequencing and assembly of Haemophilus influenzae Rd. Science, 1995
  18. Venter JC, Remington K, Heidelberg JF, et al. Environmental genome shotgun sequencing of the Sargasso Sea. Science, 2004
  19. Zetsche B, Gootenberg JS, Abudayyeh OO, et al. Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell, 2015
  20. Shmakov S, Abudayyeh OO, Makarova KS, et al. Discovery and functional characterization of diverse class 2 CRISPR-Cas systems. Molecular Cell, 2015
  21. Abudayyeh OO, Gootenberg JS, Konermann S, et al. C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector. Science, 2016
  22. Doron S, Melamed S, Ofir G, et al. Systematic discovery of antiphage defense systems in the microbial pangenome. Science, 2018
  23. Millman A, Bernheim A, Stokar-Avihail A, et al. Bacterial retrons function in anti-phage defense. Cell, 2020
  24. Altae-Tran H, Kannan S, Demircioglu FE, et al. The widespread IS200/IS605 transposon family encodes diverse programmable RNA-guided endonucleases. Science, 2021
  25. Saito M, Xu P, Faure G, et al. Fanzor is a eukaryotic programmable RNA-guided endonuclease. Nature, 2023
  26. Altae-Tran H, Kannan S, Suberski AJ, et al. Uncovering the functional diversity of rare CRISPR-Cas systems with deep terascale clustering. Science, 2023
  27. Durrant MG, Perry NT, Pai JJ, et al. Bridge RNAs direct programmable recombination of target and donor DNA. Nature, 2024
  28. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature, 2021
  29. AlphaFold Protein Structure Database (EMBL-EBI and Google DeepMind)
  30. Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 2023
  31. The Nobel Prize in Chemistry 2024, press release
  32. Chikhi R, Lemane T, Loll-Krippleber R, et al. Logan: planetary-scale genome assembly surveys life's diversity. bioRxiv preprint, 2024
  33. Anzalone AV, Randolph PB, Davis JR, et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature, 2019
  34. Chien A, Edgar DB, Trela JM. Deoxyribonucleic acid polymerase from the extreme thermophile Thermus aquaticus. Journal of Bacteriology, 1976
  35. The Nobel Prize in Physiology or Medicine 1978, press release (restriction enzymes)
  36. The Nobel Prize in Chemistry 1993, press release (PCR and site-directed mutagenesis)
  37. The Nobel Prize in Physiology or Medicine 1975, press release (reverse transcriptase)

Related Articles