top of page

Anthropic Claude Enzyme Discovery Looks Like CRISPR, but the Function Is Still Unknown

1 day ago
12 min read

Anthropic says Claude autonomously found a previously uncharacterized enzyme system after roughly 950 agent sessions searched biological data for 21 hours. The Anthropic Claude enzyme discovery resembles CRISPR in one important structural respect, but no experiment has shown that it edits genes.

That distinction separates an intriguing scientific lead from the tool implied by some early headlines. Anthropic calls the system array-associated reverse transcriptases, or ARTs. Its researchers found reverse transcriptase genes beside repeating DNA arrays in bacteriophages, which are viruses that infect bacteria.

The result is Anthropic’s first public discovery from a life sciences research group formed in spring 2026. It also tests a larger proposition: whether general-purpose AI agents can notice biological anomalies, pursue them, and hand scientists candidates worth testing. Claude passed that test once, yet failed to rediscover the defining array in ten subsequent campaign runs.

What Anthropic’s Claude Enzyme Discovery Actually Found

Claude found a new biological arrangement, not a working CRISPR replacement.

Anthropic announced the result on September 23, 2026, alongside its new Bay Area molecular biology laboratory. According to the company’s enzyme system announcement, the lab handles lower-risk BSL-1 and BSL-2 work. Human scientists perform every physical experiment.

The computational campaign began with a high-level research brief. Claude was asked to search for unusual reverse transcriptase systems. A reverse transcriptase, or RT, is an enzyme that copies RNA into DNA.

The agents assembled more than 198,000 RT clusters and examined thousands of neighboring protein families. The full campaign consumed 215.6 million tokens across 949 agent sessions, according to Anthropic’s ART technical report. It ran for 21.5 hours of elapsed time and represented 77 total agent-hours.

Anthropic’s public summary rounds those figures to roughly 200,000 RTs, 950 agents, and 210 million tokens. Those numbers describe the scale of the search. They do not establish that 950 independent AI scientists simultaneously investigated the problem.

The “agents” were model sessions coordinated through a research harness. One session could plan a task, another could review it, and either could open follow-up tasks. This structure let Claude explore several branches without a scientist directing every intermediate decision.

The original brief focused on new partner genes located beside reverse transcriptases. Claude scored 3,564 protein families as possible partners, then conducted deeper investigations. Three associations survived review as previously unreported RT relationships.

ART emerged through a more unusual route. One agent investigated a reverse transcriptase lineage associated with jumbo phages, which are bacteriophages with unusually large genomes. The expected protein association looked unconvincing, but the agent continued exploring the surrounding genomic region.

It loaded raw upstream DNA into its working context and noticed a tandem repeat array. One locus contained 14 copies of a 16-nucleotide repeat, separated by unique stretches measuring roughly 100 to 200 nucleotides.

That arrangement evoked CRISPR because CRISPR systems also use repeating DNA separated by variable sequences. Cells transcribe those arrays into smaller RNAs, which can guide CRISPR-associated proteins toward matching genetic material.

Anthropic’s laboratory work found that ART arrays are also transcribed into distinct short RNAs. In one Staphylococcus bacteriophage dataset, array-derived RNA reached as much as 8 percent of the phage’s RNA 15 minutes after infection.

The researchers identified 95 distinct ART-related RT clusters across cultured phages and predicted viral sequences. Twenty-eight had a detectable repeat array located upstream. The system usually combined three elements: a reverse transcriptase, an adjacent partner gene, and an array of evenly spaced DNA repeats.

Those observations support the claim that ART is a previously uncharacterized biological system. They do not reveal what the system does. Anthropic’s researchers have not shown that the reverse transcriptase is active, that it uses the short RNAs as substrates, or that ART targets other genetic sequences.

That leaves the most interesting possibility open rather than established. ART might represent another programmable molecular system, or its CRISPR-like architecture might perform a very different role inside phages.

Why the CRISPR Comparison Matters

The comparison matters because CRISPR also began as an unexplained repeat pattern, not because ART already performs gene editing.

Scientists first encountered what became CRISPR while examining an Escherichia coli gene in 1987. They noticed an unfamiliar series of repeated DNA sequences but did not understand its function. The original CRISPR sequence90183-4) appeared years before researchers connected those repeats to microbial immunity.

The term CRISPR arrived in 2002. Later studies established that bacteria and archaea store fragments of viral genetic material between repeats. RNA copied from those regions helps associated proteins recognize and attack matching invaders.

Researchers eventually adapted that mechanism for programmable genome editing. The historical lesson is not that every repeat array becomes a biotechnology platform. It is that an odd genomic pattern can mark machinery whose importance becomes clear only after years of experimental work.

ART shares part of that visual grammar. Its repeating units are separated by variable sequences, and the array produces short RNAs. Those features justify investigating whether the RNAs guide or program another molecular activity.

Several differences also weaken any direct equivalence. The technical report found no nearby cas genes, which encode the proteins central to established CRISPR systems. ART spacers also appear conserved among related phages, while CRISPR spacers usually vary as microbes encounter different viruses.

ART is built around reverse transcriptase rather than a known CRISPR-associated nuclease. A nuclease cuts nucleic acids, while reverse transcriptase copies RNA into DNA. That difference means ART could participate in genome modification without operating like familiar CRISPR tools.

Feng Zhang, a Broad Institute and MIT researcher who helped develop CRISPR genome editing, reviewed the preprint before Anthropic’s announcement. He described the RNA-repeat arrays associated with reverse transcriptases as genuinely intriguing and said they merited further investigation.

His reaction supports the scientific interest of the finding. It does not amount to independent experimental confirmation. No outside laboratory has yet published evidence establishing ART’s function or programmability.

The comparison therefore works best as a research map. CRISPR gives scientists a reason to ask whether ART’s repeat-derived RNAs encode targets, guide an enzyme, or record information. It does not supply the answers.

A separate research group has reported non-coding RNA arrays beside another family of reverse transcriptases. Its genome mining preprint used a purpose-built genome language model to identify related structural patterns.

That parallel matters because it suggests repeat-linked reverse transcriptases might not be a single accident. Similar architectures could have evolved more than once, perhaps because they solve a recurring biological problem.

The same evidence also makes the Anthropic claim less singular. AI-supported genome mining was already being applied to obscure non-coding elements and their associated proteins. Claude’s contribution lies partly in using a general model and agent harness rather than a system designed for one genomic task.

The Anthropic Claude enzyme discovery is compelling because the model appears to have noticed the defining array without a prompt telling it to find repeats. The research brief targeted RT partner genes, and the agent pursued the array after following an initially rejected lead.

That resembles scientific serendipity. An investigator searches for one relationship, finds an anomaly, and changes direction. Whether AI can produce that pattern reliably is now more important than whether one successful transcript sounds human.

The Real Contest Is Reliable Discovery Versus Lucky Attention

Anthropic’s strongest result is that Claude recognized the anomaly; its biggest weakness is that the system rarely looked in the right place.

The discovery campaign was autonomous in a limited but meaningful sense. Once researchers supplied the initial brief, the harness ran without human intervention for more than 21 hours. Claude organized tasks, built computational searches, rejected weak candidates, and opened follow-up investigations.

No scientist instructed the successful agent to inspect a repeat array. The model independently loaded the upstream sequence, described the pattern, wrote analysis code, compared the structure with known systems, and prepared a report for human review.

That is more than summarizing papers or proposing a familiar experiment. The system moved from a broad objective to an observation that was not specified in the original task.

However, Anthropic reran the same campaign ten times after the initial discovery. None of those reruns rediscovered the array.

Several reruns sampled ART-related sequences, and two investigated the relevant lineage. The agents still failed to load the upstream DNA region containing the repeats. Claude could recognize the feature, but the workflow did not consistently expose the feature to the model.

That distinction changes how “autonomous discovery” should be interpreted. The successful run shows that the model has useful pattern-recognition and follow-up abilities. The ten misses show that the surrounding agent system cannot yet produce dependable coverage.

Anthropic conducted more controlled tests to isolate recognition from search behavior. Researchers gave seven Claude models ART-related inputs at five levels of detail, either directly inside the prompt or through files and tools.

The strongest models identified the array in at least 90 percent of trials when the relevant sequences were placed directly in context. Performance fell when models had to locate and read the sequences through tools. In some conditions, recognition dropped as low as 32 percent.

Across 39 percent of file-based attempts, the model never read a continuous stretch of at least 200 nucleotides. It therefore saw too little raw sequence to recognize more than one repeat unit.

The bottleneck was not simply biological reasoning. It was attention management inside an open-ended workflow.

An agent has limited time, context, and computational resources. It must decide which files to open, which regions to inspect, and which anomalies deserve another task. A capable model can still miss a discovery if its tools summarize away the decisive evidence.

This creates a more practical benchmark for AI science. A useful discovery system must do more than identify a pattern when researchers place it in view. It must repeatedly navigate a large search space and decide to inspect the right evidence.

Traditional scientific software often handles this problem through deterministic pipelines. Researchers define explicit searches for repeat sequences, protein domains, conserved neighborhoods, or evolutionary relationships. Those methods are narrow, but they behave consistently.

A general AI agent offers a different advantage. It can combine literature review, code generation, sequence analysis, hypothesis formation, and written interpretation. It can also change direction when a result contradicts its initial expectation.

The tradeoff is variability. Two runs can follow different branches even when they receive the same brief. One session may inspect raw DNA, while another may accept a summary and move on.

Anthropic’s result therefore pressures both AI laboratories and scientific software developers. General agents must become more reproducible. Specialized tools must become easier for agents to select and orchestrate without losing raw evidence.

A stronger system would combine both approaches. Deterministic checks could require agents to inspect minimum sequence windows, run repeat-detection software, and preserve intermediate evidence. The language model could then focus on selecting hypotheses and connecting unexpected results.

This is similar to good knowledge work outside biology. An AI system produces better analysis when it can retain source material, inspect the underlying record, and connect findings across tasks. A structured AI knowledge base cannot replace expert judgment, but it can reduce the chance that decisive evidence disappears inside a summary.

For Anthropic, reliability now matters more than another large agent count. Turning one success across eleven campaigns into a repeatable result would strengthen the case for autonomous research far more than scaling the swarm.

The Wet Lab Narrows the Claim Without Settling It

Physical experiments confirmed that the arrays produce RNA, but they did not establish the enzyme’s activity or biological purpose.

AI biology announcements often blur three separate achievements: generating a hypothesis, validating an observable feature, and proving a useful function. Anthropic has completed parts of the first two stages for ART.

The computational campaign identified the genomic arrangement and proposed that it represents a system. Human scientists then tested whether the repeat arrays were expressed. They detected distinct short RNA products, supporting the idea that the array participates in an active biological process.

That validation matters. Genome databases contain sequencing errors, incomplete assemblies, and patterns with no functional importance. Observing RNA expression makes ART harder to dismiss as a computational artifact.

Yet transcription alone does not explain the system. DNA regions can produce RNA without controlling a programmable enzyme. Researchers still need to determine whether the ART reverse transcriptase copies those RNAs, whether the partner protein participates, and what biological outcome follows.

The current work also lacks independent replication. Anthropic published a technical report rather than a peer-reviewed journal article. Outside researchers can inspect its methods, but another laboratory has not publicly reproduced the biochemical observations.

The company deserves credit for documenting unfavorable results. The ten failed reruns and the weaker file-based benchmark scores complicate its marketing language, yet they appear in the report.

Those disclosures let readers distinguish three claims.

First, Claude found an overlooked genomic pattern during an unsupervised campaign. The available logs and methods support that account.

Second, ART is a previously uncharacterized system whose repeat arrays produce short RNAs. Anthropic’s experiments support that conclusion, subject to external replication.

Third, ART is comparable to CRISPR as a programmable gene-editing platform. Current evidence does not support that conclusion.

Anthropic itself uses more careful language than some summaries. It describes properties “reminiscent of CRISPR” and says the system might involve something analogous. It also states that ART’s primary function remains unknown.

The gap between the second and third claims could take substantial work to close. Researchers must purify or express the relevant proteins, identify molecular substrates, measure activity, and test whether changing spacer sequences changes a target.

They also need to establish ART’s natural role. Because the system appears mainly in bacteriophages, it could help viruses copy genetic information, evade bacterial defenses, manipulate host cells, or compete with other phages. Each possibility implies a different experimental program.

Safety and governance will matter if the system becomes programmable. Anthropic says its laboratory does not work with pathogens capable of infecting humans. The current experiments remain far from a clinical or high-risk gene-editing application.

The company is also exploring ways for AI agents to operate scientific instruments. Its hardware standard preview describes a common interface for microscopes, liquid handlers, robotic arms, and other programmable devices.

ART did not emerge from a fully robotic laboratory. Claude conducted the database search and assisted with analysis, while human scientists performed the wet-lab work. That separation is important because it limits both the autonomy claim and the immediate physical risk.

Combining agentic analysis with automated experiments would create a more complete discovery loop. An AI system could propose candidates, order tests, inspect results, and revise its hypotheses. It would also create additional failure modes involving instrument control, contamination, experimental design, and unsafe objectives.

Anthropic has already observed that Claude can struggle with physical constraints. In a laboratory automation test, the model retried a liquid-handling operation when bubbles caused errors, which made the problem worse. Human guidance helped it recognize the physical cause and adjust its behavior.

That example captures the central limit of present AI science. Models can reason across vast amounts of symbolic information, yet they do not possess a scientist’s accumulated physical intuition. A wet lab exposes errors that look plausible in text or code.

ART is therefore a useful early test precisely because the result remains incomplete. The AI generated a promising lead. Human experiments established that part of the lead was real. Neither side has yet explained the biology.

What Scientists and AI Builders Should Watch Next

Three signals will determine whether ART becomes an important discovery or remains an instructive AI-generated lead.

The first signal is functional validation. Researchers need evidence that the reverse transcriptase or its partner protein performs a measurable reaction involving the array-derived RNAs.

A convincing experiment would identify the enzyme’s substrate and product. Stronger evidence would show that changing a repeat or spacer changes the molecular outcome in a predictable way.

If ART proves programmable, the CRISPR comparison gains substance. If its RNAs are unrelated to the enzyme’s activity, the comparison weakens even if the system remains biologically interesting.

The second signal is independent replication. Another laboratory should confirm the RNA products, reconstruct the system, and test its components without relying on Anthropic’s internal workflow.

Independent results would separate the biology from the company’s model demonstration. Failure to reproduce the observations would force a closer examination of experimental conditions, sequence selection, and analysis.

Peer review will help, but publication alone is not the decisive threshold. Functional experiments from researchers with different incentives and equipment would carry more weight.

The third signal is campaign reproducibility. Anthropic or another group should rerun the discovery workflow with safeguards that force agents to inspect raw sequence around promising loci.

The controlled benchmarks already suggest a direct intervention. When Claude reads the relevant DNA, it usually recognizes the repeats. A better harness should make that inspection systematic without turning the workflow into a script designed specifically to find ART.

Researchers should report success across multiple runs, not only the best transcript. They should also disclose how many candidates required human review, how much compute each campaign used, and which automated decisions changed the final result.

If a revised system finds ART in most comparable runs, the initial one-in-eleven result will look like an early engineering limitation. If success remains rare, autonomous discovery will still resemble opportunistic exploration more than dependable scientific infrastructure.

That does not make the original observation worthless. Human science also depends on chance, unusual attention, and investigators who follow weak signals. The value of AI may partly come from creating more opportunities for such moments.

The economic question is whether those opportunities justify the cost of screening their output. Anthropic’s campaign produced thousands of possible associations and a much smaller set of reports worth reading. Human experts still had to judge candidates and conduct experiments.

For working scientists, the immediate lesson is narrower than “AI can replace researchers.” General agents can search across literature, genomic data, and computational tools at a scale that one person cannot sustain. They can also miss obvious evidence unless workflows preserve raw data and require adequate inspection.

For AI builders, ART offers a better benchmark than polished scientific prose. Can an agent explore an open search space, preserve evidence, recognize an anomaly, challenge its first explanation, and reproduce the path later?

For enterprise and knowledge-work teams, the same design principle applies. Agent quality depends on access to original records, traceable decisions, and repeatable retrieval. A useful searchable knowledge base should help people verify how an answer emerged, not merely present a confident conclusion.

The Anthropic Claude enzyme discovery has already cleared one meaningful threshold. Claude noticed a genomic structure that researchers had not characterized and directed human attention toward it.

It has not cleared the harder thresholds of repeatable discovery, independent validation, or demonstrated gene-editing function. Those are now the measurements that matter.

Watch the experiments, not the analogy. If ART proves programmable and other laboratories reproduce it, Anthropic will have a strong example of AI initiating a consequential biological discovery. If the function remains elusive, the project will still reveal something useful: AI can generate scientific leads, but dependable science begins when those leads survive repeated contact with evidence.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page