Researchers have identified a previously unknown enzyme system in bacteriophages, termed array-associated reverse transcriptase (ART). The discovery was notable because the analysis was assisted by Anthropic's Claude AI, which examined DNA sequence data and helped prioritize candidate systems for further study.
What was found?
The ART structure pairs a reverse transcriptase (RT) with an adjacent partner gene and a long series of evenly spaced tandem DNA repeats. Structurally, this arrangement resembles CRISPR systems, but the biological function of ART is currently unknown.
How the AI contributed
Claude worked as part of Anthropic's life-sciences-focused initiative to screen genomic data. According to the reported figures, nearly 950 agents ran for about 21 hours and consumed roughly 210 billion tokens. During processing, the agents generated several thousand proposed biological systems for downstream analysis.
The workflow collected more than 200,000 reverse transcriptase sequences and initially identified approximately 3,500 potential systems. That set was then narrowed to about 20 candidates for detailed reporting and human review.
Discovery process
One Claude agent focused on an unusual family of reverse transcriptases and examined the raw DNA surrounding those genes. In the RT gene neighborhood the agent detected a tandem-repeat array not previously documented. The AI counted the repeats, measured distances between them, compared the pattern against known RT-associated systems and surveyed the scientific literature. Finding no match, it flagged the arrangement to researchers as a novel biological system for follow-up.
Next steps
The biological role of ART remains to be determined, and laboratory experiments are underway to explore its function. The research team emphasizes that the AI-driven pre-screening significantly accelerated the discovery pipeline: Anthropic claims the automated analysis accomplishes work that could take human researchers weeks or months.
Significance
This finding illustrates how large language models and agent-based analysis pipelines can help identify new biological systems when large numbers of sequences must be scanned and contextual relationships inferred. Functional interpretation, however, still requires conventional experimental validation.



