A team from the University of Oxford and Fordham University published a paper in Nature describing something that sounds like science fiction but is very real lab work: an AI system that proposed, tested, and refined a new drug candidate for an eye disease almost entirely on its own.
The system is called Robin, and the story of what it did is worth unpacking, because it says a lot about where AI-assisted biology is actually headed, not just where it might be headed someday.
The Problem Robin Was Built to Solve
Scientific discovery usually moves through four stages: you observe something, you form a hypothesis, you run an experiment, and you analyze the results. Then you repeat the cycle with what you learned. This sounds simple, but each stage takes real human time and expertise. Reading enough papers to form a good hypothesis can take hundreds of hours. Designing the right experiment takes judgment built from years of training. Making sense of messy experimental data takes even more judgment.
Robin is a multi-agent system, meaning it’s not one AI model doing everything, but several specialized agents working together. Two of them, called Crow and Falcon, handle literature review. Crow does quick literature scans, and Falcon does deeper, more exhaustive searches to evaluate candidate ideas. A third agent, Finch, handles data analysis, writing, and running its own code to interpret experimental results such as flow cytometry readouts or RNA sequencing data.
What Robin Was Asked to Do
The researchers pointed Robin at dry age-related macular degeneration, or dAMD, the leading cause of irreversible vision loss in developed countries. There’s no approved treatment for it, and by 2050, the number of affected people in the US alone is projected to nearly triple as the population ages.
Robin started by reading about 150 papers to identify plausible disease mechanisms, then settled on a strategy: boosting the ability of retinal pigment epithelium (RPE) cells to clear cellular debris through phagocytosis. From there, Robin read roughly 400 more papers and proposed a list of existing drugs that might enhance this process. Human scientists then ran the actual lab experiments and fed the raw data back to Robin for analysis.
The Discovery
The first round pointed to a compound called Y-27632, a ROCK inhibitor that isn’t approved for use in humans. Robin then proposed a follow-up RNA sequencing experiment, which revealed that this drug was switching on a gene called ABCA1, involved in moving cholesterol and fats out of cells. That’s a meaningful finding on its own, since a related gene, APOE, has long been linked to the risk of macular degeneration.
Using those insights, Robin proposed a second round of candidates, and this time the standout was ripasudil, a ROCK inhibitor already approved in Japan for the treatment of glaucoma. In lab testing, ripasudil outperformed Y-27632, and unlike Y-27632, it already has a known human safety profile, which matters a lot for how quickly a repurposed drug could move toward clinical testing. The finding was later validated in RPE-SCs (native adult RPE cells isolated from a donor eye), cells taken from an actual elderly donor, not just an immortalized cell line, which adds some real weight to the result. Robin also flagged a second compound, KL001, a circadian clock modulator that had never been linked to phagocytosis.
Why Speed Matters
The researchers estimated that the literature review alone, which Robin completed in about 30 minutes by processing 551 papers, would have taken a human scientist roughly 294 hours. Add in hypothesis generation, experimental design, and data analysis, and the full cycle that took Robin under two hours of computed time is estimated to take a human team somewhere between 359 and 424 hours. That’s not a small efficiency gain; it’s closer to a 200-fold reduction in the cognitive labor involved.
The team also tested what happens when you strip out the specialized literature agents and replace them with a general-purpose model. Hallucinated references jumped dramatically, which is a useful reminder that not all AI research tools are equally reliable, and that the architecture behind an AI system matters as much as the underlying language model. They also compared Robin against OpenAI’s Deep Research tool on the same task, and none of Deep Research’s 17 suggested compounds turned out to be hits, and it never landed on ROCK inhibition as a strategy at all.
What This Means Going Forward
Robin’s authors are careful to note that this is still a “lab-in-the-loop” system. Humans ran every physical experiment, and every hypothesis still requires standard preclinical validation before anything gets near a clinical trial. But as a demonstration of AI systems that can move from a research question to a testable, validated hypothesis with a real drug candidate at the end, it’s a fairly significant proof of concept, and one that the authors suggest could extend well beyond drug repurposing into other areas of science entirely.
The full code for Robin and its data-analysis agent, Finch, is open-sourced on GitHub, so anyone curious about the architecture can look under the hood themselves.
Article Source: Reference Paper | Reference Article | Robin Code: GitHub | Robin Code: GitHub
Disclaimer:
The research discussed in this article was conducted and published by the authors of the referenced paper. CBIRT has no involvement in the research itself. This article is intended solely to raise awareness about recent developments and does not claim authorship or endorsement of the research.
Follow Us!
Learn More:
Anchal is a consulting scientific writing intern at CBIRT with a passion for bioinformatics and its miracles. She is pursuing an MTech in Bioinformatics from Delhi Technological University, Delhi. Through engaging prose, she invites readers to explore the captivating world of bioinformatics, showcasing its groundbreaking contributions to understanding the mysteries of life. Besides science, she enjoys reading and painting.












