Ask anyone in pharma why drugs fail, and you’ll hear the same story. Roughly nine out of ten candidates who enter Phase I never reach approval, mostly because they don’t work well enough or turn out to be unsafe. Years of effort and billions of dollars, gone. The research team based at Stanford University (the Departments of Biomedical Data Science, Computer Science, and Genetics) tries a different approach. Their idea is simple to say but hard to pull off: build a whole biotech research team from AI agents. They call it the Virtual Biotech.

Why drug discovery needs something like this

The authors point to a few complex problems. Expertise is scattered across genetics, single-cell biology, chemistry, and clinical work, and these groups often work in silos. Decisions can be swayed by office politics or personal bias. Nobody can easily trace why a call was made. And the sheer volume of biological data is now too much for any person to read through by hand.

How the Virtual Biotech is set up

It copies the structure of a real company. At the top level sits a Chief Scientific Officer (CSO) agent. It doesn’t analyze data itself. It takes your question, asks clarifying follow-ups, and hands the work to specialist “scientist” agents.

There are eleven agents in total, grouped into four divisions: target identification, target safety, modality selection, and clinical officers. Together, they can use more than 100 tools that query databases like Open Targets, CELLxGENE, ClinicalTrials.gov, and Tahoe-100M. A chief of staff agent prepares a briefing first, and a scientific reviewer agent checks the work afterward and sends it back if the reasoning is thin. Everything is logged so that a human expert can audit it.

Test 1: Reading 55,984 clinical trials

Trial outcomes are messy; they’re often missing or buried in free text. Thus, the system deployed 37,075 clinical trial agents simultaneously, each working on one trial at a time. Every agent goes to ClinicalTrials.gov, then PubMed, then press releases, and notes their sources.

The whole batch finished in about six hours. For a single agent, it would have needed roughly 77 days. When humans spot-checked 100 trials, they agreed with the agents 89.7% of the time on primary endpoints.

Then came the interesting part. The agents looked at how drug targets are expressed in single cells and compared that with trial results. Drugs aimed at cell-type-specific targets did better on several fronts:

  • 40% more likely to move from Phase I to Phase II
  • 48% more likely to reach the market (Phase IV)
  • About 32% lower adverse event rates

These links held up after adjusting for genetic evidence, so single-cell data seems to add something new. The authors are careful to say this is observational, not proof of cause and effect.

Test 2: B7-H3 in lung cancer

Next, the team asked the system whether B7-H3 (CD276) is a good target for lung cancer. Germline genetics gave little support, but the CSO reasoned that for checkpoint targets, tumor overexpression matters more than inherited variants.

From there, the agents found that B7-H3 was concentrated in cancer-associated fibroblasts, not just tumor cells. Spatial data showed immune cells were depleted around B7-H3-high spots. In TCGA data, high B7-H3 tracked with worse disease-specific survival (HR 1.82). Small molecules looked unlikely to work, so the CSO suggested an antibody–drug conjugate.

The whole analysis took under a day and cost $46 in API fees. And here’s a nice check: the model’s knowledge stopped in January 2025, but in August 2025, the FDA granted Breakthrough Therapy Designation to ifinatamab deruxtecan, a B7-H3-targeting ADC.

Test 3: Why did an ulcerative colitis trial fail?

The MOONGLOW trial tested vixarelimab, an antibody against OSMRβ, in ulcerative colitis. It was stopped for futility after 79 participants. The genetics for OSMR looked strong, so what went wrong?

The agents noticed the trial enrolled patients without checking their OSMR levels. Across five earlier UC trials, people who didn’t respond to biologics had higher baseline OSMR. In single-cell data, certain fibroblasts in non-responders showed rising OSMR alongside more JAK-STAT signaling. The suggestion: screen patients and enroll only those with high OSMR, the way HER2 or PD-L1 testing works in cancer. This one costs about $54.

The fine print

The authors are upfront about the constraints. AI agents make mistakes, and one wrong step can ripple through the chain. The case studies produce hypotheses that still need lab and clinical validation. Their trial analysis only covers registered trials with mapped targets. Humans stay in the loop.

Why it matters

The real value may be speed and reach. A question that once took a team weeks can be explored overnight, and many targets can be screened with the same rigor. The paper isn’t peer-reviewed yet, so treat the numbers as early. But it’s a clear sign of where research teams are heading: not scientists replaced, but scientists with a much bigger bench behind them.

Sources: Reference Paper | Published Abstract

Disclaimer:
The research discussed in this article was conducted and published by the authors of the referenced paper. CBIRT has no involvement in the research itself. This article is intended solely to raise awareness about recent developments and does not claim authorship or endorsement of the research.

Learn More:

Author

Anchal is a consulting scientific writing intern at CBIRT with a passion for bioinformatics and its miracles. She is pursuing an MTech in Bioinformatics from Delhi Technological University, Delhi. Through engaging prose, she invites readers to explore the captivating world of bioinformatics, showcasing its groundbreaking contributions to understanding the mysteries of life. Besides science, she enjoys reading and painting.

LEAVE A REPLY

Please enter your comment!
Please enter your name here