37,000 AI agents found a virtual biotech: Stanford predicts clinical trial success and rediscovers an FDA-validated anticancer drug
🔎 A pharma with no employees, no lab, no payroll
On September 17, 2026, the team of James Zou, associate professor of biomedical data science at Stanford Medicine, and PhD student Harrison Zhang published in Science the portrait of an improbable company: a biotech with 37,000 employees, without a single human. No lab, no payroll, no badges — only AI agents organized like the divisions of a real pharma (Stanford Medicine), with the story covered in detail by SingularityHub.
Why now? Because the industry runs things backwards. About 90% of drug candidates that enter phase I never reach approval, and failures stem mostly from insufficient efficacy or unexpected toxicity (preprint on bioRxiv). A tool capable of triaging targets before committing hundreds of millions attacks the problem at its root.
And it's not a lab demo: the agents sifted through more than 55,000 clinical trials, produced a quantified predictive signal, and designed an anticancer strategy that a major pharma independently converged on — a molecule since designated a "breakthrough therapy" by the FDA. My take after reading the sources: the agent count makes the headline, but it's the org chart that makes the result.
The Essentials
- A "virtual biotech" published in Science on September 17, 2026 by James Zou and Harrison Zhang (Stanford Medicine): up to 37,000 AI agents under a virtual CSO, zero human employees.
- One agent per late-phase trial: the CSO assigned 37,075 agents; in total, 55,984 clinical trials annotated and analyzed, including roughly 50,000 in under a week.
- A strong predictive signal: drugs targeting "switch-like" genes gain +40% odds of advancing from phase 1 to phase 2, +48% of reaching the market, with -32% adverse events.
- The B7-H3 case: an antibody-drug conjugate designed by the agents using data predating January 2025; a major pharma independently arrived at the same strategy in August 2025, and the molecule received a breakthrough therapy designation from the FDA.
- The limitations are acknowledged: retrospective comparison, no lab, no clinical pipeline — Zou announces validation in a real laboratory as the next step.
Recommended Tools
You won't be directing 37,000 agents, but you can reproduce the mechanics — a supervisor, specialized divisions, a job description for each agent — on your own projects. Here's what you need to get started.
| Tool | Main use | Price (month year) | Ideal for |
|---|---|---|---|
| Selection of the best autonomous AI agents | Set up an agent org chart (supervisor + divisions) | Free, mostly open source | Reproduce Stanford's CSO/divisions structure |
| Local AI agents with Ollama | Run agents on sensitive data, without the cloud | Free | Medical, legal, or confidential data |
| Hostinger | VPS to run a swarm 24/7 | From a few €/month (September 2026, check hostinger.com) | Swarms that run continuously |
A hierarchy of agents, not a cloud of bots
No, the 37,000 agents don't form an anonymous swarm: it's an org chart, modeled on that of a classic biotech. A "chief scientific officer" agent leads specialized divisions that work in parallel — molecular targets on one side, trial design on the other (phys.org). Stanford makes it clear from the outset: this is not a startup with employees, labs, or a pipeline, but a multi-agent system organized into divisions under a CSO agent (findarticles).
Nature provides the definition and the key figure: agents are systems that interact with LLMs or with each other on multi-step tasks, and the CSO assigned 37,075 of them — one per late-phase trial. In total, the agents annotated and analyzed the results of 55,984 clinical trials (preprint), including roughly 50,000 in under a week (Stanford Medicine).
This is exactly the hierarchical pattern we detail in The 5 AI agent patterns that work: a supervisor, specialized executors, and definitely not a group chat with 37,000 participants.
One agent per trial: granularity as a principle
The study's smartest choice is also its least spectacular: each "clinical-trialist" agent receives a single trial to structure. It extracts the results and links them to multi-omic annotations derived from single-cell RNA-seq atlases (preprint).
This granularity has three virtues. A bounded context: one trial, not 55,000. A verifiable output: a trial's record can be checked. And trivial scaling: you add agents, not per-agent complexity. It's the opposite of the common reflex, which is to ask a single overloaded agent for everything.
The real result: "switch" genes predict success — and risk
Yes, and that's the scientific core of the study, beyond the buzz around the 37,000 agents. For trials with single-cell data available, the agents built two scores: cell-type targeting specificity, and bimodality — in other words, does the gene behave like an on/off switch or like a "dimmer" that modulates continuously (Stanford Medicine).
The verdict, on the analyzed cohort:
| Metric (drugs targeting switch-like genes) | Observed gap |
|---|---|
| Chances of moving from phase 1 to phase 2 | +40% |
| Chances of reaching the market | +48% |
| Adverse events | -32% |
Figures: Stanford Medicine, September 2026.
And this isn't a side effect of a single subdomain: the patterns remain stable across cancers and brain, heart, kidney, and lung diseases (Stanford Medicine). VentureBeat adds that targets carried by these single-cell features were about 50% more likely to reach the market, and that one of the drug designs was independently confirmed by Merck.
Put these figures up against the economics of the sector: ~90% failure in phase I, with failures concentrated on efficacy and toxicity (preprint). A score computed from public data that shifts those probabilities doesn't just save time — it redraws portfolio selection. This predictive layer fits into a broader trend we're tracking closely: DeepMind is already precomputing the molecular effect of 9 billion variants with AlphaGenome (our analysis). Biology is becoming, slowly but surely, a data-reading problem.
B7-H3: the "rediscovery" that hits the mark — and what it really proves
No, the agents didn't discover an approved drug: they designed a strategy that a major pharma company independently arrived at — and that's almost stronger than a discovery.
The case: B7-H3, a target highly expressed in the fibroblasts that surround certain lung tumors. These fibroblasts suppress the activity of nearby immune cells — the tumor protects itself as if under a blister (Stanford Medicine). The agents designed an antibody-drug conjugate (ADC) against B7-H3, using only data predating January 2025.
Then the twist: in August 2025, a few months after the agents' proposal, a major established private pharma company independently arrived at the same ADC strategy against B7-H3 (phys.org). The molecule — ifinatamab deruxtecan — has since received a "breakthrough therapy" designation from the FDA. VentureBeat credits Merck with this independent confirmation, and Zou sees it as third-party validation that is "independent and consistent" (Stanford Medicine). His line in Nature sums up the ambition: "We want to see how far these agent teams of AI scientists can help us to really accelerate drug discovery and development".
The enthusiasm still needs calibrating: the B7-H3 case is a retrospective comparison, not clinical proof (findarticles). The agents didn't predict the future — they read the past faster, more exhaustively, and without the fatigue of an expert committee. For an industry that reads poorly, that's already a decisive advantage.
What this changes for pharma economics — and what it doesn't
Predictive target selection can break the cost of failure; it won't eliminate failure itself.
The reasoning is simple: the price of an approved drug funds the nine candidates that die along the way. If upstream selection raises the success rate from 10% to ~15% (+48% in relative terms), every dollar of R&D mechanically buys more probability of approval. It's the highest-return lever in the entire chain, because it acts before the massive spending of phases II and III.
But let's keep our feet on the ground: agents have neither labs nor pipelines, and pharma's bottleneck has never been a lack of target ideas — it's physical validation, slow and expensive (findarticles). Zou himself sets the next step: testing in a real laboratory how many of these results hold up (phys.org).
My bet: agent swarms will compress the "reading and hypotheses" phase from several months to a few days, and shift the scarcity toward real-world testing capabilities. Scarcity shifts; it doesn't disappear.
Swarms: 37,000 agents at Stanford, 700,000 at MiroFish
The scale of swarms has become, in 2026, a unit of measurement for agentic AI.
Just a few weeks after Stanford's virtual biotech, a single undergraduate built MiroFish, a swarm of 700,000 AI agents built in 10 days — an open source project that is exploding on GitHub. Two philosophies are clashing here.
Stanford plays for depth: a vertical domain, a business org chart, verifiable outputs, and third-party validation from Merck and the FDA. MiroFish plays for breadth: a generic, open source framework that anyone can repurpose for their own ends.
My take: it's not the headcount that impresses, it's the job description. 37,075 agents, each with a trial to structure, produce more value than 700,000 agents with no clear mission. The lesson to take away isn't "more agents", it's "more specialization".
Replicating the mechanics: your mini agent biotech in 4 steps
You have neither 55,000 trials nor a virtual CSO, but Stanford's organizational recipe can be copied almost as-is onto a business use case.
-
Draw the org chart before the prompts. One supervisor agent, two or three specialized divisions, an explicit job description for each agent. With OpenClaw, this is exactly the role of the SOUL, AGENTS, and Skills files — see our configuration guide.
-
One agent = one unit of work. That's Stanford's granularity: one agent per trial. For you: one agent per document, per support ticket, per client. Never "one agent that handles everything."
-
Choose the model per division, not per project. Heavy reasoning for the supervisor (GPT-5.5 or Gemini 3 Pro Deep Think), a good cost/quality ratio for execution (Claude Sonnet 4.6), self-hosting for sensitive data (Kimi K2.6 or GLM-5). The comparison of the best LLMs for AI agents details these trade-offs.
-
Close the loop with a human. Stanford's agents have no lab; your agents have no decision-making power. Log every output, have a human validate it, and host the swarm on a stable foundation — a Hostinger-type VPS (see table) or locally with Ollama if the data is sensitive.
The entry cost is negligible compared to what Stanford just demonstrated: value comes from structure, not size.
❌ Common mistakes
Mistake 1: Reciting "37,075 trials analyzed in 6 hours"
The most widespread confusion about this story — and it's understandable, given how quickly summaries circulate. The primary sources say: 37,075 refers to the agents, one per late-phase trial (Nature); in total, 55,984 trials were annotated and analyzed (preprint), of which about 50,000 in less than a week (Stanford Medicine). Rule: always cite the number with its verb — how many agents, how many trials, in how much time.
Mistake 2: Presenting B7-H3 as clinical proof
The story is a retrospective comparison: the agents reconstructed a strategy from old data, which a pharma company later converged on (findarticles). The FDA designation applies to the pharma's molecule, not to the agents' design. Present it as a consistency test, not as a won trial.
Mistake 3: Believing the agents replace the lab
The virtual biotech has neither a lab nor a clinical pipeline, and Zou announces validation in a real lab as the next step (phys.org). The agents compress reading and hypothesis generation — not biology. Treat them as an upstream filter, never as final proof.
Mistake 4: Launching a swarm without an org chart
37,000 agents in a group chat produce noise, not science. What works at Stanford: hierarchy, specialization, granularity — one agent, one task. Before adding agents, add structure: it's the number one criterion in our selection of autonomous agents.
❓ Frequently Asked Questions
Did the agents analyze 37,075 trials in 6 hours?
No. 37,075 is the number of agents — one per late-phase trial (Nature). In total, the agents annotated and analyzed 55,984 clinical trials (preprint), including roughly 50,000 in under a week (Stanford Medicine). The exact timeframe hardly matters: it's the massive parallelism that changes the game.
Which LLMs power these 37,000 agents?
The sources (Stanford Medicine, Nature, preprint) do not name the underlying model; they describe agents interacting with LLMs or with one another. The lesson is valuable: at scale, the architecture — CSO, divisions, one agent per task — matters more than the choice of engine.
Did the FDA approve the drug designed by the agents?
No. The agents designed an ADC strategy against B7-H3 based on data predating January 2025. A major pharma company independently arrived at the same strategy in August 2025; its molecule, ifinatamab deruxtecan, received a "breakthrough therapy" designation — an accelerated review, not an approval.
Is Stanford's system open source?
The cited sources do not mention any public code; the preprint detailing the method is available on bioRxiv. To experiment with a large-scale swarm, the open-source reference remains MiroFish and its 700,000 agents, whose architecture we have described.
Can this system be reproduced at a small scale?
Yes, the structure carries over: a supervisor, specialized divisions, one agent per unit of work, human validation at the end of the chain. A few agents are enough for a business use case, hosted on a VPS costing a few euros per month or running locally with Ollama.
✅ Conclusion
Stanford's virtual biotech doesn't prove that AI agents will replace researchers: it proves that an agent org chart reads 55,000 clinical trials where a human department reads 50, and that this reading predicts success better than committee intuition. To go from theory to your first structured swarm, start with our selection of the best autonomous AI agents.