The meeting request came from a virtual chief scientific officer. It delegated to target-discovery agents, safety-risk agents, trial-design agents - 37,000 of them in total. On September 17, 2026, Stanford Medicine published research in Science describing a virtual biotech company built from tens of thousands of AI agents spanning the drug-development pipeline.
The system analyzed 55,984 clinical trials with more than 37,000 AI agents working across the dataset, and processed that volume in less than a week. What the paper makes concrete - in a way that abstract multi-agent discussions rarely do - is that the hard problem was never the models. It was the coordination layer.
What Stanford's 37,000-agent system actually did
The system features a virtual chief scientific officer that takes a query from a human user and delegates tasks to an army of specialized "scientist" agents, each armed with their own databases and tools.
The organization mirrors an established biotech, with specialized divisions working in parallel on activities ranging from identifying molecular targets to designing clinical trials.
The most interesting part is not the scale - it is the shape. This was not a flat swarm of identical agents voting on outputs. It was a strict hierarchy with defined roles, shared artifacts passed between divisions, and humans explicitly owning irreversible decisions. Lead author Zhang and senior author Zou report that the system proposed a B7-H3 antibody-drug conjugate design using information available before January 2025, and months later a pharmaceutical company independently arrived at a similar strategy that received FDA breakthrough therapy designation
- an external consistency check that is hard to dismiss.
Stanford reports that the agents analyzed and catalogued about 50,000 clinical trials in less than a week, work the researchers said would have taken humans years.
The non-obvious takeaway for product teams: the architecture that made this tractable - specialized roles, crisp handoffs, human review before anything irreversible - is not specific to drug discovery. The virtual lab metaphor only works when agents have crisp jobs and humans own irreversible decisions. That principle holds whether you are processing clinical trials or routing support tickets.
MCP vs A2A: the coordination layer that teams are actually choosing between
Multi-agent AI coordination now has two standardized protocols, and the choice between them is architectural, not cosmetic.
Both MCP and A2A are complementary but focus on different layers: MCP standardizes model-tool interactions, while A2A enables agents to coordinate tasks and communicate across systems. In plain terms: MCP is how an agent picks up a tool; A2A is how one agent hands a task to another agent.
| MCP | A2A | |
|---|---|---|
| What it connects | Agent → tools, data, APIs | Agent → other agents |
| State management | Application layer (you build it) | Native multi-turn task lifecycle |
| Coordination complexity | Lower - fewer protocol primitives | Higher - agent discovery, capability negotiation |
| Ecosystem maturity | 97M+ monthly SDK downloads | 150+ orgs in production, April 2026 |
| When to reach for it | Single agent needing many tools | Multiple agents delegating across boundaries |
A University of York team evaluated an MCP-based and an A2A-based multi-agent implementation of the same software engineering task against requirements including agent discoverability, multi-turn conversations, asynchronous communication, and access control. MCP can support inter-agent coordination through a comparatively lightweight implementation model with lower coordination complexity, although conversational state management and task lifecycle handling must be implemented explicitly at the application layer.
The MCP version was lighter and less complex to coordinate, but conversational state and task lifecycle handling had to be built at the application layer. The A2A version got stateful, multi-turn coordination natively, at the price of greater implementation complexity.
That is the real trade-off teams are navigating. Neither answer is wrong - it depends on whether you want to pay the complexity tax in protocol or in application code.
Both protocols are now governed by the Agentic AI Foundation under the Linux Foundation, with co-founding members including Anthropic, Google, OpenAI, Microsoft, AWS, and Block. As of May 2026, the foundation has 190 member organizations - making it one of the fastest-growing projects in Linux Foundation history. No single vendor controls either standard, which matters for teams that learned the hard way from proprietary integrations.
The coordination patterns that actually work at scale
The Stanford paper and the MCP/A2A research converge on the same three patterns for multi-agent AI coordination:
Role specialization with narrow scope. Calling three APIs in sequence is not multi-agent coordination - it is a workflow. The distinction matters because multi-agent systems introduce autonomous decision-making: agents choosing what to do, not just executing a predefined script. If your "agents" are just running scripted tool calls, you have a pipeline, not a multi-agent system.
Hierarchical delegation, not flat consensus. Stanford's system had a CSO agent, division agents, and specialist agents - each layer with a defined authority. At VB Transform 2026, James Zou argued that the next frontier isn't a single more capable agent, it's tens of thousands of them collaborating. For developers and product builders, the most critical takeaway is how these massive systems are orchestrated. Flat swarms with no clear ownership are debuggable only in theory.
Human gates at irreversible points. Agentic AI can expand the amount of scientific evidence examined, allow specialized work to proceed in parallel, and surface connections that would be extremely difficult for human teams to find at the same speed. In an industry where development programs can take years and clinical trials can cost tens or hundreds of millions of dollars, that capacity matters. But the story is scientific process acceleration, not a license to ship decisions from an unsupervised agent.
A teammate like Beagle operates exactly this way inside Slack and Teams - it drafts and surfaces, a human approves before anything posts. The draft-and-approve model is a human gate at the last reversible moment.
When to actually build a multi-agent system
The biggest mistake is building a mesh of six specialized agents when a single well-prompted agent with good tools would do the job. Multi-agent systems are harder to debug, harder to monitor, and harder to explain.
The decision is simpler than the hype suggests. A single agent with MCP tool access covers the vast majority of real team workflows - reading from Notion, posting to Slack, querying a database, filing a Linear issue. MCP delivers immediate impact on coding workflows, while A2A addresses peer-to-peer task delegation between autonomous agents - a pattern that only becomes necessary once a team is coordinating multiple agents across processes or machines.
Reach for a multi-agent architecture when:
- The task is genuinely decomposable and the sub-tasks are independent enough to parallelize
- Different sub-tasks need different tools, permissions, or context windows that cannot be shared in a single call
- The output of one agent is structured enough to be the verified input to another - not just raw text passed forward
- You have the observability infrastructure to trace a decision across agent boundaries
One challenge that will grow quickly is debugging across protocol boundaries once multiple agents, runtimes, and orchestration layers interact. Understanding one agent run end-to-end is already difficult before introducing cross-protocol coordination.
Start with one agent and MCP tool connections. Add A2A when your architecture genuinely grows into agents delegating to agents. The Stanford team started with five to eight agents structured to mirror a physical Stanford lab before scaling to 37,000. The hierarchy came first; the scale followed from it.
Multi-agent AI coordination: common questions
What is the difference between MCP and A2A?
MCP connects an agent to external tools and data sources - databases, APIs, file systems. A2A connects one agent to another agent, enabling task delegation, capability discovery, and multi-turn coordination across agent boundaries. Use MCP first for most single-agent workflows; add A2A when you have genuine multi-agent delegation to manage.
How many agents do you need before multi-agent coordination matters?
Two. The moment you have one agent handing a result to another and expecting it to act, you have a coordination problem. The Stanford system started at five to eight agents with explicit role definitions before it scaled. The role structure matters more than the count.
Is A2A production-ready?
A2A reached version 1.0 in April 2026 with over 150 organizations in production and SDKs across five languages, with native support in LangGraph, CrewAI, LlamaIndex, Semantic Kernel, AutoGen, and Google's ADK. It is production-ready for teams with the engineering capacity to manage the implementation complexity. The MCP alternative is lighter but requires building conversational state yourself.
What does the Stanford virtual biotech paper mean for product teams?
It is a proof of concept that role-specialized agent hierarchies can handle work at a scale no human team can match - and that the architecture (crisp roles, human gates, shared artifacts) is what makes that tractable. The drug-discovery context is specific, but the coordination pattern is not.
What breaks first in a multi-agent system?
Observability. Debugging across protocol boundaries once multiple agents, runtimes, and orchestration layers interact is already difficult before introducing cross-protocol coordination. Build tracing and logging into the coordination layer before you build the agents that use it.