Your agent calls a tool. The Kubernetes pod that handled the handshake restarts mid-call. The session ID was opaque to the load balancer, so every client had to be pinned to one pod - and when that pod died, all its in-flight tool calls died with it. That was the design of every MCP server built before July 28, 2026.
The highlight of the 2026-07-28 release is a stateless protocol core: MCP is transforming from a bidirectional stateful protocol into a request/response stateless protocol.
The specification was developed with the MCP Transports Working Group, co-founded by Google and Hugging Face. It is the biggest structural revision since Anthropic open-sourced the protocol in November 2024, and for teams running agents in production it changes the infrastructure picture more than any feature addition would.
What the stateful design actually cost in production
MCP has worked like this since launch: a client opens a connection, sends an initialize request, gets back an initialized response, and every request in the session carries an Mcp-Session-Id header that ties it to one specific handshake. If you were running more than one server instance behind a load balancer, that load balancer needed to route every request from the same client back to the same instance every time, or nothing worked.
Keeping that session alive in production meant one of three things: sticky sessions at the load balancer, a shared session store like Redis so any instance could serve any request, or a client-side retry strategy that tolerated session loss and re-initialized when a pod restarted.
Each solution adds a moving part that fails in a specific, familiar way during a rolling deployment: a pod gets terminated, its in-memory session state goes with it, and every client pinned to that pod either gets a connection reset mid-conversation or silently starts talking to a session that no longer exists.
Sticky sessions, where the balancer pins a client to the instance that started its session, undercut the whole point of load balancing by creating hot spots - and they do nothing for resilience since a dead instance still loses the session.
The cost was quiet and constant. Teams building agent workflows on top of MCP were writing infrastructure glue - Redis session stores, sticky ALB annotations, reconnect logic - just to paper over a protocol-level assumption that every connection lands on the same machine.
What the 2026-07-28 spec actually removes
The 2026-07-28 revision is different from the previous three: its headline changes are subtractions. Protocol-level sessions are gone. The initialize handshake is gone. Stream resumability is gone. So are ping and logging/setLevel.
What is left is a protocol that behaves like an ordinary HTTP API. Every request carries everything the server needs to handle it, so any request can land on any instance behind a plain round-robin load balancer, with no shared session store and no sticky routing.
The routing story has a useful detail that most coverage skips. Method and tool names now travel in the Mcp-Method and Mcp-Name HTTP headers, so gateways can route and authorize on headers directly.
That means a load balancer, gateway, or rate-limiter can route and throttle on the operation without opening the request body. Before, the session ID was opaque to infrastructure - load balancers had no idea what kind of call was in flight.
The spec also solves long-running interactions without reintroducing state. Previously, if a tool needed user confirmation mid-execution, the server had to hold open a bidirectional stream and initiate a reverse request, which required persistent connections and made serverless deployment impossible.
Server-to-client requests for things like sampling and elicitation are now redesigned to use Multi Round-Trip Requests (MRTR), removing the need for constantly open bidirectional streams.
One non-obvious consequence of the stateless move: because tools/list, resources/list, and prompts/list no longer vary per connection, you cannot use the session to serve a different tool catalog to different clients.
If your server did that, the capability has to move into the request itself or into your authorization layer. Teams that built per-client tool filtering into session state will need to rethink that design.
Enterprise-Managed Authorization: the companion change
The stateless update landed in late July. Six weeks earlier, on June 18, a separate but complementary piece went stable: Enterprise-Managed Authorization, or EMA.
Every MCP server used to mean one more OAuth consent screen per user. Enterprise-Managed Authorization - promoted to stable on June 18, 2026 - moves that decision to the organization's identity provider: an admin approves a server once, and every authorized employee inherits access automatically.
EMA lets a company's identity provider - Okta at launch - grant access to MCP servers in bulk, scoped to user groups and roles. End users open Claude or VS Code, sign in once, and inherit every MCP connector their admin already approved without seeing an OAuth screen.
Admins enable specific MCP servers in the IdP, set group scopes, and audit usage through the same IdP logs they use for SaaS. Reduced token lifetimes let admins deprovision a leaving employee from every MCP server at once.
For teams, the offboarding case alone is significant. Previously, removing an employee's agent access meant hunting down every MCP server they'd individually authorized. Now it's one IdP action. As of August 24, 2026, EMA is generally available, with connector coverage expanded to Datadog, Notion, and Slack.
Support for Microsoft Entra ID, Google Workspace, and Ping Identity remains in beta with waitlists available.
What this means for teams building with agents now
The stateless spec is backward-incompatible, but the migration window is real. The 2026-07-28 MCP spec makes the protocol stateless, removing the initialize/initialized handshake and the Mcp-Session-Id header, so any server instance can handle any request. Existing servers keep working during a 12-month window, but transport, auth, and deprecated features all need attention.
All four Tier 1 SDKs - TypeScript, Python, Go, and C# - speak the new spec as of launch day. TypeScript ships as two new packages, @modelcontextprotocol/client and @modelcontextprotocol/server, both at 2.0, while the old @modelcontextprotocol/sdk line continues for 2025-era servers.
For teams using Beagle or any agent that calls MCP servers inside Slack or Teams, the practical upshot is reliability: fewer dropped tool calls when pods cycle, no need to keep long-lived connections warm between morning standups. The agent's calls become short, discrete HTTP requests - the same shape as everything else in your stack.
The comparison nobody is making: MCP's session model was borrowed from its stdio origins, where a single local process handled everything. The old model was a design that did not survive contact with real HTTP infrastructure. The July 2026 spec finally aligns the protocol with how production HTTP actually works - which should have been the default from the start.
MCP stateless spec: common questions
What did the MCP 2026-07-28 spec change?
The 2026-07-28 spec transforms MCP from a bidirectional stateful protocol into a request/response stateless protocol. The initialize handshake and Mcp-Session-Id header are removed. Every request now carries its own protocol version and client identity, so any server instance can handle any request behind a standard load balancer.
Do I need to migrate my existing MCP server immediately?
No. Existing servers keep working during a 12-month window, but transport, auth, and deprecated features all need attention. The new SDKs (TypeScript 2.0, Python mcp-sdk 2.0, Go v2, C# 2.x) implement the new spec. You can migrate incrementally - clients fall back to the old handshake when they reach a server running an older revision.
What is MCP Enterprise-Managed Authorization (EMA)?
EMA - promoted to stable on June 18, 2026 - moves MCP server access decisions to the organization's identity provider: an admin approves a server once, and every authorized employee inherits access automatically.
Okta is the first supported identity provider; organizations using Okta can provision MCP access through Okta's Cross App Access (XAA).
Can MCP servers now run on serverless infrastructure?
Yes. Agent servers can now run behind standard round-robin load balancing and on serverless, without sticky-session workarounds. Until the July 28, 2026 spec, remote MCP relied on long-lived sessions. Load balancers required sticky affinity. Serverless functions, which go dormant after each request, were a poor fit. Both constraints are lifted.
What happens to server-to-client requests like sampling and elicitation?
Server-to-client requests like sampling, elicitation, and roots no longer ride an open stream.
They are redesigned to use Multi Round-Trip Requests (MRTR), removing the need for constantly open bidirectional streams. This is what makes serverless deployment viable for servers that need back-and-forth interaction with the model.