Every developer who put an MCP server behind a load balancer in the last year built the same workaround: sticky sessions at the ingress layer, or a Redis store bolted on so every replica could read a shared session.
The failure mode was subtle - pod A handled initialize and issued the session ID, then the SDK's long-lived SSE stream silently hashed to pod B, which knew nothing about that session, and returned a 404. Teams spent days debugging what looked like a network issue when the real culprit was a protocol-level assumption that every connection lands on the same machine.
On July 28, 2026, that assumption left the spec.
The MCP team officially pushed the 2026-07-28 specification on July 28, and its headline change is a stateless protocol core - MCP transformed from a bidirectional stateful protocol into a request/response stateless protocol.
It was locked as a release candidate on May 21, 2026, and is the largest revision to the protocol since it launched.
Across the Tier 1 SDKs, the MCP team is seeing close to half a billion downloads a month, with both the TypeScript and Python SDKs crossing one billion cumulative downloads. At that scale, a protocol-level flaw in how sessions pinned connections to specific instances was going to cause real incidents. This spec fixes it structurally.
What the stateless core actually removes
The 2026-07-28 specification removes the initialize handshake and the Mcp-Session-Id: every request now stands on its own, and any request can land on any server instance behind a plain round-robin load balancer with no shared session store required.
The old design's cost was not theoretical. Sticky sessions undercut the whole point of load balancing by creating hot spots and did nothing for resilience once an instance died. The more robust option was externalizing session state into a shared store such as Redis - which solved scaling but added a piece of critical infrastructure to deploy, manage, and pay for, plus a network hop on every request to validate the session before any real work could start.
The 2026-07-28 specification removes the initialize/initialized handshake and the Mcp-Session-Id header that pinned a client to one server instance. Protocol version, client info, and capabilities now travel inline in a _meta field on each request. Any server instance can serve any request, so a remote MCP server can run behind a plain round-robin load balancer without sticky sessions or a shared session store.
Application state did not disappear - it moved somewhere more honest.
Tools now return explicit opaque handles (think a basket_id or job_id). The model can see these handles, reason about them, and pass them forward on subsequent calls. Your server stores actual state in Redis, PostgreSQL, or wherever makes sense.
Multi Round-Trip Requests and what they replace
The old spec's mechanism for a server to ask the client something mid-call - "confirm before deleting these files?" - required holding an SSE stream open back to the client. Delivering these requests meant the server held an SSE stream open back to the client. With the latest protocol version, server-initiated requests are now only permitted while the server is actively processing a client request.
The replacement is Multi Round-Trip Requests (MRTR).
The MRTR pattern replaces the previous approach of sending server-initiated requests such as roots/list, sampling/createMessage, or elicitation/create. Servers return an InputRequiredResult whose inputRequests field carries the requests for additional information, and clients respond with inputResponses on a retry of the original request.
Because requestState carries everything needed to resume, the retry can land on a completely different server instance and still pick up where it left off.
This matters beyond the architecture elegance. Server-initiated requests are now only permitted while the server is actively processing a client request. Every prompt a user sees traces back to something they or their agent started. That is a security property, not just a design preference.
Three things that just got better: headers, caching, and auth
Header-based routing.
Method and tool names now travel in the Mcp-Method and Mcp-Name HTTP headers, so gateways can route and authorize on headers directly.
Previously a WAF or rate limiter had to parse the JSON-RPC body to know what operation was happening. Now it is a header read.
Servers also reject requests where headers and body disagree, closing off a class of routing and security mismatches.
Cacheable list results.
Results returned by tools/list, prompts/list, resources/list, resources/read, and resources/templates/list now carry ttlMs and cacheScope fields. ttlMs is a freshness hint in milliseconds; cacheScope - "public" or "private" - controls whether shared intermediaries may cache the response.
List responses also carry cache hints and a deterministic order, so clients can cache tool catalogs and keep upstream prompt caches stable across reconnects. That second part is easy to miss: if your tool list shifts order on every reconnect, the tokenized version of the list is different each time, and you lose prompt-cache hits on the models that cache by prefix. A stable, cached tool list is quietly worth real inference cost savings.
A result marked cacheScope: "private" that gets cached and served across users is a cross-user data leak.
Check whatever sits between your clients and your MCP servers - gateways, proxies, CDNs - and confirm they understand cacheScope before you set ttlMs to anything non-zero on user-specific results.
Auth hardening. Authorization hardening includes RFC 9207 issuer validation and a formal shift away from Dynamic Client Registration (DCR) toward client metadata documents (CIMD). If your current MCP auth flow uses DCR, that is now deprecated with a twelve-month runway.
Deprecation policy. For the first time, MCP has a formal deprecation policy. Features move through a structured Active → Deprecated → Removed lifecycle with a minimum 12-month transition window.
Three features enter deprecation today: Roots (replaced by explicit tool parameters, resource URIs, or server configuration), Sampling (replaced by calling LLM provider APIs directly), and Logging (replaced by standard stderr for stdio connections, or OpenTelemetry for structured cloud observability).
Your migration checklist
This is not a preview you can wait out - 2026-07-28 is the current specification, and the Tier-1 SDKs for TypeScript, Python, Go, and C# already ship it. Here is what needs attention, roughly in order of urgency:
- Grep for
Mcp-Session-Idandinitialize- those are your migration starting points. Every handler that reads session state needs to switch to explicit handles passed in request parameters. - Add
Mcp-MethodandMcp-Nameheaders to Streamable HTTP requests. Required by the spec; servers reject requests where headers and body disagree. - Audit your elicitation flows - if you use server-initiated elicitation, migrate to
InputRequiredResultand the MRTR pattern. Legacy clients that do not advertise elicitation capability will receive a-32021error instead. - Set
cacheScopecorrectly before settingttlMs- default isttlMs: 0, cacheScope: "private", which is always safe. Only setcacheScope: "public"when the result is genuinely identical for all callers. - Plan off Roots, Sampling, and Logging - none break immediately, but new implementations should not adopt them. Budget the migration inside the twelve-month window, not at the end of it.
- Fix error code
-32002- it changed to-32602(Invalid Params, aligned with JSON-RPC). Hardcoded checks for the old code will silently pass. - Deprecate HTTP+SSE transport - Streamable HTTP is the path forward. The legacy transport has a year-long offramp.
A teammate like Beagle tracking your MCP server changelog in a shared channel can surface these migration items as they hit your SDK version - catching the cacheScope default flip before it reaches a gateway that does not understand the new field is exactly the kind of quiet problem that benefits from a logged, searchable thread.
MCP stateless spec: common questions
What does the MCP 2026-07-28 spec actually change?
The specification removes the initialize handshake and the Mcp-Session-Id header, making every request self-contained.
Any server instance behind a plain load balancer can handle any request. It also adds cacheable list results, header-based routing, Multi Round-Trip Requests to replace SSE-held streams, auth hardening, and a formal extensions framework - all shipped together on July 28, 2026.
Does migrating to 2026-07-28 mean rewriting my MCP server?
No full rewrite is needed.
Most servers need a focused checklist: upgrade the SDK, drop Mcp-Session-Id assumptions and use explicit handles for state, emit the new Mcp-Method and Mcp-Name headers, migrate experimental Tasks to the new lifecycle, add ttlMs to list responses, and harden auth.
Deprecated features have a twelve-month runway, so those moves can be scheduled separately.
What happens to my existing MCP clients built on the old spec?
The old 2025-11-25 spec keeps working, and the new deprecation policy guarantees at least a twelve-month overlap, so this is a migration you can plan rather than scramble through.
Tier-1 SDK maintainers preserved backward compatibility in their initial 2026-07-28 releases.
Legacy pre-2026-07-28 connections cannot receive an InputRequiredResult
, so client support for MRTR needs to be verified before you switch elicitation flows.
Are Roots, Sampling, and Logging gone now?
Not yet. All three are deprecated with a 12-month removal window. They still function. New implementations should not adopt them, and existing ones should plan migration: Roots → explicit tool parameters or resource URIs; Sampling → direct LLM provider API calls; Logging → stderr (stdio) or OpenTelemetry (HTTP).
What is the cacheScope field and why does it matter for security?
cacheScope on a list response is either "public" (safe to share across all callers) or "private" (belongs to one authorization context).
A result marked cacheScope: "private" that gets cached and served across users is a cross-user data leak.
The default is "private" with ttlMs: 0, so servers that do not touch these fields are safe - but any gateway or proxy caching list responses needs to understand cacheScope before ttlMs is set to anything non-zero.