Migrate Your MCP Server to the 2026-07-28 Stateless Spec

The MCP 2026-07-28 spec drops sessions and the initialize handshake entirely. Here's what breaks, what gets cheaper, and the one caching trap nobody warns you about.

Cover art for Migrate Your MCP Server to the 2026-07-28 Stateless Spec

Every developer who put an MCP server behind a load balancer in the last year built the same workaround: sticky sessions at the ingress layer, or a Redis store bolted on so every replica could read a shared session. The failure mode was subtle - pod A handled initialize and issued the session ID, then the SDK's long-lived SSE stream silently hashed to pod B, which knew nothing about that session, and returned a 404. Teams spent days debugging what looked like a network issue when the real culprit was a protocol-level assumption that every connection lands on the same machine. On July 28, 2026, that assumption left the spec.

The MCP team officially pushed the 2026-07-28 specification on July 28, and its headline change is a stateless protocol core - MCP transformed from a bidirectional stateful protocol into a request/response stateless protocol.

It was locked as a release candidate on May 21, 2026, and is the largest revision to the protocol since it launched.

Across the Tier 1 SDKs, the MCP team is seeing close to half a billion downloads a month, with both the TypeScript and Python SDKs crossing one billion cumulative downloads. At that scale, a protocol-level flaw in how sessions pinned connections to specific instances was going to cause real incidents. This spec fixes it structurally.

What the stateless core actually removes

The 2026-07-28 specification removes the initialize handshake and the Mcp-Session-Id: every request now stands on its own, and any request can land on any server instance behind a plain round-robin load balancer with no shared session store required.

The old design's cost was not theoretical. Sticky sessions undercut the whole point of load balancing by creating hot spots and did nothing for resilience once an instance died. The more robust option was externalizing session state into a shared store such as Redis - which solved scaling but added a piece of critical infrastructure to deploy, manage, and pay for, plus a network hop on every request to validate the session before any real work could start.

The 2026-07-28 specification removes the initialize/initialized handshake and the Mcp-Session-Id header that pinned a client to one server instance. Protocol version, client info, and capabilities now travel inline in a _meta field on each request. Any server instance can serve any request, so a remote MCP server can run behind a plain round-robin load balancer without sticky sessions or a shared session store.

Application state did not disappear - it moved somewhere more honest. Tools now return explicit opaque handles (think a basket_id or job_id). The model can see these handles, reason about them, and pass them forward on subsequent calls. Your server stores actual state in Redis, PostgreSQL, or wherever makes sense.

Scaling an MCP server past one replica
Without Beagle
sticky sessions at the ingress, a shared Redis session store, and a network hop per request to validate state before any tool runs
With Beagle
plain round-robin load balancing, self-contained requests with _meta inline, explicit handles for any state the tool needs to carry forward

Multi Round-Trip Requests and what they replace

The old spec's mechanism for a server to ask the client something mid-call - "confirm before deleting these files?" - required holding an SSE stream open back to the client. Delivering these requests meant the server held an SSE stream open back to the client. With the latest protocol version, server-initiated requests are now only permitted while the server is actively processing a client request.

The replacement is Multi Round-Trip Requests (MRTR). The MRTR pattern replaces the previous approach of sending server-initiated requests such as roots/list, sampling/createMessage, or elicitation/create. Servers return an InputRequiredResult whose inputRequests field carries the requests for additional information, and clients respond with inputResponses on a retry of the original request.

Because requestState carries everything needed to resume, the retry can land on a completely different server instance and still pick up where it left off.

This matters beyond the architecture elegance. Server-initiated requests are now only permitted while the server is actively processing a client request. Every prompt a user sees traces back to something they or their agent started. That is a security property, not just a design preference.

Beagle in action#engineering, reviewing an on-call incident from an MCP tool confirmation dialog
The ask
team asks why a destructive tool ran without the expected confirmation step
Beagle drafts
surfaces the server's MCP spec version, checks whether the client advertises elicitation capability, drafts a summary of the MRTR flow and what the client's missing support means
You approve
the team approves a one-paragraph explanation in the channel; the gap is documented before the next deploy
Do this in your workspace →

Three things that just got better: headers, caching, and auth

Header-based routing. Method and tool names now travel in the Mcp-Method and Mcp-Name HTTP headers, so gateways can route and authorize on headers directly. Previously a WAF or rate limiter had to parse the JSON-RPC body to know what operation was happening. Now it is a header read. Servers also reject requests where headers and body disagree, closing off a class of routing and security mismatches.

Cacheable list results. Results returned by tools/list, prompts/list, resources/list, resources/read, and resources/templates/list now carry ttlMs and cacheScope fields. ttlMs is a freshness hint in milliseconds; cacheScope - "public" or "private" - controls whether shared intermediaries may cache the response.

List responses also carry cache hints and a deterministic order, so clients can cache tool catalogs and keep upstream prompt caches stable across reconnects. That second part is easy to miss: if your tool list shifts order on every reconnect, the tokenized version of the list is different each time, and you lose prompt-cache hits on the models that cache by prefix. A stable, cached tool list is quietly worth real inference cost savings.

A result marked cacheScope: "private" that gets cached and served across users is a cross-user data leak. Check whatever sits between your clients and your MCP servers - gateways, proxies, CDNs - and confirm they understand cacheScope before you set ttlMs to anything non-zero on user-specific results.

Auth hardening. Authorization hardening includes RFC 9207 issuer validation and a formal shift away from Dynamic Client Registration (DCR) toward client metadata documents (CIMD). If your current MCP auth flow uses DCR, that is now deprecated with a twelve-month runway.

Deprecation policy. For the first time, MCP has a formal deprecation policy. Features move through a structured Active → Deprecated → Removed lifecycle with a minimum 12-month transition window.

Three features enter deprecation today: Roots (replaced by explicit tool parameters, resource URIs, or server configuration), Sampling (replaced by calling LLM provider APIs directly), and Logging (replaced by standard stderr for stdio connections, or OpenTelemetry for structured cloud observability).

~500MMCP SDK downloads/monthacross all Tier-1 SDKs at July 2026 release
1B+cumulative downloadsTypeScript and Python SDKs each, crossed at release
12 monthsminimum deprecation windowfor Roots, Sampling, Logging, and HTTP+SSE transport
4Tier-1 SDKs updatedTypeScript, Python, Go, C# - all shipped 2026-07-28 support by launch day

Your migration checklist

This is not a preview you can wait out - 2026-07-28 is the current specification, and the Tier-1 SDKs for TypeScript, Python, Go, and C# already ship it. Here is what needs attention, roughly in order of urgency:

  • Grep for Mcp-Session-Id and initialize - those are your migration starting points. Every handler that reads session state needs to switch to explicit handles passed in request parameters.
  • Add Mcp-Method and Mcp-Name headers to Streamable HTTP requests. Required by the spec; servers reject requests where headers and body disagree.
  • Audit your elicitation flows - if you use server-initiated elicitation, migrate to InputRequiredResult and the MRTR pattern. Legacy clients that do not advertise elicitation capability will receive a -32021 error instead.
  • Set cacheScope correctly before setting ttlMs - default is ttlMs: 0, cacheScope: "private", which is always safe. Only set cacheScope: "public" when the result is genuinely identical for all callers.
  • Plan off Roots, Sampling, and Logging - none break immediately, but new implementations should not adopt them. Budget the migration inside the twelve-month window, not at the end of it.
  • Fix error code -32002 - it changed to -32602 (Invalid Params, aligned with JSON-RPC). Hardcoded checks for the old code will silently pass.
  • Deprecate HTTP+SSE transport - Streamable HTTP is the path forward. The legacy transport has a year-long offramp.

A teammate like Beagle tracking your MCP server changelog in a shared channel can surface these migration items as they hit your SDK version - catching the cacheScope default flip before it reaches a gateway that does not understand the new field is exactly the kind of quiet problem that benefits from a logged, searchable thread.

Beagle in action#platform-eng, after upgrading the MCP Python SDK to 2.2
The ask
'did the list caching change anything about how our user-specific tool results work?'
Beagle drafts
reads the 2026-07-28 changelog and the cacheScope docs, drafts a reply explaining the default (ttlMs: 0, cacheScope: private), and flags that the team's gateway does not yet parse cacheScope
You approve
approved and posted; the gateway config gets a ticket before a private result leaks
Do this in your workspace →

MCP stateless spec: common questions

What does the MCP 2026-07-28 spec actually change?

The specification removes the initialize handshake and the Mcp-Session-Id header, making every request self-contained. Any server instance behind a plain load balancer can handle any request. It also adds cacheable list results, header-based routing, Multi Round-Trip Requests to replace SSE-held streams, auth hardening, and a formal extensions framework - all shipped together on July 28, 2026.

Does migrating to 2026-07-28 mean rewriting my MCP server?

No full rewrite is needed. Most servers need a focused checklist: upgrade the SDK, drop Mcp-Session-Id assumptions and use explicit handles for state, emit the new Mcp-Method and Mcp-Name headers, migrate experimental Tasks to the new lifecycle, add ttlMs to list responses, and harden auth. Deprecated features have a twelve-month runway, so those moves can be scheduled separately.

What happens to my existing MCP clients built on the old spec?

The old 2025-11-25 spec keeps working, and the new deprecation policy guarantees at least a twelve-month overlap, so this is a migration you can plan rather than scramble through. Tier-1 SDK maintainers preserved backward compatibility in their initial 2026-07-28 releases. Legacy pre-2026-07-28 connections cannot receive an InputRequiredResult , so client support for MRTR needs to be verified before you switch elicitation flows.

Are Roots, Sampling, and Logging gone now?

Not yet. All three are deprecated with a 12-month removal window. They still function. New implementations should not adopt them, and existing ones should plan migration: Roots → explicit tool parameters or resource URIs; Sampling → direct LLM provider API calls; Logging → stderr (stdio) or OpenTelemetry (HTTP).

What is the cacheScope field and why does it matter for security?

cacheScope on a list response is either "public" (safe to share across all callers) or "private" (belongs to one authorization context). A result marked cacheScope: "private" that gets cached and served across users is a cross-user data leak. The default is "private" with ttlMs: 0, so servers that do not touch these fields are safe - but any gateway or proxy caching list responses needs to understand cacheScope before ttlMs is set to anything non-zero.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle