Four MCP servers - Linear, Notion, Slack, Postgres - loaded into a single agent session. Zero messages sent. Tool definitions alone consumed over 21,000 tokens, 10.5% of a Claude 200K context window and 16.5% of GPT-4o's 128K context. That overhead is paid on every conversation turn, not just the first. The MCP ecosystem now has a success problem: the more servers you connect, the worse your agent performs.
This is the argument the 2026-07-28 stateless spec did not fix. It fixed session management and horizontal scaling. The token tax is still fully in effect, and the community is only just starting to measure it.
Why MCP tool definitions are so expensive by design
MCP requires every exposed tool to carry its full JSON Schema: name, description, parameter types, field descriptions, enum values. That is not a bug - it is how the model knows what to call. MCP tool definitions consume 5-15× more tokens than the simplest possible schema for the same tool.
For a single tool like ctx_batch_execute, the field descriptions add ~400 tokens, type definitions add ~300, and nested object structures add ~300 more - roughly 1,000 tokens for one tool that a human could describe in two sentences.
Scale that to a realistic enterprise agent. A single well-documented tool might consume 200-500 tokens. Load 50 tools - which is common in enterprise setups - and you've spent 10,000-25,000 tokens just on definitions, before any tool has been called. Some popular MCP packages make this dramatically worse: GitHub's MCP server alone costs 55,000 tokens in tool definitions.
Connect three full-featured services simultaneously and you could be burning 143,000 of your 200,000 token context on schema definitions alone - 71% of your context window gone before your agent has read a single line of actual work.
The structural inefficiency is that tool schemas are loaded eagerly.
Linear ships 42 tool definitions (~12,807 tokens), even if you only ever use get_issue and save_issue.
Some popular MCP server packages for GitHub, Notion, and Salesforce expose 40-80 tools by default. Most agents only use a fraction of those in any given session.
The stateless spec fixed the infrastructure layer, not the context layer
With MCP going stateless, the maintainers are able to handle production traffic more reliably, securely, and at scale, without impractical infrastructure workarounds. That is a real improvement. When MCP servers went remote, it transposed the stateful connection that worked well locally onto web infrastructure - building a well-behaved MCP server meant managing request routing to sticky sessions, holding open streams, and message replay. Stateless HTTP is the right call for any team that wants to put MCP behind a load balancer.
But the stateless spec does nothing to reduce the volume of tool definitions your context window has to absorb. That problem lives one layer up, in how servers expose their capabilities. The two problems are orthogonal.
One partial fix has shipped: Claude Code's Tool Search with Deferred Loading, which loads MCP tool schemas on-demand and reduces context usage by 85%+. The context bloat described above is largely addressed for users on current Claude Code versions. That is meaningful - but it is a client-side mitigation, not a protocol-level fix. Teams running other hosts (VS Code, custom agent frameworks, any non-Anthropic client) still absorb the full schema payload on every turn.
The supply chain problem that compounds the cost
While teams are busy optimising token budgets, a second problem sits underneath: most of the servers they are pulling in are unverified. With over 3,500 servers, Smithery is a popular hub that allows community contributions. A sample of 847 servers showed that only 8% carry an "official" badge. The remaining 92% are unverified, creating a large surface area for potential impersonation.
Endor Labs analyzed 2,614 MCP servers and found 82% prone to path traversal and 67% to code injection. The NSA published its own advisory in May 2026, noting that MCP's rapid proliferation has outpaced the development of its security model - much like early web protocols, MCP was released with a flexible and underspecified design, allowing implementers freedom but also introducing ambiguity for safe usage.
The typosquatting exposure illustrates the blast radius. OX Security researchers
crafted a proof-of-concept payload targeting the popular mcp-server-postgres package, creating a clone named mcp-server-postgress (double 's').
Nine out of 11 major MCP directories accepted and published the squatted payload without a single automated security review, source-code analysis, or publisher verification check. And because agents can auto-install tools during a session, had this been a genuine APT attack, the combination of AI-driven development and unverified registries would have compromised thousands of developer machines within hours.
What you can actually do about it this week
The right response is not to abandon MCP. The TypeScript SDK has over 34,700 dependent projects, and MCP is becoming invisible infrastructure for most developers. Walking away from it is not a realistic choice once your stack depends on it. The realistic choices are:
- Audit before you add. Run
mcp-tokens analyzeagainst any server before connecting it. The GitHub action wrapper makes this part of your CI pipeline. - Prefer minimal servers. The official MCP registry has ~1,000 servers with stricter review; MCP.so has 17,000+ servers that are unmoderated and high-risk. Start from the official list.
- Scope what you expose. Most agent tasks use 2-5 tools. Build or configure servers that expose only what the agent needs for a given workflow - not every tool the underlying API supports.
- Separate high-tool-count servers from general sessions. If your Salesforce integration exposes 80 tools, run it as a specialist sub-agent via A2A, not as a permanent attachment to your orchestrator's context.
- Pin and sign.
Carefully verify package names before installation - typosquatting is a common attack vector.
Treat your MCP server list the way you treat your
package-lock.json.
The steelman for keeping large server manifests is real: generic agents need broad capability, and hand-curating tool sets for every workflow adds operational overhead. If you are building a research assistant that genuinely needs 60 tools across four services, there is no clean answer. But most production agent tasks are narrow, and narrow tasks deserve narrow servers.
The non-obvious second-order problem
Here is the thing the token-cost discussion mostly misses: context bloat from tool definitions does not just raise cost. It degrades reasoning. Models lose coherence on complex tasks as the context fills up - and a context pre-loaded with 20,000 tokens of JSON schemas before the first user message arrives is not starting from a clean state.
The MCP community has been measuring tokens. Fewer people have been measuring task accuracy as tool count scales. That experiment is worth running. A hypothesis: agent error rates on multi-step tasks increase non-linearly as registered-tool count crosses ~30, independent of whether those tools are ever called. That would make tool-count a silent performance variable that no benchmark currently captures.
If that holds, the next urgent thing the spec needs is not another transport change. It is a standard for lazy discovery - a way for agents to declare "I want to know what this server can do, but only load the schemas I actually request." A2A already does something like this at the agent level;
any A2A server publishes an AgentCard at /.well-known/agent-card.json declaring its skills, supported MIME types, transport bindings, and security schemes.
MCP could learn from that pattern at the tool level.
MCP tool definitions: common questions
How many tokens does an MCP server consume?
It varies widely by server design.
A single well-documented tool might consume 200-500 tokens.
A minimal server like a single-tool Postgres wrapper runs under 100 tokens total.
GitHub's MCP server reaches 55,000 tokens in definitions alone.
Measure with mcp-tokens analyze before connecting any server to a production agent.
Does the 2026-07-28 stateless MCP spec fix the token overhead problem?
No. The chief change in the new revision removes protocol-level sessions, allowing each request to be handled independently - a stateless approach common in cloud-native infrastructure. Tool schema loading is a separate layer entirely. The stateless spec improves scalability and operational simplicity; it does not reduce what the context window must absorb.
What is the safest MCP server registry to pull from?
GitHub's curated official list contains only 57 official entries from established service providers. The threat of an attacker-controlled server slipping into this list is very low, making it a safe, though limited, reference point. For anything beyond that list, treat each server as an unreviewed dependency and audit it accordingly.
Can I reduce MCP tool overhead without changing the protocol?
Yes. Claude Code's deferred loading cuts schema overhead by 85%+ on supported clients. For other hosts, the practical options are scoping server manifests to the tools your agent actually uses, running tool-heavy servers as dedicated sub-agents, and using lazy-loading middleware that injects schemas only when the model requests a specific tool. None of these are as clean as a protocol-level fix, but all of them move the needle measurably.