Your Agent Burns 55,000 Tokens on Tool Descriptions Before You Type a Word
Connected MCP servers can eat a quarter of your context window in tool definitions alone. Here is how to measure it and what to cut.
Kemal EsensoyĀ·Modified on October 3, 2026
Anthropic published a five-server example that I keep coming back to. GitHub, Slack, Sentry, Grafana and Splunk connected to one agent: 58 tools, roughly 55,000 tokens of tool definitions. That is 27% of a 200K context window spent before the agent has read a single word of your request. Add Jira and the same post says you approach 100,000 tokens. In their own internal testing, one setup was carrying 134,000 tokens of tool definitions.
I build Claude Code skills and agents for clients, and I use these tools every working day. For a long time I never once thought about MCP token usage. I treated the servers I had connected as free. Connect a server, get capabilities, done. They are not free. They are a standing charge on every single turn of every single conversation, and most of what you are paying for is a JSON schema the model will never call.
What 55,000 Tokens Actually Buys You
The numbers vary wildly by server, which is the first thing worth internalising. A public benchmark that measured nine popular servers over stdio on 2026-07-08, using the o200k_base tokenizer, found a 25x efficiency spread. The official Notion server: 24 tools, 17,161 tokens, 8.6% of a 200K window. Firecrawl: 26 tools, 16,565 tokens. Slack: 8 tools, 679 tokens. The archived npm GitHub server: 26 tools, 3,546 tokens. Same protocol, same idea, two orders of magnitude apart in what they cost you to keep plugged in.
Other measurements of the current official GitHub server put it far higher, in the 17,000 to 42,000 token range depending on which toolsets are enabled. That spread is not sloppy measurement. It is the toolset flags doing exactly what they are supposed to do, and it tells you where the lever is.
The uncomfortable part is that this cost is not a one-time load. It is not like a large file you read once and then it scrolls out of relevance. Tool definitions sit in the request on every turn, because the model has to be able to choose a tool at any point in the conversation.
Why the Whole Schema Ships on Every Turn
Tools are not code the model calls. They are text the model reads. Every tool you connect becomes a name, a description, and a full JSON Schema for its input, serialised into the request alongside your system prompt and your conversation. The model does not have a directory it can look things up in. It has one flat list, in the prompt, every time.
So when you connect a server with 35 tools and then ask the agent to rename a variable, the agent still reads all 35 tool schemas. It reads the pagination parameters for the issue search tool. It reads the enum of valid sort orders for the pull request lister. Then it renames your variable. This is what people mean when they say MCP token usage eats your context window: not that the tools are chatty when they run, but that they are expensive when they sit still. If the mechanics of the protocol are new to you, I wrote a longer piece on Model Context Protocol (MCP): Everything you need to know that covers how the client and server actually talk.
Prompt caching softens the money side of this, since a stable tool block can be cached and reread cheaply. It does nothing for the context side. Cached or not, those 55,000 tokens are still occupying the window.
The Bill Is Not Mostly Descriptions
Here is the finding that changed how I write tools. In that same benchmark, 97% of Notion's token cost sat in inputSchema, not in the human-readable descriptions. The prose was fine. The parameter definitions were enormous.
The team redesigned it as 7 workflow-level tools instead of 24 endpoint-level ones and got the same practical coverage for 773 tokens. From 17,161 to 773, a 95.5% cut, by changing what a tool is rather than by trimming adjectives.
That is the actual lesson. Most MCP servers are a thin wrapper around a REST API, one tool per endpoint, one schema per tool, every optional query parameter faithfully reproduced. That design is honest and it is easy to generate, and it is why your context window is full. A tool that says "create a page in this database with this title and this content" costs a fraction of a tool that exposes every field the Notion API accepts, and it is the one your agent actually needed.
The Real Damage Is Tool Selection
If the cost were only tokens, you could buy your way out with a bigger window. The worse problem is that the model gets less accurate at choosing.
Anthropic's own evaluation gives the cleanest evidence, because it holds everything else constant and only changes whether tools are loaded upfront or searched for on demand. On their MCP evaluation, Opus 4 went from 49% to 74% accuracy. Opus 4.5 went from 79.5% to 88.1%. The tools did not change. The model did not change. Only the number of tool definitions sitting in context at once changed, and a quarter of the failures disappeared.
There is a much more dramatic figure circulating, from work associated with the Berkeley Function Calling Leaderboard, where accuracy on calendar scheduling tasks falls from 43% to 2% as the tool list grows from 4 to 51. I have only seen that number through secondary write-ups, not at the source, so treat it as directional rather than as a fact you quote in a meeting. The direction is not controversial though. Practitioners consistently report that quality starts sliding somewhere past 15 to 20 tools in active rotation.
What this looks like in practice is not the agent saying it is confused. It never says that. It picks a plausible neighbour. It calls the search tool when you wanted the get tool. It calls the read-only variant of the thing you asked it to write. You read the transcript afterwards and the wrong call looks almost reasonable, which is exactly why it is hard to catch. I have written before about which Claude model I actually use for coding work, and I will say this: a smaller model with five well-chosen tools beats a bigger model drowning in sixty.
How to Measure Your Own Overhead
Do not take anyone's numbers, including these. Measure your own MCP token usage, because your server versions and your toolset flags decide the answer.
In Claude Code, /context breaks your window down by system prompt, memory files, MCP tools and free space. One caveat that matters: earlier versions overstated MCP usage badly, because they measured each tool in a separate request and so counted the shared system overhead once per tool. One developer measured XcodeBuildMCP at roughly 14,000 tokens directly while /context reported about 45,000. That was fixed in January 2026, so make sure you are current before you panic at the number.
The precise method is the token counting endpoint. POST /v1/messages/count_tokens takes the same body as a normal message request, including the tools array, and returns the input token count. Send it your tool list with a one-word message, then send it again with an empty tool list, and the difference is your standing charge. Same model ID as you use in production, since counts are model-specific.
The offline method, if you just want a ranking: start each server, run the tools/list handshake, serialise the response to JSON, and count with a tokenizer. That is exactly what the nine-server benchmark did, and it is a twenty-line script.
What to Cut First
Start with the servers you have not called in a week. Not the ones you do not use, the ones you do not call. There is a difference, and the second list is longer than you expect.
Then use the toolset flags. The official GitHub server takes GITHUB_TOOLSETS to enable only the groups you need, or --dynamic-toolsets (GITHUB_DYNAMIC_TOOLSETS=1 in Docker) so the agent can turn groups on when a prompt calls for them. Enabling issues and pull requests instead of everything is the single highest-yield edit most people can make, and it takes one line of config.
After that, split by task rather than by capability. A sub-agent with four tools and a narrow brief will beat one general agent with forty, and it will cost less per run too. This is the same instinct behind most of the AI tooling that runs my one-person agency: narrow, purpose-built, disposable.
The trade-offs are real, so be honest about them. Filtering by hand means somebody maintains the allowlist, and it goes stale. Lazy loading spends a round trip to find a tool before it can call it, which costs latency. Sub-agents cost orchestration complexity and their own context. Proxying many servers behind one gateway centralises the trimming, and centralises the failure. None of these are free. They are all cheaper than the status quo.
What Shipped and What Is Still a Workaround
Real things have shipped. Anthropic's Tool Search Tool is the big one: you mark tools defer_loading: true, they stay discoverable through regex or BM25 search, and only the matched ones enter context. Their published example goes from about 77,000 tokens of upfront definitions to about 8,700, preserving 191,300 tokens of usable window instead of 122,800. That is the 85% reduction everyone quotes, and it comes with the accuracy gains above. Claude Code applies the same idea automatically once your MCP tool descriptions cross roughly 10,000 tokens.
Programmatic tool calling is the other half. Instead of tool results flowing back through the model, the model writes code that calls tools and only surfaces what matters. Anthropic measured average consumption on complex research tasks dropping from 43,588 to 27,297 tokens, a 37% cut, with accuracy on their GIA benchmark going from 46.5% to 51.2%.
The protocol itself has moved more cautiously. The 2026-07-28 specification is the largest revision since launch: a stateless core, self-describing requests, an optional discovery call, and clients now able to cache tools/list responses for as long as the server's ttlMs allows. That is a real improvement for infrastructure and for repeat cost. It is not a lazy loading mechanism. Deferred loading is still a client feature, which means it works where your client implements it and nowhere else. If you have ever wondered why the same MCP setup behaves differently in two tools, that is why. It is also part of why I keep a local model in the loop for some work, where I control the whole request and know exactly what is in it.
What I Actually Do About It
I keep three or four servers connected at a time, not fifteen. I split work into sub-agents with small toolsets. When I build a server for a client, I write workflow-level tools rather than mirroring their API endpoint for endpoint, because the Notion result convinced me that is where the money is.
I do not have a tidy rule for when a server earns its place. The honest heuristic I use is: if I cannot remember the last time the agent called it, it goes. That is not rigorous. It has been right more often than not.
What I am fairly confident about is that this stops being a problem you manage by hand within a year or two, once deferred loading is universal instead of client-specific. Until then it is a config problem you own, and it is worth twenty minutes with /context and the token counting endpoint. If you are also weighing what all this costs to run, I went through the break-even math on local hardware versus cloud APIs separately.
If you are building agents for a business and the tool sprawl has got away from you, that is the kind of thing I untangle for clients at Wunderlandmedia. Usually the fix is smaller than people expect.
Find these posts useful? Mark Wunderlandmedia as a preferred source on Google ā my articles will then show up more often in your Search results, AI Overviews and AI Mode.
Set as preferred sourceAbout the Author
Kemal Esensoy
Kemal Esensoy, founder of Wunderlandmedia, started his journey as a freelance web developer and designer. He conducted web design courses with over 3,000 students. Today, he leads an award-winning full-stack agency specializing in web development, SEO, and digital marketing.