Why we cut our MCP server from 83 tools to 8
Our MCP server used to expose one tool per REST route. We replaced it with eight tools built around the browser loop, which cut the tool definitions a model carries on every turn from about 38k tokens to about 4.5k. Here is what we measured, what we changed, and what the smaller surface cost us.
Our first hosted MCP server was generated, not designed. FastMCP read our OpenAPI schema and turned every public route into a tool. It launched in June with 59 tools and grew with the API. The same generator produces 83 today.
On September 4 we replaced it at https://api.notte.cc/mcp with eight hand-written tools. The generated server still runs at /mcp/full for developers who want the whole API.
One tool per route looked like the obvious design
The reasoning is written down in our own code. The docstring for the generated server keeps hidden routes in because they "are still legitimate public CRUD that AI agents should be able to call." If a route was useful to a developer, we assumed it was useful to a model.
That assumption was the mistake. A developer reads the docs once, writes the orchestration once, and reuses it. A model starts every conversation from nothing and re-derives the plan from whatever list it is handed. Vercel made the same case in The second wave of MCP: tools should cover a complete user intention, not mirror API operations.
Look at what our list asked a model to sort out. To read a page it could pick page_scrape, which needs a running session, or scrape_webpage, which does not. Every browser tool took a session_id, a word users do not reach for when they mean a browser. Anthropic sees the same failure in its own evals: with large tool libraries, the most common errors are wrong tool selection and incorrect parameters, and similar names make it worse.
![]()
Generated tools by area of the API. Only the two green areas are still served by the curated server.
What 83 tools cost before the first message
Most MCP clients load every tool definition into context when a conversation starts and send it again on every request. We measured what the model actually reads from each server's tools/list: name, description and input schema.
![]()
Model-visible definitions per server, in tokens. The middle bar is a single generated tool.
A typical generated tool is small, around 200 tokens. The cost came from having 83 of them, plus one outlier: page_execute, whose generated schema covers every browser action, is larger on its own than the entire curated server.
For scale, Anthropic's advanced tool use post counts 58 tools across five common servers at about 55k tokens, and suggests on-demand tool search once definitions pass 10k tokens or ten tools. One Notte server was over both lines. The curated one is under both, so it needs no search layer.
Eight tools built around the browser loop
The new surface covers one job: drive a remote browser.
| Tool | What it does |
|---|---|
manage_browsers | Start, list, check or stop a browser |
browser_action | Navigate or interact, with 19 typed action variants |
browser_observe | Read the page, tabs and element IDs |
browser_scrape | Return markdown or structured data |
browser_screenshot | Return the viewport as an image |
manage_files | Upload files, list and fetch downloads |
manage_profiles | Create, reuse or delete persistent profiles |
manage_auth | Check or reauthenticate Managed Auth connections |
Four of them take a single action parameter that picks the operation. GitHub made the same move in its own server.
The server instructions describe one loop, and every tool maps onto a step of it:
The browser loop. Files, profiles and Managed Auth attach to it at start or during actions; they are not separate workflows.
Cutting the count was half the work. The other half was deciding what the model sees:
- Vocabulary. Tools and results say
browser_id, neversession_id, and internal fields such as the CDP URL are dropped from responses. - Output size. Observations are trimmed to what a model can act on. A test feeds in a 600,000-character screenshot and requires the result to stay under 4,000 characters.
- Server instructions. MCP lets a server ship instructions that work like a system prompt. Beyond the loop, ours cover stale element IDs, profile persistence and file uploads.
The budget is held by tests, not intentions. The server must expose exactly eight tools, their combined parameter schemas must stay under 12,000 characters, and names like session_id and function_create must not appear anywhere in them. Adding a ninth tool means changing a test on purpose.
The smaller surface also picks better. We ran a benchmark of our own, and tool selection improved on the curated server. Published work points the same way: Vercel cut its internal data agent from 17 tools to a shell and a SQL tool and went from 4 of 5 test queries to 5 of 5, 3.5x faster with 37% fewer tokens, and Anthropic's tool search took Claude Opus 4 from 49% to 74% on MCP evals with large libraries.
What it cost us
Functions, vaults, personas and agents are gone from the default server. They live at /mcp/full, and clients that depended on the old /mcp list had to change their URL. We took that breaking change deliberately: the default should serve the person who wants a browser, not the person administering an account.
browser_action is still 2,500 tokens, more than half the server. We chose one tool with a typed union over nineteen separate tools.
Code execution is a different path. Cloudflare's Code Mode has the model write code instead of calling tools one at a time. We went the other way on the default server: it leaves out running JavaScript in the page, which /mcp/full still exposes through page_execute.
If you run an MCP server, measure your own tools/list before arguing about it. Count only what the model sees and sort by size: the problem may be one tool, not the total. Then design for the task the user brings, not for your route table. Our setup is in the MCP server docs.
Frequently asked questions
How many tools should an MCP server expose?
There is no universal number, but fewer than you think. Anthropic suggests adding on-demand tool search once a server passes about ten tools or 10k tokens of definitions. Our browser server exposes eight tools in about 4,500 tokens, and a test fails if anyone adds a ninth without meaning to.
How did you decide which tools to keep?
We started from the loop a model runs to finish a browser task: start, navigate, observe, act, read, stop. A tool made the cut only if it served a step of that loop or attached to it, like files, profiles and Managed Auth. Everything else stayed on /mcp/full.
Does grouping operations behind an action parameter hurt tool selection?
It moves information from the tool name into the description, so the description has to list every action clearly. GitHub uses the same pattern in its MCP server. In our own benchmark, tool selection improved after we moved to the curated server.
Does this apply to local stdio MCP servers too?
Yes. The cost comes from tool definitions sitting in the model's context, and that happens the same way whether the server runs locally over stdio or remotely over HTTP. The transport changes deployment and latency, not how many tokens the tool list takes.
