As we connected more MCP servers to our agents, we noticed a problem before the agents had even called a tool: their context was already filling up. Loading complete tool catalogs meant including every tool’s name, description and input schema upfront. Those definitions consumed input tokens even when the tools were irrelevant to the request, leaving less room for the conversation and the data the agent actually needed.
Our experience matched what we were reading in publications such as Anthropic’s articles on code execution with MCP and advanced tool use. In the latter, Anthropic described five MCP servers exposing 58 tools whose definitions occupied roughly 55,000 tokens before the conversation began. That figure comes from their configuration, but it illustrates the problem we wanted to address: making more tools available could carry a substantial context cost before any useful work happened.
There was also a practical reason to manage those connections centrally. Configuring the same MCP servers for every new agent meant repeating setup and maintaining it in several places. An MCP gateway lets us configure those integrations once and connect each new agent to a single endpoint, through which it can access the tools available to it.
Together, these concerns shaped our work on MCP discovery in Aixy. We wanted to keep a growing collection of integrations available through the gateway while bringing only the tool definitions needed for the current task into the model’s context. We explored tool selection and response encoding, then measured their effects separately.
Selection and encoding are different problems
The MCP tool contract contains the information needed to call a tool. When a client loads all those definitions upfront, every schema contributes input tokens. We saw two places to intervene: choose which definitions enter the conversation, and choose how to encode the selected information. Each change has its own baseline; their saving percentages cannot be added.
How we built discovery
Our MCP endpoint exposes four stable helpers. The agent searches for a capability and receives a small selection of action names and descriptions, with complete input schemas when they fit the response budget. Otherwise, it retrieves the chosen action’s exact schema before execution. The full catalog stays available to the system while the model carries the definitions needed for its next step.
A separate mcp-tool-search service handles retrieval and response formatting; the gateway forwards its reply over the existing MCP interface. Search uses BM25S with Lucene scoring, k1=1.2 and b=0.75. We build a local index from names, descriptions, schemas and catalog metadata, with extra weight on names. Queries normalize naming conventions, accents and case, and expand a small English and Spanish vocabulary. Discovery makes no embedding-model request.
Ranking prioritizes exact names, overlap with the requested capability and compatible read or write operations, with BM25 breaking ties. Debugging also showed us that a matching tool could leave the agent missing the identity or workspace needed to use it. When the catalog declares those dependencies, the reply can include context helpers alongside the primary action.
Returned schemas retain required fields, types, limits and accepted values. Internal identifiers and definition revisions stay in the session. Before execution, the gateway rechecks permissions, connection state and definition identity. A smaller discovery reply still has to preserve the contract needed to make a valid call.
Discovery formatting belongs to the search service
The client carries its MCP session context. The search service selects and formats discovery; the gateway checks access and executes.
Agent client → Aixy gateway
aixy_search_toolssends a capability query.Gateway → Tool search service
One
DiscoverToolsRPC for a ready catalog.Tool search service
Ranks relevant actions, selects complete schemas when useful, binds semantic names and serializes compact JSON.
Search service → Gateway → Agent client
The gateway forwards the bounded discovery text unchanged.
Agent client
Chooses a semantic name, fetches the exact schema if needed and obtains approval when required.
Agent client → Gateway
aixy_execute_toolsubmitsnameand JSONarguments.Gateway → Search service
Resolves the name against the caller's selection context and obtains the native definition.
Gateway
Rechecks permissions, connection, allowlist and definition revision. A stale or expired selection requires rediscovery.
Gateway → Connected MCP server
Runs the requested native action.
Connected MCP server → Gateway → Client
Returns the native result through the existing protocol.
Agent client
Continues the model loop.
How we measured discovery
On 6 October 2026, we captured synthetic catalogs and responses from the real DiscoverTools handler in a temporary SQLite-backed local service. On 7 October, we recaptured the same native schemas with the current JSON implementation and all four gateway helpers. Counts use tiktoken 0.14.0 with o200k_base and cl100k_base; the offline TOON control uses toon-format 1.1.0. The tables link to both dated captures and their methods.
Initial context
The current initial comparison uses two synthetic catalogs of 100 tools, with flat and nested schemas. Both requests contain the same user message. One includes all native definitions; the other includes the four literal gateway helpers. The original 6 October control contained only search and execute: its earlier “four helpers” label was incorrect. That control remains available with its original files and counts.
Initial serialized request · discovery only
Current JSON capture · 7 October 2026
Local controlled fixtures · serialized input counts · four gateway helpers
Tokenizer: o200k_base
| Initial context | Eager JSON | Four helpers | Reduction |
|---|---|---|---|
| 100 flat schemas | 10,022 | 544 | 94.57% less |
| 100 nested schemas | 18,622 | 544 | 97.08% less |
Tokenizer: cl100k_base
| Initial context | Eager JSON | Four helpers | Reduction |
|---|---|---|---|
| 100 flat schemas | 9,719 | 538 | 94.46% less |
| 100 nested schemas | 17,919 | 538 | 97% less |
Original two-helper control · 6 October 2026
Local controlled fixtures · serialized input counts · two discovery/execution helpers
Tokenizer: o200k_base
| Initial context | Eager JSON | Two helpers | Reduction |
|---|---|---|---|
| 100 flat schemas | 10,022 | 428 | 95.73% less |
| 100 nested schemas | 18,622 | 428 | 97.7% less |
Tokenizer: cl100k_base
| Initial context | Eager JSON | Two helpers | Reduction |
|---|---|---|---|
| 100 flat schemas | 9,719 | 421 | 95.67% less |
| 100 nested schemas | 17,919 | 421 | 97.65% less |
Label corrected: these frozen requests contain search and execute only. The original counts and files are unchanged.
With o200k_base, the full-catalog requests use 10,022 and 18,622 tokens, compared with 544 for the four helpers: 94.57% and 97.08% less initial context. Selected schemas, later searches and tool results still add tokens as the task progresses.
A complete simulated sequence
The current sequence comparison uses eight focused searches returning 1, 2, 5 or 10 candidates, plus two browse pages. We count every serialized input request, repeated history and schema lookup. The baseline reconstructs the previous verbose JSON path, adding the same two connection helpers to both paths. Search and execution contracts can differ. Model steps use a fixed successful result.
Complete simulated input sequences
Current JSON capture · 7 October 2026
Local controlled fixtures · serialized input counts · four gateway helpers
Tokenizer: o200k_base
| Search | Previous sequence | Current JSON | Reduction |
|---|---|---|---|
| flat · 1 candidate | 2,643 | 2,016 | 23.72% less |
| flat · 2 candidates | 3,185 | 2,244 | 29.54% less |
| flat · 5 candidates | 4,867 | 2,212 | 54.55% less |
| flat · 10 candidates | 7,708 | 2,442 | 68.32% less |
| nested · 1 candidate | 2,908 | 2,218 | 23.73% less |
| nested · 2 candidates | 3,744 | 2,648 | 29.27% less |
| nested · 5 candidates | 6,218 | 2,414 | 61.18% less |
| nested · 10 candidates | 10,391 | 2,644 | 74.55% less |
| flat · 10 browse summaries * | 7,704 | 3,418 | 55.63% less |
| nested · 10 browse summaries * | 10,387 | 3,620 | 65.15% less |
Tokenizer: cl100k_base
| Search | Previous sequence | Current JSON | Reduction |
|---|---|---|---|
| flat · 1 candidate | 2,622 | 1,995 | 23.91% less |
| flat · 2 candidates | 3,164 | 2,223 | 29.74% less |
| flat · 5 candidates | 4,846 | 2,191 | 54.79% less |
| flat · 10 candidates | 7,687 | 2,421 | 68.51% less |
| nested · 1 candidate | 2,884 | 2,197 | 23.82% less |
| nested · 2 candidates | 3,718 | 2,627 | 29.34% less |
| nested · 5 candidates | 6,190 | 2,393 | 61.34% less |
| nested · 10 candidates | 10,362 | 2,623 | 74.69% less |
| flat · 10 browse summaries * | 7,683 | 3,388 | 55.9% less |
| nested · 10 browse summaries * | 10,358 | 3,590 | 65.34% less |
* Browse cases include a fourth request for the schema; all focused cases now use three. Both paths include the same two connection helpers; search and execute contracts differ. Counts include repeated history.
Original TOON experiment · 6 October 2026
Local controlled fixtures · serialized input counts · two discovery/execution helpers
Tokenizer: o200k_base
| Search | Previous sequence | TOON experiment | Reduction |
|---|---|---|---|
| flat · 1 candidate | 2,298 | 1,702 | 25.94% less |
| flat · 2 candidates | 2,840 | 1,978 | 30.35% less |
| flat · 5 candidates | 4,522 | 1,934 | 57.23% less |
| flat · 10 candidates | 7,363 | 2,204 | 70.07% less |
| nested · 1 candidate | 2,563 | 1,942 | 24.23% less |
| nested · 2 candidates * | 3,399 | 2,673 | 21.36% less |
| nested · 5 candidates | 5,873 | 2,174 | 62.98% less |
| nested · 10 candidates | 10,046 | 2,444 | 75.67% less |
Tokenizer: cl100k_base
| Search | Previous sequence | TOON experiment | Reduction |
|---|---|---|---|
| flat · 1 candidate | 2,274 | 1,678 | 26.21% less |
| flat · 2 candidates | 2,816 | 1,954 | 30.61% less |
| flat · 5 candidates | 4,498 | 1,910 | 57.54% less |
| flat · 10 candidates | 7,339 | 2,180 | 70.3% less |
| nested · 1 candidate | 2,536 | 1,918 | 24.37% less |
| nested · 2 candidates * | 3,370 | 2,639 | 21.69% less |
| nested · 5 candidates | 5,842 | 2,150 | 63.2% less |
| nested · 10 candidates | 10,014 | 2,420 | 75.83% less |
* The original nested two-candidate case includes a fourth request for the schema. This historical control uses two helpers. Counts include repeated history.
The one- and two-candidate cases now use 23.72–29.54% less serialized input against that previous path. Favorable ten-candidate searches reach 68.32% and 74.55%, because the old path repeats every full schema while discovery carries a primary schema and summaries. These figures measure the combined changes. The table also retains the original two-helper TOON experiment; its percentages have a different denominator.
The TOON experiment
We first tried Token-Oriented Object Notation (TOON) because it can remove repeated keys from uniform lists of records. Our replies also contained nested schemas, so we compared TOON against compact JSON using the same decoded values, field order and schema constraints. We checked round-trip equality and counted both formats with the same tokenizer.
Same data · compact JSON versus TOON
Original TOON experiment · 6 October 2026
Local controlled fixtures · serialized input counts · same response data
Tokenizer: o200k_base
| Same response data | Compact JSON | TOON | Change |
|---|---|---|---|
| 1 flat schema | 96 | 116 | 20.83% more |
| 2 flat schemas | 196 | 245 | 25% more |
| 1 nested schema | 184 | 222 | 20.65% more |
| GitHub list_tags schema | 128 | 147 | 14.84% more |
| 10 uniform summaries | 225 | 189 | 16% less |
Tokenizer: cl100k_base
| Same response data | Compact JSON | TOON | Change |
|---|---|---|---|
| 1 flat schema | 93 | 116 | 24.73% more |
| 2 flat schemas | 189 | 245 | 29.63% more |
| 1 nested schema | 175 | 222 | 26.86% more |
| GitHub list_tags schema | 120 | 146 | 21.67% more |
| 10 uniform summaries | 224 | 189 | 15.62% less |
For the captured GitHub schema, compact JSON used 128 tokens and TOON used 147: 14.84% more. Ten uniform summaries went the other way, from 225 tokens as JSON to 189 as TOON: 16.00% less. The shape of the data changed the result. Comparing against pretty-printed JSON would have obscured that trade-off.
The compact JSON follow-up
We first replayed a compact JSON formatter on eleven frozen response values. The GitHub schema went from 147 TOON tokens to 128 JSON tokens, 12.93% less than the TOON reply. The ten-summary menu became larger, from 189 to 225 tokens. Native schemas and their constraints remained intact.
The fresh RPC capture then showed a behavior change: two nested schemas now fit in the search reply, so that focused sequence needs three requests instead of four. On the current decoded values, with prompts, all four helpers and rounds fixed, JSON uses 1.66–5.63% less sequence input than the offline TOON control. The two browse cases remain slightly larger. A separate synthetic GitHub case captures the declared identity helper alongside repository search. No model or native action ran.
We chose compact JSON for the revised implementation. It gives discovery one encoding path and removes the need to encode and tokenize competing formats during a request. The choice accepts a larger response for some uniform lists; it follows the schema-heavy responses we tested.
Exploring the trade-off
The explorer below compares an eager catalog with helper and selected-schema context.
Discovery context explorer
Estimated context
90 tools across 3 services24,600 tokens less in this model.
Full catalog
90 schemasAction definitions in the model context
Select a service to see its schemas.
With Aixy
4 helpersA stable set of discovery helpers
aixy_search_toolsaixy_execute_toolaixy_connectionsaixy_connect
Select a service to compare discovery.
Adjust the model assumptions
The service catalog counts are illustrative. Aixy exposes four helpers; the number of discovered action schemas is capped by the selected catalog. Formatting measurements are separate from these assumptions.
What we still need to measure
Our counts support a smaller initial context and show why encoding needs to be tested on the actual response shapes. They leave model behavior, retries, latency, provider message framing, caching and billed cost unmeasured. The next experiment is paired runs with the same model, catalog, permissions and approval policy, with caching enabled in both. We need cost per successful task as well as token counts, using the client’s real baseline: a client that already discovers tools on demand starts from a different position.