Skip to content
All articles

When MCP tools fill the context before the work begins

Growing MCP catalogs consume model context before work begins. We examine discovery, encoding and the measurements from our experiments, including what they leave unresolved.

As we connected more MCP servers to our agents, we noticed a problem before the agents had even called a tool: their context was already filling up. Loading complete tool catalogs meant including every tool’s name, description and input schema upfront. Those definitions consumed input tokens even when the tools were irrelevant to the request, leaving less room for the conversation and the data the agent actually needed.

Our experience matched what we were reading in publications such as Anthropic’s articles on code execution with MCP and advanced tool use. In the latter, Anthropic described five MCP servers exposing 58 tools whose definitions occupied roughly 55,000 tokens before the conversation began. That figure comes from their configuration, but it illustrates the problem we wanted to address: making more tools available could carry a substantial context cost before any useful work happened.

There was also a practical reason to manage those connections centrally. Configuring the same MCP servers for every new agent meant repeating setup and maintaining it in several places. An MCP gateway lets us configure those integrations once and connect each new agent to a single endpoint, through which it can access the tools available to it.

Together, these concerns shaped our work on MCP discovery in Aixy. We wanted to keep a growing collection of integrations available through the gateway while bringing only the tool definitions needed for the current task into the model’s context. We explored tool selection and response encoding, then measured their effects separately.

Selection and encoding are different problems

The MCP tool contract contains the information needed to call a tool. When a client loads all those definitions upfront, every schema contributes input tokens. We saw two places to intervene: choose which definitions enter the conversation, and choose how to encode the selected information. Each change has its own baseline; their saving percentages cannot be added.

How we built discovery

Our MCP endpoint exposes four stable helpers. The agent searches for a capability and receives a small selection of action names and descriptions, with complete input schemas when they fit the response budget. Otherwise, it retrieves the chosen action’s exact schema before execution. The full catalog stays available to the system while the model carries the definitions needed for its next step.

A separate mcp-tool-search service handles retrieval and response formatting; the gateway forwards its reply over the existing MCP interface. Search uses BM25S with Lucene scoring, k1=1.2 and b=0.75. We build a local index from names, descriptions, schemas and catalog metadata, with extra weight on names. Queries normalize naming conventions, accents and case, and expand a small English and Spanish vocabulary. Discovery makes no embedding-model request.

Ranking prioritizes exact names, overlap with the requested capability and compatible read or write operations, with BM25 breaking ties. Debugging also showed us that a matching tool could leave the agent missing the identity or workspace needed to use it. When the catalog declares those dependencies, the reply can include context helpers alongside the primary action.

Returned schemas retain required fields, types, limits and accepted values. Internal identifiers and definition revisions stay in the session. Before execution, the gateway rechecks permissions, connection state and definition identity. A smaller discovery reply still has to preserve the contract needed to make a valid call.

Discovery formatting belongs to the search service

The client carries its MCP session context. The search service selects and formats discovery; the gateway checks access and executes.

Discovery formatting belongs to the search service After initialization, the agent client sends a capability query to Aixy's gateway. For a ready catalog the gateway makes one DiscoverTools RPC to the search service, which ranks actions, selects schemas and serializes compact JSON. The gateway forwards that text unchanged to the client. The client chooses a semantic name, obtains approval when required and sends the name and JSON arguments. The search service resolves the name in the caller's selection context. The gateway checks current permissions, connection, allowlist and native definition revision, then runs the action on the connected MCP server. Native results return through the gateway. The client continues its model loop. Stale or expired selections require rediscovery. Native action results retain their original MCP contract. Agent clientModel loop + approval Aixy gatewayAccess + execution Tool searchSelection + JSON Connected MCPserverRuns the native action aixy_search_toolsCapability query DiscoverToolsOne search RPC Rank relevant actionsSelect complete schemasBind names + JSON Compact JSON text Forward unchanged Choose semantic nameFetch schema if neededObtain required approval aixy_execute_toolname + JSON arguments Resolve scoped nameExact schema lookup Current definition Check access + revisionConnection + allowlistReject stale selections Run the native action Native result Tool result Continue the model loop
  1. Agent client → Aixy gateway

    aixy_search_tools sends a capability query.

  2. Gateway → Tool search service

    One DiscoverTools RPC for a ready catalog.

  3. Tool search service

    Ranks relevant actions, selects complete schemas when useful, binds semantic names and serializes compact JSON.

  4. Search service → Gateway → Agent client

    The gateway forwards the bounded discovery text unchanged.

  5. Agent client

    Chooses a semantic name, fetches the exact schema if needed and obtains approval when required.

  6. Agent client → Gateway

    aixy_execute_tool submits name and JSON arguments.

  7. Gateway → Search service

    Resolves the name against the caller's selection context and obtains the native definition.

  8. Gateway

    Rechecks permissions, connection, allowlist and definition revision. A stale or expired selection requires rediscovery.

  9. Gateway → Connected MCP server

    Runs the requested native action.

  10. Connected MCP server → Gateway → Client

    Returns the native result through the existing protocol.

  11. Agent client

    Continues the model loop.

Simplified external MCP client flow for the revised JSON implementation. The search service ranks the authorized catalog with BM25S and returns compact JSON discovery text; the gateway forwards it unchanged. Execution arguments and native results retain their existing protocol. The client owns its model loop and approval policy. Session headers and catalog publication are omitted.

How we measured discovery

On 6 October 2026, we captured synthetic catalogs and responses from the real DiscoverTools handler in a temporary SQLite-backed local service. On 7 October, we recaptured the same native schemas with the current JSON implementation and all four gateway helpers. Counts use tiktoken 0.14.0 with o200k_base and cl100k_base; the offline TOON control uses toon-format 1.1.0. The tables link to both dated captures and their methods.

Initial context

The current initial comparison uses two synthetic catalogs of 100 tools, with flat and nested schemas. Both requests contain the same user message. One includes all native definitions; the other includes the four literal gateway helpers. The original 6 October control contained only search and execute: its earlier “four helpers” label was incorrect. That control remains available with its original files and counts.

Initial serialized request · discovery only

Current JSON capture · 7 October 2026

Local controlled fixtures · serialized input counts · four gateway helpers

Tokenizer: o200k_base
Initial serialized request · discovery only, Current JSON capture · 7 October 2026, o200k_base, tokens
Initial contextEager JSONFour helpersReduction
100 flat schemas10,02254494.57% less
100 nested schemas18,62254497.08% less
Tokenizer: cl100k_base
Initial serialized request · discovery only, Current JSON capture · 7 October 2026, cl100k_base, tokens
Initial contextEager JSONFour helpersReduction
100 flat schemas9,71953894.46% less
100 nested schemas17,91953897% less

Counts and method · Captured fixtures · Recount script

Original two-helper control · 6 October 2026

Local controlled fixtures · serialized input counts · two discovery/execution helpers

Tokenizer: o200k_base
Initial serialized request · discovery only, Original two-helper control · 6 October 2026, o200k_base, tokens
Initial contextEager JSONTwo helpersReduction
100 flat schemas10,02242895.73% less
100 nested schemas18,62242897.7% less
Tokenizer: cl100k_base
Initial serialized request · discovery only, Original two-helper control · 6 October 2026, cl100k_base, tokens
Initial contextEager JSONTwo helpersReduction
100 flat schemas9,71942195.67% less
100 nested schemas17,91942197.65% less

Label corrected: these frozen requests contain search and execute only. The original counts and files are unchanged.

Counts and method · Captured fixtures · Recount script

Measured synthetic input context. Required schemas, later discovery exchanges and execution results still need to enter the conversation; this is an initial-context comparison, not completed-task savings.

With o200k_base, the full-catalog requests use 10,022 and 18,622 tokens, compared with 544 for the four helpers: 94.57% and 97.08% less initial context. Selected schemas, later searches and tool results still add tokens as the task progresses.

A complete simulated sequence

The current sequence comparison uses eight focused searches returning 1, 2, 5 or 10 candidates, plus two browse pages. We count every serialized input request, repeated history and schema lookup. The baseline reconstructs the previous verbose JSON path, adding the same two connection helpers to both paths. Search and execution contracts can differ. Model steps use a fixed successful result.

Complete simulated input sequences

Current JSON capture · 7 October 2026

Local controlled fixtures · serialized input counts · four gateway helpers

Tokenizer: o200k_base
Complete simulated input sequences, Current JSON capture · 7 October 2026, o200k_base, tokens
SearchPrevious sequenceCurrent JSONReduction
flat · 1 candidate2,6432,01623.72% less
flat · 2 candidates3,1852,24429.54% less
flat · 5 candidates4,8672,21254.55% less
flat · 10 candidates7,7082,44268.32% less
nested · 1 candidate2,9082,21823.73% less
nested · 2 candidates3,7442,64829.27% less
nested · 5 candidates6,2182,41461.18% less
nested · 10 candidates10,3912,64474.55% less
flat · 10 browse summaries *7,7043,41855.63% less
nested · 10 browse summaries *10,3873,62065.15% less
Tokenizer: cl100k_base
Complete simulated input sequences, Current JSON capture · 7 October 2026, cl100k_base, tokens
SearchPrevious sequenceCurrent JSONReduction
flat · 1 candidate2,6221,99523.91% less
flat · 2 candidates3,1642,22329.74% less
flat · 5 candidates4,8462,19154.79% less
flat · 10 candidates7,6872,42168.51% less
nested · 1 candidate2,8842,19723.82% less
nested · 2 candidates3,7182,62729.34% less
nested · 5 candidates6,1902,39361.34% less
nested · 10 candidates10,3622,62374.69% less
flat · 10 browse summaries *7,6833,38855.9% less
nested · 10 browse summaries *10,3583,59065.34% less

* Browse cases include a fourth request for the schema; all focused cases now use three. Both paths include the same two connection helpers; search and execute contracts differ. Counts include repeated history.

Counts and method · Captured fixtures · Recount script

Original TOON experiment · 6 October 2026

Local controlled fixtures · serialized input counts · two discovery/execution helpers

Tokenizer: o200k_base
Complete simulated input sequences, Original TOON experiment · 6 October 2026, o200k_base, tokens
SearchPrevious sequenceTOON experimentReduction
flat · 1 candidate2,2981,70225.94% less
flat · 2 candidates2,8401,97830.35% less
flat · 5 candidates4,5221,93457.23% less
flat · 10 candidates7,3632,20470.07% less
nested · 1 candidate2,5631,94224.23% less
nested · 2 candidates *3,3992,67321.36% less
nested · 5 candidates5,8732,17462.98% less
nested · 10 candidates10,0462,44475.67% less
Tokenizer: cl100k_base
Complete simulated input sequences, Original TOON experiment · 6 October 2026, cl100k_base, tokens
SearchPrevious sequenceTOON experimentReduction
flat · 1 candidate2,2741,67826.21% less
flat · 2 candidates2,8161,95430.61% less
flat · 5 candidates4,4981,91057.54% less
flat · 10 candidates7,3392,18070.3% less
nested · 1 candidate2,5361,91824.37% less
nested · 2 candidates *3,3702,63921.69% less
nested · 5 candidates5,8422,15063.2% less
nested · 10 candidates10,0142,42075.83% less

* The original nested two-candidate case includes a fourth request for the schema. This historical control uses two helpers. Counts include repeated history.

Counts and method · Captured fixtures · Recount script

Current compact JSON and the original TOON control, each against its stated previous discovery path. Includes all requests and repeated history. This measures combined discovery changes, not encoding alone or an eager-catalog baseline. Fixed successful results; task quality and billed savings remain unmeasured.

The one- and two-candidate cases now use 23.72–29.54% less serialized input against that previous path. Favorable ten-candidate searches reach 68.32% and 74.55%, because the old path repeats every full schema while discovery carries a primary schema and summaries. These figures measure the combined changes. The table also retains the original two-helper TOON experiment; its percentages have a different denominator.

The TOON experiment

We first tried Token-Oriented Object Notation (TOON) because it can remove repeated keys from uniform lists of records. Our replies also contained nested schemas, so we compared TOON against compact JSON using the same decoded values, field order and schema constraints. We checked round-trip equality and counted both formats with the same tokenizer.

Same data · compact JSON versus TOON

Original TOON experiment · 6 October 2026

Local controlled fixtures · serialized input counts · same response data

Tokenizer: o200k_base
Same data · compact JSON versus TOON, Original TOON experiment · 6 October 2026, o200k_base, tokens
Same response dataCompact JSONTOONChange
1 flat schema9611620.83% more
2 flat schemas19624525% more
1 nested schema18422220.65% more
GitHub list_tags schema12814714.84% more
10 uniform summaries22518916% less
Tokenizer: cl100k_base
Same data · compact JSON versus TOON, Original TOON experiment · 6 October 2026, cl100k_base, tokens
Same response dataCompact JSONTOONChange
1 flat schema9311624.73% more
2 flat schemas18924529.63% more
1 nested schema17522226.86% more
GitHub list_tags schema12014621.67% more
10 uniform summaries22418915.62% less

Counts and method · Captured fixtures · Recount script

Formatting only: identical decoded values, field order and schema constraints. Negative savings are shown as more tokens. The uniform-summary case is particularly favorable to TOON.

For the captured GitHub schema, compact JSON used 128 tokens and TOON used 147: 14.84% more. Ten uniform summaries went the other way, from 225 tokens as JSON to 189 as TOON: 16.00% less. The shape of the data changed the result. Comparing against pretty-printed JSON would have obscured that trade-off.

The compact JSON follow-up

We first replayed a compact JSON formatter on eleven frozen response values. The GitHub schema went from 147 TOON tokens to 128 JSON tokens, 12.93% less than the TOON reply. The ten-summary menu became larger, from 189 to 225 tokens. Native schemas and their constraints remained intact.

The fresh RPC capture then showed a behavior change: two nested schemas now fit in the search reply, so that focused sequence needs three requests instead of four. On the current decoded values, with prompts, all four helpers and rounds fixed, JSON uses 1.66–5.63% less sequence input than the offline TOON control. The two browse cases remain slightly larger. A separate synthetic GitHub case captures the declared identity helper alongside repository search. No model or native action ran.

We chose compact JSON for the revised implementation. It gives discovery one encoding path and removes the need to encode and tokenize competing formats during a request. The choice accepts a larger response for some uniform lists; it follows the schema-heavy responses we tested.

Exploring the trade-off

The explorer below compares an eager catalog with helper and selected-schema context.

Discovery context explorer

Choose your MCP services

Discovery only: compare loading every action schema with selecting the actions a task needs. Tool counts below are illustrative; no TOON multiplier is applied.

Estimated context

90 tools across 3 services
Full catalog
27,000 tokens
On-demand discovery
2,400 tokens
91.1% less context

24,600 tokens less in this model.

Full catalog

90 schemas

Action definitions in the model context

GitHub40 schemas
Jira30 schemas
Notion20 schemas
Every selected action schema, upfront.

With Aixy

4 helpers

A stable set of discovery helpers

  • aixy_search_tools
  • aixy_execute_tool
  • aixy_connections
  • aixy_connect
3 action schemasDiscovered for this task
900 tokens of discovery context included.
Adjust the model assumptions

The service catalog counts are illustrative. Aixy exposes four helpers; the number of discovered action schemas is capped by the selected catalog. Formatting measurements are separate from these assumptions.

Illustrative discovery-only inputs. This estimates tool schemas, four helper definitions and discovery context. It applies no TOON reduction and does not measure model usage or billed savings.

What we still need to measure

Our counts support a smaller initial context and show why encoding needs to be tested on the actual response shapes. They leave model behavior, retries, latency, provider message framing, caching and billed cost unmeasured. The next experiment is paired runs with the same model, catalog, permissions and approval policy, with caching enabled in both. We need cost per successful task as well as token counts, using the client’s real baseline: a client that already discovers tools on demand starts from a different position.