OguzhanTekin
The MCP Cost That Starts Before the First Tool Call
Cloud & FinOpsOctober 7, 2026

The MCP Cost That Starts Before the First Tool Call

By Oguzhan Tekin•Back to Blog

MCP servers are usually described as connectors. They give an AI agent access to code repositories, databases, collaboration platforms, cloud services, and internal tools.

But the cost of those connections can begin before the agent takes an action.

The first cost is exposure. When MCP tool definitions are loaded into the active model context, they consume tokens and context capacity — even if most of those tools are never called.

Anthropic documented a five-server example in which 58 tool definitions consumed approximately 55,000 tokens before the conversation started. GitHub alone contributed about 26,000 tokens across 35 tools, while Slack contributed about 21,000 across 11 tools.

Against a task that would otherwise begin with 10,000 input tokens, adding 55,000 tokens of definitions raises the initial model input to 65,000 tokens:

  • 55,000 additional tokens
  • 550% above the original baseline
  • 6.5 times the original input

The percentage is illustrative, not universal. It depends on which definitions are loaded, how large their schemas are, and how the agent harness manages them.

Context capacity is also a budget. The same 55,000 definitions would occupy 27.5% of a 200,000-token context window, but only 5.5% of a one-million-token window.

That difference does not make the definitions free in the larger window. It means the agent has more room left for code, documents, instructions, conversation history, tool results, and output.

Comparison of the same 55,000 loaded MCP tool-definition tokens in two context windows. They occupy 27.5 percent of a 200,000-token window, leaving 145,000 tokens, and 5.5 percent of a one-million-token window, leaving 945,000 tokens. No tool has been called yet.
Loaded definitions consume context before the first tool call. Open the full PDF visual.

This is not a separate MCP surcharge. The economic treatment appears through each provider's normal usage model.

GitHub Copilot CLI identifies MCP Tools as a separate part of the active context through the /context command. Copilot usage is then affected by the model selected and the input, output, and cached tokens processed through GitHub's AI-credit system.

Anthropic generally processes loaded tool definitions as model input. Its Tool Search approach demonstrates the alternative: discover relevant tools on demand instead of loading the entire catalog upfront. In Anthropic's example, this reduced initial context consumption from approximately 77,000 tokens to 8,700 — an 85% reduction.

OpenAI states that customers pay for tokens used when importing MCP tool definitions or making MCP calls, with no additional fee for each MCP tool call. It recommends restricting imported tools with allowed_tools because large catalogs can increase both cost and latency.

Subscription products may absorb some of this overhead into included allowances rather than displaying it as a separate MCP line item. The resource is still being consumed.

The control is tool exposure. Engineers should measure not only how often tools are called, but also:

  • How many definitions are exposed to the model
  • How many exposed tools are actually used
  • How often definitions are reprocessed
  • Whether prompt caching is preserved
  • How much context tool results consume
  • Whether only relevant tools are loaded for each task

Excessive MCP tool catalogs can create avoidable cost, latency, and context pressure — and make selecting the correct tool harder.

The practical controls are familiar: inventory the tools, scope them to the task, use allowlists, defer or dynamically discover definitions, simplify verbose schemas, constrain tool output, and preserve stable cache prefixes.

The FinOps takeaway is that active tool definitions are part of the inference-cost architecture. They should be managed like any other cost-bearing resource: inventoried, scoped, measured, cached where appropriate, and loaded only when needed.

The model does not need to call a tool for its definition to have a cost. That cost can begin with the decision to place the tool in front of the model.

References