CLI vs MCP for AI agents: context cost, auth and when to use each

CLI vs MCP for AI agents: what tool definitions and results cost in Claude Code's context, how auth and sandboxes differ, and when MCP is the better pick.

Claude Code and the Codex CLI run in a terminal with a shell, and for agents like these a command line tool plus a skill file is the simpler default. Anthropic’s Claude Code best practices call CLI tools “the most context-efficient way to interact with external services.” A Model Context Protocol (MCP) server earns its place when the client has no shell or a remote service needs each person to sign in with OAuth. It also wins when the agent runs in a cloud sandbox with its network shut off, or when one integration has to serve many clients.

The usual argument against MCP, that tool definitions fill the context window, has shrunk to a few hundred tokens in Claude Code. With tool search on, which is the default, 25 Playwright tools added about 390 input tokens per request in one measured run, against about 6,710 with tool search off. The differences that remain are output shaping, authentication, sandboxing and which clients can run each route.

What each option is

An MCP server offers tools the model can call, each with a JSON Schema for its inputs, along with resources and prompt templates. The current specification, dated 2026-07-28, defines two standard transports. With stdio, the client launches the server as a local subprocess. With Streamable HTTP, each message is a POST to one endpoint on a server that can live anywhere.

A skill is a folder with a SKILL.md file holding a name, a description and instructions, plus optional reference files. Paired with a CLI, it tells the agent what --help leaves out, such as which flags belong together and what each exit code means. Agents already know git and gh from training, so common tools often need no skill.

What tool definitions cost now

In Claude Code, tool search loads only tool names and server instructions at session start. A tool’s full schema loads when Claude searches for it. Tool search needs Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 or a later model, and the ENABLE_TOOL_SEARCH environment variable controls it:

Setting What Claude Code does
Unset Defers all MCP tools. Loads them upfront when ANTHROPIC_BASE_URL points at a non-first-party host, on Google Cloud Agent Platform models older than Claude 4.5, or on a Microsoft Foundry deployment hosted on Azure
auto, auto:N Loads tools upfront while their definitions total under 10% (or N percent) of the context window, and defers all of them past that
false Loads every MCP tool upfront

Setting CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS keeps tool search off whatever ENABLE_TOOL_SEARCH says. Run /context all to see what each loaded tool costs.

How we measured it

We connected Microsoft’s Playwright MCP server, version 0.0.83, to Claude Code 2.1.292 on October 10, 2026. Its 25 default tools serialize to 20,286 characters, counted as JSON.stringify of the tools array from a tools/list call; its --caps flag adds more. We ran this twice per configuration, with only that server loaded through --strict-mcp-config:

claude -p "Reply with the single word ok. Do not use any tools." \
  --output-format json --strict-mcp-config --mcp-config pw.json \
  --no-session-persistence

Each run made one API turn on Claude Opus 5.5. We summed input_tokens, cache_creation_input_tokens and cache_read_input_tokens from the usage field, which counts everything in context whatever caching does to the bill.

Configuration Baseline, no server With Playwright MCP Added by the server
Tool search on 40,301 and 40,309 40,693 and 40,689 about 390
ENABLE_TOOL_SEARCH=false 59,747 and 59,741 66,456 and 66,463 about 6,710

With tool search off, all 25 schemas occupy context on every request of the session; with it on, only the deferred tool list grows. Our baselines include our own instruction files and skills, so only the differences carry over to your setup.

Codex defers MCP tools the same way on models that support tool search, according to its client source. Other clients vary, so check yours.

Skills have a standing cost too

A skill’s description loads every session. Claude Code truncates the combined description and when_to_use text at 1,536 characters, and Codex caps the initial skills list at 2% of the context window, or 8,000 characters when the size is unknown. The full SKILL.md loads when the agent picks the skill. Playwright’s CLI skill is 15,428 bytes, larger than any single Playwright MCP schema, which tool search would load one at a time.

Output is where the routes differ

Every result enters the context window, whether a server returned it or a command printed it, and each client caps both.

Limit Claude Code Codex
MCP tool result Warns past 10,000 tokens and caps at 25,000 by default (MAX_MCP_OUTPUT_TOKENS). A text result over 50,000 characters goes to a file, and Claude gets a reference tool_output_token_limit covers every tool; output_token_limit under [mcp_servers.<id>.tools.<tool>] overrides it for one
Shell output Inline up to about 30,000 characters, adjustable to 128,000 with bashOutputMaxChars. Past that, a file path and a 2,000-character preview. A failed command gets about 10,000 characters. BASH_MAX_OUTPUT_LENGTH sizes only the read-back window tool_output_token_limit

A shell command can pipe through head, jq or grep before anything reaches the model. An MCP result can be trimmed only through parameters the server offers.

The same task through both routes

Microsoft ships both a Playwright MCP server and a Playwright CLI with skills, and recommends the CLI to coding agent users. Our task: list the section headings of the Wikipedia article on the Model Context Protocol, which has 12.

We called the MCP server’s tools directly over stdio on October 10, 2026, headless Chromium at a 1280x720 viewport, four runs a few minutes apart, and counted the characters each call returned:

MCP call Characters returned Headings included
browser_snapshot, no arguments 87,472 to 87,968 12
browser_snapshot with depth: 8 47,323 to 47,412 12
browser_snapshot with depth: 4 1,938 to 1,958 1
browser_find with text heading " 8,660 12, with tree context
grep -n 'heading "' on the saved snapshot 779 12

The CLI’s own equivalent of browser_find is playwright-cli find, which the README says returns matching nodes with three lines of context. Its file route looks like this:

playwright-cli open https://en.wikipedia.org/wiki/Model_Context_Protocol
playwright-cli snapshot --filename=page.yml
grep -n 'heading "' page.yml

We did not install the CLI. The 779-character row is that grep run on the snapshot the MCP server saved. Claude Code’s Grep tool returns the same output on that file, so searching a file gives the small result on either route.

Treat these as one dated session on a live page. The snapshot that browser_navigate saves measured 87,710 bytes twice and 204,809 bytes once within minutes, a swing of more than 2x. Claude Code would save the 87,000-character inline snapshot to a file, since it passes the 50,000-character limit, and Claude would search that file. A client without the fallback puts all of it in context.

Permissions and authentication

Both routes go through your client’s permission system. In Claude Code, Bash(gh pr list *) allows one command pattern and mcp__playwright__* allows every tool from one server. Codex sets MCP approval per server with default_tools_approval_mode, which takes auto, prompt, writes or approve; writes prompts only for tools not marked read-only.

A CLI uses its own stored login, such as gh auth login, and the agent inherits it. The specification tells stdio servers to read credentials from the environment too. Remote servers that require sign-in use OAuth 2.1, and each token must be bound to the one server it was issued for. Authorization is optional in the spec, so many remote servers take a static token in a header, which Claude Code can generate per connection with headersHelper. You sign in with /mcp or claude mcp login <name> in Claude Code and codex mcp login <name> in Codex.

Prompt injection threatens both routes: a web page or issue comment can carry instructions aimed at the model, whichever way it arrived. Anthropic’s docs tell you to verify you trust each server and warn that servers fetching external content raise the risk.

Sandboxes

Claude Code’s sandbox wraps shell commands only. A sandboxed CLI reaches an API once you add its host to sandbox.network.allowedDomains, and credential masking shows the command a placeholder that the sandbox proxy swaps for the real value. Local MCP servers run outside the sandbox with your full access, so the sandbox does nothing to contain them.

Claude Code cloud sessions are where MCP helps. Their network access can be set to None, and MCP connectors you enable still work because their traffic travels through Anthropic’s servers, outside the session’s network.

Code execution as a third route

A third pattern has the model write code that calls tools, so only the final result enters context. Anthropic’s programmatic tool calling does this on the Claude API inside the code execution container, though tools from the API’s MCP connector are excluded. Cloudflare’s Code Mode converts MCP tools into a TypeScript API that the model codes against, run in a V8 isolate. An agent with a shell gets the same effect from a script that chains CLI calls.

When MCP is the better choice

The client has no shell. Claude on the web and in Claude Desktop reaches outside tools through custom connectors, which are remote MCP servers added under Settings > Connectors. Claude Desktop also runs local stdio servers listed in claude_desktop_config.json and asks before each action.

The service is remote and each person signs in. OAuth through MCP replaces API keys pasted into environment variables, and the vendor keeps a hosted server current. Claude Code can narrow the scopes a server requests with an oauth.scopes entry in .mcp.json.

The agent runs in a locked-down cloud session. Connectors still work with the session’s network set to None.

One integration serves many clients. The same server works in Claude Code, Claude Desktop and Codex, and OpenAI’s docs say Codex shares its MCP configuration with the ChatGPT desktop app.

The tool holds state across many calls. The Playwright README keeps MCP for long exploratory browser sessions where continuous context outweighs token cost.

When a CLI wins

The work is local. Files, repositories, build tools and databases you already query from a terminal need no server, and the agent uses your existing credentials. Agent memory kept in markdown notes belongs here too, as giving AI agents persistent memory explains.

You want filtering without waiting on a server author. Any command can pipe through grep or jq, whether or not the tool offers a search parameter.

The command should outlive the session. A line that worked for the agent runs unchanged in CI or cron, and running it yourself shows exactly what the agent saw.

Failures should surface at once. A missing binary or a nonzero exit code reaches the agent with the error text attached.

How skills fit with both

Anthropic’s feature overview gives the pattern: an MCP server connects to your database, and a skill documents your schema and query patterns. A Claude Code plugin can bundle both into one install.

Skills can also travel over MCP. The official Skills extension (io.modelcontextprotocol/skills, SEP-2640, final since September 2026) lets a server publish SKILL.md files through MCP resources, keeping instructions with the service they describe. Client support is still arriving. Instructions every session needs belong in your instruction file, as covered in what to put in CLAUDE.md.

Why Jotura ships a CLI and no MCP server

Jotura is a notes app that keeps plain .md files on your disk. The agents it supports, including Claude Code, Codex, Gemini CLI, GitHub Copilot and OpenCode, all have a shell, so the jotura command does the work. An edit carrying --if-hash is refused with exit code 2 when the note changed since the agent read it, and a --replace with more than one match exits 4.

You install the skills like any other. In Claude Code, run /plugin marketplace add adamrichardson14/jotura-agents and then /plugin install jotura@jotura-agents; other agents copy the files from github.com/adamrichardson14/jotura-agents. A chat assistant without a shell cannot use Jotura. The commands are in the CLI reference, and notes apps with a CLI compares the alternatives.

Frequently asked questions

Do I need an MCP server to use my tools with Claude Code?

No. Claude Code runs shell commands, so any tool with a command line interface works, and Anthropic’s best practices recommend CLIs such as gh and aws for external services. Add a server for a service with no CLI or an account that needs OAuth sign-in.

Do MCP servers use more tokens than CLI tools?

At session start, only slightly in Claude Code with tool search on. In one measured run, 25 Playwright tools added about 390 input tokens per request with tool search on and about 6,710 with it off. Per call, each route costs what it returns, and output that lands in a file can be searched on either route.

How do I turn off MCP tool search in Claude Code?

Set ENABLE_TOOL_SEARCH=false in your shell or in the env field of settings.json. The values auto and auto:N load tools upfront only while their definitions fit under 10%, or N percent, of the context window. Setting "alwaysLoad": true on one server in .mcp.json exempts that server alone.

How to use Claude Code with your Obsidian vault2026-10-10
Use Claude Code with your Obsidian vault: run it in the vault folder, add a CLI for hash-checked edits and link-safe renames, and set permission rules.
How to give AI agents persistent memory, and what each agent remembers2026-10-10
Why agents forget between sessions, what Claude Code, Codex, Gemini CLI and Copilot keep natively, and how to give AI agents persistent memory across tools.
How to sync CLAUDE.md and Claude Code settings between machines2026-10-10
How to sync CLAUDE.md between machines with skills, settings and plugins: what to move, what never to copy, Windows symlink traps and a setup for teams.

More in AI agents.

Download free

Jotura is free, and your notes stay plain markdown files you keep forever.