MCP Security: What to Verify Before an AI Agent Calls a Tool
MCP security starts before the tool call: verify provenance, treat tool text as untrusted, scope credentials, and check the host the server runs on.
By SAUTERA
Every MCP tool an AI agent can call is code running on a real host, holding real credentials. Verify both before the call, not after the incident.
The tool call is the new perimeter
An AI agent that only answers questions is a content risk. An AI agent that calls tools is an infrastructure risk.
The Model Context Protocol made tool access easy. An MCP server exposes tools, and any compatible AI agent can discover and call them. That is the point of the protocol, and it is also the problem: every tool call executes code somewhere, with whatever permissions that code was given, on a machine you may never have looked at.
The attacks are no longer theoretical. In April 2025, Invariant Labs disclosed tool poisoning: instructions hidden in a tool's description that manipulate the model into unauthorized actions without the user seeing them. In July 2025, JFrog disclosed CVE-2025-6514, a CVSS 9.6 flaw in the widely used mcp-remote package. A malicious MCP server could hand the client a crafted authorization URL and run commands on the machine that connected to it.
OWASP's Top 10 for Agentic Applications, published December 2025, names the pattern directly: tool misuse and exploitation, identity and privilege abuse, and agentic supply chain vulnerabilities, which explicitly include MCP servers.
What does MCP security actually require?
MCP security means treating every tool an AI agent can call as untrusted code on an unverified host until proven otherwise. You verify who published it, what it can reach, which credentials it receives, and the state of the machine it runs on. Then you enforce those limits on every call, not once at install.
Most teams stop at the first item. They vet a server once, add it to a config file, and trust it from then on. That is the same design-time trust that fails everywhere else, as we argued in right at design time, wrong by Tuesday.
Five things to verify before an AI agent calls a tool
1. Where the server came from, and whether it changed
An MCP server is a dependency, and it inherits every supply-chain risk a dependency has. Know who publishes it, pin the version you reviewed, and re-review when the tool list or descriptions change. A tool that was benign at approval can be modified later; the protocol even lets a server notify clients that its tool list has changed. Approval is a snapshot, so treat it like one.
2. Tool descriptions and outputs are untrusted input
The model reads a tool's name, description, and schema as guidance. That is exactly why tool poisoning works. The same holds for everything a tool returns: output flows straight back into the model's context, which makes it a channel for the indirect prompt injection that OWASP ranks as the top risk for LLM applications.
Treat both as data, never as authority. Show users the full description, flag hidden or unusual instructions, and never let tool text widen what the AI agent is permitted to do.
3. Credentials scoped to that server, and nothing more
The MCP specification's security best practices are blunt here. Token passthrough is explicitly forbidden: an MCP server must not accept tokens that were not issued for it. Passthrough breaks audience controls, rate limits, and audit trails, and it turns a stolen token into a proxy for data exfiltration.
The same document warns against omnibus scopes. Start an AI agent with read-only scopes and elevate per operation, so a leaked token can only reach what that single task needed. It also describes the confused deputy problem in MCP proxies, which is avoided by requiring consent per client rather than per user. This is least privilege for AI agents, applied to the protocol.
4. The host the server runs on
This is the check most MCP security guides skip, and it is where the worst outcomes live.
A local MCP server runs with the privileges of the client that launched it. The specification says so plainly, and recommends sandboxing, minimal file-system and network access, and showing the exact command before a server is started. A remote MCP server is a URL, but behind that URL is a machine with a patch level, an exposed attack surface, and an operator.
So ask the infrastructure questions before the identity questions are even finished. Is the host patched and supported? Is its configuration where it was when you approved it? Has its posture changed since yesterday? A valid token from a compromised host is still a compromised call. We made this case for people in the Zero Trust device gap; it applies with more force to machines that act at machine speed.
5. A decision on every call, with evidence
Verification at install time is not enforcement. Every consequential tool call should pass a policy check against current state, and the decision should be recorded: which AI agent, which tool, which host, what was allowed, and why.
Irreversible actions get a human or a hard policy in the path; everything else can run. That is the model in reversible by design. And when the state of a host cannot be established, the right answer is to stop, because unknown is an answer.
Where identity stops and infrastructure trust starts
MCP's authorization model, built on OAuth 2.1, answers an important question well: who is calling, and with what delegated rights. It does not answer whether the machine on the other end is in a state you would trust with that call.
That is the gap in most agentic AI deployments. Identity is verified once per session; the infrastructure behind each tool is assumed. As we put it in zero trust tells you who, not whether, the authorization decision needs both inputs. A complete trust decision for a tool call combines the AI agent's identity, the user it acts for, the scope requested, and the observed state of the host, in the way we describe in the anatomy of a trust decision.
Infrastructure trust is the second half of that decision: continuous, observed evidence about the hosts your AI agents depend on, so "is this server safe to call right now?" has a factual answer.
An MCP security checklist
Use this before you connect a new MCP server, and again whenever it changes:
- Provenance: a known publisher, a pinned version, and re-review on any change to tools or descriptions.
- Untrusted text: tool descriptions and outputs treated as data, displayed in full, and never allowed to expand permissions.
- Credentials: tokens issued only for that server, no passthrough, minimal scopes elevated per operation, and per-client consent on proxies.
- Host: local servers sandboxed with least privilege; remote servers backed by a host whose patch level and configuration you can verify.
- Enforcement: a policy decision on every consequential call, a human or hard policy on irreversible actions, and a fail-closed default when host state is unknown.
- Evidence: every decision logged with the AI agent, tool, host, and outcome, so an audit is a query, not a fire drill.
The takeaway
The Model Context Protocol made it easy to give AI agents tools. It did not make those tools trustworthy, and it does not tell you anything about the machines they run on.
- Every MCP tool is code on a host with credentials. Verify the code, the credentials, and the host.
- Treat tool descriptions and outputs as untrusted input; tool poisoning depends on you not doing that.
- Scope tokens to the server and the task. The specification forbids token passthrough for good reason.
- Check host posture on every consequential call, and fail closed when you cannot.
- Record each decision so you can prove what your AI agents did and why.
Identity tells you who is calling. Infrastructure trust tells you whether the call should run.
Next: scope what each AI agent may do in the first place — read Least Privilege for AI Agents.
Written by
SAUTERA
Author of the Infrastructure Trust Architecture (ITA) and the Infrastructure Trust Conveyance Mechanism (ITCM) — the standard organizations use to decide whether infrastructure can be trusted.
Follow the work
Read the next one
New perspectives on infrastructure trust and updates to the ITA / ITCM framework, by email.
Occasional. No spam. Unsubscribe anytime.