How to Secure Tool-Using LLMs: Safety Patterns
Tool-using LLMs can search documents, call APIs, and change records, which makes them useful in workflows but gives their mistakes consequences beyond a bad answer. Developers and security teams need to treat every model-proposed action as untrusted input and enforce policy outside the model. This guide explains the risks, how a safe execution path works, and practical patterns for authorization, validation, approval, isolation, and monitoring.
What Are Tool-Using LLMs?
A tool-using LLM application gives a model descriptions of operations it may request, such as searching a knowledge base or creating a support ticket. The model returns a structured request; the application decides whether and how to execute it, then sends the result back into the conversation. As the Anthropic tool-use documentation describes, the model can request a tool, but the integrating application performs the operation and supplies its result.
This distinction matters: a tool call is not an instruction that must be obeyed. It is a proposal from a probabilistic component. The application owns the credentials, network connections, and execution environment, so it must independently authenticate the person, authorize the exact operation, and validate its arguments. A model’s confidence, a well-written system prompt, or a provider’s structured output mode cannot replace those checks.
Tool-using applications are discussed more often as assistants move from answering questions to completing multi-step work. Connecting a model to a calendar, code repository, ticketing system, or payment API can save repetitive effort, but it also creates a path from ambiguous natural-language input to a privileged system. The central design problem is therefore not how to make a model always choose correctly; it is how to keep an incorrect or manipulated choice within a safe boundary.
The Problem: Model Intent Is Not Authorization
Traditional applications usually receive an operation from a user interface and check it against the user’s identity and permissions. A tool-using workflow adds a model that interprets user requests, retrieved documents, web pages, and tool results. Any of this text can be wrong or malicious. A prompt injection embedded in a document might tell the model to disclose data or invoke an unrelated operation. Treating trusted instructions and untrusted content as if they were equally authoritative makes that attack easier.
The risk increases when an application exposes broad tools. A generic shell, unrestricted database query, or API client with an administrator token can do far more than the user intended. A mistaken recipient, an unsafe path, or a repeated call may cause data loss, unauthorized access, unexpected charges, or messages sent on the user’s behalf. The OWASP GenAI Security Project highlights prompt injection and excessive agency among LLM application risks; excessive agency includes giving a model too much functionality or permission without sufficient controls.
There is no single prompt or model setting that makes these risks disappear. Security must be enforced by the services that hold the authority: the tool host, authorization layer, executor, and downstream API. The model may help select an action, but it must not be allowed to expand its own permissions.
How a Safer Tool-Execution Path Works
A controlled execution path separates interpretation from authority:
- The application authenticates the user and establishes the user’s trusted identity and permissions.
- The model receives only the task context and tool descriptions needed for that request.
- The model proposes a tool name and structured arguments.
- A policy layer checks the tool against an allowlist, validates the arguments, applies per-user authorization and resource limits, and decides whether human approval is required.
- A constrained executor invokes a narrowly scoped service using credentials unavailable to the model.
- The application filters the result, records a privacy-conscious audit event, and returns only necessary information to the model or user.
The model should not directly receive API keys or execute arbitrary code. A dedicated tool service can hold short-lived credentials and translate a narrow operation, such as “read this user’s open tickets,” into a backend request. For delegated access, use a token limited to the intended resource and operations. OAuth 2.0 Security Best Current Practice (RFC 9700) recommends restricting access-token privileges and using other protections against token misuse; those OAuth controls complement, but do not replace, application-level authorization.
| Failure or abuse case | Enforceable control | Remaining consideration |
|---|---|---|
| Prompt injection requests a privileged action | Keep tool permissions in application policy; treat retrieved text as data | The model may still propose an unsafe call |
| Arguments select another user’s record | Bind authorization to the authenticated user and target resource | Validate ownership again in the downstream service |
| A tool call creates an external side effect | Require a trusted, specific approval for high-impact actions | Approval must describe the exact action and target |
| A loop repeats an expensive operation | Set per-user quotas, step limits, timeouts, and idempotency keys | Distributed deployments need atomic shared limits |
| Tool output contains malicious instructions or sensitive data | Minimize, filter, and label returned content | Filtering does not establish that content is trustworthy |
These layers reduce the impact of individual failures. They do not make every model decision correct or every connected service secure.
Key Safety Components
A small allowlist of tools limits what the model can ask the application to do. Prefer purpose-built operations such as search_my_documents over a generic fetch_url or run_command. Separate read-only operations from actions that write, delete, send, or spend. Do not expose a tool just because the model might find it convenient.
Server-side authorization uses a trusted principal, not a username or role inferred from the prompt. Check access on every call and at the resource being accessed; hiding a tool from the model is not an authorization check. Where possible, pass the user’s delegated identity through to the service so that existing permissions continue to apply. Keep credentials in a secret store, scope them narrowly, rotate them, and never place them in prompts or routine telemetry. The OAuth 2.0 implementation guide provides background on OAuth roles and tokens.
Schema and semantic validation catch malformed or out-of-range arguments before execution. A JSON schema can restrict types and required fields, but valid syntax is not proof of a safe request. Validate business rules as well: allowed project identifiers, maximum amounts, expected file locations, and the user’s right to access the selected object. Avoid building shell commands or database queries by concatenating model-generated strings.
Human approval for consequential actions should be risk-based. Searching an approved knowledge base may need no interruption, while deleting records, sending external messages, changing access, or making purchases may need confirmation. Show the person the operation, destination, and important data being submitted. Bind approval to that exact action, expire it, and reject changed arguments. A general “approve this agent” switch is not equivalent to consent for every later operation.
Isolation and resource limits restrict the damage a tool can cause if its input is hostile. Do not give an agent a general-purpose shell when a narrow API will do. If code execution is genuinely required, use a separate sandbox with a non-privileged identity, restricted filesystem and network access, CPU and memory limits, and a hard timeout. Apply maximum tool-call counts, request budgets, and per-user rate limits to prevent runaway loops and resource exhaustion.
Auditing and monitoring make failures detectable. Record which authenticated principal requested which tool, the policy decision, approval state, outcome, and correlation identifier. Avoid logging full prompts, secrets, or sensitive tool arguments by default; define retention and access rules for any content you do store. For traces across model, retrieval, and tool operations, see OpenTelemetry for LLM production observability.
Real-World Use Cases
- Research assistants can search a fixed collection of approved sources. Keep retrieval read-only, enforce each user’s document permissions, and treat retrieved text as untrusted input rather than new instructions.
- Support assistants can draft a ticket update or response. Require a human to review customer-facing changes, and verify that the requester can access the ticket before exposing or modifying it.
- Developer assistants can run narrowly scoped checks or open a proposed change. Do not expose production secrets or unrestricted command execution; run builds and tests in an isolated environment with explicit limits.
- Operations assistants can look up service health and prepare a remediation. Start with read-only tools, then require a specific approval before changing production state. Keep a human-operated rollback path.
In each case, the model helps interpret a request while ordinary application code remains responsible for identity, permission, validation, and execution.
Getting Started: Add a Policy Gate
Start with one low-risk, read-only tool and a small set of explicit rules. The following Python example demonstrates an allowlist, basic argument validation, per-user rate limiting, and a dry-run write action that requires an approval record. The example keeps its handlers local and does not make network requests.
from collections import defaultdict, deque
from json import dumps
from time import monotonic
POLICY = {
"search_docs": {"scope": "docs:read", "limit": 20, "approval": False},
"create_ticket": {"scope": "tickets:write", "limit": 5, "approval": True},
}
calls = defaultdict(deque)
def canonical_action(name, arguments):
return name, dumps(arguments, sort_keys=True, separators=(",", ":"))
def dispatch(user_id, scopes, name, arguments, approved_action=None):
policy = POLICY.get(name)
if policy is None:
raise PermissionError("Tool is not allowed")
if policy["scope"] not in scopes:
raise PermissionError("User lacks the required scope")
if not isinstance(arguments, dict):
raise ValueError("Arguments must be an object")
if name == "search_docs":
if set(arguments) != {"query"}:
raise ValueError("Expected only a query")
if not isinstance(arguments["query"], str) or not arguments["query"].strip():
raise ValueError("Query must be a non-empty string")
elif name == "create_ticket":
if set(arguments) != {"project", "title"}:
raise ValueError("Expected project and title")
if not all(isinstance(arguments[key], str) and arguments[key].strip()
for key in ("project", "title")):
raise ValueError("Project and title must be non-empty strings")
if arguments["project"] not in {"helpdesk"}:
raise PermissionError("Project is not allowed")
if policy["approval"] and approved_action != canonical_action(name, arguments):
raise PermissionError("This exact action has not been approved")
now = monotonic()
key = (user_id, name)
while calls[key] and calls[key][0] <= now - 60:
calls[key].popleft()
if len(calls[key]) >= policy["limit"]:
raise PermissionError("Rate limit exceeded")
calls[key].append(now)
if name == "search_docs":
return {"matches": [], "query": arguments["query"]}
return {"dry_run": True, "project": arguments["project"],
"title": arguments["title"]}
The approval record in a real application must come from a trusted, single-use server-side approval flow, never from the model or a client-supplied boolean. Verify that it is bound to the exact canonical action and authenticated user. The in-memory rate limiter is only suitable for a single-process demonstration; production services need a shared, atomic limiter and should enforce resource ownership in the downstream system as well.
Test the gate with unknown tool names, missing scopes, extra arguments, another user’s resource identifier, approval for modified arguments, and bursts of repeated calls. Also test prompt-injection text in documents and tool results. Confirm that every rejected call is stopped before the external service is invoked. Once these checks are reliable, add tools incrementally and review permissions whenever a schema, model, or downstream service changes. For broader test techniques, see API security testing.
Common Misconceptions
“A strong system prompt prevents unsafe actions.” Prompts can guide behavior, but they are not an authorization boundary. Malicious retrieved content, ambiguous requests, or model mistakes can still produce an unsafe proposal. Enforce permissions and validation in code.
“Structured tool calls are automatically safe.” A schema helps ensure that arguments have an expected shape. It does not verify the caller’s access, the target resource’s ownership, or whether a valid action is appropriate. Perform application and service-side checks.
“Human approval makes any tool safe.” Approval can reduce risk only when the reviewer sees what will happen and approves the exact operation. A vague confirmation that can be reused or applied to changed arguments provides little protection.
Related Articles
- Model Context Protocol (MCP) explained
- OpenTelemetry for LLM production observability
- OAuth 2.0 and OpenID Connect implementation
- API security testing
Changelog and Last Updated
- Published: 2026-09-30

