Zero-trust AI¶
Zero trust grants no implicit trust because of network location, component type, model reputation or earlier interactions. For AI, that includes the model itself: a model can be steered by any text in its context, so it cannot be the component that decides what is allowed.
Principles¶
- Verify the identity of every user, workload, service, model, agent, tool and data source.
- Authorize the purpose and the action at each consequential trust boundary.
- Grant least privilege through short-lived, task-bound credentials.
- Treat prompts, retrieved content, model output and tool responses as untrusted.
- Keep control instructions apart from data, and carry provenance with every piece of content.
- Assume breach. Limit the blast radius by tenant, task, data class and transaction.
- Evaluate identity, device, behavior, policy and system health signals continuously.
Enforce outside the model¶
Put every security decision in deterministic code that the model cannot rewrite.
| Decision | Where to enforce it |
|---|---|
| Who may see a document | The retrieval service, using the calling user's identity, before content reaches the model |
| Whether a tool call is allowed | A policy check between the agent and the tool, using the task's scope |
| Whether an action needs approval | A gate that pauses the workflow and records the approver |
| Whether output is safe to use | A validator for the destination: HTML encoding, SQL parameters, schema checks |
A system prompt that says "never reveal salary data" is a hint, not a control.
In practice
AEGIS enforces this with a deterministic hook: code agents cannot write outside the active work package's declared paths. See the AEGIS reference implementation. The framework does not require any particular tool.
Agent and tool controls¶
| Control | What it prevents |
|---|---|
| Allowlisted, typed tools with narrow schemas | The agent calling tools it does not need, or passing arbitrary input |
| Deny-by-default scopes, one credential per tool and task | One compromised call reaching everything the agent can reach |
| Transaction and rate limits | Runaway loops, bulk exfiltration, cost attacks |
| Egress allowlists | Data sent to attacker-controlled endpoints |
| Sandboxes with no standing credentials for generated code | Code execution escaping into the host or network |
| Human approval for consequential or irreversible actions | Injected instructions completing high-impact actions alone |
| Idempotency keys and rollback | Duplicate or unrecoverable actions after retries |
| Correlation IDs across user, agent, model and tool | Investigations that cannot reconstruct what happened |
Never put credentials in prompts or anywhere the model can read them.
Agent identity¶
Give each agent its own workload identity, separate from the user it acts for. Record both identities on every action, so the audit trail shows who asked and what acted. When an agent acts for a user, pass a token scoped to that user and task, and never a broad service credential.
In multi-agent systems, authenticate each agent to the others. Validate every message against a schema before acting on it, and do not let one agent grant another more privilege than it holds.
MCP authorization¶
The Model Context Protocol (MCP) specification, version 2026-07-28, bases authorization on OAuth 2.1. Its security best practices name specific attacks. Check each MCP server you run or connect to against them.
| Requirement | Why |
|---|---|
| An MCP server must not accept tokens that were not issued for it | Token passthrough lets a client reuse a token meant for another service and bypass its controls |
| Validate the token audience, and use resource indicators (RFC 8707) | Tokens stay bound to one server |
| A proxy server that uses a static client ID must get per-client consent before forwarding to a third-party authorization server | Prevents the confused deputy problem |
Exact-match redirect URI validation, and a single-use state value | Prevents authorization code theft |
| Request the minimum scopes | Limits what a stolen token can do |
Tool descriptions are part of the attack surface as well. The threat model covers tool poisoning, rug pulls and tool shadowing.
Sources¶
Checked 2026-09-28.