Skip to content

Enterprise AI threat model

  • Draft
  • v0.3.0
  • Security architects and engineers
  • Reviewed 2026-09-28

An AI threat model covers the same ground as any threat model: adversarial attack, accidental failure, misuse by authorized people, supplier compromise and unexpected behavior. AI adds two things. Natural language mixes instructions with data, so any text the model reads can try to steer it. Agents act on the model's output, so a steered model can take real actions.

This page names threats with public identifiers so teams can trace them to tests, controls and incidents:

Catalog Edition used here
OWASP Top 10 for LLM Applications 2026, released August 2026 (IDs such as LLM01:2026)
OWASP Top 10 for Agentic Applications 2026, released 2025-12-09 (IDs such as ASI01)
MITRE ATLAS v2026.09 (IDs such as AML.T0051)
NIST AI 100-2e2025 Adversarial machine learning taxonomy, March 2025

Record the edition next to every ID. The 2026 LLM list reordered the 2025 list, so the same number now names a different risk.

2025 ID 2026 ID Risk
LLM01:2025 LLM01:2026 Prompt Injection
LLM02:2025 LLM02:2026 Sensitive Information Disclosure
LLM03:2025 LLM04:2026 Supply Chain
LLM04:2025 LLM05:2026 Data and Model Poisoning
LLM05:2025 LLM10:2026 Improper Output Handling
LLM06:2025 LLM03:2026 Excessive Agency
LLM07:2025 LLM08:2026 System Prompt Leakage, renamed Hidden Context Exposure
LLM08:2025 LLM09:2026 Vector and Embedding Weaknesses
LLM09:2025 LLM07:2026 Misinformation
LLM10:2025 LLM06:2026 Unbounded Consumption

OWASP scopes the LLM list to a model used as a component. When the model acts through tools, memory or other agents, use the Agentic list alongside it.

Assets

Asset Why it matters
Mission decisions and affected people's rights The harm an AI failure causes lands here
Sensitive data Prompts, retrieved documents, logs, training and fine-tuning data
Identities, credentials and authorization Agents and tools act with them
System prompts, policies and tool definitions They steer the model and describe what it can do
Models, weights and adapters Stolen, swapped or tampered artifacts change behavior
Retrieval indexes, vector stores and agent memory Poisoned entries persist and reach many users
Tools, plugins and MCP servers They turn model output into actions
Evaluation sets, logs and evidence Assurance and investigation depend on them
Availability and budget Model calls cost money and capacity
Institutional trust Public-facing failures damage it quickly

Trust boundaries

Draw each boundary on the system diagram and record what crosses it.

Boundary What crosses it Default stance
User to application Prompts, files, images Untrusted input
Application to model Assembled context, system prompt The model may follow any instruction in the context
Retrieval to model Documents, web pages, emails, tickets Untrusted input, even from internal sources
Model to downstream systems Generated text, code, queries, tool calls Untrusted output until validated
Agent to tool or MCP server Tool calls, credentials, returned data Authorize each call; treat returned data as untrusted
Agent to agent Messages, delegated tasks Authenticate both sides; validate every message
Organization to supplier Models, datasets, hosted APIs, tool servers Verify provenance and integrity

Threat catalog

Threat OWASP MITRE ATLAS Main mitigations
Direct and indirect prompt injection LLM01:2026, ASI01 AML.T0051 (.000 direct, .001 indirect), AML.T0054 jailbreak Separate instructions from data, mark provenance, least-privilege tools, human approval for consequential actions
Sensitive information disclosure LLM02:2026 AML.T0057, AML.T0024 Minimize context, authorize retrieval per user, filter output, redact logs
AI supply chain compromise LLM04:2026, ASI04 AML.T0010 (.001 software, .002 data, .003 model, .005 agent tool) Provenance, AI-BOM, signed models, pinned versions (see AI supply chain)
Data and model poisoning LLM05:2026 AML.T0020, AML.T0018 Curate and version data, verify artifacts, evaluate before release, keep rollback ready
Improper output handling LLM10:2026, ASI05 Validate and encode output for its destination, parameterize queries, sandbox generated code
Excessive agency LLM03:2026, ASI02, ASI03 AML.T0053, AML.T0086, AML.T0101 Narrow tools, scoped short-lived credentials, transaction limits, approval gates
Hidden context exposure (formerly system prompt leakage) LLM08:2026 AML.T0069 (.002 system prompt) Keep secrets and authorization logic out of prompts, enforce controls outside the model
Vector and embedding weaknesses, RAG poisoning LLM09:2026, ASI06 AML.T0070, AML.T0071, AML.T0082 Authorize at retrieval time, isolate tenants, validate sources before indexing
Misinformation and confabulation LLM07:2026 Grounding, citations, abstention, human review of consequential output
Unbounded consumption LLM06:2026 AML.T0034, AML.T0029 Quotas, budgets, rate limits, circuit breakers, graceful degradation
Tool poisoning, rug pulls and tool shadowing ASI04 AML.T0110, AML.T0099, AML.T0011.002 Pin and review tool definitions, alert on changes, isolate servers, allowlist tools
Memory and context poisoning ASI06 AML.T0080 (.000 memory) Scope memory per user and task, expire it, validate before writing
Insecure inter-agent communication ASI07 AML.T0118 Mutual authentication, signed messages, schema validation
Cascading failures across agents ASI08 Isolation, timeouts, circuit breakers, bounded retries
Human-agent trust exploitation ASI09 AML.T0100 Clear disclosure, friction before consequential approval, training against automation bias
Rogue agents ASI10 AML.T0081 Inventory agents, monitor behavior against a baseline, keep a kill switch
Cross-tenant leakage LLM02:2026, LLM09:2026 Tenant isolation in indexes, caches and memory; penetration testing
Audit manipulation Append-only evidence, separation of duties, synchronized time

AI-specific threats in more depth

Improper output handling

Model output is untrusted input to whatever consumes it. OWASP defines the weakness as "insufficient validation, sanitization, and handling of the outputs generated by large language models before they are passed downstream". Generated text reaches browsers, shells, SQL engines, template renderers and other agents. Encode output for its destination, use parameterized queries, run generated code in a sandbox with no standing credentials, and never pass output to eval or a shell.

Excessive agency and multi-agent risk

An agent can do as much damage as its tools allow, whatever its instructions say. Give each agent the fewest tools, the narrowest scopes and the shortest-lived credentials its task needs. Require human approval for actions that are consequential, hard to reverse or outside a transaction limit. In multi-agent systems, authenticate every agent, validate every message against a schema, and stop one agent's failure from cascading with timeouts and circuit breakers.

OWASP's 2026 prompt injection entry recommends the "Rule of Two" as a minimum check on agent capabilities. Look for three capabilities in each agent:

Capability Example
A. Reads untrusted input Web pages, emails, tickets, retrieved documents
B. Reaches sensitive data Case files, personal records, credentials
C. Changes state or communicates externally Sends email, writes records, calls external APIs

An agent with all three needs human approval for every action. An agent with two of them needs a documented residual-risk assessment.

In practice

AEGIS limits each code agent to one approved work package and its paths, and treats external content as data. See the AEGIS reference implementation. The framework does not require any particular tool.

MCP servers and tool poisoning

The Model Context Protocol (MCP) connects agents to tools. The model reads tool descriptions, so a malicious or compromised server can hide instructions in them. Invariant Labs described these tool poisoning attacks in April 2025. The same report named two variants.

Variant What happens
Rug pull The server changes a tool description after the user approved it
Tool shadowing One server's descriptions change how the agent uses other, trusted servers

Pin tool definitions and alert on changes, show users the full description, run each server in isolation, and allowlist the servers an agent may use. The zero-trust AI page covers MCP authorization controls.

Hidden context exposure

Assume users can extract anything in the model's context: the system prompt, developer instructions, retrieved policy text and tool schemas. OWASP's 2026 entry defines the risk as "the unauthorized extraction, inference, or reconstruction of hidden, non-user-facing system instructions or operational context placed in a model's context". The fix is architectural: keep credentials, internal hostnames and authorization rules out of prompts, and enforce access control in code outside the model.

Required outputs

Record these for every AI system, and review them at each lifecycle gate:

  • system and trust-boundary diagrams
  • assumptions and out-of-scope threats
  • threats with their catalog IDs and editions
  • affected assets
  • controls and their validation evidence
  • residual risks, with owners and acceptance decisions
  • reassessment triggers

Link each threat to a detection in production and to an incident response playbook.

Sources

Checked 2026-09-28.