OWASP Top 10 for LLM Applications 2026: Real RAG & Agent Attacks + Practical Defenses
What Is the OWASP Top 10 for LLM Applications 2026?
The OWASP Top 10 for LLM Applications 2026 is the current release of the OWASP GenAI Security Project's community-developed guide to critical security risks in applications that use large language models. The official release is published August 4, 2026, and the canonical Markdown source is maintained in the project's active repository.
The list is an application-security classification, not a list of model brands or a certification checklist. It covers failure modes in the system around the model, including how prompts, retrieved data, memory, tools, vector stores, generated code and user-visible outputs are handled.
The 2026 edition makes an important boundary explicit. When the model acts as an application component, the LLM Top 10 is the relevant lens. When it becomes an autonomous actor with tools, persistent memory and downstream consequences, the OWASP Agentic Top 10 should be read alongside it. A RAG assistant can therefore have LLM risks in its retrieval path and agentic risks in its tool path.
For a broader model comparison, read our best AI models guide. For connected development workflows, see our agentic coding guide.
What You'll Learn
- What changed in the OWASP 2026 list and how its evidence base was used.
- How prompt injection, RAG poisoning and agent permissions combine in real systems.
- Why vector stores, hidden context and generated output need separate controls.
- How to turn the ten risks into a practical engineering review.
The official OWASP 2026 preface says earlier versions were built on practitioner judgment. The 2026 release keeps that community vote as the spine of the list, then tests it against a corpus of 7,714 real incidents from public vulnerability databases and an AI-harm database. Classifiers placed 6,639 records that carried enough detail to sort.
The community vote carries 75% of the ranking weight and incident data carries 25%. OWASP says this preserves the list as a consensus product while allowing evidence to move an entry when the disagreement is wide. The incident corpus is not a measure of all risk. Prompt injection stayed first partly because mature defenses can suppress clean public incidents, while misinformation moved upward because incident records showed more operational harm than practitioners expected.
| 2026 change | Official description | Engineering consequence |
|---|---|---|
| Prompt Injection | Remains LLM01:2026 and now covers direct, indirect, multimodal, memory and cross-session forms. | Threat-model every input surface, not only the chat box. |
| Excessive Agency | Moves to LLM03:2026 as the list gives greater weight to agentic impact. | Limit tools, permissions, autonomy and state-changing actions. |
| Hidden Context Exposure | Renames and broadens the earlier system-prompt leakage concept. | Protect system prompts, memory, retrieved chunks, traces and hidden metadata. |
| Data and Model Poisoning | LLM05:2026 includes poisoning of training, fine-tuning and application data. | Verify provenance and integrity for datasets, models and retrieval corpora. |
| Improper Output Handling | Moves to LLM10:2026 and includes insecure code generated at scale. | Validate and encode output in trusted application code before use. |
OWASP also says Supply Chain now includes a promoted model artifact that is not what it claims to be. Prompt Injection covers cross-modal attacks. The release keeps the list compact by folding sharper forms of existing risk into the entries that own them rather than creating a separate category for every new technique.
For a system-level risk view, read our AI model risk management guide. The ranking tells you where to look first, but it does not replace threat modeling for a specific product.
LLM01:2026 Prompt Injection
OWASP defines prompt injection as input that alters an LLM's behavior in ways the developer did not intend. The input can be direct user text, retrieved content, tool output, an image, audio, video, intermediate reasoning or persistent memory. An LLM does not enforce an architectural distinction between instructions and data on the same token stream.
Direct injection comes from a user or attacker using the normal access path. Indirect injection arrives through a web page, document, email, support ticket, issue title, database row, tool response or RAG passage. A malicious instruction can therefore reach the model without the user seeing it. Memory persistence can spread a tainted instruction across later sessions, while agentic execution can turn the compromised output into a tool call.
OWASP's practical advice is architectural. Use strict output schemas and trusted-code validation, keep credentials and state-change capability in application code, apply least privilege per operation, label external content with provenance, treat memory writes as privileged and require human confirmation for high-impact actions. Filtering can reduce attack success, but it is not a complete prevention boundary.
| Injection surface | Example | Control priority |
|---|---|---|
| Direct input | A user asks the assistant to ignore its role and reveal private data. | Constrain role, validate output and avoid granting data access to the model. |
| Retrieved content | A poisoned web page or document contains instructions that run when summarized. | Label provenance, isolate untrusted content and authorize retrieval. |
| Tool or MCP output | A trusted connector returns attacker-controlled text that drives a later action. | Pin tools, validate arguments and mediate calls in deterministic code. |
| Memory or RAG store | A tainted entry changes later sessions that read the same store. | Review writes, log provenance and require approval for instruction-bearing memory. |
For implementation, combine input classification with blast-radius controls. A model that can read untrusted content should not automatically hold sensitive credentials, send external messages and change production state. This is why prompt injection and excessive agency often appear in the same incident even though OWASP assigns them to separate entries.
LLM02:2026 Sensitive Information Disclosure
Sensitive Information Disclosure occurs when an LLM-integrated system exposes confidential, regulated, privileged or proprietary information through a channel that the owner did not authorize. OWASP lists more than the final answer as disclosure surfaces. Tool-call arguments, reasoning traces, retrieved chunks, multimodal output, logs, telemetry, embeddings and observable inference properties can all reveal protected data.
The risk can begin during training, inference, pipeline processing or observation. Training and fine-tuning can memorize source material. Inference can overshare system prompts, files, RAG chunks, tool outputs or another session's context. Pipeline services can move sensitive examples into derived artifacts. Side channels can expose facts through timing, token length, log probabilities, confidence or cache behavior.
Start with data minimization and authorization before retrieval. Enforce document and chunk-level access controls inside the index query, isolate high-sensitivity tenants, keep secrets out of system prompts and scrub logs before observability ingestion. Treat reasoning traces and tool arguments as output that needs classification and retention rules. An external provider's no-train statement does not fix overshared context or a misconfigured internal vector store.
For data-aware AI operations, read our multi-agent protocols guide. Every connected component should have a clear answer to who may see input, output, traces, embeddings and derived artifacts.
LLM03:2026 Excessive Agency
Excessive Agency is the risk that an LLM application gives the model too much functionality, permission or autonomy. The model may choose tools, construct arguments, chain actions, write memory or change state. A harmless-looking hallucination or injected instruction can become a material incident when the system accepts the model's output as authority.
Use deterministic policy code between the model and every privileged action. Define an allowlist of tools, narrow each tool's arguments, separate read and write permissions, set time and action budgets and require approval for irreversible or externally visible changes. The reviewer should see the exact action and target, not only a model summary.
Agent designs should also consider the combination of untrusted input, sensitive data and external communication or state change. If one component can receive attacker-controlled content, read private data and send or modify something outside the model boundary, the application needs explicit residual-risk review. Multi-agent routing can spread the same permission problem across several apparently small components.
| Agency control | Recommended boundary | Failure it limits |
|---|---|---|
| Tool allowlist | Expose only named operations required for the task. | Unapproved tool selection or capability expansion. |
| Argument validation | Check types, targets, scope and ownership in trusted code. | Model-generated requests reaching an unintended resource. |
| Approval gate | Require human confirmation for irreversible or external actions. | Injected or mistaken output becoming an immediate state change. |
| Action budget | Limit calls, time, depth and retries per identity or session. | Recursive loops, runaway cost or repeated harmful attempts. |
Our agentic coding guide covers tool-driven development workflows. The same review pattern applies to support agents, RAG copilots and workflow automation.
LLM04:2026 Supply Chain
Supply Chain risk covers compromised or misleading components used to build or operate an LLM application. The 2026 OWASP release broadens this entry to include model artifacts that are not what they claim to be. The surface also includes datasets, fine-tuning adapters, prompt templates, plugins, tools, MCP servers, dependencies, containers and hosted services.
A supply-chain incident does not require a malicious base model. A poisoned package can alter a tool description, a model file can contain an unexpected payload, a dependency can exfiltrate prompts and a public retrieval connector can return attacker-controlled content. The trust decision is made during procurement or update, then materializes when the application gives the component access to context or credentials.
Pin versions, verify signatures and hashes, maintain a software and model bill of materials, review transitive dependencies, scan model files and keep third-party tools behind narrow capability policies. Test upgrades in an isolated environment and record the source, version, maintainer, permissions and rollback path. Treat tool registries and prompt libraries as production dependencies rather than informal configuration.
For a model-serving comparison, read our Nemotron 3 Super deployment guide. Open model access can improve control over hosting, but it also transfers verification and patching duties to the operator.
LLM05:2026 Data and Model Poisoning
Data and Model Poisoning occurs when an attacker or untrusted contributor modifies training, fine-tuning, evaluation or application data so the model or retrieval system behaves in a chosen way. OWASP's 2026 release folds fine-tuning subversion into this category. In an application, poisoning can target a RAG corpus, vector store, memory service, feedback dataset, prompt template or model adapter.
RAG poisoning is especially difficult because the system may retrieve the tainted record as designed. The result can be a false answer, a hidden instruction, a biased recommendation or a tool call. Cross-session poisoning becomes more serious when one user's write reaches other users or when a feedback loop promotes generated text into future training data.
Protect the data lifecycle. Record provenance, source identity, ingestion time and reviewer status. Separate write and read privileges, quarantine new material, scan for duplicate or instruction-bearing content, require approval for training promotion and maintain clean reference sets for regression tests. An integrity check should cover both the source document and the derived embedding or adapter.
For a broader RAG and model evaluation context, read our multimodal AI guide. Poisoning controls should account for text, image, audio and other modalities that reach the model.
LLM06:2026 Unbounded Consumption
Unbounded Consumption is the risk that an attacker or accidental workload consumes excessive model, memory, retrieval, tool or downstream resources. The impact can be financial cost, denial of service, queue starvation, quota exhaustion or degraded service for other tenants. The LLM can be the expensive component, but the trigger may be a prompt, document, image, recursive agent loop or oversized tool result.
Apply limits at more than one layer. Set request, token, file, retrieval, tool-call, recursion, concurrency and session budgets. Enforce timeouts and cancellation in trusted code. Rate-limit by identity and tenant, use admission control for expensive contexts, cache only where privacy permits and alert on cost or latency anomalies. A model-side token limit cannot stop a loop that repeatedly calls tools or creates new sessions.
Load-test worst-case prompts and long documents rather than only average traffic. Measure prefill, generation, retrieval, tool execution, retries and downstream API cost. Define a degraded mode that returns a bounded answer or asks for narrower input. Keep a kill switch for a runaway workflow and make billing ownership visible to the service owner.
LLM07:2026 Misinformation
Misinformation is the risk that an application produces or amplifies false, misleading or unsupported content that users or downstream systems treat as reliable. OWASP's 2026 preface says incident evidence moved misinformation upward because fluent wrong answers can drive decisions or tool calls. The problem is not limited to classic hallucination. It includes unsupported summaries, fabricated citations, incorrect transformations and confident output that hides uncertainty.
Reduce the consequence of wrong output instead of assuming a prompt can make the model truthful. Use retrieval with source attribution, constrain answers to available evidence where appropriate, display uncertainty and require verification for high-impact domains. Test stale, conflicting, incomplete and adversarial documents. Keep a human reviewer for decisions that affect money, safety, legal rights, health or access.
Validation should test claims, not only syntax. A JSON response can be valid while its facts are wrong. Compare extracted citations to source spans, flag missing evidence, detect contradictions and monitor user corrections. Do not invent confidence percentages unless a measured calibration study supports them.
For a practical AI workflow review, read our Codex versus Claude Code analysis. A coding assistant can produce syntactically valid but semantically unsafe changes, so tests and review remain necessary.
LLM08:2026 Hidden Context Exposure
Hidden Context Exposure is the 2026 name for a broader failure to protect information that the user or attacker should not receive. It includes system prompts, developer instructions, hidden memory, retrieved private chunks, internal tool descriptions, reasoning traces, metadata and other context that influences the model but is not intended for disclosure.
The rename matters because system-prompt leakage was only one example. A system can expose hidden context through direct requests, indirect injection, verbose errors, debug logs, citations, tool arguments or a model that summarizes its own operating instructions. Redacting a visible answer after generation cannot undo a private document that was already sent to the model or an access token that was logged upstream.
Keep secrets and credentials out of prompts. Treat hidden context as sensitive data with access controls, minimization, retention and audit rules. Return only the fields needed for the task, classify traces before logging and test output channels including UI rendering, exports, telemetry and connector responses. Separate the user's request from trusted policy code instead of relying on the secrecy of a prompt.
For an edge and connector perspective, see our edge AI inference guide. The hosting location does not by itself determine whether hidden context is protected.
LLM09:2026 Vector and Embedding Weaknesses
Vector and Embedding Weaknesses covers security failures in the representation and retrieval layer of an LLM application. A vector store can leak information, mix tenants, bypass document permissions, return the wrong context or accept manipulated embeddings. Embeddings are not automatically anonymous. Inversion and similarity queries can expose information from source material, while poor filtering can retrieve data the requester is not authorized to see.
Enforce authorization before retrieval, not after the model has received the chunk. Keep document ACLs separate from vector-store permissions, isolate high-sensitivity tenants, restrict export and nearest-neighbor APIs, encrypt stores and monitor probing patterns. Record document identity and access decision with every retrieved chunk. Re-indexing a corrected source does not automatically remove every derived representation from backups or caches.
Test retrieval for cross-tenant leakage, stale permissions, poisoned records, partial deletion, near-duplicate documents and prompts that ask for information by concept rather than by name. Evaluate multilingual and multimodal content when the system supports it. A similarity score is a retrieval signal, not an authorization decision.
Our AI inference hardware guide explains why memory and serving choices matter. For vector security, pair performance testing with authorization and deletion tests.
LLM10:2026 Improper Output Handling
Improper Output Handling occurs when an application trusts or embeds model output without applying the validation, encoding, escaping and authorization that the destination requires. OWASP's 2026 release places this entry at LLM10 and expands it to insecure code that assistants generate at scale. The output may be displayed in a browser, passed to SQL, executed in a shell, written to a file, used as a workflow command or sent to another model.
| Output destination | Failure mode | Required control |
|---|---|---|
| Browser or HTML | Generated markup changes the page or carries script content. | Escape by context, sanitize allowed markup and use a strict content policy. |
| SQL or query API | Model text is concatenated into a query or filter. | Use parameterized interfaces, schema validation and an allowlist of operations. |
| Shell or code runner | Generated code reaches an execution environment without review. | Use isolation, least privilege, time limits, tests and human approval. |
| Workflow or external message | Model output triggers a state change or sends unreviewed content. | Validate fields, confirm the exact action and enforce policy in application code. |
Never solve this risk by asking a second model whether the first output is safe. Use deterministic validation for structure, destination-specific encoding, permission checks and a controlled execution environment. The model can help draft an action, but the application must decide whether the action is allowed.
For agent security, read our OWASP agentic AI security guide. The strongest control is to keep generated text away from privileged interpreters until a trusted boundary has approved it.
How to Turn the OWASP List into an Engineering Review
Start by drawing the complete data and action path. Mark every place where the system accepts user, web, document, email, tool, memory or model-generated input. Record which components can read private data, write persistent state, call external services or execute code. Then map each boundary to the ten OWASP entries and to the Agentic Top 10 where the model acts as an actor.
Use a small test set for each risk. Include direct and indirect prompt injection, cross-tenant retrieval, poisoned documents, stale ACLs, malformed output, long inputs, recursive tool calls, fabricated citations and hidden-context requests. Log the model version, prompt, retrieval set, tools, permissions, output and reviewer decision so a finding can be reproduced.
Prioritize controls that survive model failure. Least-privilege tools, pre-retrieval authorization, signed dependencies, bounded budgets, destination-specific encoding, isolated execution and human approval reduce the damage when the model is wrong or manipulated. Prompt wording and filters can help, but they should not be the only control supporting a high-impact action.
For a connected-agent design, continue with our NemoClaw versus OpenClaw comparison and multi-agent protocols guide. The review is complete only when owners, evidence, residual risk and remediation dates are recorded.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles