ChecklistAI Agents
AI agent security checklist: what to verify before an agent gets tools and credentials
A chatbot that says something wrong is embarrassing; an agent that does something wrong can move money, leak records or change systems. This checklist covers the controls to verify before an agent receives tools and credentials: untrusted inputs, tool validation, identity, execution and memory, oversight and third-party components. A closing table maps each area to OWASP's risk categories for LLM and agentic applications.
On this page
- How tools, credentials and memory change the threat model
- How untrusted content reaches an agent's tools
- Inputs: treat everything the agent reads as untrusted
- Tools: allow-lists, schemas and validation outside the model
- Identity and credentials for each agent
- Execution, memory and stored state
- Oversight, tracing and the kill switch
- Supply chain: tools, plug-ins and MCP servers
- Mapping the checklist to OWASP risk categories
- Questions and answers
- Sources
How tools, credentials and memory change the threat model
A model without tools can only produce text, so the worst outcome is misleading or leaked text. Give it tools and its output becomes a command: a database query, an email, a payment instruction, a file write. Give it credentials and those commands run with real authority. Give it memory and one malicious input can shape decisions days later.
The underlying problem is that a language model cannot reliably separate instructions from data. Text inside a retrieved document, a support email or a tool response can be read as an instruction, which is why OWASP places prompt injection first among risks to LLM applications and agent goal hijack first among risks to agentic applications12. No wording in a prompt fully prevents it, so the controls below assume injection will sometimes succeed and limit what an injected agent can do.
Apply the checklist to the agent's design before choosing or hardening a framework or host; host-level hardening for open-source frameworks is covered on the OpenClaw and Hermes page.
How untrusted content reaches an agent's tools
- Untrusted content
Emails, uploads, web pages, tickets and tool outputs written by someone outside your control.
- Retrieval or tool read
The agent fetches the content as part of a legitimate task.
- Model context
Instructions hidden in that content now sit beside your own system prompt.
- Proposed tool call
The model may propose the action an attacker wanted, with plausible arguments.
- Validation and approval
Schema checks, policy rules and approval points outside the model decide whether it runs.
- System of record
Only validated, permitted actions reach real data and services.
Inputs: treat everything the agent reads as untrusted
Tools: allow-lists, schemas and validation outside the model
Identity and credentials for each agent
Execution, memory and stored state
Oversight, tracing and the kill switch
Supply chain: tools, plug-ins and MCP servers
Mapping the checklist to OWASP risk categories
Category names come from the OWASP Top 10 for LLM Applications 2025 and the OWASP Top 10 for Agentic Applications for 2026, published in December 202512. The mapping is our reading of where each control area mainly applies; most controls reduce more than one risk.
| Checklist area | OWASP LLM category | OWASP agentic category |
|---|---|---|
| Inputs | LLM01:2025 Prompt Injection | ASI01: Agent Goal Hijack |
| Tools | LLM06:2025 Excessive Agency; LLM05:2025 Improper Output Handling | ASI02: Tool Misuse and Exploitation |
| Identity and credentials | LLM06:2025 Excessive Agency; LLM02:2025 Sensitive Information Disclosure | ASI03: Identity and Privilege Abuse |
| Execution | LLM05:2025 Improper Output Handling | ASI05: Unexpected Code Execution (RCE) |
| Memory and state | LLM04:2025 Data and Model Poisoning; LLM08:2025 Vector and Embedding Weaknesses | ASI06: Memory & Context Poisoning |
| Oversight and tracing | LLM10:2025 Unbounded Consumption | ASI09: Human-Agent Trust Exploitation; ASI10: Rogue Agents |
| Supply chain | LLM03:2025 Supply Chain | ASI04: Agentic Supply Chain Vulnerabilities |
ASI07 (Insecure Inter-Agent Communication) and ASI08 (Cascading Failures) mainly concern systems of several cooperating agents, covered on the agent swarms page.
Questions and answers
Can prompt injection be fully prevented in an AI agent?
Not with current models. Filters, delimiters and careful instructions reduce how often injection works, but a determined attacker can usually find wording that gets through. Design on the assumption that it will sometimes succeed: limit the tools and data an agent can reach, validate actions outside the model and require approval for anything with serious consequences.
How do we red-team an AI agent before launch?
Write attack scenarios against the agent's real tools and data: hidden instructions in documents and emails, requests to reveal its system prompt, attempts to call tools outside its role or with manipulated arguments, and inputs designed to make it loop. Run them in a sandbox with realistic permissions, record which control stopped each attempt and keep the successful ones as permanent test cases.
What should be logged for a tool-using AI agent?
Enough to reconstruct any action: the triggering input, sources retrieved, model and prompt versions, every proposed and executed tool call with arguments and response, validation failures, approvals with the reviewer's identity and the final outcome. Protect the logs themselves, because they may hold personal data and secrets, and set retention to match your obligations.
Are system prompt instructions enough to keep an agent safe?
No. Instructions such as 'only use approved tools' are useful guidance, but a model can be talked out of them, and the system prompt itself may leak. Anything that must never happen has to be enforced by permissions, argument validation, sandboxing or approval, in code the model cannot change.
Is it safe to connect third-party MCP servers to an agent?
It can be, with the diligence you would apply to any dependency that touches your data. Check who maintains the server, what it can reach, how it handles authorization and whether it passes tokens through to other services. Pin the version, run it with minimal privileges and treat its tool descriptions as untrusted text the model will read.
Sources
- OWASP Top 10 for LLM Applications 2025 — OWASP GenAI Security Project · checked 10 October 2026
- OWASP Top 10 for Agentic Applications for 2026 (Agentic Security Initiative, December 2025) — OWASP GenAI Security Project · checked 10 October 2026
- Model Context Protocol: Security Best Practices — Model Context Protocol · checked 10 October 2026