ChecklistAI Agents

AI agent security checklist: what to verify before an agent gets tools and credentials

A chatbot that says something wrong is embarrassing; an agent that does something wrong can move money, leak records or change systems. This checklist covers the controls to verify before an agent receives tools and credentials: untrusted inputs, tool validation, identity, execution and memory, oversight and third-party components. A closing table maps each area to OWASP's risk categories for LLM and agentic applications.

Reviewed 7 min read

On this page
  1. How tools, credentials and memory change the threat model
  2. How untrusted content reaches an agent's tools
  3. Inputs: treat everything the agent reads as untrusted
  4. Tools: allow-lists, schemas and validation outside the model
  5. Identity and credentials for each agent
  6. Execution, memory and stored state
  7. Oversight, tracing and the kill switch
  8. Supply chain: tools, plug-ins and MCP servers
  9. Mapping the checklist to OWASP risk categories
  10. Questions and answers
  11. Sources

How tools, credentials and memory change the threat model

A model without tools can only produce text, so the worst outcome is misleading or leaked text. Give it tools and its output becomes a command: a database query, an email, a payment instruction, a file write. Give it credentials and those commands run with real authority. Give it memory and one malicious input can shape decisions days later.

The underlying problem is that a language model cannot reliably separate instructions from data. Text inside a retrieved document, a support email or a tool response can be read as an instruction, which is why OWASP places prompt injection first among risks to LLM applications and agent goal hijack first among risks to agentic applications12. No wording in a prompt fully prevents it, so the controls below assume injection will sometimes succeed and limit what an injected agent can do.

Apply the checklist to the agent's design before choosing or hardening a framework or host; host-level hardening for open-source frameworks is covered on the OpenClaw and Hermes page.

How untrusted content reaches an agent's tools

fetchedinsertedinfluencescheckedonly if allowed01Untrusted content02Retrieval or tool read03Model context04Proposed tool call05Validation and approval06System of record
  1. Untrusted content

    Emails, uploads, web pages, tickets and tool outputs written by someone outside your control.

  2. Retrieval or tool read

    The agent fetches the content as part of a legitimate task.

  3. Model context

    Instructions hidden in that content now sit beside your own system prompt.

  4. Proposed tool call

    The model may propose the action an attacker wanted, with plausible arguments.

  5. Validation and approval

    Schema checks, policy rules and approval points outside the model decide whether it runs.

  6. System of record

    Only validated, permitted actions reach real data and services.

Conceptual path of an indirect prompt injection. Controls must hold at the validation step, because none of the earlier steps can be fully trusted.

Inputs: treat everything the agent reads as untrusted

0 of 5 checked

Tools: allow-lists, schemas and validation outside the model

0 of 5 checked

Identity and credentials for each agent

0 of 5 checked

Execution, memory and stored state

0 of 5 checked

Oversight, tracing and the kill switch

0 of 5 checked

Supply chain: tools, plug-ins and MCP servers

0 of 5 checked

Mapping the checklist to OWASP risk categories

Category names come from the OWASP Top 10 for LLM Applications 2025 and the OWASP Top 10 for Agentic Applications for 2026, published in December 202512. The mapping is our reading of where each control area mainly applies; most controls reduce more than one risk.

Checklist areaOWASP LLM categoryOWASP agentic category
InputsLLM01:2025 Prompt InjectionASI01: Agent Goal Hijack
ToolsLLM06:2025 Excessive Agency; LLM05:2025 Improper Output HandlingASI02: Tool Misuse and Exploitation
Identity and credentialsLLM06:2025 Excessive Agency; LLM02:2025 Sensitive Information DisclosureASI03: Identity and Privilege Abuse
ExecutionLLM05:2025 Improper Output HandlingASI05: Unexpected Code Execution (RCE)
Memory and stateLLM04:2025 Data and Model Poisoning; LLM08:2025 Vector and Embedding WeaknessesASI06: Memory & Context Poisoning
Oversight and tracingLLM10:2025 Unbounded ConsumptionASI09: Human-Agent Trust Exploitation; ASI10: Rogue Agents
Supply chainLLM03:2025 Supply ChainASI04: Agentic Supply Chain Vulnerabilities

ASI07 (Insecure Inter-Agent Communication) and ASI08 (Cascading Failures) mainly concern systems of several cooperating agents, covered on the agent swarms page.

Questions and answers

Can prompt injection be fully prevented in an AI agent?

Not with current models. Filters, delimiters and careful instructions reduce how often injection works, but a determined attacker can usually find wording that gets through. Design on the assumption that it will sometimes succeed: limit the tools and data an agent can reach, validate actions outside the model and require approval for anything with serious consequences.

How do we red-team an AI agent before launch?

Write attack scenarios against the agent's real tools and data: hidden instructions in documents and emails, requests to reveal its system prompt, attempts to call tools outside its role or with manipulated arguments, and inputs designed to make it loop. Run them in a sandbox with realistic permissions, record which control stopped each attempt and keep the successful ones as permanent test cases.

What should be logged for a tool-using AI agent?

Enough to reconstruct any action: the triggering input, sources retrieved, model and prompt versions, every proposed and executed tool call with arguments and response, validation failures, approvals with the reviewer's identity and the final outcome. Protect the logs themselves, because they may hold personal data and secrets, and set retention to match your obligations.

Are system prompt instructions enough to keep an agent safe?

No. Instructions such as 'only use approved tools' are useful guidance, but a model can be talked out of them, and the system prompt itself may leak. Anything that must never happen has to be enforced by permissions, argument validation, sandboxing or approval, in code the model cannot change.

Is it safe to connect third-party MCP servers to an agent?

It can be, with the diligence you would apply to any dependency that touches your data. Check who maintains the server, what it can reach, how it handles authorization and whether it passes tokens through to other services. Pin the version, run it with minimal privileges and treat its tool descriptions as untrusted text the model will read.

Sources

  1. OWASP Top 10 for LLM Applications 2025 — OWASP GenAI Security Project · checked 10 October 2026
  2. OWASP Top 10 for Agentic Applications for 2026 (Agentic Security Initiative, December 2025) — OWASP GenAI Security Project · checked 10 October 2026
  3. Model Context Protocol: Security Best Practices — Model Context Protocol · checked 10 October 2026

More in AI Agents

Back to AI Agents

Next step

Have your agent's design reviewed before it reaches production

Send the agent's role, its tool list and the credentials it will hold. We will work through this checklist with your security team and flag the controls to add before launch.

Request an agent security review