AI agent safety is the engineering and governance discipline that ensures autonomous agents act within prescribed, auditable constraints. The three immediate priorities are identity and least privilege, runtime enforcement at the point of action, and continuous adversarial testing. Everything else, from guardrails to compliance mapping, builds on these three foundations.
TL;DR:
- Runtime gates should check actions before, during, and after execution; controlled experiments found this architecture reduced soft window overages by 89.5% versus policy as code alone.
- Give each agent a short lived, task scoped token, and require named human approval for privilege expansions while allowing permission tightening automatically.
- A 2025 red team competition tested agents across 44 scenarios; some injections reached 100% success, so convert every failure into a regression test.
- Inventory agents, tools, and credentials first; by day 90, add runtime gates and approval workflows, then establish recurring red team exercises by day 180.
- For irreversible payments or deletions, deny actions until a human approves them; egress allowlists can also block prompt injection leaks to unfamiliar destinations.
Table of Contents
- Why agentic AI changes the security model
- Top agentic AI threats you need to triage first
- Runtime governance and enforcement architectures
- Locking down agent privileges with per-action authorization
- Defensive controls that reduce attack surface and contain failures
- Testing, monitoring, and incident response for agent-specific failures
- Mapping agent safety to governance and audit requirements
- How we apply these controls in production agent deployments
- A 90 to 180 day priority list for operationalizing agent safety
- Get a Dogtooth assessment of your agent deployment
- FAQ
- Sources
Why agentic AI changes the security model
Traditional application security assumes a fixed set of inputs and outputs you can validate at the edge. Agentic systems break that assumption. An agent plans multi-step tasks, calls tools, writes to memory, and persists state across sessions, which means the attack surface now includes the agent’s reasoning process itself, not just its interfaces.

Because agent reasoning is probabilistic, you cannot rely on the model to police itself. Enforcement has to sit outside the model, at runtime, where it can deterministically allow or block an action regardless of how the agent arrived at its decision.
Three failure modes dominate real incidents:
- Goal hijacking: an attacker injects instructions that redirect the agent’s objective mid-task.
- Memory poisoning: corrupted or adversarial data persists into an agent’s long-term memory and influences future decisions.
- Tool misuse: an agent calls a legitimate tool in an unintended way, often because its permissions are broader than the task requires.
The OWASP Top 10 for Agentic Applications formalizes these patterns into a ranked taxonomy, and the EchoLeak incident, detailed below, shows how quickly a single injection path can turn into data exfiltration with no user interaction at all.
Top agentic AI threats you need to triage first
Ranking threats by operational impact helps security and development teams decide where to spend limited review time. The following sequence follows the themes the OWASP Top 10 for Agentic Applications consolidates from community-reported incidents.
- Goal manipulation: injected content redirects an agent’s objective. Watch for sudden shifts in stated intent mid-session; mitigate with input partitioning and goal-state checkpoints.
- Tool misuse: an agent invokes a tool outside its intended scope. Watch for anomalous tool call sequences; mitigate with per-action authorization.
- Identity and privilege abuse: an agent accumulates permissions beyond its task. Watch for privilege-expansion events in access logs; mitigate with deny-by-default policy updates.
- Memory poisoning: adversarial data persists and skews future reasoning. Watch for inconsistent outputs tied to specific memory entries; mitigate with memory diffing and provenance tags.
- Inter-agent attacks: one agent manipulates another in a multi-agent pipeline. Watch for message content that does not match the sender’s declared role; mitigate with signed inter-agent messages.
- Cascading failures: a single bad action propagates through downstream automated steps. Watch for rapid-fire repeated actions; mitigate with rate limits and reversibility windows.
- Human-agent trust exploitation: an agent’s confident tone masks an incorrect or harmful action. Watch for low-confidence actions presented as certain; mitigate with mandatory human review on high-impact steps.
- Rogue agent behavior: an agent deviates from its assigned task entirely. Watch for actions outside its declared capability set; mitigate with hard capability boundaries.
- Unauthenticated data exfiltration: a prompt injection triggers data leakage with zero user interaction, as seen in the EchoLeak exploit. Watch for outbound requests to unfamiliar domains; mitigate with egress controls.
- Supply-chain compromise: a poisoned tool, plugin, or model update introduces malicious behavior. Watch for unsigned or unpinned dependencies; mitigate with signed manifests.
Pro Tip: Instrument detection around tool-call anomalies and privilege-expansion events first. Those two signal families catch the majority of the threat categories above before they escalate.
Runtime governance and enforcement architectures
Policy documents do not stop an agent mid-action. That gap between written compliance and live behavior is exactly what runtime governance architectures close. The SARC architecture compiles legal and contractual obligations directly into runtime constraints, checked at three distinct points in an agent’s execution.
- Pre-Action Gate (PAG): evaluates a proposed action against policy before execution begins.
- Action-Time Monitor (ATM): observes the action as it executes and can interrupt it mid-flight.
- Post-Action Auditor (PAA): verifies the completed action against expected outcomes and logs the result.
- Escalation Router: routes ambiguous or high-impact decisions to a human reviewer instead of auto-approving or auto-denying.
Each action type maps to a different combination of these checkpoints. A low-risk read operation might only pass through the PAG, while a payment or data-deletion action triggers all four.
In controlled experiments, compiling obligations into these runtime checkpoints reduced soft-window overages by 89.5% compared to a policy-as-code-only baseline, a meaningful gap between documentation-based compliance and enforcement that actually holds at execution time.
The operational tradeoff is latency. Every checkpoint adds milliseconds to an action, and a fail-closed default, where an unverifiable action is blocked rather than allowed, adds friction on the margin. For high-impact or irreversible actions, that friction is the point: a short reversibility window, where an action can still be rolled back after execution, matters more than shaving latency.
Locking down agent privileges with per-action authorization
Least privilege sounds simple until you try to enforce it on an agent that calls dozens of tools across a session. The Progent framework addresses this by encoding privileges as symbolic rules over tool names and arguments, then using a deterministic SMT solver to classify every proposed policy update as either a narrowing change or an expansion.
- Narrowing updates (tightening what an agent can do) are allowed automatically.
- Expansion updates (granting new capability) require explicit human approval before they take effect.
This distinction matters because Progent’s SMT-based check prevents silent privilege escalation: an agent cannot quietly widen its own access just because a task seems to require it. The approval step is the control, not a suggestion layered on top of it.
Pairing this with identity-first controls closes the remaining gap. Give each agent a short-lived, scoped token tied to a specific task rather than a standing credential, and tie token issuance to a per-agent identity lifecycle that mirrors how you manage human service accounts, including rotation and revocation.
Pro Tip: Configure your approval flow to deny privilege expansions by default and route them to a named human reviewer, never a general queue. Specificity in the approval chain is what keeps expansion requests from piling up unreviewed.
Defensive controls that reduce attack surface and contain failures
Beyond identity and runtime gates, a set of concrete engineering patterns limits how far a compromised agent can get.
- Input and output guardrails: partition untrusted content (web pages, documents, emails) from trusted instructions, and sanitize both directions of the interaction.
- RAG hygiene: apply content disarm and reconstruction (CDR) to uploaded documents before they enter a retrieval pipeline, so embedded instructions cannot reach the model unfiltered.
- Provenance controls: sign prompts and tool definitions, maintain an AIBOM or SBOM for every component in the agent’s pipeline, and pin dependencies by hash rather than by version tag.
- Staged rollout: deploy agent updates to a small percentage of traffic first, with automatic rollback triggers if error or escalation rates spike.
- Per-session sandboxing: isolate each agent session with its own syscall and network limits, so a compromised session cannot reach other sessions or persistent infrastructure.
- Egress controls: restrict outbound network calls to an allowlist of known destinations, which is precisely the control that would have blunted the EchoLeak exfiltration path.
OpenAI’s agent SDK documentation describes a related pattern worth adopting directly: guardrails applied separately at the input, tool, and output layers, combined with an approval interruption flow that records the interruption, produces a resumable state, and only continues the run once a human approves it. None of these controls work in isolation. Provenance without sandboxing still lets a poisoned tool run unchecked, and sandboxing without egress controls still lets data leave the boundary.
Testing, monitoring, and incident response for agent-specific failures
Static review catches configuration mistakes, not the dynamic ways an agent can be manipulated in production. Continuous red-teaming closes that gap. A 2025 competition that ran adversarial prompts against agents across 44 real-world scenarios found that some prompt injections achieved 100% attack success rates, with participants submitting over 1.8 million adversarial prompts in aggregate. Every failure that testing surfaces should become a permanent regression test, not a one-off fix.
- Collect telemetry on the full reasoning chain: tool calls, goal state changes, memory diffs, inter-agent messages, and every approval decision.
- Preserve forensics artifacts the moment an incident is detected: the decision-chain log, the exact prompt and context window, and the token or session identifier involved.
- Contain immediately: revoke the affected token, trigger the kill switch for that agent instance, and isolate the session before investigating root cause.
- Remediate and regress: patch the gap, then add the exact attack pattern to your adversarial test suite so it cannot recur silently.
Mapping agent safety to governance and audit requirements
Security controls only satisfy a compliance function when they produce evidence. The NIST AI RMF profile for Generative AI organizes that evidence around four functions: govern, map, measure, and manage. Runtime traces from your PAG, ATM, and PAA checkpoints map directly onto the “measure” and “manage” functions, giving auditors a live record instead of a point-in-time attestation.
Specific artifacts worth producing on an ongoing basis:
- Decision-chain logs showing every action an agent proposed, whether it was allowed, and why.
- Approval records for every privilege expansion or high-impact action routed to a human.
- SBOM/AIBOM attestations covering every model, tool, and dependency in the pipeline.
- Red-team reports documenting what was tested, what failed, and what changed as a result.
Risk tiering determines how strict enforcement needs to be. Actions with a short or nonexistent reversibility window, like an irreversible financial transfer or a permanent data deletion, warrant a default-deny posture until a human signs off. Related trace-based enforcement work shows the payoff of this rigor: a runtime monitor checking GDPR-style predicates over execution traces achieved attack-success rates at or below 12% under noisy extraction conditions, and zero under perfect extraction, evidence that compliance predicates enforced at runtime hold up even under adversarial pressure. For teams building out a formal governance playbook, the AI Agent Governance Framework from Everything Cloud offers templates for structuring oversight across cloud and AI workloads.
How we apply these controls in production agent deployments
We build agents that act inside real workflows rather than staying confined to a chat window, which means the controls above are not theoretical for us. Our AI booking system for three combat sports brands books sessions, charges cards, and sends invites across three separate brands, and every payment action runs through scoped, per-action authorization rather than a standing credential with broad access.
Our practitioner portal with Stripe Connect payouts applies the same pattern to a different guarded interaction: payout actions are logged, scoped, and routed through approval before funds move. We log the full decision chain on both systems, not just the outcome, because that is what turns an incident response from guesswork into a five-minute lookup.
A 90 to 180 day priority list for operationalizing agent safety
Start with inventory: list every agent, every tool it can call, and every credential it holds. In the first 30 days, isolate the highest-impact agent-tool bindings, such as payment or data-deletion actions, and enforce least privilege on those first, then run an initial red-team sweep against your highest-traffic agent.

By day 90, implement runtime gates (PAG, ATM, PAA) and automated approval workflows for privilege expansions, and complete an SBOM/AIBOM inventory across your agent pipeline. By day 180, track detection time and response time as standing metrics, measure the percentage of agent actions running under enforced policy rather than ad hoc review, and set a recurring cadence for red-team exercises so testing never lapses into a one-time audit.
Teams preparing for a formal audit may find the audit-ready AI risk assessment framework from Project JTH useful for structuring that evidence collection in parallel.
— Chase Weir
Get a Dogtooth assessment of your agent deployment
We design and ship the same agent systems this article describes, inside real client workflows rather than as a proof of concept. Our services cover AI agents and automation, platform builds and migrations, e-commerce engineering, and secure integrations, with a senior team that stays on after launch rather than handing off and disappearing.

Our AI booking system case study and practitioner portal with Stripe Connect payouts show how we apply per-action authorization and guarded payment flows in production, not just in a design document. If you are deploying agents into workflows that touch payments, customer data, or scheduling, request an assessment and we will map the controls your specific deployment needs.
FAQ
What is the 30% rule for AI?
The phrase sometimes appears informally in discussions about limiting how much of a workflow an agent automates without human review, but it is not a defined standard from OWASP, NIST, or another named framework.
Is ChatGPT an AI agent?
ChatGPT functions as an AI agent when it is given tools, memory, and the ability to take multi-step actions on a user’s behalf, rather than only generating text responses. On its own, a conversational chat interface without tool access or persistent task execution is closer to an assistant than a fully autonomous agent.
What are 5 risks of AI agent deployment?
Based on the OWASP Top 10 for Agentic Applications, the highest-impact risks include goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, and unauthenticated data exfiltration through prompt injection. Each requires a different control, from per-action authorization to egress restrictions, rather than a single blanket fix.
What are 7 types of AI?
Common categorizations include reactive machines, limited memory systems, theory-of-mind systems, self-aware systems, narrow AI, general AI, and superintelligent AI, though these categories describe theoretical capability levels rather than a formal industry standard. Agentic AI, the focus of this article, is best understood as a narrow AI system given tools, memory, and multi-step planning ability rather than as a separate category on that list.
How do I start building AI agent safety into an existing deployment?
Begin with an inventory of every agent, tool binding, and credential in your environment, then enforce least privilege on the highest-impact bindings first. From there, add runtime enforcement checkpoints and a continuous red-team cadence, following the prioritized sequence covered earlier in this guide.



