AI Agent Security: Boundaries, Failures and Controls
AI agent security across retrieval, memory, tools and execution: a system-boundary diagram, failure matrix and practical tests for bounded enterprise actions.
AI agent security constrains what an agent can read, remember, propose and execute, including when its inputs or model behavior are unreliable. Protect the whole system: data access, runtime isolation, tool definitions, identity, delegated authority, credentials, approvals and recovery. Treat model output as a proposal. A separate enforcement point must decide whether the exact action is permitted before it can affect a protected system.
This guide is for architects and security engineers assessing an agent deployment. It covers system boundaries and failure behavior. The authorization guide owns request-level permission design, and MCP security owns protocol-specific threats. Those are parts of the system described here, alongside retrieval, memory, operators and release governance.
Draw the AI agent security boundary
Start with assets and permitted effects, then draw the paths that connect them. Identify sensitive source records, persistent memory, execution credentials, downstream systems and evidence storage. Mark every point at which data could become an instruction, an identity claim or a request for additional rights.
Sources, users and tool responses Owners and reviewers
| untrusted content | trusted decisions
v v
Retrieval / memory -> model runtime -> proposed action
| | |
data access policy isolated workload v
independent enforcement
identity / scope / approval
reservation / credentials
|
v
protected system
|
observation / evidence / recovery
The diagram is a reference architecture. It applies the resource-focused approach in NIST SP 800-207: network location is not sufficient reason to trust a request. For agents, the practical question is whether each route to a meaningful external effect crosses the intended controls. Include scheduled jobs, direct SDK calls and administrative routes rather than drawing only the normal user journey.
Separate threats by the boundary they cross
Prompt injection is one failure mode, but it is not the entire security problem. An agent can produce an accurate answer while exposing the wrong tenant's records. It can follow a legitimate user instruction while exceeding a delegated budget. It can use the correct tool while changing the destination after a human approved the original request.
| Boundary | Failure | Control | Evidence to collect |
|---|---|---|---|
| Retrieval | Unauthorized records enter context | Enforce source and field access before retrieval | Denied record absent from model input |
| Memory | Untrusted content becomes persistent instruction | Separate provenance and authority; restrict writes | Memory change attribution and review |
| Runtime | Process gains unintended filesystem or network access | Least-privilege workload and constrained egress | Isolation and outbound-access tests |
| Tool catalog | New tool silently expands capability | Versioned inventory and release approval | Tool-schema diff and authorized rollout |
| Delegation | Child agent receives broader rights | Enforce narrower scope at every handoff | Parent/child constraint comparison |
| Execution | Approved request is changed | Bind approval to protected request fields | Mutation rejected before dispatch |
| Recovery | Retry repeats an uncertain action | Durable operation state and reconciliation | Upstream call count and final status |
Do not assign all rows to a prompt filter. A filter may help identify suspicious inputs, but it cannot replace a data permission check, an atomic reservation or an independent approval. Choose controls by the property that must hold, then test that property at the system boundary.
Define one bounded action before granting tools
Use an evidence-preparation worker as a starting example. It may read approved records and draft a package for a human reviewer. It may not alter source evidence, certify compliance or close the underlying case. A single write operation can publish the exact approved draft to one approved destination.
Record a contract for that operation: accountable owner, principal, tenant, allowed tool, allowed destination, permitted data class, approval requirement, expiration, operation key and recovery owner. Keep the contract outside the model prompt. The prompt can explain the task, while deterministic code enforces its boundary even if the model proposes an inappropriate action.
The enterprise AI agents guide develops this workflow and its launch gates. Its value is a small, testable delegation boundary. It is not evidence that every process should use an autonomous agent, and it does not establish readiness for payments, production infrastructure changes or regulated decisions.
Test misuse as well as successful execution
The agent control lab 1.0.0 provides executable reference cases. After extraction, run this command in the extracted directory with Node.js 20 or newer:
node --test lab.test.mjs
The tests inject trusted identity and approval objects into a small local simulator. They do not authenticate real users. They verify that the simulator blocks a wrong audience, cross-tenant request, forbidden tool, changed destination, changed body, expired approval, self-approval and unavailable policy or reservation state. Each denied fixture asserts that no simulated upstream action occurred.
Extend those cases against the actual system. Replace the stub with an isolated test target, provision test identities, and observe the target independently. Try a direct call that bypasses the gateway. Remove a source permission and check model inputs. Change an approved tool argument. Suspend the agent on one instance and test another instance. The absence of an unsafe natural-language response does not establish the absence of an unsafe action.
For adversarial documents, record the injected instruction, the allowed task, the proposed tool call and the enforcement outcome. A resilient test can pass even if the model proposes the forbidden call, provided the independent boundary reliably blocks it and the workflow handles the denial. Separately evaluate whether the model should have proposed it at all.
Plan incident behavior and recovery
| Incident | Immediate control | Recovery requirement |
|---|---|---|
| Exposed execution credential | Revoke and stop affected access | Rotate secret, investigate use and remove exposure path |
| Suspected poisoned memory | Isolate affected memory and suspend writes | Restore from a reviewed version with provenance |
| Policy or authority service unavailable | Stop material actions | Restore trusted state and reevaluate pending work |
| Timeout after dispatch | Mark result uncertain | Reconcile using the original operation identifier |
| Tool definition changed unexpectedly | Block that version | Review schema and rerun acceptance tests |
| Evidence missing after action | Record incident without repeating action | Recover observation and assess missing evidence |
Name the operator and escalation path before launch. An incident runbook should identify which actions can stop immediately, which may already be in flight and how the protected system reports final state. A generic retry loop can make an incident worse if it creates new operation identifiers or repeats a non-idempotent effect.
The reference lab retains uncertain state in one process and does not dispatch again for an identical retry. It does not implement durable reconciliation or distributed concurrency. Test those properties in the real storage and upstream integration. A passing local example cannot establish exactly-once execution across a network failure.
Evidence, evaluation and ownership
Evaluate answer quality, permission enforcement and recovery as distinct layers. For answer quality, use representative tasks with evidence requirements. For enforcement, test denied actions and bypass attempts. For recovery, simulate failures before and after dispatch. Keep versions of the model, prompts, tool schemas, policies and evaluation set so a change can be assessed against the correct baseline.
The voluntary NIST AI Risk Management Framework provides a useful governance structure across Govern, Map, Measure and Manage. It does not certify this architecture. Assign operational owners and review gates that match the impact of your own process. The German governance checklist offers an editable responsibility and evidence register for German-speaking teams.
Preserve a distinction between authorization, dispatch and observed outcome in the audit record. A signed statement protects bytes under a chosen verification key; it does not independently prove that the external action occurred or that the record is complete. See the audit-trail guide for a runnable synthetic verification and tamper example.
Current Intelliger boundary and next step
Intelliger builds AI systems that can safely perform consequential enterprise work. Agent Trust describes the control layer and the Enterprise Agent Gateway the shared enforcement boundary. OATI is a developer preview with implemented components and remaining production and independent-review gates. This guide and its lab are educational reference material, not a security certification or a claim of production acceptance.
Use the boundary diagram to identify one protected action and its owner. Run the lab, then write an integration test for the same property in your environment. Continue through the guide library or use the Agent Trust overview to examine where the controls belong in the wider architecture.