Agentic AI systems can retrieve information, use tools, make decisions, and change enterprise systems with limited human intervention. That capability expands the attack surface beyond the model: identity, APIs, memory, data stores, orchestration logic, network paths, and human approvals all become security boundaries.
The most practical guidance currently points to a simple principle: do not attempt to secure an agent through prompts alone. Enforce authority in the surrounding system—the “harness”—with identity, least privilege, deterministic policy checks, isolation, logging, and rapid containment.
What the Whitepaper Establishes
ISACA’s Cybersecurity Recommendations for Securing AI Agents identifies 11 control areas, including governance, secure development, identity, sandboxing, prompt-injection defense, memory security, tool integration, human oversight, monitoring, supply-chain security, and resilience. The paper specifically addresses agents that ingest untrusted content, reason over enterprise data, and act through tools and APIs. [1]
A scope caveat matters: the paper states that its recommendations are for securing AI agents generally, rather than “agentic AI” as a separate technical category. Organizations should therefore use it as a control baseline, then add testing for autonomy, multi-agent delegation, persistent memory, and agent-to-agent communication. [1]
NIST’s AI Risk Management Framework remains useful for governance through its Govern, Map, Measure, and Manage functions, but its Generative AI Profile was not designed as a complete autonomous-agent control standard. [2] Google’s Secure AI Framework, or SAIF, is likewise a framework for integrating security into machine-learning applications and managing AI/ML model risk—not a substitute for runtime authorization and enterprise access controls. [3]
1. Build an Agent Inventory Before Granting Access
Maintain a continuously updated inventory of:
- Agents and sub-agents
- Models and model endpoints
- Tools, plugins, and MCP servers
- Vector databases and memory stores
- Data sources and retrieval pipelines
- External providers
- Owners, environments, permissions, and business purpose
The inventory should record what each agent can read, write, execute, delegate, and reach over the network. It should also identify trust boundaries between the user, agent runtime, model provider, orchestration layer, tool runner, enterprise systems, and approval interface. [1]
Established control: every deployed agent needs an owner, an approved purpose, defined capabilities, and a decommissioning process.
Reported industry direction: Google’s Agent Gateway uses a central registry for approved agents, tools, and MCP servers, allowing policy decisions to be made against registered destinations. Unregistered remote agents, tools, and MCP servers are blocked by default in the documented configuration. [4]
2. Give Every Agent a Distinct Identity
Do not run agents under a shared administrator account, a reused human session, or a broadly privileged service credential.
Use:
- A distinct identity for each agent or workload
- Short-lived credentials with automatic rotation
- Federated or workload identity instead of static keys
- RBAC or ABAC for tools, data, and APIs
- Just-in-time elevation for sensitive functions
- Separate identities for the human, agent runtime, tool runner, and administrator
Authorization should preserve both identities: which user initiated the task and which agent performed the action. A user’s authority should not automatically become the agent’s unrestricted authority.
Google’s documented Agent Gateway model assigns agents unique, trackable identities and uses those identities in authorization decisions. Its gateway can apply IAM policies to restrict which agents may reach specific tools or other agents. [4]
3. Enforce Least Privilege at the Tool Layer
The model should not decide whether its own tool call is authorized. Put tools behind an API gateway, action broker, or policy enforcement point that validates:
- Agent identity
- User or delegated authority
- Target resource
- Requested operation
- Parameters and schema
- Business rules
- Risk level
- Required approval
Prefer narrow tools such as look_up_order_status over general-purpose tools such as “run arbitrary SQL.” Separate read and write functions, restrict file paths and repositories, and block unrestricted shell access, arbitrary URL fetching, and access to cloud metadata services. [1]
Recommended default: read-only access first; write access only when the business case is documented and tested.
For high-impact operations—deleting records, changing permissions, moving money, sending external communications, or modifying production—require a deterministic approval gate outside the model’s reasoning loop. ISACA recommends human approval for destructive, financial, legal, regulated, or irreversible actions. [1]
4. Treat Retrieved Content as Untrusted
Prompt injection can arrive through a user request, webpage, email, PDF, code comment, retrieved document, memory entry, or tool response. The content does not need to exploit a software vulnerability; it can attempt to redirect the agent’s objective or induce an unsafe tool call.
Controls should include:
- Clear separation between instructions and retrieved data
- Content labeling and provenance tracking
- Input and output screening
- Tool allowlists
- Independent authorization before action
- Validation of tool outputs before reuse
- Source trust levels that do not grant authority to untrusted content
- Adversarial tests using real documents, websites, emails, and tools
ISACA explicitly advises treating external content and tool output as untrusted and preventing retrieved content from directly triggering actions without a separate policy decision. [1]
Anthropic’s Claude Code documentation provides a practical example of layered defense: manual permission mode, approval for sensitive operations, network-request approval, command-injection detection, isolated web-fetch contexts, and trust verification for new codebases and MCP servers. Anthropic also warns that no system is completely immune to attack. [5]
Uncertainty: filtering and classifiers can reduce exposure, but they should not be treated as a complete defense. The durable control is limiting what a manipulated agent can access and do.
5. Isolate Execution and Restrict Egress
Run browser automation, code interpretation, file parsing, and tool execution in containers, virtual machines, or microVMs. Use read-only filesystems where practical, prohibit privileged containers, restrict system calls, and make execution environments ephemeral.
Network controls should:
- Deny outbound access by default
- Permit only approved domains, APIs, and destinations
- Route requests through inspected proxies
- Block internal administration interfaces
- Block cloud metadata services
- Separate development, test, and production environments
- Prevent the agent from directly reaching unrestricted execution systems
The ISACA whitepaper specifically recommends segmented networks, sandboxed execution, no default egress, and separation between reasoning and execution. [1]
Google’s Agent Gateway provides another implementation pattern: network-layer enforcement for agent interactions, centralized authorization, content inspection through Model Armor, and telemetry exported to Cloud Logging and Cloud Trace. [4]
6. Protect Memory, Context, and Secrets
Persistent memory turns a one-time manipulation into a potentially recurring influence. Secure memory as an enterprise data store, not as an informal extension of the prompt.
Required controls include:
- Never place secrets in prompts or persistent memory
- Store credentials in a secrets manager
- Isolate memory by tenant, user, session, and use case
- Apply retention limits and time-to-live values
- Authorize memory retrieval
- Record provenance for memory writes
- Validate memory integrity before reuse
- Encrypt prompts, traces, logs, memory, and outputs
- Redact sensitive information from telemetry
ISACA recommends authorization checks for stored memory, cross-session and cross-tenant isolation, retention controls, and provenance validation to reduce poisoning risk. [1]
7. Secure the Agent Supply Chain
An agent’s effective capability depends on more than its model. Plugins, MCP servers, libraries, containers, embeddings, APIs, prompts, policies, and external providers can all introduce risk.
Enterprise controls should include:
- Pinning model, tool, plugin, and dependency versions
- Artifact signing and provenance verification
- SBOM and AI-BOM records
- Dependency and container scanning
- Restricted publishing rights for tools and policies
- Third-party provider assessments
- Contractual controls for data use, retention, residency, and breach notification
- Monitoring for endpoint spoofing and provider-side behavior changes
- Rollback capability
These controls are established software-supply-chain practices adapted to AI systems; the agent-specific uncertainty is how rapidly tools and capabilities can change at runtime. ISACA recommends applying formal change management to models, prompts, policies, tools, memory, retrieval sources, providers, and action scope. [1]
8. Log Actions, Decisions, and Delegation
A conventional application log is insufficient if it records only the final API request. Investigators need the chain of authority and context.
Log, with appropriate redaction:
- User and agent identities
- Session and parent-agent identifiers
- Prompts and responses where permitted
- Retrieved sources and content hashes
- Memory reads and writes
- Tool names and arguments
- Authorization decisions
- Human approvals and denials
- Policy violations
- Network destinations
- Model, prompt, tool, and policy versions
Use centralized, tamper-resistant storage and connect agent telemetry to existing SIEM and incident-response processes. ISACA recommends monitoring unusual tool use, excessive retrieval, repeated bypass attempts, anomalous outbound traffic, behavioral changes, and guardrail failures. [1]
9. Add Rate Limits, Circuit Breakers, and Kill Switches
Agent failures can be fast, repetitive, and expensive. Set limits on:
- Tool-call frequency
- Token and compute consumption
- Workflow duration
- Retry counts
- Data volume retrieved or exported
- Number of delegated agents
- Financial or transactional value
- Concurrent actions
Implement per-capability and global shutdown controls. A safe response may be to revoke credentials, disable a tool, quarantine the agent, shift it to read-only mode, or revert the workflow to manual operation.
ISACA recommends quotas, timeouts, capped retries, circuit breakers, rollback, kill switches, and safe fallback modes such as recommendation-only or manual-approval operation. [1]
10. Test the Live Agent, Not Only the Model
Predeployment model evaluation is not enough. Test the complete workflow with its actual prompts, retrieval corpus, memory, tools, permissions, network paths, and approval process.
Minimum test scenarios should include:
- Direct and indirect prompt injection
- Malicious documents and tool responses
- Unauthorized data retrieval
- Cross-tenant and cross-user access
- Unsafe tool arguments
- SSRF and internal-network discovery
- Memory poisoning
- Credential exposure
- Malicious or altered MCP servers
- Excessive delegation
- Cascading tool or agent failures
- Guardrail bypass
- Kill-switch and rollback behavior
Retest after changes to models, tools, permissions, memory, retrieval sources, providers, prompts, orchestration, or deployment environments. [1]
A Practical Enterprise Baseline
Before an agent receives production access, require evidence that it has:
- A named owner and current inventory record
- A documented business purpose and risk assessment
- A distinct, least-privileged identity
- Short-lived credentials and revocation procedures
- An allowlisted tool set
- A policy enforcement point before every consequential action
- Sandboxed execution and restricted egress
- Tenant- and session-isolated memory
- Human approval for irreversible actions
- Tamper-resistant action logs
- Adversarial test results
- A tested kill switch and rollback path
- A documented supplier and dependency review
The central lesson from the current whitepaper and platform guidance is established: agent security must be enforced outside the model. Prompts can describe intended behavior, but identity systems, policy engines, sandboxes, network controls, approval gates, and incident-response mechanisms determine what the agent can actually do. [1] [4] [5]
Agentic AI standards and implementation patterns are still developing. Enterprises should therefore avoid waiting for a single definitive standard. Use established cybersecurity controls now, map them to the NIST AI RMF and ISO/IEC 42001 programs where relevant, and continuously test whether the live agent remains within its approved authority.
Sources
- White Papers 2026 Cybersecurity Recommendations for Securing AI Agents
- Agentic AI Security: 8 Critical Risks and 5 Best Practices
- 6 Agentic AI Security Risks to Monitor in 2026
- AI Agent Security Checklist (2026): Agentic Risks & Controls
- Agent Gateway overview
- Is Claude Code Safe to Use? Risks and Best Practices
- AI Governance Framework: ISO 42001 and NIST AI RMF…
- Securing AI agents: Key controls and best practices
- Above the Stack: AI Controls for MSPs
- A manufacturing blueprint for secure agentic AI