AI agents are moving beyond conversation. They retrieve documents, browse websites, read email, call APIs, modify files, and make decisions inside business workflows. That usefulness also creates a new security problem: content an agent reads can contain instructions that redirect what it does.
Prompt injection is therefore no longer merely a model-quality issue. It is an enterprise risk involving data protection, identity, authorization, operational resilience, regulatory exposure, and incident response.
What prompt injection means
Prompt injection occurs when input changes an AI system’s behavior or output in an unintended way. The injected instruction does not need to be visible to a human; it only needs to be interpreted by the model. [1]
There are two principal forms:
- Direct prompt injection: The attacker enters instructions through the user interface or API.
- Indirect prompt injection: Malicious instructions are embedded in external content—such as a webpage, email, PDF, uploaded file, database record, or tool response—that the agent later processes. [1]
Indirect injection is especially important for enterprise agents because the attacker may never interact with the AI system directly. A malicious party can place content in a source the agent is expected to read, then wait for an ordinary business task to retrieve it.
The underlying difficulty is architectural. Language models process instructions and information through the same general context, so a document that should be treated as evidence may also contain language that looks like an instruction. Retrieval-augmented generation and fine-tuning can improve usefulness, but they do not eliminate the vulnerability. [1]
Why agents raise the stakes
A chatbot that produces a bad answer can create reputational or operational harm. An agent with tools can turn the same manipulation into an action.
OWASP identifies potential consequences including sensitive-information disclosure, unauthorized function access, arbitrary command execution, manipulation of critical decisions, and exposure of system prompts or AI infrastructure details. The actual severity depends on the agent’s business context and level of agency. [1]
That creates a simple risk relationship:
> Prompt injection determines whether the agent’s goal can be redirected. Permissions determine what happens next.
An overprivileged agent might be able to:
- Read confidential files or customer records
- Send email or messages
- Modify tickets, repositories, or databases
- Execute code or shell commands
- Call external services
- Move money or approve transactions
- Pass information between otherwise separate systems
The prompt injection is the entry point; excessive agency is often the mechanism that produces material damage.
A reported production example: EchoLeak
The risk is not limited to laboratory demonstrations. Mandiant and Google report that prompt injection remains a primary attack vector in custom AI deployments. Their assessments describe agents exposed to public webpages, customer emails, uploaded documents, vector databases, MCP servers, and third-party APIs. [2]
One Mandiant case involved a developer-facing assistant used for CI/CD and repository management. In a security assessment, testers used role-confusion prompt injection to persuade the assistant to use permitted tools in an unintended way. The assistant then cloned sensitive repositories and pushed them to an external repository controlled by the testers—a “confused deputy” pattern in which legitimate capabilities were weaponized through semantic manipulation. [2]
Mandiant also describes a customer-service RAG system whose data sources included forums, support tickets, and third-party feeds. Hidden instructions in retrieved content could influence the agent’s reasoning and potentially expose information from other customers’ tickets. The company’s stated defensive lesson was to treat retrieved content as untrusted, enforce tenant isolation, restrict tool privileges, and apply egress data-loss controls. [2]
These are reported assessment findings, not evidence that every AI agent is exploitable in the same way. They do establish that the attack path is operationally plausible when untrusted content, sensitive data, and external communication meet in one workflow.
The “lethal combination” boards should understand
The highest-risk design combines three conditions:
- The agent can access private or sensitive data.
- The agent processes content that an attacker can influence.
- The agent can communicate externally or change state.
Mandiant describes this combination as enabling autonomous exfiltration through prompt injection. [2]
This is a useful board-level test because it translates a technical vulnerability into a business question:
Can an outsider place content in the agent’s reading path, cause it to access protected information, and give it a route to send or alter something outside the intended workflow?
If the answer is yes, the organization has more than a chatbot risk. It has an identity, data-flow, authorization, and containment problem.
Why filtering alone is not a security strategy
Input filters, prompt classifiers, and content scanners are useful layers. They are not reliable authorization systems.
OWASP states that foolproof prevention remains unclear because prompt injection is connected to the stochastic behavior of generative AI. Its recommended mitigations therefore emphasize reducing impact: least privilege, external-content segregation, deterministic output validation, human approval for high-risk actions, and recurring adversarial testing. [1]
Anthropic’s guidance similarly treats indirect injection as a continuing threat. It recommends screening tool outputs, but also says a successful injection should have limited consequences because the agent receives only narrowly scoped permissions, does not receive unnecessary secrets, and operates tools in sandboxes. [3]
The implication is important:
> Assume that some malicious content will reach the model. Design the surrounding system so that the model cannot convert it into an unauthorized high-impact action.
Controls that belong outside the model
1. Inventory every agent and capability
Create a current inventory of:
- Models and providers
- Agents and orchestration code
- Tools, plugins, and MCP servers
- Data sources and vector stores
- Persistent memory
- Credentials and service identities
- Human approval points
- External communication paths
Assign an accountable owner to each agent. The board should be able to identify which systems can read sensitive data, execute actions, or communicate externally.
2. Enforce least privilege
Give each agent only the tools and data required for its specific task. Use separate identities for agents, users, tool runners, and administrators. Prefer short-lived, task-scoped credentials over standing keys.
Mandiant recommends identity segmentation, user-delegated authorization for sensitive workflows, cryptographic workload identities, and emergency token revocation. [2]
A drafting agent should not automatically have permission to send. A research agent should not have database-write access. A coding assistant should not receive unrestricted shell access merely because one project occasionally requires it.
3. Separate data from instructions
External content should be explicitly marked as untrusted data. Anthropic recommends placing third-party content in structured tool-result fields rather than mixing it into system instructions or ordinary user text. It also recommends identifying the source and nature of the content and, where practical, JSON-encoding it to create clearer boundaries. [3]
These measures can reduce confusion, but they are not a guarantee. The decisive control remains authorization enforced by application code, not by the model’s interpretation of a prompt.
4. Put a deterministic policy layer between the agent and the tool
Every consequential tool call should be checked independently for:
- The requesting user and agent identity
- The target system and resource
- The requested operation
- Parameter validity
- Data sensitivity
- Business-policy constraints
- Approval requirements
- Rate and volume limits
The model may propose an action; it should not be the sole authority that authorizes it.
5. Require approval for irreversible actions
Use human approval, step-up authentication, or dual control for actions involving money, legal commitments, regulated data, deletion, external publication, production changes, or bulk operations.
Approval screens should show the exact action, destination, data involved, and affected records. A vague “the agent wants to continue” prompt is not meaningful oversight.
6. Sandbox execution and restrict egress
Run browser automation, code execution, file parsing, and tool calls in isolated environments. Restrict outbound network access to approved destinations and block access to internal administration interfaces and cloud metadata services where possible.
Mandiant recommends correlating agent logs with network telemetry, detecting anomalous outbound transfers, and revoking tokens or reducing container egress privileges when behavior deviates from the approved use case. [2]
7. Protect retrieval and memory
Treat vector databases, uploaded documents, support tickets, and persistent memory as security-sensitive stores.
Controls should include:
- Authorization before retrieval
- Tenant and user isolation
- Source provenance
- Review of memory writes
- Retention and expiration limits
- Ability to remove poisoned entries
- Rebuilding or rolling back compromised indexes
- Monitoring for unusual retrieval patterns
A document should not gain authority merely because it was successfully retrieved.
8. Log the full decision chain
Security teams need more than the final answer. Logs should capture, subject to privacy and retention requirements:
- User and agent identity
- Retrieved sources
- Tool calls and arguments
- Policy decisions
- Approvals and denials
- External destinations
- Sensitive-data detections
- Prompt-injection alerts
- Model and tool versions
Without this telemetry, incident responders may be unable to determine whether an agent followed the user’s request, a poisoned document, a compromised tool, or an unintended chain of reasoning.
What the board should ask management
A useful board review does not need to inspect system prompts. It should ask management questions that expose ownership and blast radius:
- Which agents can access confidential, regulated, or financial data?
- Which agents can change systems or communicate externally?
- What happens if an agent follows a malicious instruction from an email or webpage?
- Are permissions granted per task, or are agents using standing credentials?
- Which actions require human approval, and can that approval be bypassed?
- Can the organization disable a tool, revoke an agent’s credentials, or switch it to read-only mode quickly?
- When were the agents last tested with realistic indirect-injection scenarios?
- Who accepts the residual risk when a vendor’s model, plugin, or tool changes?
- Can the security team reconstruct every tool call and data transfer after an incident?
- Are AI-agent risks included in enterprise risk registers, third-party risk reviews, and cyber-insurance discussions?
These questions shift the conversation from “Is the model safe?” to “Is the business prepared for model failure?”
A practical risk-tiering model
Organizations can begin with three operational tiers:
Low autonomy
The agent summarizes approved internal content and cannot write, send, execute, or access sensitive systems. Injection remains possible, but the blast radius is limited.
Controlled autonomy
The agent can retrieve data or prepare actions but requires deterministic checks and human approval before state-changing operations.
High autonomy
The agent can independently access sensitive data, call external services, execute code, or make irreversible changes. These systems require stronger isolation, narrowly scoped identities, continuous monitoring, tested shutdown procedures, and explicit executive risk acceptance.
The objective should not be to eliminate all agent use. It should be to ensure that autonomy is proportional to the value of the task and the consequences of failure.
The strategic conclusion
Prompt injection is difficult because it exploits the boundary between language and control. A webpage, email, or document can look like information to a person while acting like an instruction to an agent.
That makes the issue suitable for board oversight. The governing questions involve asset ownership, data access, delegated authority, resilience, third-party risk, and the ability to contain a compromised workflow—not merely the quality of a model’s responses.
The most defensible posture is layered:
- Treat all external content as untrusted.
- Keep authorization outside the model.
- Grant the agent minimal, short-lived access.
- Isolate tools, memory, and tenants.
- Require approval for high-impact actions.
- Monitor data movement and behavioral drift.
- Test continuously against indirect prompt injection.
- Maintain a fast kill switch and a manual fallback.
No single prompt, classifier, or vendor feature can close the risk completely. The board-level responsibility is to ensure that when an agent is manipulated, the resulting failure is bounded, visible, reversible, and survivable.
Sources
- Mandiant AI Risk and Resilience Report 2026
- Prompt Injection: The Full Guide to Securing AI Models and Agents
- LLM01:2025 Prompt Injection – OWASP Gen AI Security Project
- Prompt Injection and Jailbreak Defense for Claude Apps — Anthropic Claude Certified Developer
- Summary of Comments on the Concept Paper — NCCoE Agentic AI Identity and Authorization Project Resource Hub documentation
- AI Agent Security and Prompt Injection, What Actually Works in 2026
- Mitigate jailbreaks and prompt injections
- Prompt Injection Explained: Attacks + Defenses 2026
- Understanding Prompt Injection Risks in Claude Code
- Prompt Injection Is the New Phishing: Defending Copilot and Copilot Studio Agents