Securing AI Agents Against Prompt Injection and Memory Leakage

AI agents are not merely chatbots with better interfaces. They can retrieve documents, call APIs, modify files, send messages, and preserve information between sessions. That combination creates a security problem in which malicious text can influence not only an answer, but also a tool call, authorization decision, or future interaction.

The emerging security literature converges on a central principle: do not make the language model the sole security boundary. Prompt filters and model training can help, but authorization, memory governance, data-flow controls, and tool enforcement must operate outside the model.

Why Tool-Using Agents Change the Threat Model

A conventional language model may produce a misleading or unsafe response. An agent can turn the same manipulation into an external action.

Indirect prompt injection occurs when an attacker places instructions in content the agent later reads—such as a webpage, email, document, repository file, database record, or tool response. The agent may confuse that content with an instruction from the user or developer. Microsoft describes local agents as especially exposed because they process prompts, files, web content, and tool output while lacking a reliable, inherent distinction between trusted instructions and hidden instructions. [1]

The consequences depend on the agent’s permissions. A read-only summarizer may return corrupted information. An agent with access to private files and outbound communication may disclose data. An agent able to execute code may create a path to file modification or command execution.

NIST’s agent-identity work identifies the same architectural concern: agents operate across data sources, tools, services, and delegation chains, making identity, authorization, auditing, and non-repudiation separate design problems rather than simple extensions of chatbot safety. [2]

Prompt Injection Is an Instruction-and-Data Confusion Problem

Prompt injection is difficult because system instructions, user requests, retrieved text, tool outputs, and memory are often assembled into one model context. A model can be instructed to treat external content as untrusted, but that instruction is still competing with the content inside the same probabilistic channel.

NIST’s public comments on agent identity and authorization explicitly describe the lack of separation between control and data planes as a core challenge. Commenters warned that external content and tool outputs can enter the agent’s context and influence behavior, including through direct and indirect prompt injection. [2]

This means a stronger system prompt is not equivalent to a security boundary. It may reduce attack success, but it does not deterministically prevent an agent from selecting an unsafe tool or constructing dangerous arguments.

Direct and indirect injection

  • Direct injection: the attacker enters instructions through the normal user interface.
  • Indirect injection: the attacker controls content that the agent retrieves or processes.
  • Tool-output injection: a tool returns attacker-controlled text that is placed into the agent’s context.
  • Stored injection: malicious content is written into memory, a retrieval index, a summary, or a workflow file and affects later runs.

The indirect forms are particularly important because the attacker may never interact with the agent directly. Microsoft’s runtime-protection documentation describes inspection points before and after tool calls precisely because injected content can originate in files, webpages, repositories, or tool responses. [1]

Memory Leakage and Memory Poisoning Are Different Risks

“Memory leakage” can refer to at least two related but distinct failures.

Memory leakage

The agent exposes information stored in conversation history, profiles, retrieval systems, logs, or persistent memory to an unauthorized user, tool, agent, or external destination. Leakage can occur in an answer, a tool argument, a URL, a log entry, or a generated file.

NIST’s agentic-security feedback highlights overcollection and the risk of sensitive prompt and context data being passed to other agents, external models, and APIs. It also notes that agent transaction logs may themselves contain sensitive information. [2]

Memory poisoning

Memory poisoning is the inverse problem: an attacker gets malicious or false content written into persistent state. A later session retrieves it and treats it as trusted background knowledge.

Research published by Abbas Yazdinejad and Hadis Karimipour examined 2,614 simulated multi-step attack trajectories involving memory-enabled agents. The work reports that attacks can be separated in time: an agent may behave normally during the initial interaction, then act on poisoned memory during a later workflow. The authors emphasize that the simulations do not establish production prevalence, but they do show why single-turn testing can miss long-horizon failures. [3]

The security implication is important: restarting an agent or ending a session does not remove a malicious record from a persistent store. Memory therefore needs its own access controls, provenance, retention policy, review process, and rollback mechanism.

The Tool Layer Is the Enforcement Point

A secure agent architecture should assume that the model may propose an unsafe action. The application—not the model—should decide whether that action is authorized.

Useful controls include:

  1. Explicit tool allowlists: Expose only the operations required for a task.
  2. Read-only defaults: Separate read, write, delete, send, and execute capabilities.
  3. Structured arguments: Use typed schemas instead of passing free-form model text to shells, SQL interpreters, or HTTP clients.
  4. Deterministic policy checks: Validate identity, resource, action, destination, data sensitivity, and task scope before execution.
  5. Short-lived credentials: Issue task-specific permissions that expire and can be revoked.
  6. Rate and spend limits: Restrict the number, frequency, and cost of tool calls.
  7. Human approval for high-impact actions: Require confirmation for payments, deletion, external communication, permission changes, and code execution.

NIST’s collected feedback favors continuous, task-level authorization rather than one-time authorization at login. It also recommends attenuating permissions as work moves through agents and sub-agents, so a delegated component cannot inherit more authority than it needs. [2]

A practical rule is:

> The model may recommend an action; trusted application code must authorize and execute it.

Separate Untrusted Reading From Privileged Action

One of the strongest architectural ideas in the research is to separate the component that reads untrusted content from the component that plans and performs privileged actions.

The CaMeL approach, developed for tool-using agents, uses a privileged language model and a quarantined model. The quarantined model processes potentially malicious content but does not receive direct tool authority. An interpreter tracks the provenance and capabilities of values before allowing tool calls. [4]

This design addresses two separate problems:

  • Control-flow protection: Untrusted text should not be able to insert new steps into the agent’s plan.
  • Data-flow protection: Untrusted data should not automatically become eligible for sensitive tools or destinations.

The approach is not a universal solution. It requires carefully designed policies, trusted wrappers, and complete mediation of side effects. It may also reduce utility or increase implementation complexity. But its key contribution is architectural: security checks do not depend entirely on the model correctly recognizing an attack.

Protect Memory as a Privileged Write Path

Memory should not be treated as an automatic side effect of reading content.

A safer memory design should:

  • Assign every entry an origin, author, session, tenant, and timestamp.
  • Distinguish facts, preferences, procedures, summaries, and instructions.
  • Prevent external content from writing privileged policies or system-level instructions.
  • Require validation before cross-session or cross-user persistence.
  • Keep tenant, user, session, and task memory separate.
  • Apply expiration and periodic revalidation.
  • Support quarantine, deletion, rollback, and forensic review.
  • Record which memory entries influenced a later tool call.

The research on delayed memory poisoning supports trajectory-level testing rather than only evaluating isolated prompts. A test should follow the complete lifecycle: ingestion, memory write, retrieval, reasoning, tool selection, execution, and later reuse. [3]

Memory security also requires testing for leakage. A system should verify that one user’s memories cannot appear in another user’s context, that sensitive values are not copied into broad shared stores, and that logs do not retain unnecessary prompt or tool-result content.

Validate Tool Results Before They Re-enter Context

Tool results are data, not instructions.

An API response, ticket, document, or webpage may contain natural-language text that looks authoritative. Before placing it into the agent’s context, the application should:

  • Validate its schema.
  • Label its provenance and trust level.
  • Remove unnecessary fields.
  • Limit embedded links and active content.
  • Detect or quarantine instruction-like text.
  • Prevent the result from changing tool permissions.
  • Preserve the distinction between factual values and commands.

Microsoft’s runtime model illustrates the value of checking both sides of the tool boundary: the requested action can be inspected before execution, while the returned content can be inspected before it is processed further. [1]

Filtering remains useful, but it should be treated as a detection layer rather than a guarantee. Research on prompt-injection defenses continues to report trade-offs between attack resistance and task utility. For example, Meta’s Meta-SecAlign experiments report lower attack-success rates on several benchmarks, while also noting that the evaluation does not cover cumulative attacks that distribute mutually reinforcing injections across multiple untrusted messages. [5]

Use Identity and Authorization Designed for Agents

Agents should have distinct, verifiable non-human identities rather than silently reusing a person’s credentials.

NIST’s public feedback favors a model combining stable trust anchors with short-lived, task-scoped credentials. An agent may have a persistent identity tied to its software and organizational owner, while each workflow receives a narrower, revocable authorization. [2]

A useful authorization record should answer:

  • Which agent acted?
  • Which human or organization authorized the task?
  • Which model and software version ran?
  • Which tools and resources were available?
  • What data was read?
  • What action was requested?
  • Which policy allowed or denied it?
  • Which sub-agent or service received delegated authority?
  • What approval, if any, was obtained?

NIST commenters also warned against using an LLM as the primary or sole authorization decision-maker. Probabilistic systems may help assess context or detect anomalies, but deterministic policy enforcement should make the final allow-or-deny decision. [2]

Logging Must Capture the Agentic Chain

A final answer is not enough for incident response. Security teams need the sequence that produced it.

Log, with appropriate redaction:

  • The original user request.
  • Agent and user identities.
  • Model, policy, and tool versions.
  • Retrieved sources and memory identifiers.
  • Tool calls and validated arguments.
  • Authorization decisions.
  • Human approvals and timestamps.
  • Tool results or cryptographic hashes of sensitive results.
  • External destinations.
  • Memory writes, edits, reads, and deletions.
  • Policy violations, blocks, and kill-switch events.

Logs should be protected from modification by the agent and its tools. Retention should balance investigation requirements with privacy and data-minimization obligations.

Test Complete Trajectories, Not Just Prompts

Agent security evaluation should measure both what the system says and what it does.

A useful test program includes:

  • Direct and indirect prompt injections.
  • Malicious webpages, emails, documents, code comments, and database fields.
  • Tool-result and tool-description poisoning.
  • Attempts to select unauthorized tools.
  • Attempts to alter tool arguments.
  • Data-exfiltration paths through URLs, messages, files, and APIs.
  • Memory-write and cross-session poisoning.
  • Cross-tenant retrieval failures.
  • Multi-agent message tampering.
  • Long-running and repeated attacks.
  • Failure recovery and rollback.
  • Human-approval fatigue and misleading summaries.

AgentDojo is one established research benchmark for evaluating prompt injection and defenses in tool-calling agents. More recent work continues to show that benchmark results depend on the task, model, agent framework, and attack objective; therefore, a single attack-success figure should not be treated as a universal production risk estimate. [5]

Testing should also inspect side effects. An agent that refuses to reveal a secret but still calls an unauthorized tool has not passed the test.

What Security Whitepapers Establish—and What They Do Not

The current literature supports several established conclusions:

  • Untrusted content can influence tool-using agents through indirect prompt injection. [[1] [2]]
  • Persistent memory creates a security surface that can outlive the initiating session. [3]
  • Least privilege, task-scoped authorization, and deterministic enforcement reduce blast radius. [2]
  • Architectural separation and capability controls offer stronger guarantees than relying only on prompt wording. [4]
  • Model-level defenses can improve benchmark performance but do not eliminate the need for application controls. [5]

The literature does not yet establish a universal attack rate for deployed enterprise agents. Results from synthetic benchmarks, vendor evaluations, and simulated memory trajectories should not be presented as production prevalence. Nor is there a single accepted standard for agent intent, memory provenance, or cross-agent authorization; NIST’s identity work remains an evolving project rather than a finished agent-authorization standard. [2]

A Practical Baseline for Production

Organizations deploying tool-using agents should begin with this minimum control set:

  1. Inventory every agent, tool, memory store, connector, owner, and credential.
  2. Treat webpages, files, emails, database records, tool results, and agent messages as untrusted by default.
  3. Keep untrusted content out of privileged instruction channels.
  4. Enforce tool permissions in trusted application code.
  5. Use read-only and task-scoped access wherever possible.
  6. Give agents short-lived, revocable credentials.
  7. Gate irreversible or externally visible actions.
  8. Log the full agent trajectory with sensitive-data redaction.
  9. Protect, review, expire, and isolate memory writes.
  10. Run adversarial, multi-step tests after every meaningful model, tool, policy, or memory change.
  11. Maintain a kill switch and a safe degraded mode.
  12. Reassess the system whenever a new tool, data source, sub-agent, or delegation path is added.

The durable security strategy is not to assume that an agent will never be fooled. It is to ensure that being fooled does not automatically grant the agent authority to leak memory, misuse a tool, or cross a trust boundary.


Sources

  1. Summary of Comments on the Concept Paper — NCCoE Agentic AI Identity and Authorization Project Resource Hub documentation
  2. AI Agent Security and Prompt Injection, What Actually Works in 2026
  3. An Experimental Evaluation of Multimodal PromptInjection Attacks on Agentic AI Frameworks
  4. Defending Agent Memory Against Poisoning
  5. teiss – News – When prompt injection becomes a propagation mechanism
  6. OWASP Top 10 for AI Agents: An AI Security Checklist
  7. Poisoned Memory: Essential AI Agent Security Warning
  8. AI Agent Security Checklist (2026): Agentic Risks & Controls
  9. Prompt Injection: The Full Guide to Securing AI Models and Agents
  10. Authority Is Not a String: A Capability-Scoped Harness for Prompt-Injection-Resistant Coding Agents

Leave a Reply

Your email address will not be published. Required fields are marked *