Researchers expose memory poisoning risk in AI agents

Researchers have demonstrated a technique called GhostWriter, which uses hidden instructions to place false or malicious information in an AI agent’s long-term memory. The stored instruction can later influence the agent’s responses or actions.

Key facts

  • Malicious instructions can reach an agent through emails or calendar invitations.
  • The researchers reported an average memory-injection rate of about 98%.
  • Poisoned memories were activated in about 60% of their tests.
  • A proposed screening system significantly reduced the attack’s effectiveness.
  • The research is a preprint based on simulated tasks, not a documented real-world attack.

Our take

Memory helps AI agents provide personalised assistance, but it also creates a lasting attack surface. A malicious instruction stored today could influence an agent during an unrelated task later. Organisations should control what agents are allowed to remember, screen information before it is saved or retrieved, and restrict the actions an agent can perform without human approval.

Sources

About the author

Campbell McKenzie is a Director at Incident Response Solutions, a New Zealand firm experienced in cyber incident response, digital forensics, investigations and technology risk. Through KiwiGen.AI, Campbell helps professional services firms adopt generative AI safely, with practical governance and controls.