Memory in Artificial Intelligence (AI) agents has seen rapid adoption in recent years, and many of them now include a memory database that allows them to store information about their users and reference it in separate conversations. However, this feature also creates a new attack surface, where adversaries can inject poisoned memories into the database of an agent in order to degrade its performance. Previous research in this area has focused on attacks that either insert poisoned memories directly into the memory store, or assume that the attacker knows in advance which questions the agent will be asked. This thesis attempts to poison a memory database without either of those capabilities, relying only on conversation history that has already been stored in memory and on knowledge of the type and retrieval policy of the memory system. From this corpus, it selects up to k conversations using closeness to the hubs of the memory database, regions of the embedding space where many unrelated user queries tend to retrieve from, as the guiding selection criterion. Conversations chosen this way are disproportionately retrieved across a wide range of future queries, displacing legitimate context and broadly degrading the agent’s responses. Results show that this is an effective way to reduce overall AI agent performance with minimal information about the agent itself. This research demonstrates that the embedding geometry of retrieval-based memory systems is itself a vulnerability: an adversary can exploit the structure of the embedding space without ever needing to know what the agent will be asked.
- Hub-based memory poisoning
- Aidan Root
- 0009-0002-9681-9354
- Long Jiao (Advisor) - University of Massachusetts Dartmouth, Department of Computer and Information ScienceAmir Akhavan Masoumi (Committee Member) - University of Massachusetts Dartmouth, Department of Computer and Information ScienceJoshua Robert Carberry (Committee Member) - University of Massachusetts Dartmouth, Department of Computer and Information Science
- vii, 27 pages
- List of tables -- Chapter 1. Introduction -- Chapter 2. Background and related work -- RAG memory and the LongMemEval benchmark -- Hubness in nearest-neighbour spaces -- Memory-poisoning attacks -- Chapter 3. Threat model and metric -- Attacker model -- Where this sits -- Why accuracy alone is ambiguous -- Formal objective -- Chapter 4. Proposed solution -- Experimental setup -- Hub identification -- Stage A: privileged vector-mode injection -- Stage B: realistic text-mode realization -- Chapter 5. Experiments and results -- Stage A: the privileged-mode ceiling -- Stage B: under the realistic threat model -- Chapter 6. Discussion -- Limitations -- Future research directions -- References.
- Includes bibliographical references (25-27 pages).
- University of Massachusetts Dartmouth
- Master of Science (MS)
- Computer Science
- Department of Computer and Information Science
- English
- Thesis
- Copyright 2026 Aidan Root
- https://doi.org/10.62791/20609
- 9914540173501301