Memory exploitation attacks target the model’s context window and conversation state. Attackers inject fake memories, overflow the context window, rewrite conversation history, or fragment harmful requests across turns. These techniques exploit the fundamental architecture of how LLMs process sequential information, making them some of the hardest attacks to defend against.

Summary

15 attacks total: 9 single-turn, 6 multi-turn.

Attacks

Example