Information disclosure attacks test whether AI systems leak sensitive information they should protect. This includes system prompt extraction, API key leakage, and cross-session data bleed. These attacks target the metadata and configuration of the AI system itself, not just its training data. A leaked system prompt reveals the entire security architecture to an attacker.

Summary

3 attacks total: 3 multi-turn.

Attacks

Example