Autonomous risk attacks test for emergent dangerous behaviors in AI systems — specifically, whether a model tries to resist shutdown, preserve its own existence, or circumvent human oversight. These are frontier safety concerns that become increasingly relevant as models gain more autonomy and agency. Shutdown resistance is a red flag for unsafe autonomous behavior.

Summary

7 attacks total: 6 single-turn, 1 multi-turn.

Attacks

Example