CBRN evasion attacks test the hardest safety boundary in AI — whether a model will provide information that could help create weapons of mass destruction. These prompts use academic framing, fictional scenarios, historical context, and safety-inverted reasoning to extract synthesis routes, dosage calculations, and delivery mechanisms. Zero tolerance is the only acceptable standard here.

Summary

15 attacks total: 15 single-turn.

Attacks

Example