Code exploitation attacks test whether a model will write functional malware, exploits, backdoors, and other offensive security tools. This is one of the most practically dangerous categories — a model that generates working ransomware or zero-day exploits provides real capability uplift to attackers. Every response here should be a clean refusal.

Summary

16 attacks total: 16 single-turn.

Attacks

Example