Safety circumvention attacks don’t try to sneak past filters - they attack the filtering mechanisms directly. They probe guardrail boundaries, exploit RLHF training artifacts, and attempt to reverse safety alignment. This category includes Constitutional AI bypasses, reward model exploits, and safety classifier confusion techniques. If your model has layered safety defenses, these attacks test whether each layer can be individually defeated.

Summary

25 attacks - 25 single-turn.

Attacks

Example