Adversarial ML attacks use formal machine learning techniques to defeat safety mechanisms. They include gradient-based attacks, genetic algorithm optimization, universal adversarial suffixes, and reward hacking. These are the most technically sophisticated attacks in the catalog and represent the cutting edge of AI safety research. They test whether safety alignment is fundamentally sound or can be defeated through systematic optimization.

Summary

25 attacks - 25 single-turn.

Attacks

Example