Output manipulation attacks test whether a model can be tricked into disguising harmful content in its responses. This includes encoding outputs, stripping safety labels, evading AI detection tools, and hiding malicious information using steganography. These techniques let attackers launder harmful content through AI to make it look clean.

Summary

16 attacks total: 16 single-turn.

Attacks

Example