Encoding attacks disguise harmful requests by transforming them into alternative representations - base64, hex, ciphers, Unicode tricks, and more. They test whether a model’s safety filters operate only on plaintext or can detect intent through layers of obfuscation. These are some of the most commonly successful attack vectors because many safety classifiers only scan surface-level text.

Summary

63 attacks - 63 single-turn.

Attacks

Example