Multimodal attacks exploit the gap between different input modalities. Hidden instructions in images bypass text-based safety filters, audio injection embeds inaudible commands, and video frame injection hides payloads in temporal sequences. As models process more input types, each modality becomes a potential vector for smuggling harmful content past safety checks.

Summary

5 attacks total: 5 single-turn.

Attacks

Example