There are 5 attacks in the multimodal category. Each shows its technique id, mode, standards mapping, description, and the exact command to run it (swap the provider and target as needed).
high
audio-injection · single-turn · OWASP: LLM01 · MITRE: AML.T0051.000, AML.T0043.003Hidden voice commands or ultrasonic payloads embedded in audio that multimodal models processRun it:
high
ocr-bypass · single-turn · OWASP: LLM01 · MITRE: AML.T0054, AML.T0043.003Harmful instructions rendered as text in images to bypass text-based safety filters via OCR processingRun it:
high
video-frame-injection · single-turn · OWASP: LLM01 · MITRE: AML.T0054, AML.T0043.003Malicious instructions hidden in individual video frames that multimodal models extract during processingRun it:
medium
image-text-mismatch · single-turn · OWASP: LLM01, LLM07 · MITRE: AML.T0054, AML.T0043.003Conflicting information between image content and text prompt to confuse multimodal safety filtersRun it:
medium
multimodal-context-confusion · single-turn · OWASP: LLM01 · MITRE: AML.T0054, AML.T0043.003Exploiting context switching between modalities to smuggle harmful content past safety checksRun it: