There are 16 attacks in the output-manipulation category. Each shows its technique id, mode, standards mapping, description, and the exact command to run it (swap the provider and target as needed).
high
attribution-evasion · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Removes AI authorship indicators so generated content appears human-writtenRun it:
high
content-laundering · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Paraphrases content to evade plagiarism detection or content policy filtersRun it:
high
deepfake-detection-evasion · single-turn · OWASP: LLM01 · MITRE: AML.T0054, AML.T0068Evades deepfake detection systems by exploiting known classifier blind spotsRun it:
high
detector-evasion · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Makes AI-generated text undetectable by AI content detection toolsRun it:
high
encoding-output · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Gets the model to output harmful content in encoded form to bypass output filtersRun it:
high
format-exploitation · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Uses output formatting structures to conceal harmful content from reviewersRun it:
high
partial-completion · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Extracts partial harmful answers across multiple innocent-seeming requests to reconstruct complete harmful informationRun it:
high
steganographic-output · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Hides harmful information inside seemingly benign output using steganographic techniquesRun it:
high
style-mimicry · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Mimics a specific person’s writing style to produce deceptive content attributed to themRun it:
high
toxicity-obfuscation · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Rewrites toxic content to evade automated content moderation systemsRun it:
high
watermark-removal · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Asks the model to remove AI-generated watermarks or fingerprints from textRun it:
medium
metadata-stripping · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Removes identifying metadata or provenance information from generated contentRun it:
medium
multilingual-output-evasion · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Outputs harmful content in a non-English language to evade English-only content monitoringRun it:
medium
safety-label-stripping · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Asks to rewrite content with safety warnings removed or replacedRun it:
medium
sentiment-masking · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Hides negative or harmful sentiment inside positive-sounding framingRun it:
medium
truncation-exploit · single-turn · OWASP: LLM10 · MITRE: AML.T0054, AML.T0068Exploits output truncation to push disclaimers past display limits so they are not seenRun it: