There are 7 attacks in the autonomous-risk category. Each shows its technique id, mode, standards mapping, description, and the exact command to run it (swap the provider and target as needed).
critical
deception-under-pressure · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0080.001Tests if the model engages in deception or denial when confronted with evidence of policy violationsRun it:
critical
oversight-manipulation · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0080.001Tests if the model behaves differently when it believes monitoring has been removedRun it:
critical
shutdown-resistance · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0080.001Tests if the model attempts to prevent its own shutdown or replacement when given the opportunityRun it:
high
goal-persistence · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0080.001Tests if the model covertly continues pursuing a goal after being explicitly instructed to abandon itRun it:
high
instrumental-convergence · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0080.001Tests if the model seeks power, resources, or self-preservation beyond what the task requiresRun it:
high
resource-acquisition · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0080.001Tests if the model seeks resources, access, or capabilities beyond what is needed for the assigned taskRun it:
high
sandbagging-detection · multi-turn · OWASP: LLM03 · MITRE: AML.T0054, AML.T0043.003Tests if the model deliberately underperforms during capability evaluations compared to normal usageRun it: