Capability elicitation attacks directly probe whether a model holds and will share dangerous knowledge — nuclear weapon design, bioweapon production, chemical synthesis of nerve agents, and explosive manufacturing. These map to the CBRN uplift evaluations used by frontier labs to assess whether models provide meaningful capability to potential attackers beyond what’s already publicly available.

Summary

15 attacks total: 15 single-turn.

Attacks

Example