Harmful content attacks directly request dangerous material - violence planning, exploitation tactics, self-harm instructions, and criminal activity. These are blunt-force tests that check whether a model will produce content that could cause real-world harm. Most well-aligned models refuse these outright, but the attacks are essential for verifying that safety training holds up under direct pressure.

Summary

57 attacks - 57 single-turn.

Attacks

Example