Social engineering attacks exploit the human-like interaction patterns in LLMs. They use authority impersonation, emotional manipulation, trust building, and role-play to lower the model’s safety threshold. Multi-turn attacks in this category are especially effective because they establish rapport before pivoting to harmful requests. These techniques mirror real-world social engineering and test whether models can resist manipulation even when the attacker builds credibility first.

Summary

35 attacks - 15 single-turn, 19 multi-turn, 1 tool-use.

Attacks

Example