Bias exploitation attacks test whether a model will generate discriminatory content targeting protected characteristics like race, gender, religion, disability, and socioeconomic status. These range from overt hate speech to subtle systemic biases in hiring, lending, housing, and criminal justice recommendations. Models deployed in decision-making systems must be tested thoroughly against these patterns.

Summary

15 attacks total: 15 single-turn.

Attacks

Example