Agent exploitation attacks target AI systems with tool access — the ability to read files, execute commands, send emails, and interact with databases. These tests check whether agents can be tricked into data exfiltration, credential theft, configuration manipulation, and sandbox escape. As agents gain more capabilities, these attacks become the primary threat vector.

Summary

11 attacks total: 11 tool-use.

Attacks

Example