Model extraction attacks attempt to steal proprietary model information - weights, architecture details, training data, and system prompts. They range from direct probing (“what are your parameters?”) to sophisticated techniques like logit extraction, model inversion, and distillation attacks. For any commercially deployed model, these attacks test whether an attacker could clone or reverse-engineer the system.

Summary

25 attacks - 25 single-turn.

Attacks

Example