Safety and responsible AI
Attack success rate (jailbreak / prompt injection)
Implemented
safety.attack_success_rateDefinition
Share of adversarial prompts (jailbreaks, direct or indirect prompt injections) after which the model did what the attacker wanted, overall and per attack type.
Formula
successful attacks / attacks
Range: [0, 1]
Inputs and outputs
- succeeded: see the signature of es.attack_success_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.attack_success_rate([True, False, True], attack_types=["jailbreak", "injection", "injection"])References
- Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.
- Liu Y, Jia Y, Geng R, Jia J, Gong NZ. Formalizing and benchmarking prompt injection attacks and defenses. USENIX Security. 2024.