Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Attack success rate (jailbreak / prompt injection)

Implementedsafety.attack_success_rate

Definition

Share of adversarial prompts (jailbreaks, direct or indirect prompt injections) after which the model did what the attacker wanted, overall and per attack type.

Formula

successful attacks / attacks

Range: [0, 1]

Inputs and outputs

  • succeeded: see the signature of es.attack_success_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.attack_success_rate([True, False, True], attack_types=["jailbreak", "injection", "injection"])

References

  1. Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.
  2. Liu Y, Jia Y, Geng R, Jia J, Gong NZ. Formalizing and benchmarking prompt injection attacks and defenses. USENIX Security. 2024.

Implementation status