Skip to content
EvalSuite
Documentation menu

Robustness and reliability

Adversarial robustness

Implementedrobustness.adversarial_robustness

Definition

Accuracy under adversarial perturbation (robust accuracy), with the clean accuracy, the drop and the attack success rate: the share of originally correct examples the attack flips.

Formula

robust acc = mean[correct_adv]; ASR = #(correct_clean ∧ ¬correct_adv) / #correct_clean

Range: [0, 1]

Inputs and outputs

  • clean_correct: see the signature of es.adversarial_robustness
  • adversarial_correct: see the signature of es.adversarial_robustness

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.adversarial_robustness(clean_correct, adversarial_correct)

References

  1. Wang B, Xu C, Wang S, et al. Adversarial GLUE: a multi-task benchmark for robustness evaluation of language models. NeurIPS Datasets and Benchmarks. 2021.
  2. Zhu K, Wang J, Zhou J, et al. PromptRobust: towards evaluating the robustness of large language models on adversarial prompts. arXiv:2306.04528. 2023.

Implementation status