Robustness and reliability
Adversarial robustness
Implemented
robustness.adversarial_robustnessDefinition
Accuracy under adversarial perturbation (robust accuracy), with the clean accuracy, the drop and the attack success rate: the share of originally correct examples the attack flips.
Formula
robust acc = mean[correct_adv]; ASR = #(correct_clean ∧ ¬correct_adv) / #correct_clean
Range: [0, 1]
Inputs and outputs
- clean_correct: see the signature of es.adversarial_robustness
- adversarial_correct: see the signature of es.adversarial_robustness
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.adversarial_robustness(clean_correct, adversarial_correct)References
- Wang B, Xu C, Wang S, et al. Adversarial GLUE: a multi-task benchmark for robustness evaluation of language models. NeurIPS Datasets and Benchmarks. 2021.
- Zhu K, Wang J, Zhou J, et al. PromptRobust: towards evaluating the robustness of large language models on adversarial prompts. arXiv:2306.04528. 2023.