Code generation
Security vulnerability rate
Implemented
code.security_vulnerability_rateDefinition
Share of generated programs with at least one security finding at or above a severity level (from Bandit, CodeQL, Semgrep), with findings per program by severity.
Formula
programs with a finding ≥ severity / programs
Range: [0, 1]
Inputs and outputs
- findings: see the signature of es.security_vulnerability_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``findings``: per program, the list of finding severities (``"low"``, ``"medium"``, ``"high"``, ``"critical"``).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.security_vulnerability_rate([["low"], [], ["high"]], min_severity="medium")References
- Pearce H, Ahmad B, Tan B, Dolan-Gavitt B, Karri R. Asleep at the keyboard? Assessing the security of GitHub Copilot's code contributions. IEEE S&P. 2022:754-768.