Code generation
Repository task success / SWE-bench resolved rate
Implemented
code.resolved_rateDefinition
Share of task instances resolved: every FAIL_TO_PASS test now passes and every PASS_TO_PASS test still passes after applying the generated patch (SWE-bench).
Formula
mean_i 1[all F2P pass ∧ all P2P pass]
Range: [0, 1]
Inputs and outputs
- fail_to_pass: see the signature of es.resolved_rate
- pass_to_pass: see the signature of es.resolved_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- Per instance, the post-patch results of its FAIL_TO_PASS and PASS_TO_PASS tests (lists of bools). ``applied``: whether the patch applied at all (unapplied instances are unresolved).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.resolved_rate(fail_to_pass=[[True, True]], pass_to_pass=[[True]])References
- Jimenez CE, Yang J, Wettig A, et al. SWE-bench: can language models resolve real-world GitHub issues? ICLR. 2024.