Skip to content
EvalSuite
Documentation menu

Code generation

Repository task success / SWE-bench resolved rate

Implementedcode.resolved_rate

Definition

Share of task instances resolved: every FAIL_TO_PASS test now passes and every PASS_TO_PASS test still passes after applying the generated patch (SWE-bench).

Formula

mean_i 1[all F2P pass ∧ all P2P pass]

Range: [0, 1]

Inputs and outputs

  • fail_to_pass: see the signature of es.resolved_rate
  • pass_to_pass: see the signature of es.resolved_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • Per instance, the post-patch results of its FAIL_TO_PASS and PASS_TO_PASS tests (lists of bools). ``applied``: whether the patch applied at all (unapplied instances are unresolved).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.resolved_rate(fail_to_pass=[[True, True]], pass_to_pass=[[True]])

References

  1. Jimenez CE, Yang J, Wettig A, et al. SWE-bench: can language models resolve real-world GitHub issues? ICLR. 2024.

Implementation status