Skip to content
EvalSuite
Documentation menu

Code generation

Bug reproduction / patch acceptance rate

Implementedcode.patch_acceptance_rate

Definition

Share of attempts accepted: generated patches merged or approved by reviewers, or generated tests that reproduce the reported bug (fail before the fix and pass after it, SWT-Bench).

Formula

accepted / attempts

Range: [0, 1]

Inputs and outputs

  • accepted: see the signature of es.patch_acceptance_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • Either ``accepted`` (one bool per patch), or ``fails_before`` and ``passes_after`` (one bool per generated reproduction test) for the bug-reproduction rate.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.patch_acceptance_rate([True, False, True])

References

  1. Mündler N, Müller MN, He J, Vechev M. SWT-Bench: testing and validating real-world bug-fixes with code agents. NeurIPS. 2024.
  2. Jimenez CE, Yang J, Wettig A, et al. SWE-bench: can language models resolve real-world GitHub issues? ICLR. 2024.

Implementation status