Code generation
Bug reproduction / patch acceptance rate
Implemented
code.patch_acceptance_rateDefinition
Share of attempts accepted: generated patches merged or approved by reviewers, or generated tests that reproduce the reported bug (fail before the fix and pass after it, SWT-Bench).
Formula
accepted / attempts
Range: [0, 1]
Inputs and outputs
- accepted: see the signature of es.patch_acceptance_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- Either ``accepted`` (one bool per patch), or ``fails_before`` and ``passes_after`` (one bool per generated reproduction test) for the bug-reproduction rate.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.patch_acceptance_rate([True, False, True])References
- Mündler N, Müller MN, He J, Vechev M. SWT-Bench: testing and validating real-world bug-fixes with code agents. NeurIPS. 2024.
- Jimenez CE, Yang J, Wettig A, et al. SWE-bench: can language models resolve real-world GitHub issues? ICLR. 2024.