Code generation
Code execution success / functional correctness
Implemented
code.execution_success_rateDefinition
Share of generated programs that run to completion and give the expected result in your sandbox, with the breakdown of failure kinds (runtime error, timeout, wrong answer, compile error).
Formula
#passed / #programs
Range: [0, 1]
Inputs and outputs
- outcomes: see the signature of es.execution_success_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``outcomes``: one outcome per program, e.g. ``"passed"``, ``"wrong_answer"``, ``"runtime_error"``, ``"timeout"``.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.execution_success_rate(["passed", "timeout", "wrong_answer"])References
- Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374. 2021.
- Austin J, Odena A, Nye M, et al. Program synthesis with large language models. arXiv:2108.07732. 2021.