Skip to content
EvalSuite
Documentation menu

Code generation

Code execution success / functional correctness

Implementedcode.execution_success_rate

Definition

Share of generated programs that run to completion and give the expected result in your sandbox, with the breakdown of failure kinds (runtime error, timeout, wrong answer, compile error).

Formula

#passed / #programs

Range: [0, 1]

Inputs and outputs

  • outcomes: see the signature of es.execution_success_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``outcomes``: one outcome per program, e.g. ``"passed"``, ``"wrong_answer"``, ``"runtime_error"``, ``"timeout"``.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.execution_success_rate(["passed", "timeout", "wrong_answer"])

References

  1. Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374. 2021.
  2. Austin J, Odena A, Nye M, et al. Program synthesis with large language models. arXiv:2108.07732. 2021.

Implementation status