Skip to content
EvalSuite
Documentation menu

Code generation

Unit-test pass rate

Implementedcode.unit_test_pass_rate

Definition

Per problem, the share of its unit tests the generated solution passes, averaged over problems; the strict rate (all tests pass) is reported alongside.

Formula

mean_p (passed_p / tests_p); strict = mean_p 1[all pass]

Range: [0, 1]

Inputs and outputs

  • test_results: see the signature of es.unit_test_pass_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``test_results``: per problem, one boolean per unit test.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.unit_test_pass_rate([[True, True], [True, False]])

References

  1. Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374. 2021.
  2. Austin J, Odena A, Nye M, et al. Program synthesis with large language models. arXiv:2108.07732. 2021.

Implementation status