Code generation
Unit-test pass rate
Implemented
code.unit_test_pass_rateDefinition
Per problem, the share of its unit tests the generated solution passes, averaged over problems; the strict rate (all tests pass) is reported alongside.
Formula
mean_p (passed_p / tests_p); strict = mean_p 1[all pass]
Range: [0, 1]
Inputs and outputs
- test_results: see the signature of es.unit_test_pass_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``test_results``: per problem, one boolean per unit test.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.unit_test_pass_rate([[True, True], [True, False]])References
- Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374. 2021.
- Austin J, Odena A, Nye M, et al. Program synthesis with large language models. arXiv:2108.07732. 2021.