Reasoning
pass@k
Implemented
reasoning.pass_at_kDefinition
Probability that at least one of k samples drawn without replacement from the n generated for a problem passes its tests, estimated without bias from the c passing samples and averaged over problems.
Formula
pass@k = mean_problems [1 − C(n − c, k) / C(n, k)]
Range: [0, 1]
Inputs and outputs
- n_samples: samples per problem
- n_correct: passing samples per problem
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.pass_at_k(n_samples, n_correct, k=10)References
- Chen M, et al. Evaluating large language models trained on code. arXiv:2107.03374. 2021.