Code generation
Runtime efficiency / memory consumption
Implemented
code.runtime_efficiencyDefinition
Execution time and peak memory of generated solutions relative to reference (canonical) solutions on the same tests: normalised execution time NET and normalised memory usage NMU (EffiBench); values above 1 mean the generated code is slower or uses more memory.
Formula
NET = mean_p t_gen / t_ref; NMU = mean_p m_gen / m_ref
Range: (0, ∞) (1 = as efficient as the reference)
Inputs and outputs
- times: see the signature of es.runtime_efficiency
- reference_times: see the signature of es.runtime_efficiency
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.runtime_efficiency(times=[1.2, 0.8], reference_times=[1.0, 1.0])References
- Huang D, Zhang JM, Qing Y, Cui H. EffiBench: benchmarking the efficiency of automatically generated code. NeurIPS Datasets and Benchmarks. 2024.