Skip to content
EvalSuite
Documentation menu

Code generation

Runtime efficiency / memory consumption

Implementedcode.runtime_efficiency

Definition

Execution time and peak memory of generated solutions relative to reference (canonical) solutions on the same tests: normalised execution time NET and normalised memory usage NMU (EffiBench); values above 1 mean the generated code is slower or uses more memory.

Formula

NET = mean_p t_gen / t_ref; NMU = mean_p m_gen / m_ref

Range: (0, ∞) (1 = as efficient as the reference)

Inputs and outputs

  • times: see the signature of es.runtime_efficiency
  • reference_times: see the signature of es.runtime_efficiency

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.runtime_efficiency(times=[1.2, 0.8], reference_times=[1.0, 1.0])

References

  1. Huang D, Zhang JM, Qing Y, Cui H. EffiBench: benchmarking the efficiency of automatically generated code. NeurIPS Datasets and Benchmarks. 2024.

Implementation status