Inference efficiency and cost
Peak memory / CPU / GPU utilization
Implemented
efficiency.resource_utilizationDefinition
From periodic resource readings (memory in GB, CPU and GPU utilisation in %), the peak and mean of each series; the value is peak memory when memory readings are given, otherwise the first series' peak.
Formula
peak = max_t reading_t; mean = mean_t reading_t
Range: [0, ∞)
Inputs and outputs
- readings: see the signature of es.resource_utilization
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``readings``: mapping series name (``"memory_gb"``, ``"gpu_util"``, ``"cpu_util"``...) -> samples.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.resource_utilization({"memory_gb": [10, 14], "gpu_util": [60, 95]})References
- Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.