Inference efficiency and cost
Input / output token counts
Implemented
efficiency.token_usageDefinition
Token usage per request: mean input and output tokens (the value is mean total tokens), with percentiles and totals.
Formula
mean_r (input_r + output_r)
Range: [0, ∞)
Inputs and outputs
- input_tokens: see the signature of es.token_usage
- output_tokens: see the signature of es.token_usage
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.token_usage(input_tokens=[1200, 800], output_tokens=[300, 150])References
- Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.