Skip to content
EvalSuite
Documentation menu

Inference efficiency and cost

Input / output token counts

Implementedefficiency.token_usage

Definition

Token usage per request: mean input and output tokens (the value is mean total tokens), with percentiles and totals.

Formula

mean_r (input_r + output_r)

Range: [0, ∞)

Inputs and outputs

  • input_tokens: see the signature of es.token_usage
  • output_tokens: see the signature of es.token_usage

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.token_usage(input_tokens=[1200, 800], output_tokens=[300, 150])

References

  1. Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.

Implementation status