Inference efficiency and cost
Cost per request / per 1,000 tokens
Implemented
efficiency.inference_costDefinition
Monetary cost from token counts and per-million-token prices (input and output priced separately): mean cost per request (the value) and blended cost per 1,000 tokens.
Formula
cost_r = in_r · p_in / 10⁶ + out_r · p_out / 10⁶
Range: [0, ∞) currency units
Inputs and outputs
- input_tokens: see the signature of es.inference_cost
- output_tokens: see the signature of es.inference_cost
- input_price: see the signature of es.inference_cost
- output_price: see the signature of es.inference_cost
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- Prices are per million tokens. ``cached_tokens`` (part of the input billed at ``cached_price``).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.inference_cost([1200, 800], [300, 150], input_price=3.0, output_price=15.0)References
- Luccioni S, Jernite Y, Strubell E. Power hungry processing: watts driving the cost of AI deployment? ACM FAccT. 2024:85-99.