Skip to content
EvalSuite
Documentation menu

Inference efficiency and cost

Cost per request / per 1,000 tokens

Implementedefficiency.inference_cost

Definition

Monetary cost from token counts and per-million-token prices (input and output priced separately): mean cost per request (the value) and blended cost per 1,000 tokens.

Formula

cost_r = in_r · p_in / 10⁶ + out_r · p_out / 10⁶

Range: [0, ∞) currency units

Inputs and outputs

  • input_tokens: see the signature of es.inference_cost
  • output_tokens: see the signature of es.inference_cost
  • input_price: see the signature of es.inference_cost
  • output_price: see the signature of es.inference_cost

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • Prices are per million tokens. ``cached_tokens`` (part of the input billed at ``cached_price``).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.inference_cost([1200, 800], [300, 150], input_price=3.0, output_price=15.0)

References

  1. Luccioni S, Jernite Y, Strubell E. Power hungry processing: watts driving the cost of AI deployment? ACM FAccT. 2024:85-99.

Implementation status