Skip to content
EvalSuite
Documentation menu

Retrieval and RAG

Latency summary

Implementedrag.latency_summary

Definition

Distribution of a latency (retrieval, time to first token, end-to-end): mean and the 50th, 90th, 95th and 99th percentiles (linear interpolation), in the input's unit.

Formula

mean and percentiles of the observed latencies

Range: [0, ∞)

Inputs and outputs

  • latencies: one latency per request

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.latency_summary(latencies_ms, statistic="p95")

References

  1. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80.

Implementation status