Retrieval and RAG
Latency summary
Implemented
rag.latency_summaryDefinition
Distribution of a latency (retrieval, time to first token, end-to-end): mean and the 50th, 90th, 95th and 99th percentiles (linear interpolation), in the input's unit.
Formula
mean and percentiles of the observed latencies
Range: [0, ∞)
Inputs and outputs
- latencies: one latency per request
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.latency_summary(latencies_ms, statistic="p95")References
- Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80.