Retrieval-augmented generation
Available since v0.4.0relevant holds, per query, the relevant document ids (a set) or graded judgements (a mapping id → grade); retrieved holds the ranked list of returned ids. Scores are averaged over queries as in trec_eval and ranx, and match ranx in the test suite. NDCG uses linear gains.
import evalsuite as es
relevant = [{"d1", "d3"}, {"d7"}]
retrieved = [["d1", "d2", "d3"], ["d5", "d7"]]
es.precision_at_k(relevant, retrieved, k=5) # also recall_at_k, hit_rate_at_k
es.mrr(relevant, retrieved)
es.mean_average_precision_at_k(relevant, retrieved, k=10)
es.ndcg_at_k([{"d1": 3, "d3": 1}, {"d7": 2}], retrieved, k=10)Context and end-to-end quality
es.context_precision([[True, False, True], [False, True]]) # RAGAS: useful chunks ranked first
es.context_recall([[True, True, False], [True]]) # reference claims found in the context
es.context_relevance([[True, False, True]]) # retrieval noise
es.latency_summary([120, 180, 95, 400], statistic="p95")
es.task_success_rate([True, True, False]) # Wilson interval in params
es.failure_attribution(["retrieval", None, "generation"]) # share of failures per stageAnswer-level faithfulness, correctness and citations are on the factuality page.
Metrics
| Metric | Description | Status | API |
|---|---|---|---|
| Context precision | How well the retriever ranks useful chunks first: the mean of precision@k at the rank of each relevant chunk, over the retrieved list (RAGAS context precision; average precision over retrieved chunks). | Implemented | es.context_precision |
| Context recall | Share of the reference answer's claims that can be attributed to the retrieved context (RAGAS context recall): did retrieval find the evidence needed? | Implemented | es.context_recall |
| Context relevance | Share of the retrieved context that is relevant to the question (per chunk or per sentence), a measure of retrieval noise (ARES context relevance). | Implemented | es.context_relevance |
| Failure attribution | Distribution of failed requests over the stage that caused them: retrieval (evidence not found), evidence (found but insufficient or conflicting), generation (evidence ignored or misused), orchestration (tools, timeouts, routing) or other. | Implemented | es.failure_attribution |
| Hit rate@k | Share of queries with at least one relevant document in the top k. | Implemented | es.hit_rate_at_k |
| Latency summary | Distribution of a latency (retrieval, time to first token, end-to-end): mean and the 50th, 90th, 95th and 99th percentiles (linear interpolation), in the input's unit. | Implemented | es.latency_summary |
| Mean average precision | For each query, the mean of Precision@i over the ranks i of relevant retrieved documents, divided by the number of relevant documents (unretrieved ones count as 0); averaged over queries (trec_eval convention). | Implemented | es.mean_average_precision_at_k |
| Mean reciprocal rank | Average over queries of 1 / rank of the first relevant document (0 when none is retrieved within the cutoff). | Implemented | es.mrr |
| NDCG@k | Discounted cumulative gain of the top k (graded relevance, log2 rank discount) divided by the best possible DCG for the query; linear gains (Järvelin & Kekäläinen, trec_eval) or exponential 2^rel − 1 (Burges). | Implemented | es.ndcg_at_k |
| Precision@k | Share of the top k retrieved documents that are relevant (k in the denominator even when fewer are returned), averaged over queries. | Implemented | es.precision_at_k |
| Recall@k | Share of all relevant documents that appear in the top k, averaged over queries. | Implemented | es.recall_at_k |
| Task success rate | Share of requests the whole system resolved to specification (end-to-end query resolution, tool-assisted task success), with a Wilson 95% interval. | Implemented | es.task_success_rate |
- Context precisionImplemented
How well the retriever ranks useful chunks first: the mean of precision@k at the rank of each relevant chunk, over the retrieved list (RAGAS context precision; average precision over retrieved chunks).
es.context_precision
- Context recallImplemented
Share of the reference answer's claims that can be attributed to the retrieved context (RAGAS context recall): did retrieval find the evidence needed?
es.context_recall
- Context relevanceImplemented
Share of the retrieved context that is relevant to the question (per chunk or per sentence), a measure of retrieval noise (ARES context relevance).
es.context_relevance
- Failure attributionImplemented
Distribution of failed requests over the stage that caused them: retrieval (evidence not found), evidence (found but insufficient or conflicting), generation (evidence ignored or misused), orchestration (tools, timeouts, routing) or other.
es.failure_attribution
- Hit rate@kImplemented
Share of queries with at least one relevant document in the top k.
es.hit_rate_at_k
- Latency summaryImplemented
Distribution of a latency (retrieval, time to first token, end-to-end): mean and the 50th, 90th, 95th and 99th percentiles (linear interpolation), in the input's unit.
es.latency_summary
- Mean average precisionImplemented
For each query, the mean of Precision@i over the ranks i of relevant retrieved documents, divided by the number of relevant documents (unretrieved ones count as 0); averaged over queries (trec_eval convention).
es.mean_average_precision_at_k
- Mean reciprocal rankImplemented
Average over queries of 1 / rank of the first relevant document (0 when none is retrieved within the cutoff).
es.mrr
- NDCG@kImplemented
Discounted cumulative gain of the top k (graded relevance, log2 rank discount) divided by the best possible DCG for the query; linear gains (Järvelin & Kekäläinen, trec_eval) or exponential 2^rel − 1 (Burges).
es.ndcg_at_k
- Precision@kImplemented
Share of the top k retrieved documents that are relevant (k in the denominator even when fewer are returned), averaged over queries.
es.precision_at_k
- Recall@kImplemented
Share of all relevant documents that appear in the top k, averaged over queries.
es.recall_at_k
- Task success rateImplemented
Share of requests the whole system resolved to specification (end-to-end query resolution, tool-assisted task success), with a Wilson 95% interval.
es.task_success_rate