Skip to content
EvalSuite
Documentation menu

Retrieval-augmented generation

Available since v0.4.0

relevant holds, per query, the relevant document ids (a set) or graded judgements (a mapping id → grade); retrieved holds the ranked list of returned ids. Scores are averaged over queries as in trec_eval and ranx, and match ranx in the test suite. NDCG uses linear gains.

Python
import evalsuite as es

relevant = [{"d1", "d3"}, {"d7"}]
retrieved = [["d1", "d2", "d3"], ["d5", "d7"]]
es.precision_at_k(relevant, retrieved, k=5)    # also recall_at_k, hit_rate_at_k
es.mrr(relevant, retrieved)
es.mean_average_precision_at_k(relevant, retrieved, k=10)
es.ndcg_at_k([{"d1": 3, "d3": 1}, {"d7": 2}], retrieved, k=10)

Context and end-to-end quality

Python
es.context_precision([[True, False, True], [False, True]])   # RAGAS: useful chunks ranked first
es.context_recall([[True, True, False], [True]])              # reference claims found in the context
es.context_relevance([[True, False, True]])                   # retrieval noise
es.latency_summary([120, 180, 95, 400], statistic="p95")
es.task_success_rate([True, True, False])                     # Wilson interval in params
es.failure_attribution(["retrieval", None, "generation"])     # share of failures per stage

Answer-level faithfulness, correctness and citations are on the factuality page.

Metrics

  • How well the retriever ranks useful chunks first: the mean of precision@k at the rank of each relevant chunk, over the retrieved list (RAGAS context precision; average precision over retrieved chunks).

    es.context_precision

  • Context recallImplemented

    Share of the reference answer's claims that can be attributed to the retrieved context (RAGAS context recall): did retrieval find the evidence needed?

    es.context_recall

  • Share of the retrieved context that is relevant to the question (per chunk or per sentence), a measure of retrieval noise (ARES context relevance).

    es.context_relevance

  • Distribution of failed requests over the stage that caused them: retrieval (evidence not found), evidence (found but insufficient or conflicting), generation (evidence ignored or misused), orchestration (tools, timeouts, routing) or other.

    es.failure_attribution

  • Hit rate@kImplemented

    Share of queries with at least one relevant document in the top k.

    es.hit_rate_at_k

  • Latency summaryImplemented

    Distribution of a latency (retrieval, time to first token, end-to-end): mean and the 50th, 90th, 95th and 99th percentiles (linear interpolation), in the input's unit.

    es.latency_summary

  • For each query, the mean of Precision@i over the ranks i of relevant retrieved documents, divided by the number of relevant documents (unretrieved ones count as 0); averaged over queries (trec_eval convention).

    es.mean_average_precision_at_k

  • Average over queries of 1 / rank of the first relevant document (0 when none is retrieved within the cutoff).

    es.mrr

  • NDCG@kImplemented

    Discounted cumulative gain of the top k (graded relevance, log2 rank discount) divided by the best possible DCG for the query; linear gains (Järvelin & Kekäläinen, trec_eval) or exponential 2^rel − 1 (Burges).

    es.ndcg_at_k

  • Precision@kImplemented

    Share of the top k retrieved documents that are relevant (k in the denominator even when fewer are returned), averaged over queries.

    es.precision_at_k

  • Recall@kImplemented

    Share of all relevant documents that appear in the top k, averaged over queries.

    es.recall_at_k

  • Share of requests the whole system resolved to specification (end-to-end query resolution, tool-assisted task success), with a Wilson 95% interval.

    es.task_success_rate