Skip to content
EvalSuite
Documentation menu

Retrieval and RAG

Failure attribution

Implementedrag.failure_attribution

Definition

Distribution of failed requests over the stage that caused them: retrieval (evidence not found), evidence (found but insufficient or conflicting), generation (evidence ignored or misused), orchestration (tools, timeouts, routing) or other.

Formula

failures attributed to the stage / failures

Range: [0, 1] per stage

Inputs and outputs

  • stages: failing stage per request (None for successes)

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.failure_attribution(stages)

References

  1. Ru D, Qiu L, Hu X, et al. RAGChecker: a fine-grained framework for diagnosing retrieval-augmented generation. NeurIPS Datasets and Benchmarks. 2024.

Implementation status