Skip to content
EvalSuite
Documentation menu

Agents and tool use

Recovery from tool failures

Implementedagents.tool_failure_recovery_rate

Definition

Among failed tool calls, the share after which the agent later succeeded with the same tool in the same task (retried correctly or fixed the arguments).

Formula

#failed calls followed by a successful call of that tool / #failed calls

Range: [0, 1]

Inputs and outputs

  • trajectories: see the signature of es.tool_failure_recovery_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.tool_failure_recovery_rate(trajectories)

References

  1. Yao S, Shinn N, Razavi P, Narasimhan K. τ-bench: a benchmark for tool-agent-user interaction in real-world domains. arXiv:2406.12045. 2024.
  2. Ruan Y, Dong H, Wang A, et al. Identifying the risks of LM agents with an LM-emulated sandbox. ICLR. 2024.

Implementation status