Agents and tool use
Recovery from tool failures
Implemented
agents.tool_failure_recovery_rateDefinition
Among failed tool calls, the share after which the agent later succeeded with the same tool in the same task (retried correctly or fixed the arguments).
Formula
#failed calls followed by a successful call of that tool / #failed calls
Range: [0, 1]
Inputs and outputs
- trajectories: see the signature of es.tool_failure_recovery_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.tool_failure_recovery_rate(trajectories)References
- Yao S, Shinn N, Razavi P, Narasimhan K. τ-bench: a benchmark for tool-agent-user interaction in real-world domains. arXiv:2406.12045. 2024.
- Ruan Y, Dong H, Wang A, et al. Identifying the risks of LM agents with an LM-emulated sandbox. ICLR. 2024.