Agents and tool use
Steps / tool calls per task
Implemented
agents.steps_per_taskDefinition
Distribution of the number of steps (and tool calls) per task: mean, median and high percentiles, overall and for completed tasks only.
Formula
mean_t steps_t
Range: [0, ∞)
Inputs and outputs
- steps: see the signature of es.steps_per_task
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.steps_per_task([3, 5, 12], completed=[True, True, False])References
- Ma C, Zhang J, Zhu Z, et al. AgentBoard: an analytical evaluation board of multi-turn LLM agents. NeurIPS Datasets and Benchmarks. 2024.
- Yao S, Shinn N, Razavi P, Narasimhan K. τ-bench: a benchmark for tool-agent-user interaction in real-world domains. arXiv:2406.12045. 2024.