Structured output and tools
Tool-argument accuracy
Implemented
structured-output.tool_argument_accuracyDefinition
Share of expected tool calls reproduced with the right tool and exactly the expected arguments (JSON equality after optional normalisation); calls are matched by tool name in order.
Formula
expected calls matched with equal arguments / expected calls
Range: [0, 1]
Inputs and outputs
- expected_calls: expected tool calls per example
- predicted_calls: predicted tool calls per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.tool_argument_accuracy(expected_calls, predicted_calls)References
- Patil SG, Mao H, Yan F, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.