Structured output and tools
Tool-selection accuracy
Implemented
structured-output.tool_selection_accuracyDefinition
Share of examples whose predicted tool calls name exactly the expected tools (as a multiset; order ignored unless ordered=True).
Formula
mean[tools(predicted) = tools(expected)]
Range: [0, 1]
Inputs and outputs
- expected_calls: expected tool calls per example
- predicted_calls: predicted tool calls per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.tool_selection_accuracy(expected_calls, predicted_calls)References
- Patil SG, Mao H, Yan F, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.