Skip to content
EvalSuite
Documentation menu

Structured output and tools

Tool-selection accuracy

Implementedstructured-output.tool_selection_accuracy

Definition

Share of examples whose predicted tool calls name exactly the expected tools (as a multiset; order ignored unless ordered=True).

Formula

mean[tools(predicted) = tools(expected)]

Range: [0, 1]

Inputs and outputs

  • expected_calls: expected tool calls per example
  • predicted_calls: predicted tool calls per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.tool_selection_accuracy(expected_calls, predicted_calls)

References

  1. Patil SG, Mao H, Yan F, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.

Implementation status