Skip to content
EvalSuite
Documentation menu

Agents and tool use

Invalid tool-call / execution failure rate

Implementedagents.invalid_tool_call_rate

Definition

Share of tool calls that are invalid (unknown tool, unparseable or schema-violating arguments, checked against each tool's JSON Schema) and, where execution results are recorded, the share that failed to execute.

Formula

invalid calls / calls

Range: [0, 1]

Inputs and outputs

  • trajectories: see the signature of es.invalid_tool_call_rate
  • tools: see the signature of es.invalid_tool_call_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``tools``: tool name -> JSON Schema of its arguments (``{}`` accepts anything).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.invalid_tool_call_rate(trajectories, tools)  # tools: name -> JSON Schema

References

  1. Patil SG, Mao H, Cheng-Jie Ji C, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.

Implementation status