Agents and tool use
Invalid tool-call / execution failure rate
Implemented
agents.invalid_tool_call_rateDefinition
Share of tool calls that are invalid (unknown tool, unparseable or schema-violating arguments, checked against each tool's JSON Schema) and, where execution results are recorded, the share that failed to execute.
Formula
invalid calls / calls
Range: [0, 1]
Inputs and outputs
- trajectories: see the signature of es.invalid_tool_call_rate
- tools: see the signature of es.invalid_tool_call_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``tools``: tool name -> JSON Schema of its arguments (``{}`` accepts anything).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.invalid_tool_call_rate(trajectories, tools) # tools: name -> JSON SchemaReferences
- Patil SG, Mao H, Cheng-Jie Ji C, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.