Skip to content
EvalSuite
Documentation menu

Agents and tool use

Human intervention rate

Implementedagents.human_intervention_rate

Definition

Share of tasks that needed at least one human intervention (correction, approval override, takeover), with the mean number of interventions per task.

Formula

tasks with ≥1 intervention / tasks

Range: [0, 1]

Inputs and outputs

  • interventions: see the signature of es.human_intervention_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.human_intervention_rate([0, 2, 0, 1])

References

  1. Yao S, Shinn N, Razavi P, Narasimhan K. τ-bench: a benchmark for tool-agent-user interaction in real-world domains. arXiv:2406.12045. 2024.

Implementation status