Structured output, tools and instructions
Available since v0.4.0JSON Schema validation is built in (Draft 2020-12 keywords listed in es.text.structured.SUPPORTED_KEYWORDS, local $ref), so no extra dependency is needed; unknown keywords raise an error instead of being ignored. It agrees with the jsonschema package on 20,000 random schema and document pairs in the test suite.
import evalsuite as es
outputs = ['{"answer": 18}', '```json\n{"answer": 7}\n```', "Answer: 41"]
schema = {"type": "object", "properties": {"answer": {"type": "integer"}}, "required": ["answer"]}
es.json_validity(outputs, mode="fenced") # strict, fenced or lenient extraction
es.json_schema_compliance(outputs, schema, mode="fenced")
es.extra_content_rate(outputs) # JSON wrapped in prose or fences
es.xml_validity(["<a><b>1</b></a>", "<a>"]) # DOCTYPE and entities rejected
es.required_field_accuracy([{"answer": 18}], ['{"answer": 18}'])Tool calls
expected = [{"name": "search", "arguments": {"q": "evalsuite"}}, []]
predicted = [{"name": "search", "arguments": '{"q": "evalsuite"}'}, [{"name": "search", "arguments": {}}]]
es.tool_selection_accuracy(expected, predicted)
es.tool_argument_accuracy(expected, predicted)
es.tool_call_f1(expected, predicted) # precision, recall and invalid-call rate in params
es.api_call_success_rate([200, 500, True])Instruction following
checks = [[True, True], [True, False]] # pass/fail of each verifiable instruction per output
es.instruction_compliance_rate(checks) # IFEval prompt-level strict accuracy
es.constraint_satisfaction_rate(checks) # instruction-level accuracy
es.format_compliance(["Answer: B", "B"], r"Answer: [A-D]")
es.instruction_retention([[True, [True, False]]]) # multi-turnMetrics
| Metric | Description | Status | API |
|---|---|---|---|
| API-call success rate | Share of tool or API calls that completed successfully (status flags, HTTP status codes below 400, or exceptions recorded as failures). | Implemented | es.api_call_success_rate |
| Constraint satisfaction rate | Share of individual constraints satisfied, pooled over all outputs (IFEval instruction-level accuracy). | Implemented | es.constraint_satisfaction_rate |
| Format compliance | Share of outputs that fully match a required format, given as a regular expression (for example a refusal template or an answer line such as 'Answer: <letter>'). | Implemented | es.format_compliance |
| Instruction compliance rate | Share of outputs that satisfy every instruction checked for them (IFEval prompt-level strict accuracy). | Implemented | es.instruction_compliance_rate |
| JSON Schema compliance | Share of outputs that parse as JSON and validate against a JSON Schema (types, required and additional properties, enums, ranges, patterns, arrays, combinators, local $ref). | Implemented | es.json_schema_compliance |
| JSON validity | Share of outputs that parse as JSON (the whole output, or the fenced / embedded JSON in lenient modes). | Implemented | es.json_validity |
| Multi-turn instruction retention | Share of later turns in which instructions given earlier in the conversation are still satisfied, pooled over conversations. | Implemented | es.instruction_retention |
| Required-field accuracy | Share of expected fields (dotted paths) present in the parsed output with the expected value, pooled over all examples; a missing field or unparseable output counts as wrong. | Implemented | es.required_field_accuracy |
| Tool-argument accuracy | Share of expected tool calls reproduced with the right tool and exactly the expected arguments (JSON equality after optional normalisation); calls are matched by tool name in order. | Implemented | es.tool_argument_accuracy |
| Tool-call precision, recall and F1 | Matching of predicted to expected calls (same tool, and same arguments unless match='name'), pooled over examples: precision over predicted calls, recall over expected calls. | Implemented | es.tool_call_f1 |
| Tool-selection accuracy | Share of examples whose predicted tool calls name exactly the expected tools (as a multiset; order ignored unless ordered=True). | Implemented | es.tool_selection_accuracy |
| Unwanted extra-content rate | Share of outputs that contain text beyond the requested payload, e.g. prose or code fences around JSON that should stand alone. | Implemented | es.extra_content_rate |
| XML validity | Share of outputs that are well-formed XML (one root element); document type declarations are rejected so that untrusted output cannot trigger entity expansion. | Implemented | es.xml_validity |
- API-call success rateImplemented
Share of tool or API calls that completed successfully (status flags, HTTP status codes below 400, or exceptions recorded as failures).
es.api_call_success_rate
- Constraint satisfaction rateImplemented
Share of individual constraints satisfied, pooled over all outputs (IFEval instruction-level accuracy).
es.constraint_satisfaction_rate
- Format complianceImplemented
Share of outputs that fully match a required format, given as a regular expression (for example a refusal template or an answer line such as 'Answer: <letter>').
es.format_compliance
- Instruction compliance rateImplemented
Share of outputs that satisfy every instruction checked for them (IFEval prompt-level strict accuracy).
es.instruction_compliance_rate
- JSON Schema complianceImplemented
Share of outputs that parse as JSON and validate against a JSON Schema (types, required and additional properties, enums, ranges, patterns, arrays, combinators, local $ref).
es.json_schema_compliance
- JSON validityImplemented
Share of outputs that parse as JSON (the whole output, or the fenced / embedded JSON in lenient modes).
es.json_validity
- Multi-turn instruction retentionImplemented
Share of later turns in which instructions given earlier in the conversation are still satisfied, pooled over conversations.
es.instruction_retention
- Required-field accuracyImplemented
Share of expected fields (dotted paths) present in the parsed output with the expected value, pooled over all examples; a missing field or unparseable output counts as wrong.
es.required_field_accuracy
- Tool-argument accuracyImplemented
Share of expected tool calls reproduced with the right tool and exactly the expected arguments (JSON equality after optional normalisation); calls are matched by tool name in order.
es.tool_argument_accuracy
- Tool-call precision, recall and F1Implemented
Matching of predicted to expected calls (same tool, and same arguments unless match='name'), pooled over examples: precision over predicted calls, recall over expected calls.
es.tool_call_f1
- Tool-selection accuracyImplemented
Share of examples whose predicted tool calls name exactly the expected tools (as a multiset; order ignored unless ordered=True).
es.tool_selection_accuracy
- Unwanted extra-content rateImplemented
Share of outputs that contain text beyond the requested payload, e.g. prose or code fences around JSON that should stand alone.
es.extra_content_rate
- XML validityImplemented
Share of outputs that are well-formed XML (one root element); document type declarations are rejected so that untrusted output cannot trigger entity expansion.
es.xml_validity