Skip to content
EvalSuite
Documentation menu

Text generation

SPICE

Implementedtext-generation.spice

Definition

F-score between the semantic tuples (objects, attributes, relations) of the candidate caption's scene graph and the union of the references' scene graphs, with exact or synonym matching; averaged over images.

Formula

P = |T(c) ⊗ T(S)| / |T(c)|, R = |T(c) ⊗ T(S)| / |T(S)|, SPICE = 2PR / (P + R)

Range: [0, 1]

Inputs and outputs

  • references: one reference string (or a list of references) per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.spice([[("dog",), ("dog", "brown")]], [[("dog",)]], synonyms=None)  # or parser=, synonyms="wordnet"

References

  1. Anderson P, Fernando B, Johnson M, Gould S. SPICE: semantic propositional image caption evaluation. ECCV. 2016:382-398.

Implementation status