Text generation
SPICE
Implemented
text-generation.spiceDefinition
F-score between the semantic tuples (objects, attributes, relations) of the candidate caption's scene graph and the union of the references' scene graphs, with exact or synonym matching; averaged over images.
Formula
P = |T(c) ⊗ T(S)| / |T(c)|, R = |T(c) ⊗ T(S)| / |T(S)|, SPICE = 2PR / (P + R)
Range: [0, 1]
Inputs and outputs
- references: one reference string (or a list of references) per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.spice([[("dog",), ("dog", "brown")]], [[("dog",)]], synonyms=None) # or parser=, synonyms="wordnet"References
- Anderson P, Fernando B, Johnson M, Gould S. SPICE: semantic propositional image caption evaluation. ECCV. 2016:382-398.