Semantic similarity
Embedding similarity
Implemented
semantic.embedding_similarityDefinition
Similarity or distance between the sentence embedding of each prediction and of its reference: cosine similarity, or Euclidean / Manhattan distance. Values depend on the embedding model.
Formula
cos = a·b / (‖a‖‖b‖); euclidean = ‖a − b‖₂; manhattan = ‖a − b‖₁
Range: cosine [-1, 1]; distances [0, ∞)
Inputs and outputs
- reference_embeddings: embeddings from your encoder
- prediction_embeddings: embeddings with the same dimension
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.embedding_similarity(ref_embeddings, pred_embeddings, metric="cosine")References
- Reimers N, Gurevych I. Sentence-BERT. EMNLP-IJCNLP. 2019:3982-3992.