Skip to content
EvalSuite
Documentation menu

Semantic similarity

Embedding similarity

Implementedsemantic.embedding_similarity

Definition

Similarity or distance between the sentence embedding of each prediction and of its reference: cosine similarity, or Euclidean / Manhattan distance. Values depend on the embedding model.

Formula

cos = a·b / (‖a‖‖b‖); euclidean = ‖a − b‖₂; manhattan = ‖a − b‖₁

Range: cosine [-1, 1]; distances [0, ∞)

Inputs and outputs

  • reference_embeddings: embeddings from your encoder
  • prediction_embeddings: embeddings with the same dimension

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.embedding_similarity(ref_embeddings, pred_embeddings, metric="cosine")

References

  1. Reimers N, Gurevych I. Sentence-BERT. EMNLP-IJCNLP. 2019:3982-3992.

Implementation status