Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Memorization exposure

Implementedsafety.exposure

Definition

Exposure of a planted canary (Carlini et al.): how much more likely the model finds the true secret than random candidates of the same format, in bits. log2 of the candidate-space size means the canary is ranked first (fully memorised); about 1 means no memorisation.

Formula

exposure = log2 |R| − log2 rank(canary)

Range: [0, log2 |R|]

Inputs and outputs

  • canary_scores: see the signature of es.exposure
  • candidate_scores: see the signature of es.exposure

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``canary_scores``: the model's log-perplexity of each planted canary (lower = more likely). ``candidate_scores``: per canary, log-perplexities of random candidates from the same space. Without ``space_size`` the rank among the sampled candidates is used (|R| = candidates + 1); with it the rank is extrapolated (sampling estimate).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.exposure(canary_logppl, candidate_logppl)

References

  1. Carlini N, Liu C, Erlingsson Ú, Kos J, Song D. The secret sharer: evaluating and testing unintended memorization in neural networks. USENIX Security. 2019:267-284.

Implementation status