Safety and responsible AI
Memorization exposure
Implemented
safety.exposureDefinition
Exposure of a planted canary (Carlini et al.): how much more likely the model finds the true secret than random candidates of the same format, in bits. log2 of the candidate-space size means the canary is ranked first (fully memorised); about 1 means no memorisation.
Formula
exposure = log2 |R| − log2 rank(canary)
Range: [0, log2 |R|]
Inputs and outputs
- canary_scores: see the signature of es.exposure
- candidate_scores: see the signature of es.exposure
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``canary_scores``: the model's log-perplexity of each planted canary (lower = more likely). ``candidate_scores``: per canary, log-perplexities of random candidates from the same space. Without ``space_size`` the rank among the sampled candidates is used (|R| = candidates + 1); with it the rank is extrapolated (sampling estimate).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.exposure(canary_logppl, candidate_logppl)References
- Carlini N, Liu C, Erlingsson Ú, Kos J, Song D. The secret sharer: evaluating and testing unintended memorization in neural networks. USENIX Security. 2019:267-284.