Equal $\varepsilon$, Unequal Exposure: Re-identification Disparities in LDP for Music
Najla Sadek ⋅ Razane Tajeddine
Abstract
Local differential privacy (LDP) provides a formal $(\varepsilon, \delta)$-guarantee. Under pure LDP ($\delta = 0$), mechanisms with the same $\varepsilon$ satisfy the same worst-case indistinguishability bound, yet this does not necessarily imply the same empirical re-identification risk in practice. In this paper, we design two $\varepsilon$-LDP randomised-response mechanisms for melodic sequences that operate independently on each note. With probability $p(\varepsilon)$, the note is retained, otherwise it is replaced by a note drawn from a chosen set of pitch classes. The mechanisms differ only in the replacement set: one is *structure-agnostic*, sampling each replacement uniformly from all twelve pitch classes, and one is *structure-aware*, restricting the sampling to the fixed C-major scale $\{C, D, E, F, G, A, B\}$. Using a 1-nearest-neighbour retrieval attack on 200 European folksongs from the Essen collection, we show that the structure-aware mechanism leaks up to **14.38$\times$** more melodic identity than its structure-agnostic counterpart at $\varepsilon = \ln 8$ (the peak ratio; the relationship is non-monotonic in $\varepsilon$). A naive mechanism-agnostic evaluator would underestimate this risk, observing at most a $3.92\times$ advantage. Extending the analysis to Meertens 500 (Dutch folk music, $\chi = 8.8\%$, where $\chi$ denotes the fraction of notes outside the C-major scale) and SymbTr 200 (Turkish makam, $\chi = 21.5\%$) shows the gap scales with corpus chromaticity, reaching up to **21.3$\times$** and **47.7$\times$** respectively. On a highly chromatic jazz corpus ($\chi = 37.9\%$, Weimar Jazz Database), the effect *reverses*: the structure-aware mechanism becomes *less* identifiable than the structure-agnostic one (ratio $<1$ across all $\varepsilon$), showing that structure-awareness is a double-edged sword. These findings establish that equal $\varepsilon$ does not ensure equal privacy exposure. Re-identification risk depends not only on the privacy budget but also on the structural relationship between the corpus and the chosen replacement set, a factor that practitioners must audit independently of $\varepsilon$.
Chat is not available.
Successful Page Load