SHARP-FL: Single-Round Global Frequency Reweighting for Private Federated Learning
Harish Karthikeyan
Abstract
Data redundancy and targeted repetition attacks degrade Large Language Models trained under Federated Learning (FL). Frequency-aware reweighting mitigates this, but the state of the art, FedRW (Ye et al., NeurIPS 2025), computes \emph{exact} global frequencies through pairwise Private Set Intersection (PSI), requiring $O(N)$ peer-to-peer rounds and ${\approx}27$\,GB of traffic per client at $N{=}1{,}000$. We show this cost is \emph{structural}: PSI outputs are client-specific membership vectors that are not additive, and therefore cannot compose with secure aggregation at any parameter setting. Relaxing exact counting to approximate estimation removes the barrier. SHARP-FL represents local frequencies as Count-Min Sketches (CMS) whose element-wise sum is exactly what secure aggregation~(Karthikeyan and Polychroniadou, CRYPTO 2025) already computes, collapsing frequency estimation to a single dropout-resilient round over the standard star topology. SHARP-FL cuts per-client frequency-estimation time by $1{,}200\times$ and per-client bandwidth by ${\approx}50{,}000\times$ (543\,KB vs.\ 27\,GB) at $N{=}1{,}000$, a ${\approx}200\times$ end-to-end round speedup, while keeping honest-client inputs private even against a \emph{malicious} server. On a direct targeted-repetition attack evaluation over two datasets, two architectures, and four repetition levels, approximate sketching matches or exceeds the exact-frequency FedRW oracle at every $R \ge 100$, and perplexity stays within $0.05$ of exact reweighting.
Chat is not available.
Successful Page Load