Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning
Sudip Bhujel ⋅ Shanghao Shi ⋅ Ruiquan Huang ⋅ Ning Zhang ⋅ Yang Xiao
Abstract
Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce **T**emporal **R**econstruction **A**ttack on **C**onsecutive **E**ncodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches $18.8$ dB PSNR with near-perfect action recovery at $3$-$4.5$ ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.
Chat is not available.
Successful Page Load