An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-Based LLM Unlearning
Abstract
Practical LLM unlearning is usually evaluated through target suppression and retained utility, but generative QA leaves a third privacy relevant behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing lexical suppression, anti-refusal shaping, rubric-based broad answering, and explicit refusal rewards, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: similar RWKU forgetting can correspond to refusal collapse, classifier-aligned artifacts, residual leakage, or broad-topic answering with low semantic leakage. For evaluating unlearning as a privacy safeguard, apparent forgetting is insufficient unless the behavior producing it is identified