Skip to yearly menu bar Skip to main content


An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-Based LLM Unlearning

Rubén Balbastre ⋅ Juan M Orduña ⋅ Mariano P Martínez

Abstract

Chat is not available.