An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过四种奖励设计解决LLM遗忘过程中目标相关知识抑制与非目标功能保留之间的平衡问题,探讨了优化成功与行为遗忘之间的差异。
📝 Abstract
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.
Problem

Research questions and friction points this paper is trying to address.

LLM unlearning
reward specification
behavioral unlearning
benchmark reliability
GRPO
Innovation

Methods, ideas, or system contributions that make the work stand out.

reward specification
GRPO-based LLM unlearning
benchmark reliability
R
Rubén Balbastre
University of Valencia
J
Juan Manuel Orduña
University of Valencia
M
Mariano Pérez
University of Valencia