Grounding as a Reward: Reducing Confabulation in Natural Language Autoencoders
Seong Pyo Hong ⋅ Taejoon Eo ⋅ Hye S Lee ⋅ Daemyung Youn
Abstract
Natural language autoencoders (NLAs) interpret the activations of large language models (LLMs) as natural language explanations. However, these interpretations may contain confabulations involving specific details, such as names, numbers, and quotations. We propose a grounding reward for NLA training and show that it improves grounding while preserving the fraction of variance explained (FVE). We also compare sequential training, consisting of reconstruction training followed by joint reconstruction and grounding training, with joint training. Among the evaluated checkpoints, sequential training achieves higher test FVE and exact match grounding than joint training after the same number of grounding reward updates.
Chat is not available.
Successful Page Load