Skip to yearly menu bar Skip to main content


Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

Ying Zhu ⋅ Xingwei Gan

Abstract

Chat is not available.