Strengthening LLMs for Tabular Prediction with Structural Priors
Abstract
Tabular prediction has long been dominated by gradient-boosted decision trees and specialized deep tabular models, while large language models (LLMs) remain difficult to make competitive despite their cross-task adaptability and transparent reasoning traces. Existing LLM-for-tabular methods mainly adapt tables into textual inputs, but often fail to capture key structural properties of tabular data. In this work, we propose Permutation Relative Policy Optimization (PRPO), a reinforcement learning post-training method that strengthens LLMs for tabular prediction by injecting column-permutation invariance as a structural prior. By constructing label-preserving column permutations and estimating advantages both within and across them, PRPO converts sparse outcome rewards into denser and more stable optimization signals. Extensive experiments on 139 OpenML datasets show that our 8B model reaches a genuinely competitive regime against strong specialized tabular baselines. It achieves strong fully supervised performance, dominates cross-dataset zero-shot settings, and performs on par with 32-shot strong baselines. Moreover, it substantially outperforms much larger general-purpose and reasoning LLMs, including up to a 53.17% improvement over DeepSeek-R1 (685B). These results show that structural-prior RL post-training is an effective route for making LLMs competitive in tabular prediction.