Gap-Driven Post-Training: How SFT, Distillation, and RL Close Different Gaps in Web UI Generation
Zhunxuan Wang ⋅ Wei Liu ⋅ Abhishek Tripathi ⋅ Murat Sensoy
Abstract
Post-training pipelines combine supervised fine-tuning (SFT), distillation, and reinforcement learning (RL), yet benchmark gains alone do not reveal what each stage changes, which failures remain, or when an intermediate stage helps the next one. We study these transitions in web UI generation, where outputs must render correctly, look designed, and behave interactively. We propose \emph{gap-driven post-training}: evaluation failures are distilled into a root-cause taxonomy that guides synthetic data and reward design across three checkpoints. SFT learns from teacher-authored demonstrations; Rejection Fine-Tuning with Distillation (RFT-D) minimally repairs the model's own rollouts; and RL optimises verifier-scored artifacts. On OpenDesign, SFT supplies $21.64$ of the final $29.71$-point static gain while moving the aggregate interactive score by only $+0.11$. RFT-D adds $+1.37$ static directly, but a matched branch that applies the same RL recipe to SFT instead finishes $5.24$ / $5.67$ / $4.96$ static points lower across three benchmarks, consistent with RFT-D improving the checkpoint from which RL begins; it does not measurably raise the interactivity ceiling. RL contributes the largest interactive gain on every benchmark. The final Qwen3-8B checkpoint reaches the level of Claude Sonnet 4.5 on OpenDesign ($81.9$ vs. $80.9$ static; $1.81$ vs. $1.70$ interactive), WebDev Arena ($84.7$ vs. $84.7$; $1.75$ vs. $1.75$), and Design Arena ($82.5$ vs. $81.9$; $1.74$ vs. $1.50$), with margins interpreted against the reported uncertainty. The result is a stage-resolved account of post-training: the procedures are complementary, and failure analysis identifies which learning signal is useful next.
Chat is not available.
Successful Page Load