Hierarchical TTT: Efficient Test-Time Learning through Trainable Guidance
Abstract
Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Building on this feedback loop, test-time training (TTT) further adapts the model itself by updating its parameters on the target problem. Existing TTT methods typically adapt the same LLM that generates complete solutions. This becomes expensive when reliable execution requires a large model, since training must maintain gradients, optimizer states, and policy statistics while repeatedly generating long, structured outputs. It also complicates credit assignment: outcome-level verifier feedback must jointly evaluate the high-level strategy and its low-level implementation. In this work, we introduce \emph{Hierarchical TTT}, which separates these roles. A compact \emph{guidance model} is trained at test time to propose high-level strategic changes, while a frozen \emph{execution model} implements them as complete executable solutions. At each step, the system selects a promising previously discovered solution, proposes a change, executes and verifies it, and updates only the guidance model using an adaptive group-relative RL objective. This concentrates test-time learning on short strategic decisions while retaining the implementation capability of a substantially stronger model without adapting it. Without web access, Hierarchical TTT achieves the best reported web-free result on Polyomino, surpassing even a web-enabled multi-LLM baseline, and is the strongest evaluated automated-discovery result on TriMul, comparable to the top public GPUMode submission. Code is available at \url{https://anonymous.4open.science/r/hierarchical-ttt}.