Session
Atlanta Poster Session 4
Hall C1
\$OneMillion-Bench: How Far are Language Agents from Human Experts?
Yang Liu ⋅ Jiaqi Li ⋅ Jun Bai ⋅ Qianyu Yang ⋅ Zixia Jia ⋅ Tiliang Duan ⋅ Chun Zhang ⋅ Jiayun Dong ⋅ Lingyue Yin ⋅ Jianpeng Jiao ⋅ Yanglihong Xiao ⋅ Zaiyuan Wang ⋅ Tao Peng ⋅ Xiaobo Hu ⋅ Kaiyuan Chen ⋅ Ge Zhang ⋅ Gang Yao ⋅ Hao Chen ⋅ Zilong Zheng
As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce **\$OneMillion-Bench** **(\$1M-Bench)**, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focusing on expert-level problems to ensure meaningful differentiation across agents. Together, \$1M-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.
3D Point Splatting for mmWave Radar Novel View Synthesis
Adnan Armouti ⋅ Yixuan Gao ⋅ Rajalakshmi Nandakumar
Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differentiable Monte Carlo (MC) ray tracers implement the radar forward model directly with explicit material modeling and complex outputs, but do not scale to the multi-view optimization NVS demands. Optical-NVS ports of NeRF, hash grids, and 3D Gaussians train fast but discard phase and replace explicit material modeling with opaque learned features, restricting them to power-only range-azimuth (RA) magnitudes. We propose 3D Point Splatting (3DPS), the first differentiable point renderer for radar, derived directly from the standard solid-angle form of the radar equation. Each oriented 3D point carries an ITU-R P.2040 material model, evaluated in closed form, with the resulting complex phasor splatted into range bins through a precomputed point spread function (PSF). The complex-valued output makes the renderer product-agnostic. The same optimized scene yields analog-to-digital converter (ADC), complex range profile (CRP), and RA outputs through standard fast Fourier transform (FFT) pipelines without retraining for each format. On six outdoor ColoRadar scenes, 3DPS reaches 0.587 mean Pearson correlation on held-out RA images. This is between $1.7\times$ and $5.2\times$ the three optical-NVS baselines (RadarSplat, Radar Fields, DART). Training takes approximately 3 minutes per scene on a single RTX 4090.
Acting without Knowing: Planning, Prediction, and Transfer Dissociate in Interactive Visual Physics
Alison Gopnik ⋅ Eunice Yiu ⋅ Kevin Murphy ⋅ Kelsey Allen ⋅ Shiry Ginosar
Evaluations of large multimodal models (LMMs) often conflate behavioral task success with true physical understanding. However, classical accounts of ``world models'' dictate that physical understanding requires the unified operation of three core capacities: planning, prediction, and positive transfer. Because current benchmarks typically test these three abilities in isolation, they often obscure whether models actually possess a cohesive internal world representation. In this paper, we propose a joint evaluation protocol to test whether LMMs acquire a more complete and consistent understanding of the physical world. Using a suite of virtual physics-based puzzles, models are tasked with placing a tool in a scene to alter its dynamics, iteratively revising their placements based solely on visual feedback. By testing the same models on the exact same tasks, we uncover a stark dissociation between acting and knowing. While LMMs can sometimes solve diverse physical goals under an interactive action-feedback protocol, their success is not accompanied by robust feedback-guided replanning, accurate prediction, or human-like predictive transfer to new actions. Models show weak and inconsistent prediction of the consequences of their own proposed actions, and their predictive transfer to new actions is partial, inconsistent, and less selectively goal-directed than in humans. Ultimately, we demonstrate that the capacities a world model should wholly unify remain only loosely connected. We therefore argue that benchmarks for LMMs should not measure task success alone, but should jointly evaluate planning, prediction, and transfer.
Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces
Ji Gao ⋅ Caleb Ju ⋅ Guanghui Lan ⋅ Zhaohui Tong
Policy Dual Averaging (PDA) offers a principled Policy Mirror Descent (PMD) framework that more naturally admits value function approximation than standard PMD, enabling the use of approximate advantage (or Q-) functions while retaining strong convergence guarantees. However, applying PDA in continuous state and action spaces remains computationally challenging, since action selection involves solving an optimization sub-problem at each decision step. In this paper, we propose actor-accelerated PDA, which uses a learned policy network to approximate the solution of the optimization sub-problems, significantly reducing runtime. We provide a theoretical analysis that quantifies how actor approximation error impacts the convergence of PDA under suitable assumptions. We then evaluate its performance on several benchmarks in robotics, control, and operations research problems. Actor-accelerated PDA achieves superior performance compared to popular on-policy baselines such as Proximal Policy Optimization (PPO). Overall, our results take a significant step toward bridging the gap between the theoretical advantages of PDA and its practical deployment in continuous-action problems with function approximation.
A Curvature Phase Transition Governs Coherence Penalty Efficiency Against Feature Absorption in SAEs
Hak Hyun Kim ⋅ Yash Raj ⋅ Peter Chin ⋅ Soroush Vosoughi
Feature absorption is a structural failure mode of sparse autoencoders, where a single dictionary atom captures multiple distinct concepts rather than one, reducing interpretability. While prior work has shown that absorption configurations are spurious local minima of the sparse dictionary learning loss, little is known about the geometric mechanisms that determine whether a coherence penalty can escape them. We study the curvature geometry of pairwise coherence penalties at absorption configurations on the unit sphere. We show that (1) every coherence penalty in the power family $\mathcal{R}p = \sum{l 0$, yet the efficiency of this force depends critically on the penalty's curvature; (2) a phase transition at $p = 2$ governs this efficiency: the Riemannian Hessian in the splitting direction equals $p(p-2) \cdot k^{1-p/2}$ in closed form, so concave penalties ($p < 2$) receive curvature assistance and convert absorption into a strict saddle point, while convex penalties ($p > 2$) face curvature resistance and deepen the local minimum; (3) the boundary at $p = 2$ corresponds precisely to a Parseval invariance that renders the quadratic penalty geometrically neutral; and (4) under feature correlation $\mu^{}$, convex penalties reverse above $\mu^{}_{crit} = (p-2)/(p+k-2)$, giving a closed-form threshold for when resistance can be overcome. Experiments on Pythia-160M and Gemma-2-2B confirm the predicted $\lambda$-efficiency ordering for competitive-activation SAEs, with absorption reduction verified as feature-level improvement rather than feature suppression. More broadly, our results point toward a principled framework for selecting coherence penalties based on the curvature geometry of the loss landscape rather than empirical tuning.
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
Aritra Mazumder ⋅ Shubhashis Roy Dipta ⋅ Nusrat Jahan Lia ⋅ Tanzila Khan ⋅ Kainat R Hossain ⋅ Nehaa Shri ⋅ Humayra Tasnim ⋅ Gour G Shawon ⋅ Shubhrangshu Debsarkar ⋅ Sumaiya A Rani ⋅ Al Jami Islam Anik ⋅ Debjoty Mitra ⋅ Al Nafeu Khan
Multi-agent systems achieve state-of-the-art outcomes through peer collaboration. However, when an agent in the pipeline silently drops a constraint, the system's final output may look correct even though the reasoning chain was quietly corrupted, and existing outcome-based evaluations are blind to such multi-hop process failures. To make these vulnerabilities measurable before deployment, we introduce AgentCollabBench, a diagnostic benchmark of 900 human-validated tasks spanning software engineering, DevOps, and data engineering. Each task isolates one of four behavioral risks: instruction decay (does a constraint survive peer pressure?), false-belief contagion (does a falsehood spread through consensus?), context leakage (does information bleed between tasks?), and tracer durability (does marked data reach the final agent?). Evaluating four modern LLMs (GPT 4.1 mini, Gemini 2.5 Flash Lite, Qwen-3.5-35B-A3B, and Llama 3.1 8B Instruct), we expose model-specific vulnerability profiles invisible to outcome-only evaluation; Qwen-3.5-35B-A3B, for example, leads on tracer durability and instruction stability, while GPT 4.1 mini leads on leakage containment and false-belief resistance. Beyond per-model differences, communication topology emerges as a primary risk factor that explains 7-40% of the variance in multi-hop information survival. The effect traces to a synthesis bottleneck specific to converging-DAG nodes: an agent weighing competing parent inputs discards constraints carried by a minority branch, a bottleneck structurally absent from linear chains. AgentCollabBench demonstrates that suboptimal topology can silently erase the safeguards of highly capable models, arguing that multi-agent reliability is fundamentally a structural problem and that scaling model intelligence alone is no substitute for architecture.
AI Governance Should Prioritize Control and Knowledge Boundaries Over Limiting Intelligence
Hendrik Luuk
The rapid proliferation of autonomously acting AI agents makes the governance of agentic systems urgent. We propose a competence circuit, summarized as IKC + KC + C, that decomposes an agent's real-world impact into three pathways: intelligence-mediated discovery, procedural know-how, and direct control. Using the canonical fast-takeoff debate as a stress test, we map catastrophic takeoff to a conjunction of six necessary conditions, showing that most rely on rapid knowledge acquisition or control attainment rather than intelligence alone. We operationalize epistemic intelligence as hypothesis-selection efficiency and validate this framework via stochastic simulation on NK fitness landscapes across six scenarios spanning the conjunctive chain. An analytic bound supported by the simulations shows that intelligence-driven speedup is capped by the ratio of competing intelligence levels regardless of task difficulty. The simulations also expose a fast-takeoff dilemma: a large intelligence-driven speedup multiplier materializes only when the hypothesis space is noisy, which is precisely the regime where each test incurs irreducible time cost, so a large multiplier and a short absolute timescale cannot both obtain. Simulations also indicate that institutional chokepoints disproportionately impact superintelligent agents, leaving them less unproductive search time to absorb mandatory authorization delays. We argue that the dominant safety levers are restricting control and enforcing knowledge boundaries rather than limiting intelligence alone, and this applies to agentic systems at any capability level. Purely digital domains weaken these bottlenecks, strengthening the case for architectures that intentionally "put physics in the loop."
A Latent-Load Framework for Reliability Analysis and Intervention Design in LLM Pipelines
Yuanjie Shi ⋅ Yan Yan
Multi-step LLM systems increasingly combine generation, retrieval, tool use, code execution, and verification into pipelines whose components consume and transform an evolving semantic state; end-to-end reliability is therefore governed by pipeline dynamics rather than by component-level accuracy alone. In such pipelines, an early ambiguity, hallucination, or reasoning error may be amplified, masked, partially corrected, or reintroduced by later components. Existing evaluations and failure taxonomies identify where LLM pipelines fail, but they do not provide a predictive theory of how semantic corruption propagates across dependent steps or how limited interventions should be allocated. To fill this gap, we introduce a latent-load framework for reliability analysis and intervention design in black-box LLM pipelines. Specifically, each intermediate state carries an unobserved semantic error load that evolves through inherited-error amplification, context dependence, and fresh error injection. The resulting recurrence decomposes final error, identifies stable, critical, and unstable regimes, and quantifies system-level risk under parameter uncertainty. We prove that the checkpoint-placement objective is monotone submodular, yielding near-optimal budgeted interventions, and provide a proxy-based partial-identification procedure for black-box estimation. Experiments show that proxy-scale propagation estimates predict relative downstream risk and select checkpoints that reduce accumulated proxy load and directionally lower failure rates under fixed budgets.
AlgoPilot: Cross-Paradigm Reasoning in Language Models via Strategy Selection and Guidance
Yilun Hao ⋅ Ruixiao Yang ⋅ Hudson Hilal ⋅ Yongchao Chen ⋅ Kaizhi Qian ⋅ Yada Zhu ⋅ Yang Zhang ⋅ Chuchu Fan
Real-world problems span diverse domains, including program synthesis, symbolic reasoning, planning, and optimization, and often require fundamentally different solution paradigms. A central challenge for LLM-based reasoning is twofold: identifying an appropriate problem-solving approach and formulating it correctly. In this paper, we propose AlgoPilot, a cross-paradigm reasoning framework that enables adaptive selection of the solving approaches and structured problem formulation prior to execution. The AlgoPilot framework contains three components: 1) a SteerLM, trained via supervised fine-tuning (SFT) and reinforcement learning (RL), that adaptively selects the solving strategy (paradigm and modes) based on input problems, 2) the corresponding expert GuideLM, trained via SFT, that generates structured formulation guidance tailored to the selected strategy, and 3) an ExecutionLM that generates solutions conditioned on both the selected strategy and its formulation. Unlike prompting-based methods that rely on fixed reasoning templates, our approach learns to adaptively choose and structure problem-solving strategies based on the input. We build AlgoPilot on Qwen3-8B and evaluate across 5 problem categories and 13 benchmarks with 66 unique domains, comparing against both the latest similarly scaled small models and large LLMs. AlgoPilot significantly improves average accuracy for the same model from 38.4\% to 72.9\%, outperforming the best large LLM prompting baseline GLM-5 by 11.7\%. Ablation studies confirm the effectiveness of each component. Notably, we show that the learned steering and formulation modules can transfer across execution models, suggesting their potential to improve other LLMs.
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
Byeongju Woo ⋅ Zilin Wang ⋅ Byeonghyun Pak ⋅ Sangwoo Mo ⋅ Stella X. Yu
Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidence. We propose a hierarchical vision-language learning principle for understanding scenes as part-to-whole compositions: before forming a whole-scene representation, a model should uncover what semantic parts appear where in the image. To this end, we propose CAFT (Cross-domain Alignment of Forests and Trees), a vision-language model that jointly learns local text-region alignment at intermediate representations and global image-text alignment at the final representation. Exploiting the organization of long captions, where local descriptions often correspond to scene parts, CAFT employs a fine-to-coarse image encoder and a part-whole text encoder to discover localized part semantics and progressively compose them into a global image-text representation. Trained on 30M image-text pairs, CAFT achieves state-of-the-art performance on six long-text retrieval benchmarks and exhibits strong scaling behavior. Experiments show that CAFT learns fine-grained representations that localize textual semantics in image regions without explicit region-level supervision.
All-in-one Adverse Weather Removal via Prior-modulated and Velocity-constrained Rectified Flow
Wei Dong ⋅ Han Zhou ⋅ Terry Ji ⋅ Guanhua Zhao ⋅ Shahab Asoodeh ⋅ Yulun Zhang ⋅ Guangtao Zhai ⋅ Jun Chen ⋅ Xiaohong Liu
Adverse weather removal (AWR) in real-world images remains challenging due to heterogeneous and unseen degradations, while distortion-driven training often yields overly smooth results. We propose PVRF, a unified framework that integrates zero-shot soft weather Perceptions with Velocity-constrained Rectified-Flow refinement. PVRF introduces an AWR-specific question answering module (AWR-QA) that uses frozen vision--language models (VLMs) to estimate soft probabilities of weather types and low-level attribute scores. These perceptions condition restoration networks via attribute-modulated normalization (AMN) and weather-weighted adapters (WWA), producing an anchor estimate for refinement. We then learn a terminal-consistent residual rectified flow with perception-adaptive source perturbation and a terminal-consistent velocity parameterization to stabilize learning near the terminal regime. Extensive experiments show that PVRF improves both fidelity and perceptual quality over state-of-the-art baselines, with strong cross-dataset generalization on single and combined degradations.
A low-rank decoder bottleneck bounds reliability in foundation-model perturbation prediction
Zheyu Zhang ⋅ Yuanhao Huang ⋅ Fan Feng ⋅ Jie Liu
Foundation-model-based predictors of perturbation response are evaluated primarily by aggregate accuracy, yet lack principled per-prediction reliability criteria. We audit the state-transition (ST) decoder~\citep{adduri2025predicting} shared by recent predictors across three encoders and four cell lines under CRISPRi perturbation, and identify a rank-$1$ attention bottleneck whose leading direction maps in gene space to the same cycling/p53 stress-response program in every cell line tested. The bottleneck implies a distance-defined reliability boundary: per-prediction quality degrades smoothly with the gene-space distance from a model's predicted perturbation response ($\Delta$, the predicted change relative to control) to its nearest training $\Delta$, and biological-family identity, not training-pool size, drives held-out quality. We propose \emph{predicted-NN-distance}, this same distance computed from model output alone, as a per-prediction trust signal that tracks per-perturbation quality at Spearman $|\rho| \approx 0.81$ (oracle $0.88$) and admits a single pooled threshold that transfers across cells and encoders under leave-one-out evaluation. The signal outperforms a simpler predicted-magnitude baseline in most cases. Per-prediction reliability is bounded by a low-rank decoder geometry and by gene-space coverage of training perturbations, not by aggregate benchmark accuracy.
AmbientFM: A Foundation Model for Ambient Sensing
Guozhen Zhu ⋅ Yuqian Hu ⋅ Sakila S Jayaweera ⋅ Wei-Hsiang Wang ⋅ Weihang Gao ⋅ Jiaxuan Zhang ⋅ Beibei Wang ⋅ Chenshu Wu ⋅ K. Liu
Ambient intelligence seeks to continuously understand human presence, activity, and physiology in physical spaces, enabling smart environments, health monitoring, and human-computer interaction. WiFi infrastructure offers a ubiquitous, always-on, and privacy-preserving sensing substrate across billions of IoT devices. Yet ambient sensing with WiFi remains largely fragmented, with most systems relying on task-specific models, labeled data collection, and customized training pipelines. We present AmbientFM, the first foundation model for ambient sensing through ubiquitous WiFi signals. AmbientFM is pre-trained on 9.2 million unlabeled Channel State Information (CSI) samples collected over 439 days from 20 commercial device types deployed in real-world environments. It learns transferable wireless representations through contrastive learning, masked reconstruction, and physics-informed objectives tailored to wireless signals. With lightweight adaptation, a single backbone transfers to 9 downstream tasks, achieving above 0.90 AUROC on all classification benchmarks and supporting dense spatial reconstruction. AmbientFM reduces the need for labeled data and task-specific design, enabling scalable ambient intelligence on existing wireless infrastructure.
AnaDiffusion: Anatomically Compositional Latent Diffusion for Controllable 3D Brain MRI Generation
Tracy Han ⋅ Lulin Liu ⋅ Bangya Liu ⋅ Yuanhao Cai ⋅ Nuo Chen ⋅ Jade Wang ⋅ Ziqian Xie ⋅ Chenyu You ⋅ Shuiwang Ji ⋅ Degui Zhi ⋅ Zhiwen Fan
3D brain MRI generation has made significant advances for medical imaging, simulation, and controllable anatomical analysis. However, existing generative models typically synthesize 3D volumes monolithically, often overlooking regional anatomical structures and limiting local controllability. To address these limitations, we introduce AnaDiffusion, an anatomically compositional latent diffusion framework that factorizes the generation process into distinct, anatomically meaningful regions followed by part-to-whole assembly and global refinement. Our approach trains part diffusion models first to capture local structural priors. Then, we inject assembled anatomical composite of parts into whole-brain latent and continue denoising. This mechanism enables the model to resolve global context while preserving the injected anatomy. As a result, AnaDiffusion outputs both explicit part assets and a globally coherent volume, enabling controllable part editing without requiring additional dense segmentation masks at inference time while maintaining coherent part-to-whole brain structure. On ADNI, AnaDiffusion achieves the lowest FID across the whole brain, left/right hemispheres, cerebellar-brainstem complex, and seam regions, while also reducing Cohen's d for ventricles, cerebellum, and brainstem. In localized editing experiments, paired MS-SSIM shows high target transfer and off-target preservation, supporting controllable part replacement with limited non-target anatomical drift.
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
Chen Si ⋅ Yulin Liu ⋅ Bo Ai ⋅ Jianwen Xie ⋅ Rolandos Alexandros Potamias ⋅ Chuanxia Zheng ⋅ Hao Su
We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation. While recent works with foundation approaches have shown that scaling training data markedly improves hand pose estimation, existing real-world datasets are limited in coverage, and prior synthetic datasets rarely provide occlusions, arm details, and aligned depth together at scale. To address this bottleneck, our proposed AnyHand contains 2.5M single-hand and 4.1M hand-object interaction RGB-D images, with rich geometric annotations. We show that extending the original training data recipes of existing RGB baselines with AnyHand yields significant gains on multiple benchmarks (FreiHAND and HO-3D), even when keeping the architectures and training schemes fixed. Together with extensive ablations on the scale and composition of the training data setups, these results suggest that training data diversity and quality are as critical as scale for advancing hand pose estimation. We further examine the utility of AnyHand's aligned depth maps in the appendix, showing that scaling RGB-D supervision with AnyHand allows a lightweight depth-fusion variant of existing RGB baselines to outperform prior RGB-D methods.
Approximate Matrix–Vectors Under a Bounded $\ell_1$ Assumption and Applications to Kernel Matrices
Rikhav Shah ⋅ Sandeep Silwal ⋅ Tony C Wang
Matrix-vector products (MVPs) are a key primitive in numerical linear algebra and machine learning. However, the naive quadratic running time for exact computation is prohibitive for large matrices, motivating approximate methods. In this paper, we study approximate MVPs under a natural ``lightness'' assumption, which bounds the total $\ell_1$-mass of a $n \times n$ input matrix $A$ by $\gamma n$. For general matrices, we give an algorithm for approximating $Ax$ up to additive $\epsilon \|x\|_2$ error in $O(\gamma n^{1.5}/\epsilon)$ query time and polynomial preprocessing. We extend this algorithm to kernel matrices by establishing a black-box reduction to Kernel Density Estimation data structures and analyzing a noisy entry sampling scheme. For kernel matrices, our algorithm does not require any preprocessing and improves the best-known running times for Gaussian kernel matrices of [Indyk, Kapralov, Sheth, Wagner; ICLR `25] by polynomial factors in $n$ and $1/\epsilon$, and also provides the first algorithms leveraging the lightness assumption for other kernel functions as well.
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
Viet Nguyen ⋅ Thinh Cao ⋅ Tuan M Pham ⋅ Tan Dinh ⋅ Huy Nguyen ⋅ Nhat Ho ⋅ Alessandro Rinaldo
Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to improve performance, a gated attention model that leverages a gating mechanism within the multi-head self-attention has recently been proposed as a promising alternative. Gated attention has been empirically demonstrated to increase the expressiveness of low-rank mapping in standard attention and even to eliminate the attention sink phenomenon. Despite its efficacy, a clear theoretical understanding of gated attention's benefits remains lacking in the literature. To close this gap, we rigorously show that each entry in a gated attention matrix or a multi-head self-attention matrix can be written as a hierarchical mixture of experts. By recasting learning as an expert estimation problem, we demonstrate that gated attention is more sample-efficient than multi-head self-attention. In particular, while the former needs only a polynomial number of data points to estimate an expert, the latter requires exponentially many data points to achieve the same estimation error. Furthermore, our analysis also provides a theoretical justification for why gated attention yields higher performance when a gate is placed at the output of the scaled dot product attention or the value map rather than at other positions in the multi-head self-attention architecture.
A Steerable Deep Network for Model-Free Diffusion MRI Registration
Gianfranco Cortés ⋅ Xiaoda Qu ⋅ Baba C Vemuri
Nonrigid registration is vital to medical image analysis but remains challenging for diffusion MRI (dMRI) due to its high-dimensional, spatio-angular dependence. We present a novel, geometric deep learning framework for {\it model-free}, nonrigid registration of raw dMRI data. The dMRI registration problem is formulated in the native spatio-angular acquisition space, which exhibits a natural symmetry to the group of 3D roto-translations, denoted by $\mathrm{SE}(3)$. A by-product of this design choice is freedom from having to augment the data with roto-translated versions of itself. Our second novelty is the loss function formulation, based on the maximum mean discrepancy (MMD) loss used to compare two probability density functions. We apply this loss in the Fourier space, where it becomes the well-known weighted sum-of-squared differences (SSD) loss with the weights being the Fourier transform of a reproducing kernel Hilbert space (RKHS) kernel. Experimental results on HCP and OASIS-3 clinical-grade dMRI data demonstrate competitive performance compared to SOTA approaches, with the added advantage of bypassing the overhead for estimating derived representations. This work establishes a foundation for data-driven, geometry-aware dMRI registration directly in the acquisition space.
We introduce the problem of *adversary-directed* online learning, in which the learner is aware of the set of instances in advance, and an adversary adaptively determines their ordering during the learning process. Surprisingly, this formulation has not been previously studied, perhaps due to its conceptual resemblance to the traditional adversarial online learning problem, despite related models, including the transductive, self-directed, and best-order, having been extensively explored. However, by utilizing novel techniques, we demonstrate that the landscape of adversary-directed online learning significantly diverges from that of traditional adversarial online learning. In the realizable setting, we establish a trichotomy of possible rates of the minimax number of mistakes. Specifically, for a learning horizon $\operatorname{T}$, the minimax number of mistakes can only be of the orders $\Theta(\operatorname{T})$, $\Theta(\log \operatorname{T})$, or $\Theta(1)$. To prove this, we introduce a new combinatorial complexity parameter, termed the perfect Littlestone dimension, whose finiteness distinguishes the $\Theta(\log \operatorname{T})$ rate from the $\Theta(\operatorname{T})$ rate. On the other hand, in the agnostic setting, we essentially show a dichotomy of possible rates of the minimax expected regret. In particular, if the learner plays for $\operatorname{T} \in \mathbb{N}$ rounds, its minimax expected regret can only be of the orders $\Theta(\operatorname{T})$, or $\widetilde{\Theta}(\sqrt{\operatorname{T}})$, which is also characterized by the finiteness of the perfect Littlestone dimension. Technically, a key ingredient in the proof of our $\mathcal{O}(\log \operatorname{T})$ and $\widetilde{\mathcal{O}}(\sqrt{\operatorname{T}})$ upper bounds is a novel online learning algorithm that leverages a new notion of shattering based on the perfect Littlestone dimension, which exploits the adaptive adversarial nature of the problem.
NeurIPS now publishes a broad mix of algorithmic, theoretical, empirical, dataset, benchmark, systems, and evaluation-infrastructure work. Yet we lack systematic evidence on how the composition of this mix has changed over time. We audit titles, abstracts, and metadata for 19,361 accepted NeurIPS papers from 2015--2024. We define semantic axes for formal-theoretical evidence, benchmark/resource evidence, empirical-performance evidence, artifact-release orientation, scale/capability framing, claim strength, and hedging; score papers with contrastive embedding measures; and validate the constructs with blinded LLM-agent ratings, lexical probes, encoder triangulation, length checks, and topic controls. In raw decade trends, accepted papers shift toward benchmark/resource evidence (+0.842 SD), empirical-performance evidence (+0.750 SD), and scale/capability framing (+0.606 SD), and away from formal-theoretical framing (-0.872 SD). Topic composition explains much of this movement, but embedding-cluster residuals remain +0.246, +0.264, +0.247, and -0.300 SD, respectively. Independent lexical LDA controls preserve the same signs at smaller magnitudes. Artifact-release framing is a useful negative result: it rises modestly in the raw corpus (+0.226 SD) but is near zero or negative after topic adjustment. We release a reusable audit package that includes metadata, semantic and lexical features, validation artifacts, robustness scripts, and an interactive Evidence Norms Atlas for inspecting aggregate evidence-framing patterns by year, topic neighborhood, and track.
A Unified Spectral Theory of Multimodal Losses
Yu-Ang Cheng ⋅ Sixuan Chen ⋅ Zhouyang Lu ⋅ Xizheng Yu ⋅ Grégoire Dhimoïla ⋅ Thomas Serre
Modern vision--language models are trained with very different objectives: CLIP-style contrastive alignment, SigLIP-style match/no-match prediction, caption-style token generation (Flamingo, LLaVA), and latent-space prediction (VL-JEPA). Each has been claimed as best-in-class for its target task. Yet practitioners lack clear guidelines for choosing between them, and existing identifiability theory analyzes only the contrastive case. In this work, we unveil the core statistical mechanisms that distinguish each objective family using a unified spectral analysis. By leveraging closed-form solutions for the linearized losses under a latent structure in which modality-specific nuisances may spuriously carry shared semantic content, we characterize what each objective recovers from the data. We demonstrate that contrastive alignment, match/no-match prediction, and latent-space prediction collapse to the same cross-modal solution, while directed token-space generation recovers a different one. Our findings indicate that in scenarios where one modality is heavily contaminated by spurious nuisance, token-space generation is preferable because it imposes a strictly weaker, one-sided alignment condition. In scenarios where two modalities are both clean, retrieval is maximized by the latent-prediction objective whose target matches the query modality. These results clarify the trade-offs between the four objective families. We validate our theoretical predictions on numerical, controlled synthetic, and real (Flickr30k, COCO, CC3M) data.
A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
Fred Xu ⋅ Thomas Markovich ⋅ florence regol ⋅ Yizhou Sun
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods typically address these requirements with separate models and objectives. A key obstacle is the lack of a single learned representation whose structure can serve all three tasks through different readouts. Inspired by spectral approaches in graph signal processing and polynomial chaos theory, we model uncertainty-bearing node embeddings as random graph signals: graph Fourier filters capture structural variation, while a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. This yields a \emph{doubly-spectral stochastic} (DSS) expansion in which the mean coefficient carries evidence for classification and energy-based OOD scoring, higher-order coefficients encode structured logit variation, and quadrature averaging turns that variation into a calibration-sensitive predictive distribution. Calibration thus uses the quadrature-averaged predictive distribution, OOD detection uses the mean-logit energy score, and distribution-shift robustness uses the corresponding DSS branch as a regularized spectral residual when paired with a deterministic encoder. Under mild conditions this representation universally approximates Gaussian-latent random graph signals with exponentially decaying truncation error. The resulting model, \emph{DSS-GNN} (Doubly-Spectral Stochastic GNN), can be used standalone or as a residual branch alongside a deterministic encoder (\emph{DSS-Hybrid}). In standalone form, DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; in hybrid form, DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD performance, and the strongest shifted accuracy among compared baselines on all 7 GOOD concept-shift benchmarks under standard ERM training.
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model’s general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model’s benign performance as well as the adapter’s abilities learned during downstream task adaptation.
Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning
Jiahao Zhang ⋅ Lujing Zhang ⋅ Keltin Grimes ⋅ Zhuohao Yu ⋅ Gokul Swamy ⋅ Steven Wu
A recurring challenge in preference fine-tuning (PFT) is handling *intransitive* (i.e., cyclic) preferences. Intransitive preferences often stem from either *(i)* inconsistent rankings along a single objective or *(ii)* scalarizing multiple objectives into a single metric. Regardless of their source, the downstream implication of intransitive preferences is the same: there is no well-defined optimal policy, breaking a core assumption of the standard PFT pipeline. In response, we propose a novel, game-theoretic solution concept, the *Maximum Entropy Blackwell Winner* (*MaxEntBW*), that is well-defined under multi-objective intransitive preferences. To enable computing MaxEntBWs at scale, we derive $\texttt{PROSPER}$: a provably efficient PFT algorithm. Unlike prior self-play techniques, $\texttt{PROSPER}$ directly handles multiple objectives without requiring scalarization. We then apply $\texttt{PROSPER}$ to the problem of fine-tuning large language models (LLMs) from multi-objective LLM-as-a-Judge feedback (e.g., rubric-based judges), a setting where both sources of intransitivity arise. We find that $\texttt{PROSPER}$ outperforms all baselines considered across both instruction following and general chat benchmarks.
Distribution regression, where the goal is to predict a scalar response from a distribution-valued predictor, arises naturally in settings where observations are grouped and outcomes depend on group-level characteristics rather than on individual measurements. We introduce DistBART, a Bayesian nonparametric approach to distribution regression that models the regression function as a linear functional with the Riesz representer assigned a Bayesian additive regression trees (BART) prior. We argue that shallow decision tree ensembles encode reasonable inductive biases for tabular data, making them appropriate in settings where the functional depends primarily on low-dimensional marginals of the distributions. We show this both empirically on synthetic and real data and theoretically through an adaptive posterior concentration result. We also establish connections to kernel methods, and use this connection to motivate variants of DistBART that can learn nonlinear functionals. To enable scalability to large datasets, we develop a random-feature approximation that samples trees from the BART prior and reduces inference to sparse Bayesian linear regression, achieving computational efficiency while retaining uncertainty quantification.
Benchmarks as Measurement Instruments: Quantifying Signal and Noise for More Efficient AI Evaluations Under Distribution Shift
Michael Hardy ⋅ Anka Reuel-Lamparth ⋅ Jodi Casabianca ⋅ Hansol Lee ⋅ Benjamin Domingue ⋅ Sanmi Koyejo ⋅ Mykel J Kochenderfer
AI progress is increasingly tracked through evaluations, such as AI benchmarks, yet small perturbations in evaluation pipelines can produce unstable scores and model rankings. We recast benchmark reliability as a measurement problem and formalize it as a signal-to-noise ratio derived from a crossed random-effects decomposition. This framework yields a direct mapping between variance components and expected rank stability, linking theoretical reliability to observable leaderboard concordance. Modern AI benchmarks operate in a high-dimensional regime with many items and relatively few evaluated models, where classical item-level reliability measures are ill-posed. We define a lower bound for benchmark item-level reliability classically, and we further introduce a tractable proxy, $\lambda^\bigstar_6$, that preserves the variance structure underlying reliability without modifying the estimand through sparsity or dimensionality reduction. Across diverse benchmarks, we show that selecting items via $\lambda^\bigstar_6$ produces smaller subsets that achieve higher ranking reliability than random subsampling at fixed evaluation budgets. Finally, we analyze reliability under positive distribution shift--as often observed in the current AI ecosystem--by partitioning models into lower- and higher-performing cohorts. We show that reliability is population-dependent and frequently declines as between-model variance shifts among stronger systems. Our results establish a principled framework for diagnosing and improving benchmark reliability and demonstrate that ranking stability is neither intrinsic nor static, but a measurable and optimizable property of the benchmark–population pairing.
Bernini: Latent Semantic Planning for Video Diffusion
Chenchen Liu ⋅ Junyi Chen ⋅ Lei Li ⋅ Lu Chi ⋅ Mingzhen Sun ⋅ Zhuoying Li ⋅ Yi Fu ⋅ Ruoyu Guo ⋅ Yiheng Wu ⋅ Ge Bai ⋅ Zehuan Yuan
Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong semantic grounding, while diffusion models synthesize images and videos with photorealistic fidelity. We argue that these two families can be unified through a simple division of labor: MLLMs perform semantic planning, while diffusion models render pixels from high-level semantic guidance and low-level visual features. Building on this idea, we propose \textbf{Bernini}, a unified framework for video generation and editing. An MLLM-based \textbf{planner} predicts the target semantic representation directly in the ViT embedding space, and a DiT-based \textbf{renderer} synthesizes pixels conditioned on this plan, augmented by text features and, for editing, source VAE features for detail preservation. Because semantics serve as the interface, the planner and renderer can be trained separately and only lightly co-trained, preserving the pretrained strengths of both components while keeping training efficient. To better handle multiple visual inputs, we introduce Segment-Aware 3D Rotary Positional Embedding (SA-3D RoPE), and further incorporate chain-of-thought reasoning in the planner to better transfer understanding into generation. Bernini achieves state-of-the-art performance across a wide range of video generation and editing benchmarks, with the MLLM's pretrained understanding translating into strong generalization on challenging editing tasks.
Beyond 3 Million Tokens: A Multi-Modal Foundation Model for Full-Resolution Heliophysics
Sujit Roy ⋅ Johannes Schmude ⋅ Ata A Asanjan ⋅ Thorsten Kurth ⋅ Rohit Lal ⋅ Kshitiz Mandal ⋅ Vishal Gaur ⋅ Harris Abdul Majid ⋅ Nikolaos Dionelis ⋅ Berkay Aydin ⋅ Himanshu Patil ⋅ Andres Munoz-Jaramillo ⋅ Campbell Watson ⋅ Juan Moreno ⋅ Manil Maskey ⋅ Rahul Ramachandran
Context length remains a fundamental bottleneck for vision foundation models operating on high-resolution imagery. Current architectures rarely exceed 1M tokens, forcing practitioners to downsample inputs at the cost of fine-grained spatial information. This limitation is particularly acute in heliophysics, where satellite instruments record full-disk solar observations at native $4096 \times 4096$ resolution across $13$ channels, and downsampling discards the small-scale magnetic structures and localized dynamics critical to understanding solar phenomena. In this work, we present the first vision foundation model for heliophysics trained at native 4K resolution on approximately $15$ years of multimodal solar data ($257$ TB) spanning eight extreme ultraviolet channels and five magnetic field and velocity products from the Solar Dynamics Observatory (SDO). We adapt MultiMAE with a dual-view formulation: 25% of tokens are observed directly, while the remaining 75% are replaced with fixed Gaussian random Fourier projections that act as structured, non-invertible frequency views of the masked content. Using two-way parallelism (feature + sequence) with FP8 mixed precision, we scale the context window to $>3$ million tokens at $8 \times 8$ patch size, a $3\times$ advance over prior work. We demonstrate strong scaling efficiency up to $1024$ GPUs across five hardware configurations, with an architecture capable of processing sequences exceeding $10$M tokens. The learned representations achieve strong zero-shot reconstruction across masking ratios up to 90% and full missing-modality scenarios. On downstream tasks (AR segmentation and EVE irradiance prediction) with a frozen encoder, we outperform SOTA by 9.5% and 35.3% respectively.
Beyond Copy-Paste: How Well Do Subject-Driven Video Models Understand Their Subjects?
Zun Wang ⋅ Kenan Deng ⋅ Daniel Blackburn ⋅ Linlin Lu ⋅ Junbang Liang ⋅ Yu Lou ⋅ Mohit Bansal ⋅ Shan Yang
Subject-driven video generation aims to produce videos that faithfully incorporate user-provided reference subjects. Current evaluation relies on free generation followed by reference-similarity scoring, which rewards a fundamental shortcut: models can obtain high scores by reproducing visible cues from the references while avoiding transformations that would test whether the subject is truly preserved beyond the conditioning input. We design anchored evaluation: controlled settings in which generation is anchored to held-out targets of the same subject rather than left unconstrained. Each probe withholds target content from the model's conditioning, so success requires using the references to infer content not directly visible, rather than simply copy-pasting what was given. We instantiate two probes. (1) Viewpoint control specifies target geometry via depth-based warping, leaving occluded regions that the model must complete using subject appearance from the references. (2) Degradation--recovery corrupts held-out videos showing the subject in real-world contexts that differ from the reference images, then recovers them conditioned on those references; recovery requires cross-context generalization of identity and inference of state dynamics. Both probes require only lightweight modifications to the inference loop of flow-matching models, need no retraining, and generalize across architectures. Evaluation across seven diverse, strong models reveals that rankings shift substantially: VACE-14B ranks first conventionally but Phantom-14B, ranked last, leads under anchored evaluation. Moreover, conventional scores strongly correlate with the improvement from providing ground-truth references (Pearson r = 0.91), suggesting they largely measure copying ability. Fine-grained analysis further reveals distinct failure modes: copy-paste models drift at unseen viewpoints, and no model adaptively balances copying and generalization.
Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation
Haozhan Tang ⋅ Xiuqi Zhu ⋅ Xinyin Zhang ⋅ Boxun Li ⋅ Virginia Smith ⋅ Kevin Kuo
Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides the representational plasticity required for high-entropy knowledge injection, Low-Rank Adaptation (LoRA) can match or surpass FFT performance because many tasks only require updates in a low-rank space and benefit from LoRA's additional regularization. Through empirical evaluation across diverse tasks (SQL, Medical QA, and Counterfactual Knowledge) and varying language models (Gemma-3-1B, Qwen2.5-1.5B, and Qwen2.5-3B), we verify both trends and demonstrate that relying solely on either static architecture is structurally limited. To address this challenge, we propose a Mixture of LoRA and Full (MoLF) Fine-Tuning, a unified framework that enables continuous navigation between both training regimes. MoLF dynamically routes updates between FFT and LoRA at the optimizer level to ensure that exact gradient signals are available to both experts throughout training, yielding stable training dynamics. For memory-constrained environments, we also introduce MoLF-Efficient, which freezes base weights and only routes updates among a pair of LoRA experts of potentially varying rank. Our evaluations show that MoLF either improves on or stays within $1.5$% of the better of FFT and LoRA across all settings, while MoLF-Efficient outperforms prior adaptive LoRA approaches by up to $20$% on Fact and $9$% on Med and SQL.
Beyond Marginal Coverage: Efficient Localized Conformal Prediction via Residual Rank Calibration
Xiangshi Li ⋅ Wenqing He
Conformal prediction provides finite-sample marginal coverage guarantees under exchangeability, but marginal validity alone does not ensure that prediction intervals adapt to local structure in the conditional distribution. Existing adaptive methods such as conformalized quantile regression (CQR) and conformal histogram regression (CHR) improve conditional coverage by incorporating estimated quantiles or conditional densities into the conformity score, but doing so requires direct estimation of the conditional distribution of response $Y$ given covariates $X$, which can be statistically and computationally demanding. We propose RLR-CHR, a residual-based local rank conformal histogram regression method that sidesteps full conditional distribution estimation by decoupling location and scale effects via a learned noise proxy and constructing histogram-based conformity scores on the standardized residuals. Local calibration is then performed through rank comparisons rather than weighted conformal quantiles, yielding a procedure that is both localized and computationally efficient. We establish finite-sample marginal coverage for RLR-CHR under exchangeability and asymptotic conditional coverage under local stability of the residual score distribution. We further prove that RLR-CHR is asymptotically equivalent to its weighted-localization counterpart RBC-CHR under suitable regularity conditions, providing a precise theoretical justification for the reduction in per-query computational cost without sacrificing asymptotic accuracy. Experiments on synthetic and real datasets demonstrate that RLR-CHR consistently improves worst-slab coverage and produces more compact prediction intervals than existing conformal approaches.
We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized with respect to a fixed reference policy through KL regularization. Prior work relies on pessimistic value estimation to handle distribution shift, yielding only $\widetilde{\mathcal{O}}(1/\sqrt n)$ statistical rates. We develop a new pessimism-free algorithm and analytical framework for KL-regularized games, built on the smoothness of KL-regularized best responses and a stability property of the Nash equilibrium induced by skew symmetry. This yields, to our knowledge, the first pessimism-free offline learning guarantee for KL-regularized games, with a fast $\widetilde{\mathcal{O}}(1/n)$ sample complexity bound. We further propose an efficient self-play policy optimization algorithm that replaces exact equilibrium computation with iterative KL-regularized policy updates, and prove that its last iterate preserves the same pessimism-free statistical guarantee up to a controlled optimization error.
Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings
Shubhang Bhatnagar ⋅ Dheeraj Baiju ⋅ Narendra Ahuja
Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched. A multimodal large language model (MLLM), shown the same pair, can articulate those attributes and use them to predict whether the images share a class. We propose \textbf{SAGA}, a framework that turns this language-grounded, attribute-aware perception into a training signal for the encoder itself. Specifically, we use Group Relative Policy Optimization (GRPO) to reward the MLLM for correct predictions on the vision encoder's tokens. Since correct predictions require those tokens to expose the specific attributes that differ or match between the pair, the gradient pushes the encoder to encode them, replacing the uniform pair-level scalar with attribute-resolved supervision. An auxiliary attention-distillation loss anchors the encoder's embedding to tokens the MLLM attended to, and a standard metric-learning loss shapes the embedding geometry for nearest-neighbour retrieval. The MLLM is frozen throughout and discarded at inference, matching the deployment cost of a metric-learning baseline. SAGA improves Recall@1 by 3 to 6 points over state-of-the-art baselines on CUB-200-2011, Cars-196, FGVC-Aircraft, and iNaturalist Aves on zero-shot image retrieval.
Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search
Jingdong Zhang ⋅ Yizhou Wang ⋅ Zhengzhong Tu ⋅ Xin Li ⋅ Wenping Wang ⋅ Xiaohang Zhan
Humanoid Visual Search (HVS) requires agents to actively explore immersive 360$^\circ$ environments. While prior methods treat this as a monolithic task relying on cumulative, multi-turn Chain-of-Thought (CoT) reasoning, they impose heavy cognitive burdens and require expensive trajectory-level annotations. In this paper, we propose Imagining in 360$^\circ$, a novel framework that decouples the exploration process into a specialized Imaginator and an Actor. The Imaginator functions as a probabilistic predictor of spatial priors; instead of maintaining a cumulative reasoning chain, it infers the semantic layout of both observed and unobserved regions in a single step. By sampling multiple hypotheses within this semantic space, we provide the Actor with a distribution of effective spatial information, offering robust guidance that hedges against uncertainty during active search. This decoupled architecture significantly lowers data engineering costs by eliminating the need for full-trajectory CoT annotations, enabling the generation of over 1.96 million curated training samples. Extensive experiments demonstrate that explicitly modeling semantic spatial priors drastically improves search efficiency and success rates in complex, in-the-wild environments.
Bridging Textual Profiles and Latent User Embeddings for Personalization
Zhaoxuan Tan ⋅ Xiang Zhai ⋅ Yan Zhu ⋅ Meng Jiang ⋅ Mohamed Hammad
Personalized systems rely on user representations to connect behavioral history with downstream recommendation applications. Existing methods typically employ either supervised latent user embeddings, which are effective for retrieval but difficult to interpret, or textual user profiles, which are interpretable but challenging to optimize for downstream utility due to lack of direct supervision. To bridge this gap, we present BLUE, a reinforcement learning framework that unifies these two forms of user representation by aligning language-based user profiles with embedding-based recommendation objectives. Given a user’s interaction history, BLUE leverages a profiler Large Language Model (LLM) to generate textual profiles, while an embedding model provides reward signals. This encourages the resulting textual representations to move closer to positive items and farther from negative ones in the embedding space. We further introduce a text-space supervision signal based on next-item prediction, ensuring the learned profiles remain both semantically meaningful and highly effective for downstream retrieval. Experiments on Amazon Reviews 2023 and Google Local Reviews in zero-shot sequential recommendation settings demonstrate that BLUE consistently outperforms strong baselines under both frozen and trainable embedding conditions. Notably, BLUE achieves clear gains in cross-domain transfer, highlighting the strong generalization ability of the learned user profiles. Furthermore, these generated profiles provide superior personalized context for question answering compared to raw user histories or alternative profile optimization methods. Overall, these results show that BLUE provides an effective way to unify interpretable textual profiling with discriminative latent embeddings for personalization.
Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion
Xinyu Wang ⋅ Jinbo Bi ⋅ minghu song
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions and delivers them through a trust budget set by the frozen sampler's own step norm: quotient geometry chooses the direction, while sampler motion sets the scale. We derive the horizontal lift, closed-form sampler-scaled update, KL/kinetic interpretation around a frozen reverse step, equivariance conditions, and a product-budget split for separately normalized section and residual controls. Controlled quotient tasks confirm that sampler-relative delivery activates signals that raw local quotient gradients leave dormant. On frozen TargetDiff backbones, official seed-0 CBGBench ligand-generation/editing sweeps show practical quality--runtime gains: CheapMain improves validity from 0.815 to 0.864 on fragment growing, 0.664 to 0.707 on scaffold hopping, and 0.681 to 0.712 on linker design, while PCMain improves fragment/scaffold and remains near-neutral on linker. Novelty remains 1.000 and diversity is preserved in the matched multi-seed molecular slice, giving task-dependent improvements without sampler retraining or backbone modification. Overall, QRG provides a lightweight route to quotient-aware inference for frozen molecular samplers with explicit runtime accounting.
BuresTomFlow: Bures-Geometric Flow Matching for Posterior Quantum Tomography
Lu Wei ⋅ Yufeng Wang ⋅ Chenfeng Cao ⋅ Haibin Ling
Finite-shot quantum tomography often leaves a posterior over many density matrices compatible with the same measurement record, but fast reconstruction methods collapse this uncertainty to a single state. We frame this setting as amortized posterior sampling and introduce BuresTomFlow, a measurement-conditioned Flow Matching sampler for full-rank density matrices. The method transports base density matrices along Bures-Wasserstein paths on the positive matrix cone, projects the path back to the trace-one state space, and trains the velocity field with a Bures-aligned tangent loss. Thus the sampler is trained in the same fidelity-based geometry used to judge posterior quality. We evaluate BuresTomFlow in dense four-qubit simulations with count, shadow, and mixed measurement records. Against Cholesky, log-density, and generic Riemannian Flow Matching controls at comparable learned-sampler runtime, BuresTomFlow improves Bures distance, credible-ball coverage, and held-out observable prediction. In the main comparison across multiple seeds, mean Bures distance decreases from $0.618$ for the Cholesky baseline to $0.367$ for BuresTomFlow, and mean coverage error decreases from $0.203$ to $0.100$. These results show that the value of Bures geometry is not merely enforcing physical states, but producing calibrated posterior uncertainty when finite-record tomography cannot be reduced to one reconstruction.
Cardinality-Decomposed Loss for Heterogeneous GNNs
Parul Maheshwari ⋅ Amulya Paruchuri ⋅ Alireza Sahami Shirazi
Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems. These graphs often express relations that vary in cardinality, for example, user-item preferences are one-to-many and user-attribute features are one-to-one. Traditionally, a unique loss function is applied for all of the network components which is often Bayesian Personalized Ranking (BPR). While BPR works well for the recommendation task, we find that it causes attribute embeddings to collapse to near-random geometry — a silent failure that leaves standard ranking metrics largely unaffected and therefore invisible to conventional evaluation. This in turn pollutes user node embeddings, which are shaped by both edge types simultaneously, hurting downstream tasks like personalization, segmentation, etc. Here we propose a Cardinality-Decomposed Loss (CDL) that combines both Cross Entropy (CE) and BPR to enable the model to collectively optimize for relations across cardinalities. As we implement this loss, we also confirm the conflict between CE and BPR by showing that the two losses compete against each other in the shared encoder's parameter space. We evaluate CDL on five datasets spanning two structural configurations — one-to-one attributes on user nodes (MovieLens-1M, Last.fm-360K, PayPal Audience Factory, BookCrossing) and on item nodes (Yelp) — and find that CDL consistently improves discriminability in attribute embeddings. We also show that whenever these attributes contain meaningful preference signal, we also see improvement in the ranking task (measured by NDCG). On the other hand, when attributes are weakly correlated with preferences, there is an inherent tension between the two objectives. We use a lambda parameter to navigate this trade-off, and a lambda-sweep reveals that dataset behavior is governed by two graph properties — semantic alignment and topology leakage. Semantic alignment captures whether the one-to-one attribute is predictive of user preferences, while topology leakage captures whether message passing already encodes attribute structure implicitly through the graph's connectivity.
Causal Effects with Unobserved Unit Types in Interacting Human–AI Systems
William Overman ⋅ Mohamad Sadegh Shirani Faradonbeh ⋅ Mohsen Bayati
We study experiments on interacting populations of humans and AI agents, where both unit types and the interaction network remain unobserved. Although causal effects propagate throughout the system, the goal is to estimate effects on humans. Examples include online platforms where human users interact alongside AI-driven accounts. We assume a human–AI prior that gives each unit a probability of being human. While humans cannot be distinguished at the unit level, the prior allows us to compute the average human composition within large subpopulations. We then model outcome dynamics through a causal message passing (CMP) framework and analyze sample-mean outcomes across subpopulations. We show that by constructing subpopulations that vary in expected human composition and treatment exposure, one can consistently recover human-specific causal effects. Our results characterize when distributional knowledge of population composition (without observing unit types or the interaction network) is sufficient for identification. We validate the approach on a simulated human–AI platform driven by behaviorally differentiated LLM agents. Together, these results provide a theoretical and practical framework for experimentation in emerging human–AI systems.
CCDiff: Inverse Canonical Correlation Analysis for Discovering Visual Differences in Natural Language
Neelesh Bisht ⋅ Xingjian Li ⋅ Zihan Li ⋅ Bo Jiang ⋅ Runmin Jiang ⋅ Mostofa Rafid Uddin ⋅ Yang Liu ⋅ Min Xu
Set-level visual difference discovery is increasingly important for dataset auditing and for understanding model behavior under distribution shift, yet manually inspecting thousands of images is impractical. This motivates the emerging task of *set difference captioning*, i.e., describing concepts that are more often true for one image set than another. Different from existing approaches which heavily rely on large language models (LLMs), incurring substantial computational cost in terms of time and tokens, we introduce **CCDiff**, a lightweight, statistically grounded, and training-free framework for set difference captioning that significantly reduces LLM dependence during core difference discovery, achieving over **$2\times$** speedup and reducing token usage to zero. CCDiff operates in three stages: constructing a shared pool of candidate concepts from domain vocabularies or lightweight captioning models, filtering candidates to obtain set-specific concept pools, and performing inverse canonical correlation analysis (CCA) across sets to identify low-correlation directions that correspond to set differences. Candidate concepts are then ranked using a CCDiff score that balances inverse correlation with within-set representativeness. We evaluate CCDiff on VisDiffBench and further validate it on additional domains including MetaShift and MIMIC-CXR. Across benchmarks and real-world settings, CCDiff matches or exceeds prior LLM-based approaches while being substantially more efficient and scalable.
ChanSFormer: A Channel Agnostic Vision Transformer for Multi-Channel Cell Painting Images
Jingwei Zhang ⋅ Srinivasan Sivanandan ⋅ Dimitris Samaras
High-content multichannel imaging, from Cell Painting assays to remote sensing, is central to modern scientific pipelines. Because these channels are highly heterogeneous and experimental configurations frequently evolve, a practical vision backbone must be fundamentally channel-agnostic. Current channel-adaptive vision transformers rely on global self-attention with learnable channel embeddings, which can be computed only for previously known channels. Attempts to overcome this limitation of new channels either discard the embeddings, losing channel identity, or process each channel independently, thereby eliminating critical cross-channel interactions. To address this, we introduce ChanSFormer, an efficient Vision Transformer that replaces rigid embeddings with disentangled spatial-channel attention and individual channel CLS tokens. This architecture natively preserves both distinct channel identities and cross-channel information flow. Furthermore, the individual channel class token design unlocks more informative feature representations for self-supervised learning and enables representation-based channel sampling while remaining channel agnostic. We evaluate ChanSFormer on two biology multi-channel datasets, CHAMMI, JUMP-CP and a satellite dataset, So2Sat. Experimental results show that ChanSFormer outperforms state-of-the-art methods by up to 4.27% in classification accuracy. It also demonstrates exceptional cross-dataset transferability under a self-supervised setting, outperforming previous methods by up to 2.72% in accuracy in the same classification tasks. Furthermore, the disentangled attention reduces the quadratic complexity of channel $\times$ spatial sequence length of the global attention baseline, thus improving throughput by 55%--313%.
CipherFlow: Hardware-Aware Compiler Framework for Low-Latency Hybrid Secure Inference
Hedong Zhang ⋅ Mengxin Zheng ⋅ Qian Lou
Secure inference enables clients to use cloud neural networks without revealing private inputs or intermediate values. Homomorphic encryption (HE) and secure multi-party computation (MPC) provide complementary cryptographic mechanisms for this goal, and recent hybrid HE-MPC secure inference systems have shown better performance than using either primitive alone. However, existing hybrid secure inference systems typically rely on fixed execution strategies that are brittle across deployments. The best hybrid plan depends on both the target execution environment and the evolving cryptographic state of the computation, making manual or static mappings insufficient. We present \textsc{CipherFlow}, a hardware-aware compiler for hybrid HE-MPC secure inference. Given a plaintext neural-network graph and a target deployment profile, \textsc{CipherFlow} automatically generates a deployment-specialized, latency-optimized secure execution program. It profiles deployment-specific primitive costs and performs state-aware graph optimization to jointly decide operator placement, cross-domain conversion, and cryptographic maintenance actions. We implement \textsc{CipherFlow} on real HE and MPC backends and evaluate end-to-end BERT-base and ViT inference across diverse CPU/GPU and network settings. \textsc{CipherFlow} achieves up to 16.5$\times$ speedup on BERT-base and over 22$\times$ on ViT over state-of-the-art static hybrid baselines. These results show that compiler support can make hybrid cryptographic inference portable across heterogeneous deployments, turning manual secure-inference engineering into an automatic deployment-specific optimization process.
CLAMP: A Sim-to-Real Benchmark for Closed-Loop Kinematic Pose Estimation and Assembly Reasoning
Kevin Murray ⋅ Randolph B Robert ⋅ Petar Z Duric ⋅ Zoran Duric
Existing benchmarks for known-object pose estimation either treat objects as rigid bodies or restrict articulation to serial-chain robot arms with revolute joints, and all assume that every part is always present. Real-world industrial and consumer equipment, however, exhibits closed-loop kinematic chains, prismatic joints, and structural assembly perturbations—challenges that current datasets and methods do not address. We introduce CLAMP (Closed-Loop Assembly and Mechanism Perception), a sim-to-real benchmark for kinematic pose estimation and assembly reasoning on mechanically complex known objects. The dataset spans seven diverse pieces of equipment—ranging from a delta 3D printer with 27 coupled DOFs to a hydraulic jack with closed kinematic loops—comprising over two million synthetic training images and approximately 21,000 real test images across 210 scenes, each annotated with ground-truth joint states and part-level assembly labels. To produce this data we develop (i) a graph-based kinematic representation that extends beyond URDF's tree structure to support closed-loop constraints and dynamic topology changes induced by missing parts; (ii) a synthetic data generator that samples valid poses on the kinematic constraint manifold and stochastically removes parts with automatic kinematic restructuring; and (iii) a real-image labeling pipeline that jointly recovers camera alignment and equipment pose via multi-view constrained optimization. Baseline evaluation with RoboPEPP—the current state of the art on the DREAM robot pose benchmark—shows that our constrained pose sampling is critical for closed-loop equipment, yet performance on the combined task of kinematic estimation under assembly perturbation remains low, highlighting an open challenge for the community. We release the full dataset, the synthetic data generation pipeline, evaluation code, and baseline implementation.
Clustering with Weak Distance Oracles
Aryan Esmailpour ⋅ Rahul Raychaudhury ⋅ Sainyam Galhotra ⋅ Stavros Sintos
Clustering is a fundamental task in unsupervised learning, serving as a core tool for data mining and exploratory data analysis. Classical $k$-clustering objectives such as $k$-median and $k$-means formalize this task, but typically assume access to exact pairwise distances. This assumption is increasingly unrealistic in modern applications where distances are estimated through proxy embedding models, learned similarity functions, human feedback, or other low-cost procedures that may be noisy. We study $k$-clustering in an unknown metric space where the algorithm has access only to a weak distance oracle: for each pair of objects, the oracle returns the true distance with probability of $1/2+\varepsilon$ , where $\varepsilon$ is a constant, and otherwise may return an arbitrary corrupted value. We present randomized algorithms that, given a metric space with $n$ vertices and a parameter $k$, use only weak distance-oracle queries to compute a set of representative centers together with a mapping from every input object to one of these centers, such that the total mapping cost is within a constant factor of the optimal $k$-clustering cost. Our first algorithm returns $O(k\cdot\mathsf{polylog}(n))$ centers using $O(n\cdot k\cdot \mathsf{polylog}(n))$ weak-oracle queries and runs in quasi-polynomial time. We also give a polynomial-time algorithm that returns $O(k^2\cdot\mathsf{polylog}(n))$ centers using $O(n\cdot k\cdot \mathsf{polylog}(n))$ weak-oracle queries. Preliminary experiments for $k$-means clustering show that our approach remains close to the optimum under substantial oracle noise, while noise-oblivious baselines degrade sharply.
CMPQ: Compensatory Quantization via Input-Aware Hessian Damping
DENG PAN ⋅ Grigorii Khvatskii ⋅ Xiaobao Huang ⋅ Peiyu Li ⋅ Shangqian Gao ⋅ Ting Hua ⋅ Nitesh Chawla
With the recent advancement of large language models (LLMs) scaling, post-training weight quantization (PTQ) has become one of the standard tools for making models deployable on commodity hardware with minimal performance degradation. However, existing quantization solvers such as GPTQ calibrate each weight matrix against a corrupted signal: the calibration input observed at any deep layer is the output of the cumulative quantized chain, not a clean full-precision reference. We propose \textbf{CMPQ (Compensatory Quantization)}, which reframes the corruption as evidence about \emph{which input directions have become unreliable} and tells the solver to stop spending its compensation budget on them. CMPQ measures the divergence between the quantized and full-precision input and uses it to reshape the curvature that GPTQ already relies on, redirecting compensation onto weights attached to clean input subspaces and cutting the propagation path of upstream noise. On the Qwen3 family at three sizes (4B / 8B / 14B), CMPQ matches GPTQ-family baselines at 3-bit and substantially outperforms them at 2-bit, where standard solvers collapse to near-random accuracy.
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
Han Guo ⋅ Jack Zhang ⋅ Arjun Menon ⋅ Driss Guessous ⋅ Vijay Thakkar ⋅ Yoon Kim ⋅ Tri Dao
Transformer training systems are built around dense linear algebra, yet a nontrivial fraction of end-to-end time is spent in the memory-bound operators that surround it. Normalization, activations, residual updates, reductions, and related computations repeatedly move large intermediate tensors through global memory while performing little arithmetic, making data movement an increasingly important bottleneck in otherwise highly optimized training stacks. We introduce CODA, a GPU kernel abstraction that expresses these computations as GEMM-plus-epilogue programs. CODA is based on the observation that many Transformer operators exposed as separate framework kernels can be algebraically reparameterized so that their work executes while a GEMM output tile remains on chip, before it is written to memory. The abstraction fixes the GEMM mainloop and exposes a small set of composable epilogue primitives for scaling, reductions, pairwise transformations, and accumulation. This constrained interface preserves the performance structure of expert-written GEMMs while remaining expressive enough to cover nearly all non-attention computation in the forward and backward pass of a standard Transformer block. Across representative Transformer workloads, both human- and LLM-authored CODA kernels achieve high performance, suggesting that GEMM-plus-epilogue programming offers a practical path toward combining framework-level productivity with hardware-level efficiency.
Coherent Routing in Decision Trees: Phase-Interference Learning for Interpretable Tabular Prediction
David Li ⋅ Angela Li
We introduce Phase-Interference Decision Trees (PIDT), a differentiable routing model for tabular prediction in which each node propagates a bounded complex state rather than only a scalar routing probability. The main model uses two-state norm-preserving branch maps: a gate controls how much amplitude is sent to the left and right children, while learned unitary maps rotate the latent state before subsequent splits. This formulation contains ordinary soft decision trees as the phase-free scalar case, and two-state routing creates explicit score-level cross terms among latent-state trajectories. We give a self-contained PIDT definition with exact mass conservation, prove scalar phase cancellation, establish a scoped containment relation with soft trees, derive an algebraic interference decomposition, give a companion construction for shared two-state parity routing, and derive a fixed-architecture empirical Rademacher upper bound for bounded two-state PIDT scores. We evaluate PIDT on a ten-dataset OpenML suite of low-class-count tabular tasks under matched differentiable-tree controls. Across ten seeds per dataset, the main summaries track accuracy as context and report small descriptive NLL and Brier deltas relative to the scalar PIDT-1S, phase-frozen PFB, and SDT controls. The empirical study is framed as a mechanism-level study for compact differentiable trees, not as a broad tabular-performance claim.
CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support
Yichen WU ⋅ Yujin Oh ⋅ Sangjoon Park ⋅ Kailong Fan ⋅ Yuhan Liu ⋅ Zhiyi Shi ⋅ Sekeun Kim ⋅ Dania Daye ⋅ Hana Farzaneh ⋅ Wei Liu ⋅ Xiang Li ⋅ Raul N Uppot ⋅ Quanzheng Li
Recent multi-agent frameworks have shown promise for oncology decision support, yet most assume centralized data access and rely on prompt-based assignment, limiting their applicability in privacy-sensitive clinical settings. We propose Contribution-Aware Medical Multi-Agents (CoMMa), a decentralized LLM-agent framework where specialists operate on partitioned clinical data streams. Unlike prior approaches that share inputs across agents, CoMMa enforces data decentralization to include stronger role specialization and further enhances this via agent-specific finetuning. To enable reliable and interpretable coordination, we introduce a contribution-aware aggregation mechanism that replaces stochastic, narrative-based reasoning with deterministic embedding projections to approximate each agent's marginal utility. This yields explicit credit assignment over agents, providing a stable and interpretable decision pathway aligned with clinical requirements. We evaluate CoMMa on multiple oncology benchmarks, including real-world multidisciplinary tumor board datasets from large academic hospitals in North America and East Asia, as well as public datasets, demonstrating strong performance and generalization across heterogeneous clinical settings.
Scaling test-time compute has emerged as a powerful mechanism for enhancing Large Language Model (LLM) performance. However, standard post-training paradigms, Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), optimize the likelihood of individual samples under a base policy, creating a misalignment with test time procedures that rely on aggregated or filtered outputs. In this work, we propose Compute Aligned Training, which aligns training objectives with test-time strategies. By conceptualizing inference strategies as operators on the base policy, we derive new loss functions that maximize performance when said strategies are applied. We instantiate such loss functions for SFT and RL across common test time strategies. Finally, we provide empirical evidence that this training method substantially improves test time scaling over standard training.
Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction--Concealment Tradeoff in MLLMs
Md Farhamdur Reza ⋅ Richeng Jin ⋅ Tianfu Wu ⋅ Huaiyu Dai
Intent-obfuscation-based jailbreak attacks on multimodal large language models (MLLMs) transform a harmful query into a concealed multimodal input to bypass safety mechanisms. We show that such attacks are governed by a \emph{reconstruction--concealment tradeoff}: the transformed input must hide harmful intent from safety filters while remaining recoverable enough for the victim model to reconstruct the original request. Through a reconstruction analysis of three representative black-box methods, we find that existing transformations struggle to balance this tradeoff, limiting their effectiveness. In contrast, we show that character-removed variants achieve a better balance. Building on this, we propose \emph{concealment-aware variant construction}, which greedily selects character-removed variants that are low in harmful-keyword alignment and mutually diverse, and instantiates them through five modality-aware prompting strategies. We further introduce \emph{keyword-related distractor images} that depict the harmful keyword in diverse contexts, providing more effective auxiliary visual context than generic distractor images. Experiments across closed-source and open-source MLLMs show the proposed strategies outperform strong baselines, revealing an underexplored vulnerability: a model's own reconstruction ability can be exploited to recover hidden harmful intent and produce unsafe responses.
Concept Modulation Models: A Unified Framework for Identifiability and Extrapolation
Soheun Yi ⋅ Yizhou Lu ⋅ Chandler Squires ⋅ Pradeep Ravikumar
Reliable generalization in conditional latent-variable models requires understanding both identifiability and extrapolation: whether observed variation across attributes determines latent structure, and whether that structure determines distributions at unseen attributes. However, existing identifiability and extrapolation guarantees are largely model-specific, with separate analyses in nonlinear ICA, causal representation learning, perturbation modeling, and related conditional latent-variable models. We introduce *concept modulation models (CMMs)*, an attribute-indexed class of conditional generative models with structure $A\to \Lambda \to C\to X$, where attributes select modulators, modulators induce latent concept laws, and concepts generate observed features. CMMs lift transition-based identifiability to conditional settings by showing that feature agreement on observed attributes induces a latent concept transition constrained by the CMM class. We express these constraints through *attribute potentials*, log-density ratios between attribute-conditioned concept laws, separating the generic lifting step from model-specific rigidity arguments. The same potentials control extrapolation: agreement at unseen attributes holds exactly when the transported attribute-potential identities extend to those attributes. This yields algebraic extrapolation criteria, identifies the common potential-based proof objects behind several existing identifiability and extrapolation results, and, when combined with the model-specific rigidity arguments in those works, recovers their stated conclusions.
Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification
Jiayi Xin ⋅ Evan Qiang ⋅ Zihan Zhu ⋅ Xiang Li ⋅ Weijie Su ⋅ Qi Long
Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods often fail to remain stable under simple, meaning-preserving perturbations. We identify an important source of instability: semantically equivalent paraphrases of the same input can induce substantial variability in predictive confidence, even for methods with formal guarantees such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework that explicitly enforces invariance to semantic rewordings. Instead of relying on a single input, our approach constructs uncertainty scores by aggregating predictions across a set of its paraphrases, using a lightweight proxy model to produce calibrated and comparable outputs. This aggregation yields uncertainty estimates that are both more stable and more informative, while remaining compatible with conformal calibration techniques to retain coverage guarantees. Across multiple UQ benchmarks and model families, our method achieves nominal coverage with smaller and more stable prediction sets under paraphrase perturbations. Code is available at https://anonymous.4open.science/r/PA_Score-8C0D.
Constrained Code Generation with Discrete Diffusion
Lize Shao ⋅ Michael Cardei ⋅ Zichen Xie ⋅ Ferdinando Fioretto ⋅ Wenxi Wang
Discrete diffusion models are a powerful, emerging, paradigm for code generation. They construct programs through iterative refinement of partially corrupted token sequences and enable parallel token refinement. Importantly, this paradigm exposes a global program state at each denoising step, which provides a natural intervention point for enforcing program-level functionality and security constraints and guide generation before the final code is committed. Building on this observation, the paper introduces Constrained Diffusion for Code(CDC), a training-free neurosymbolic inference framework that integrates constraint satisfaction directly into the reverse denoising process. CDC augments the base discrete diffusion sampler with constraint-aware denoising operators that combine mathematical optimization with program analysis to identify constraint-relevant regions of the intermediate program state and locally adjust the denoising trajectory, steering generation toward feasible programs while remaining close to the base diffusion model. Across code generation benchmarks, CDC consistently improves constraint satisfaction in functional correctness, security, and even syntax, outperforming discrete diffusion and autoregressive baselines with less corrective computation and more localized edits.
Constraint-Aware Influence Estimation in Deep Constrained Learning
Xin Wang ⋅ R.Tyrrell Rockafellar ⋅ Xuegang Ban
Constrained learning has been increasingly applied to various domains to ensure explicit feasibility requirements due to fairness, safety, robustness, regularization, and physics or logic constraints. Understanding how training samples influence the solution (e.g., learned parameters) of constrained learning is crucial for interpretability and robustness. The classical influence function may become unreliable in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a new estimator that explicitly incorporates the constraints into influence estimation. DIF formulates the optimality conditions of constrained learning as a variational inequality (VI) and analyzes how perturbing training data affects this VI. We validate DIF in constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.
Consumer Search and Social Learning in Agentic Markets
Brendan Lucier ⋅ Nicole Immorlica ⋅ Markus Mobius ⋅ Aleksandrs Slivkins ⋅ Daniel G Goldstein ⋅ Jake M Hofman ⋅ Sonia Jaffe ⋅ David Rothschild
Motivated by agentic markets -- two-sided markets in which AI tools facilitate search -- we propose a model that incorporates individual consumer search and market-wide learning, solve it to understand the long-run market outcomes, and then study the impact of improved search technologies on these outcomes. A sequence of consumers engage in costly search to acquire progressively refined signals of product fit prior to purchase. The market observes acquired signals and post-transaction feedback thereby impacting future searches. We solve for the per-consumer optimal search policy and characterize the steady-state of the learning dynamics. As search technologies reduce the cost and/or increase the informativeness of search, they cause an improvement in learning and consumer surplus. Using numerical simulations, we also find that these technologies improve the rate of learning and the trajectory of consumer utility. Notably, we show that if the market is unable to observe acquired signals, search technologies can degrade market outcomes, highlighting the importance of transparency in agentic markets.
Controllable and Content Based Recommendations
Firat Oncel ⋅ Jihoon Jeong ⋅ Emiliano Penaloza ⋅ Mirco Ravanelli ⋅ Laurent Charlin ⋅ Cem Subakan
Traditional recommendation systems rely on latent (dense) representations, making them difficult to interpret and control. We propose the Controllable and Content-Based Recommendations (CCBR) framework, which builds its recommendations from textual user profile representations. CCBR plugs into collaborative filtering models and introduces controllability via text bottlenecks. We show that CCBR enables text-based and multimodal interventions, allowing users to steer the model towards the directions they prefer. Different from existing controllable recommendation systems, CCBR infers the text summaries directly from item contents (images, audio or video). Across image-, audio-, and video-based datasets, we demonstrate that the proposed framework obtains competitive model performance with standard (latent-representation) models while providing controllable model summaries via text. The model also outperforms TEARS, a recent baseline for controllable recommendation systems. Through systematic interventions, we demonstrate the efficacy of the user steering mechanism.
We propose Coreset-Induced Conditional Velocity Flow Matching (CCVFM), a generative model that augments hierarchical rectified flow with a data-informed source distribution. Hierarchical flow matching models the full conditional velocity law in velocity space, but its inner flow is asked to transport isotropic Gaussian noise to a multimodal target velocity distribution from scratch. Our key observation is that this inner source can be replaced by a closed-form surrogate built from a coreset of the target. CCVFM first compresses the target into weighted atoms using an entropic Sinkhorn coreset and lifts them to a Gaussian mixture. The induced conditional velocity law is then a closed-form Gaussian mixture that can be sampled without a learned neural sampler. A lightweight correction flow, trained from this exact surrogate source, then refines the remaining surrogate-to-target residual rather than learning an entire noise-to-data map. We prove that the surrogate transport cost equals the target--surrogate Wasserstein gap under an explicit compression assumption, whereas the noise-source analogue has a dimension-scale lower bound. We further characterize the conditional second moment of the direct surrogate-source training target and show that its source-dependent excess is small when the surrogate conditional law is close to the true conditional velocity law in mean and covariance. Empirically, on MNIST, CIFAR-10, ImageNet-32, and CelebA-HQ, the proposed method reaches competitive few-step generation under matched architectures. Codes are available at \url{https://anonymous.4open.science/r/ccvfm-code-11D3/}.
Counterfactual Distillation: Internalizing Reflective Experience into LLM Agent
Rui Ge ⋅ Yichao Fu ⋅ Yu-Yang Qian ⋅ Junda Su ⋅ Yiming Zhao ⋅ Hao Zhang
Large language models are increasingly deployed as autonomous agents that plan and act over long-horizon interactions, learning from both final outcomes and rich environmental feedback. However, optimizing from such interaction experience remains challenging, as feedback is heterogeneous and rarely carries explicit reward information. Existing post-training approaches rely on outcome-driven methods, such as reinforcement learning with verifiable rewards (RLVR), which primarily exploit final success signals while leaving interaction experience and feedback underutilized. In sparse-reward, long-horizon tasks, this often results in **distribution sharpening**: the policy reinforces a narrow set of already-successful behaviors, without substantially improving the feedback-grounded agency needed for broader problem-solving capacity (e.g., Pass@$k$). We propose ***Counterfactual Distillation***, a framework that converts exploration-derived experience into trainable policy supervision. Specifically, we organize exploration as a tree-structured search process, where the agent reflects on past decision points and makes new experience-guided decisions to form alternative branches. These corrections are then distilled as counterfactual targets: we train the model to predict the revised action from the original history without explicitly providing the experience, thereby introducing out-of-distribution behavior patterns and expanding exploration capacity. Across diverse interactive coding and agentic tasks, our method outperforms outcome-driven baselines such as GRPO and experience-based methods such as Early Experience, with gains of up to **14%** on Pass@128. When interleaved with reinforcement learning updates, it further raises the performance ceiling, yielding over **10%** improvement in Pass@1. We provide an [anonymized implementation](https://anonymous.4open.science/r/Counterfactual-Distillation-0E71).
Covariate-Adjusted Deep Causal Learning for Heterogeneous Panel Data Models
Guanhao Zhou ⋅ Yuefeng Han ⋅ Xiufan Yu
This paper studies the task of estimating heterogeneous treatment effects in causal panel data models with covariate effects. We propose a novel Covariate-Adjusted DEep CAusal Learning (CoDEAL) framework, that cohesively deal with the underlying heterogeneity and nonlinearity of both panel units and covariate effects. CoDEAL integrates nonlinear covariate effect components (parameterized by a feed-forward neural network) with nonlinear factor structures (modeled by a multi-output autoencoder) to form a heterogeneous causal panel model. The nonlinear covariate component flexibly captures complex covariate influences on outcomes, and the nonlinear factor decomposition enables CoDEAL to effectively capture both cross-sectional and temporal dependencies inherent in the data panel. This latent structural information is subsequently integrated into a customized matrix completion algorithm, thereby facilitating more accurate counterfactual imputation. Moreover, the use of a multi-output autoencoder explicitly accounts for heterogeneity across units and enhances the model interpretability of the latent factors. We establish theoretical guarantees on the convergence of the estimated counterfactuals and demonstrate the compelling performance of the proposed method using extensive simulation studies and real-world data applications.
CupOFMoCA: Coupled Objective-Guided Discrete Flows for Molecular Conjugate Assembly
Ruoxi Zhang ⋅ Ziang Li ⋅ Jiatao Gu ⋅ Pranam Chatterjee
Molecular conjugates, including PROTACs and peptide-drug conjugates (PDCs), derive their function from the joint behavior of multiple coupled components, yet most generative approaches design these components independently and combine them only after generation. Such staged pipelines ignore cross-component dependencies and often produce conjugates that are chemically invalid, fall outside empirical conjugate distributions, or lose function upon assembly. We introduce $\textbf{C}$o$\textbf{up}$led $\textbf{O}$bjective-Guided Discrete $\textbf{F}$lows for $\textbf{Mo}$lecular $\textbf{C}$onjugate $\textbf{A}$ssembly $(\textbf{CupOFMoCA})$, a discrete generative framework that formulates conjugate design as a constrained, coupled generation problem. CupOFMoCA restricts generative trajectories to a chemically feasible conjugate manifold and biases local transitions using target-specific activity predictors, ensuring all components remain mutually compatible throughout generation. We show that coupling constraints and objective guidance enable anticipatory design that preserves post-assembly predicted activity and produces structurally realistic conjugates across both PDC and PROTAC settings, outperforming staged baselines, including LinkerNet, DiffLinker, and DiffPROTACs for PROTACs, across assembly validity, predicted activity, and physicochemical property ranges. These results demonstrate that explicit coupling and constraint enforcement are sufficient to recover functional conjugates across conjugate classes, and provide a principled foundation for generative modeling where function emerges only at the level of the assembled system. Our anonymous code repository can be found at https://anonymous.4open.science/r/Cupofmoca-Neurips.
Formal verification provides strong correctness guarantees, but proof development remains difficult because developers must diagnose coarse verifier errors and repair failed proof obligations. Recent approaches leverage language models (LMs) to improve proof automation through verifier-in-the-loop repair, but they often rely on hand-engineered prompts, retrieved examples, or sparse pass/fail feedback. We present \tool, an LM-based framework for Verus proof repair that internalizes counterexample-guided diagnosis as a model capability. Our key insight is to train counterexample generation not as a standalone validation task, but as a data-driven diagnostic objective: a counterexample is useful when conditioning on it helps the model produce a safe, verifier-passing proof. Given a failed Verus proof and verifier error, \tool generates a structured source-level counterexample, repair rationale, and repaired proof in one trajectory, optimized with reinforcement learning from Verus feedback and a proof-code safety gate. To address sparse and order-sensitive credit assignment, \tool introduces a topological graph reward that derives invariant dependencies from inductiveness tests and encourages counterexample generation following the same topological order as the invariant dependencies. Across seven Verus benchmarks with 471 problems, \tool achieves 61.6\%/72.2\% weighted-average Safe-Pass@1/Safe-Pass@3, outperforming the strongest baseline, Claude Sonnet 4.5, by 5.7/6.3 points. We also show that \tool's learned counterexample can directly augment the reasoning of existing LM-based proof repair approaches, improving their success rate by 64.3\% compared to using the chain-of-thought diagnostics.
DD-CAM: Minimal Sufficient Explanations for Vision Models Using Delta Debugging
Krishna Khadka ⋅ Yu Lei ⋅ Raghu Kacker ⋅ D. R Kuhn
Class Activation Mapping (CAM) and related saliency techniques are widely used to explain predictions of deep vision models, producing post-hoc heatmaps that highlight image regions a model relied on. However, existing CAM methods aggregate weighted contributions over the full set of internal units, including units whose contribution may be incidental. The resulting heatmaps are often diffuse and obscure which units are actually sufficient to preserve the prediction. We reframe visual explanation as the selection of a 1-minimal sufficient subset of internal representational units, feature maps in CNNs and patch tokens in ViTs, whose joint activation preserves the model's prediction. A subset is 1-minimal sufficient if removing any single unit changes the prediction. We propose \textbf{DD-CAM}, a gradient-free procedure adapted from delta debugging in software engineering that instantiates this formulation. DD-CAM exploits classifier-head structure with two regimes: iterated single-unit removal for non-interacting heads (GAP+FC), and recursive partition-and-reduce for interacting heads (multi-layer FC stacks and ViT self-attention). On 2{,}000 ImageNet validation images across eight CNN and ViT architectures, DD-CAM is the strongest method on 13 of 18 faithfulness metric-group cells against thirteen baselines. On the NIH ChestX-ray14 bounding-box set, DD-CAM improves IoU by 45\% and F1 by 22\% over the strongest baseline while producing single-region explanations. These results suggest that necessity-grounded internal-unit selection can produce more focused and faithful explanations than weighted aggregation over all units, with particular promise for safety-critical settings such as medical imaging.
DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents
Yun-Shiuan Chuang ⋅ Ruixuan Tu ⋅ Chengtao Dai ⋅ You Li ⋅ Binwei Yao ⋅ Michael Tessler ⋅ Sijia Yang ⋅ Dhavan Shah ⋅ Robert Hawkins ⋅ Junjie Hu ⋅ Timothy T Rogers
Modeling opinion change through social interactions is important for understanding polarization, misinformation, and societal conflict. Recent work simulates opinion dynamics with role-playing LLM agents (RPLAs), but multi-agent simulations often display unnatural group behavior, such as premature convergence, and lack empirical benchmarks for assessing alignment with real human group interactions. We introduce DEBATE, a benchmark for evaluating the authenticity of opinion dynamics in multi-agent RPLA simulations. DEBATE contains multi-round public messages and private Likert-scale beliefs from U.S.-based participants across 107 topics; the cleaned benchmark used in our experiments contains 2,788 participants in 697 groups, enabling evaluation at the utterance and group levels. We instantiate “digital twin” RPLAs with seven LLMs and evaluate across two settings: next-message prediction and full dynamics simulation, using stance-based opinion-dynamics metrics. In zero-shot settings, RPLA groups exhibit strong opinion convergence relative to human groups. On the held-out group split, supervised fine-tuning (SFT) for Llama-3.1-8B-Instruct improves auxiliary stance alignment and reduces group-level convergence error, though discrepancies in opinion change and belief updating remain. DEBATE provides a benchmark for simulated opinion dynamics; reviewer-accessible code and dataset URLs are provided at submission.
Decision Focused Scenario Learning for Contextual Stochastic Programming
Jonathan Hornewall ⋅ Tito Homem-De-Mello ⋅ Vincent Leclere
Decision-focused scenario learning for contextual two-stage stochastic programming has emerged as a promising research direction, but open questions remain. A general theory for when scenario generators can yield optimal decision policies is currently lacking. Many existing methods require strict assumptions on which second-stage data are allowed to be random. Generally, the non-differentiability of the decision-focused objective is a source of difficulty. This paper makes three theoretical contributions, each aiming to remedy one of these challenges. First, we present a comprehensive theory for when scenario generators can induce optimal policies, showing that learning a single out-of-support scenario component is often sufficient. Second, we demonstrate a theoretical and methodological equivalence result between learning different second-stage data components, relaxing assumptions required for applying existing methods. Third, we study log-barrier smoothing enabled training, presenting representability and optimization results. Finally, leveraging the theory, we propose a novel method inspired by interior-point methods from classical optimization, and validate it on standard benchmarks.
Demystifying Classifier-Free Guidance for Auto-Regressive Image Generation
Zhiling Zhou ⋅ Jiachun Pan ⋅ Fengzhuo Zhang ⋅ Dirk Bergemann ⋅ Zhuoran Yang
Classifier-free guidance (CFG) has been widely adopted in autoregressive (AR) models for high-quality image generation. Despite its strong empirical performance, its mechanism in AR models remains unclear. This paper demystifies the mechanism of CFG through both empirical and theoretical studies. By examining the top-ranked semantics of different components in CFG, we show that texture information is primarily encoded in the difference between the conditional and unconditional logits, whereas both logits share previous-token repetition as their leading semantics. This strong repetition semantics obscures the desired texture information, causing greedy decoding from the conditional logits alone to produce nearly pure-color images. Through a training-dynamics analysis of shallow transformers, we prove that this shared repetition semantics does not arise from limitations or failures of pretraining, but instead originates from the texture sparsity of images. We further show that CFG improves generation quality by rectifying this repetition bias during inference. Motivated by the shared semantics between conditional and unconditional logits, we propose Attention Weight Reuse (AttnReuse), which reuses intermediate attention computations from the conditional-logit forward pass to accelerate the unconditional-logit computation. AttnReuse reduces about $25\%$ of attention computation with little performance degradation across different models.
Dense optical flow estimates correspondence from every pixel to the next frame. It must be precise and robust, which couples two choices: what units to match, and what evidence to match them. Recent work has improved the evidence with stronger features, cost volumes, attention, refinement, and generative priors. The unit, however, usually remains fixed: motion is still estimated at dense grid points. We propose Adaptive Correspondence Tokens (ACT), which adapts correspondence units through a differentiable pixel-to-token assignment. ACT learns soft maps that group pixels into tokens, build descriptors, match and refine motion in token space, and project token motion back to pixels for dense flow. Trained end to end with the flow loss, grouping and correspondence are discovered jointly: grouping makes matching easier, while matching shapes grouping. Support can expand over coherent regions and contract near boundaries or fine detail. In zero-shot cross-dataset evaluation, ACT achieves the best KITTI EPE, second-best KITTI Fl, and second-best Sintel Final EPE among published methods, while remaining competitive under benchmark fine-tuning. These results show that adaptive units make dense flow easier: matching where the image gives reliable evidence improves transfer while preserving dense, boundary-aware prediction.
DICE: Decoupling Capability from Intervention Necessity in LLM Tutoring
Sayantan Pal ⋅ Kaiyi Ji ⋅ Rohini Srihari
Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student turn warrants a response. However, our experiments indicate that this conflates tutoring capability (what to say) with intervention necessity (whether to say it). We introduce DICE, a framework that decouples intervention decisions from response generation by first selecting an explicit pedagogical action. To calibrate this action selection policy, we define Intervention Value (IV), a rollout-grounded counterfactual metric that compares each action against non-intervention. IV shows that many prescribed interventions provide little or no marginal benefit. We further introduce DICE-Bench, a multi-variant math tutoring benchmark with skill-preserving problem variants for session-level evaluation. Using IV-weighted and KL-regularized policy optimization, DICE learns to intervene selectively while preserving tutoring effectiveness. In simulated tutoring sessions, DICE reduces the over-intervention rate to near zero while guiding students to correct solutions in approximately 3-4 fewer turns on average than existing Socratic tutoring baselines.
DiffATS: Diffusion in Aligned Tensor Space
Jinhua Lyu ⋅ Tianmin Yu ⋅ Brian Kim ⋅ Lizhuo Zhou ⋅ Chanwook Park ⋅ Naichen Shi
Direct diffusion modeling of high-resolution spatiotemporal fields is computationally challenging. Parameter-efficient primitives address this by representing high-dimensional data with a compact set of parameters. In this paper, we construct data-dependent tensor primitives without pretrained compression autoencoders. Our construction starts from Tucker decomposition, which captures low-rank multilinear structure through a core tensor and mode-wise factors. However, Tucker factors are non-unique: the same tensor can be represented by different rotated factors, which complicates generative modeling. We address this issue with orthogonal Procrustes (OP) alignment. Specifically, we select medoid anchor matrices from the data and align the factor matrices to resolve the gauge ambiguity. This yields matrix Grassmannian primitives and tensor Grassmannian primitives that are compact, data-adaptive, and directly decodable by explicit multilinear reconstruction. Theoretically, we prove that the proposed primitive maps are homeomorphisms between low-rank tensors and their corresponding primitive spaces, certifying that the representations are non-degenerate and topologically faithful. Building on these primitives, we propose *Diffusion in Aligned Tensor Space* (DiffATS), a generative framework that trains diffusion models directly on aligned tensor primitives. Across images, videos, and PDE solutions, DiffATS achieves strong unconditional and conditional generation performance while compressing original data by $3.9\times$ to $210\times$, without relying on any pretrained deep compression autoencoders.
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
Sayeh Sharify ⋅ Mahsa Salmani ⋅ Hesham Mostafa
Diffusion Transformers (DiTs) achieve state-of-the-art image generation quality but incur substantial memory and computational costs at inference. While aggressive Post-Training Quantization (PTQ) to 4-bit precision offers significant efficiency gains, it typically results in severe quality degradation. Existing approaches, including smoothing-based methods, mixed-precision schemes, rotation techniques, and low-rank residual methods, partially mitigate this issue but still leave a noticeable gap to FP16/BF16 performance. In this work, we introduce *DiRotQ*, a W4A4 PTQ framework that mitigates this degradation through rotation-aware activation quantization. DiRotQ identifies a low-rank subspace capturing dominant activation variance via Principal Component Analysis (PCA), preserving coefficients in this subspace at higher precision while quantizing the remaining components to 4-bit. Activations are rotated into the PCA basis at inference time using calibration-derived orthogonal transformations, while the inverse rotation is fused into the layer weights offline. Combined with GPTQ-based weight quantization, DiRotQ achieves an FID $(\downarrow)$ of $15.9$ and PSNR $(\uparrow)$ of $19.1$ dB on PixArt-$\Sigma$ over the MJHQ-30K dataset, outperforming the prior state-of-the-art SVDQuant (FID $18.9$, PSNR $17.6$) under the same INT W4A4 setting. Beyond standard metrics, we introduce a *VLM-as-a-Judge* evaluation protocol for diffusion model quantization, the first such evaluation in this setting, providing a more holistic assessment of perceptual quality and prompt alignment under aggressive compression. On the systems side, we implement a *Triton-based custom kernel* to enable efficient end-to-end inference, reducing memory usage of the 12B FLUX.1-dev model by $2.1\times$ and delivering $2.3\times$ speedup over the BF16 baseline, on a 24GB RTX 4090 GPU. Our anonymous codebase is available [here](https://anonymous.4open.science/r/DiRotQ-5BF6/).
DoFP-Aligned Lookup Tables for Real-Time Polarization Demosaicking
Yidong Luo ⋅ Chenggong Li ⋅ Yunfeng Song ⋅ Mengyuan Liu ⋅ Junchao Zhang ⋅ Xin Yuan
Real-time polarization imaging with off-the-shelf color division-of-focal-plane (DoFP) cameras requires demosaicking that is both accurate and deployment-friendly. Interpolation methods are lightweight but limited in polarization fidelity, whereas learning-based methods improve quality at much higher online cost. We propose DoFP-LUT, a sensor-structured LUT framework and, to our knowledge, the first LUT-based framework for real-time color DoFP polarization demosaicking. It leverages the deterministic and memory-light nature of LUT inference while aligning lookup corrections with DoFP-specific residual structures. Unlike generic LUTs designed for RGB restoration, DoFP-LUT targets post-initialization residuals that are coupled across analyzer orientations, Stokes relations, and local spatial structures. It factorizes teacher-guided correction into shared analyzer correction, polarization redistribution, and high-frequency compensation, compiling each component into compact low-dimensional LUTs offline. Online inference then requires only nearest-neighbor lookup and fixed write-back. Experiments show that DoFP-LUT achieves a favorable quality--efficiency trade-off, improving polarization reconstruction while preserving deterministic, memory-light inference for real-time polarization imaging.
DopplerWild: A Doppler Dataset and Benchmark for Human Kinematic Understanding in the Wild
Shubo Yang ⋅ Aaron Wang ⋅ Narayan Schuetz ⋅ Jida Zhang ⋅ Edison R Altamirano ⋅ Kaitlyn Leitherer ⋅ Ehsan Adeli ⋅ Amin Arbabian
Understanding human kinematics in outdoor, real-world environments is critical for assistive robots and intelligent infrastructure. Many perception systems rely on cameras and LiDAR, which infer motion indirectly from spatial observations, making estimates sensitive to discontinuities and ambiguity under low lighting or at distance. In contrast, mmWave radar is robust to lighting and weather, operates at long range, and directly measures motion through Doppler signatures. However, Doppler-based kinematic understanding remains largely unexplored in the wild because existing datasets are primarily controlled and scripted. We introduce DopplerWild, a dataset and benchmark for evaluating Doppler-based kinematic understanding in the wild. DopplerWild comprises 447k radar frames across four outdoor locations, including an unlabeled subset for self-supervised representation learning and a labeled subset of 539 subjects with nearby-person interference (56\%) and occlusion (16\%). The benchmark spans coarse and fine-grained motion classification, as well as velocity estimation. Beyond aggregate metrics on controlled datasets, DopplerWild evaluation isolates real-world factors: subject variation, location shift, multi-person interference, occlusion, sensing geometry, and low-label regimes. DopplerWild reveals three findings. First, self-supervised pretraining matches supervised performance with less than half the labeled data and improves performance under interference, occlusion, and lateral motion. Second, we identify key challenges for single-view Doppler sensing, including physical limits in lateral motion, subtle asymmetric limb motion, and location distribution shifts. Third, we observe bidirectional transfer asymmetry between DopplerWild and an external controlled Doppler dataset, suggesting that real-world coverage is important for transferable Doppler representations. DopplerWild provides an evaluation framework for characterizing the generalization, limits, and transferability of Doppler-based kinematic representations in the wild.
Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also generally believed that this more "deliberative" mode should improve alignment and safety, by providing the model a safe space to consider whether its planned answer to the request violates its safety principles. We present evidence that this intuition is not always correct. Across frontier open-weight reasoning models spanning GPT-OSS, Phi, OLMo, and Qwen model families, we find that the model's decision is already strongly encoded at the beginning of thinking, with a probe on the first token's hidden representation predicting refusal/compliance with $\ge0.85$ AUROC and $\sim90$\% balanced accuracy. We also find little evidence that current models can use their thinking trace to deliberate about safety, as additional thinking after the first 20\% of the trace rarely moves the final decision. While sentence-level inspection of thinking traces may show signs of oscillation between refusal- and compliance-leaning rationales, we find that in $\geq 85$\% of thinking traces, such oscillations exert limited to no influence on the final response. We also examine the effect of existing inference-time and training-based safety interventions and find that they largely alter thinking behavior by shifting models toward more refusal-leaning reasoning while substantially reducing helpfulness on benign prompts. Our results suggest that safety behavior in current reasoning models is much less deliberative than commonly assumed, highlighting the need for training methods that more effectively utilize thinking traces for safety-critical decision making.
ECHO: Terminal Agents Learn World Models for Free
Vaishnavi Shrivastava ⋅ Ahmed Awadallah ⋅ Dimitris Papailiopoulos
CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned stream---stdout, errors, file contents---records the consequences of its actions. We argue that this stream is a supervision signal, but standard agent RL largely discards it: GRPO-style training updates action tokens while ignoring environment responses, even though rollouts already contain them. We introduce ECHO (Environment Cross-entropy Hybrid Objective), an auxiliary loss that trains the policy to predict observation tokens alongside the policy gradient. ECHO reuses the same forward pass, requires no additional rollouts, and turns terminal feedback into a dense signal that remains informative even on failed trajectories. This objective shapes the policy toward a world model of terminal dynamics, plausibly improving its priors over which actions are likely to succeed. On TerminalBench 2, ECHO nearly doubles pass@1: Qwen3-8B improves from 2.70% to 5.17%, and Qwen3-14B from 5.17% to 10.79%. From base Qwen3-8B, ECHO+GRPO matches expert-SFT-then-GRPO on held-out terminal tasks without the expert demonstrations required for SFT, and closes roughly half this gap on TerminalBench 2. Without verifier rewards, ECHO can lead to self-improvement through exploration alone, suggesting environment observations are not merely context for actions but supervision for the policy itself.
Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts
Seungmin Oh ⋅ Seunghun Kang ⋅ Jongbin Ryu
Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose **Test-Time Correction (TTC)**, a hypothesis-based correction framework guided by a simple principle: *hypothesize, reconstruct, correct*. Given a test feature and its top-$k$ candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting **divergence shift**. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This parameter-free candidate verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to $2\times$ speedup and over $3\times$ lower memory usage.
Elastic Representations via Hyperbolic Geometry
Arjun Ramesh Kaushik ⋅ Rudrasis Chakraborty ⋅ Nalini Ratha ⋅ Venu Govindaraju
Learned representations (or embeddings) are the foundation of modern machine learning systems, yet they are typically trained as fixed-size embeddings, without accounting for varying downstream resource or task constraints. Matryoshka Representation Learning (MRL) addresses this limitation by learning nested representations, but requires costly end-to-end retraining and does not support post hoc expansion, while Contrastive Sparse Representations (CSR) reduce training overhead via lightweight adaptors but operate at a 4x larger dimensionality. Additionally, all prior works restrict their framework towards compression only. In this work, we introduce \textbf{H}yperbolic \textbf{E}lastic \textbf{R}epresentation \textbf{L}earning (HERL), a framework for learning dimension-adaptive embeddings (compression and expansion) through a lightweight adaptor without retraining the base encoder. Our key idea is to exploit the exponential capacity of hyperbolic geometry by learning a non-linear mapping from the encoder space to hyperbolic space. We train a simple MLP to downsample the representations before projecting them onto a Lorentzian and learn adaptive representations that respect the Lorentzian geometry. Our approach is model-agnostic and effective in both supervised and unsupervised settings. Through extensive experiments across image and text modalities, we show that HERL preserves downstream performance across a wide range of embedding sizes, exhibits smooth interpolation between dimensions, and achieves these benefits at a fraction of the computational cost.
Embedding Security Properties into AI-Enabled Cyber-Physical Systems
Ziyan An ⋅ John Stankovic ⋅ Meiyi Ma
AI-enabled Cyber-Physical Systems (CPS) are highly vulnerable to adversarial and anomalous inputs, where small perturbations can induce cascading errors and unsafe control actions. Existing approaches—such as rule-based filtering, training-time regularization, or diffusion-based reconstruction—either operate outside the model or lack mechanisms to incorporate formal security specifications into the prediction process. In this paper, we develop the first work toward embedding security properties directly into AI-enabled CPS, enabling predictive models to enforce system-level constraints during inference rather than relying on external defenses. We introduce a logic-conditioned bi-stage diffusion framework that integrates Signal Temporal Logic (STL) specifications into forecasting. STL serves as a first-class conditioning signal that guides both an input repair stage and an output refinement stage, allowing the model to jointly mitigate adversarial perturbations and enforce desired temporal behaviors to satisfy security-critical properties. We evaluate our approach on a real-world multivariate CPS forecasting task under both physical sensor and cyber attacks. Our method improves robustness and specification compliance, degrades more gracefully as attack strength increases, and generalizes better to unseen attacks. Ablation studies further show that embedding logical security properties yields gains unattainable by reconstruction-based methods alone, highlighting a new direction for integrating formal methods with generative models in secure CPS.
In autoregressive language models, each token is sampled by conditioning on all the past tokens; the overall string has thus been sampled from the correct underlying joint distribution represented by the model. In contrast, masked diffusion language models generate text by unmasking tokens out of order and potentially in parallel. Generating an overall string sampled from the correct underlying joint distribution would (again) require exactly one token unmasking in every full-model forward pass. The more tokens unmasked in parallel, the further away the string is from the true joint; this can be seen in the resulting drop in accuracy (but, increase in speed). In this paper we devise a way to {\em approximately} sample multiple tokens from the joint distribution in a single full-model forward pass; we do so by developing a new lightweight single-layer ``sampler" on top of an existing large diffusion LM. One forward pass of the full model can now be followed by multiple forward passes of only this sampler layer, to yield multiple unmasked tokens. Our sampler is trained to mimic exact joint sampling from the (frozen) full model. We show the effectiveness of our approximate joint sampling for both pretrained-only (Dream-7B-Base, Llada-7B-Base) and instruction-tuned (Dream-7B-Instruct, Dream-7B-Coder) models on downstream tasks (GSM8k, MATH, MBPP, HEval) and on unconditional language modeling. When eight tokens are unmasked for each full-model denoising step, our sampling algorithm achieves a MAUVE score of 0.84 (vs marginal baseline of 0.19) with respect to the true joint distribution.
Selective classification (SC) improves model reliability by allowing a classifier to abstain on low-confidence inputs. In the practical post-hoc setting, where the base classifier is fixed and often accessed as a black box, many confidence scores have been proposed, yet no single score is consistently best across models, datasets, and deployment shifts. We address this limitation by proposing, to our knowledge, the first ensemble framework for post-hoc SC, which learns to combine a diverse set of existing confidence scores using a small labeled calibration set drawn from the deployment distribution. We cast this problem as learning to separate correct from incorrect predictions of the base classifier, and train the ensemble selector with SoftRank-AURC, a differentiable soft-rank plug-in objective for the area under the risk-coverage curve (AURC), the standard evaluation metric for SC. We also study hinge and SELE losses as surrogate training variants, and provide an oracle-style guarantee showing that the SoftRank-AURC ensemble is competitive with the best single transformed base score in hindsight. Across 15 image and NLP classification tasks, with additional LLM multiple-choice experiments, the proposed method is consistently strong and often outperforms the best individual score, particularly under distribution shift, establishing ensemble post-hoc SC as a simple and effective way to improve reliability without retraining the base classifier.
Entropy Distribution as a Fingerprint for Hallucinations in Generative Models
Mattia Jacopo Villani ⋅ Pranav Deshpande ⋅ Akshay Seshadri ⋅ Romina Yalovetzky ⋅ Niraj Kumar
Large Language Models (LLMs) often generate factually incorrect outputs, commonly termed hallucinations, that undermine trust and limit deployment in high-stakes settings. Existing hallucination detection methods typically require multiple forward passes, or access to model internals. In this work, we provide theoretical background and empirical evidence that the distribution of token-level entropies, beyond the mean captured by perplexity or length-normalised entropy, serves as a fingerprint of hallucination, with distributional shape and tail behaviour carrying substantial independent signal. We formalize hallucination detection as a statistical hypothesis test and propose the Calibrated Entropy Score (CES), a lightweight algorithm requiring only a single forward pass and black-box access to token logits. CES combines the mean signal with the maximum signal of the generated entropy through a calibrated reference CDF, producing scores that are directly comparable across models and tasks. We establish finite-sample calibration guarantees via a novel random-length Dvoretzky--Kiefer--Wolfowitz inequality, and also prove that CES detects hallucinations with probability converging to one exponentially fast in the generation length. Across eight QA benchmarks and ten generator models spanning open-source and API access models, CES achieves the highest detection performance among all single-pass black-box methods while providing formal error guarantees that existing heuristics lack. Remarkably, CES is statistically indistinguishable from multi-sample methods that require far greater computational cost, closing the gap between lightweight and expensive detection and making it suitable for real-time, large-scale deployment.
Inference-time search (e.g., Tree of Thoughts, ToT) with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored—a failure mode we call \textit{reasoning basin collapse}. We introduce \textsc{BASIN}, a training-free selection method that penalizes repeated visits to the same reasoning basin, defined symbolically for arithmetic tasks or semantically via hypothesis clustering, thereby reallocating search across diverse reasoning strategies. Under matched inference budgets, \textsc{BASIN} improves over ToT by up to $+22$pp on Game of 24. To explain why this quality-agnostic penalty is effective, we introduce the \emph{redundancy gap} $\Delta$, which measures how differently search concentrates for correct versus incorrect predictions: standard ToT often operates near $\Delta \approx 0$, failing to distinguish correct from incorrect searches by their visit patterns, while \textsc{BASIN} consistently induces $\Delta > 0$, concentrating search on correct basins while dispersing incorrect ones. More broadly, \textsc{BASIN} suggests structure-aware selection as a simple and general approach to improving inference-time reasoning.
Evaluating an Evaluation: Membership Inference Attacks as Machine Unlearning Diagnostics
Umid Suleymanov ⋅ Laman Aliyeva ⋅ Nihat Abdullayev ⋅ Saida Zarbiyeva ⋅ Murat Kantarcioglu
Evaluation practices in machine learning are increasingly objects of scientific study in their own right: their assumptions, statistical properties, and failure modes determine which research conclusions can be trusted. We audit one such practice - the use of aggregate-subset Membership Inference Attacks (MIAs) for evaluating machine unlearning - across 8 dataset-architecture pairs, 4 unlearning methods, and over 500 experimental runs, identifying five failure modes: (1) limited and regime-dependent identifiability between baseline and retrained models, with KS-test indistinguishability ($p > 0.05$) in $\approx$42\% of experiment--task pairs; (2) regime-dependent signal-to-noise collapse; (3) monotonicity violations in 5 of 8 experiments under graded forgetting; (4) inter-task rank inconsistency, and (5) regime confounding by model training quality. To quantify these failures, we develop the Membership Inference Attack Unlearning Score (MIAU), a normalized diagnostic framework integrating three MIA comparisons as gap closure fractions between baseline and retrained references. We formalize four validity properties any unlearning metric should satisfy. Our findings indicate that aggregate-subset MIAs are reliable unlearning diagnostics only in narrow high-memorization regimes, and motivate complementary evaluation paradigms, alongside MIA-based protocols.
FaceProbe: Recovering HDR Environment Map via Masked Diffusion and Physical Preference
Peng Zhao ⋅ William Jing ⋅ Juyan Ba ⋅ Kairui Feng ⋅ Xuanhong Chen
Recovering $360^\circ$ HDR environment maps from unconstrained portraits is a fundamental yet ill-posed challenge in inverse rendering. Existing approaches typically suffer from a critical trade-off: traditional facial reflection-based methods yield low-fidelity, blurry results that lack practical utility, while recent generative models often resort to unconstrained panoramic hallucination (i.e., synthesizing environments with no physical grounding), which leads to severe overfitting and inaccurate lighting. In this paper, we present FaceProbe, an inpainting-driven framework that treats the human face as a reliable, physically-grounded light probe. Instead of synthesizing the environment from scratch, we reformulate the task as a masked completion and outpainting problem. By preserving the peripheral context of the input image and utilizing the portrait as a structural condition, our architecture leverages the generative priors of Diffusion Transformers (DiT) to extend sparse observations into high-fidelity, sharp $360^\circ$ HDR panoramas. To further eliminate color shifts and highlight inaccuracies, we introduce OLAT-GRPO, a post-training alignment mechanism based on Group Relative Policy Optimization. By subjecting the generated maps to image-based relighting on One-Light-At-a-Time (OLAT) datasets, we evaluate their physical validity through relighting consistency. This allows us to explicitly prune inconsistent denoising paths and align the model with multi-objective rewards, ensuring energy conservation and spectral accuracy. Extensive evaluations demonstrate that FaceProbe achieves a paradigm-shifting improvement. Most notably, in the critical portrait relighting task, our method reduces estimation errors by 54.4\% in Angular Error and 48.5\% in si-RMSE over state-of-the-art methods, delivering a robust and highly usable solution for photorealistic rendering and world model applications.
Fair Division of Work in Collaborative Mean Estimation via Bargaining
Michael O. Harding ⋅ Alex Clinton ⋅ Kirthevasan Kandasamy
Data collection is a critical component of modern machine learning pipelines. In many settings, multiple agents, such as hospitals or research labs, can collaborate to share the burden of data collection rather than each bearing it alone. But how should the work be divided among these agents fairly? Studying this question is complicated by heterogeneity in agents' data collection costs and data quality, and by the fact that fairness itself is subjective and use-case-dependent. We study these challenges in the context of PAC mean estimation, where agents wish to estimate an unknown scalar $\mu$ to within accuracy $\epsilon$ with probability at least $1-\delta$, via samples from distributions with common mean $\mu\,$. Agents may incur different costs to collect their samples, and their distributions may have different variances (data qualities). Our first contribution is a framework that casts collaborative data collection as a bargaining problem, where agents specify a notion of utility (e.g., negative cost incurred, or cost savings relative to working alone) and a welfare function over agent utilities. By maximizing the welfare over the feasible set of data collection amounts (i.e. there is enough data to satisfy $(\epsilon,\delta)$-PAC mean estimation), we obtain a fair division of work. Drawing on classical fairness axioms from the microeconomics literature, the family of admissible welfare functions takes a specific one-parameter form that includes utilitarian, egalitarian, and Nash bargaining as special cases. When agents' noise variances (data qualities) are known, this yields a clean characterization of the fair division. Our second contribution addresses the more challenging setting where variances are unknown. As the feasible set of data collection amounts itself depends on the variances, it is impossible to specify a fair division of work upfront. We propose an online algorithm that combines plug-in variance estimates with carefully calibrated forced sampling, and show that it is asymptotically optimal: the ratio of the achieved welfare to the welfare of the true fair division converges to 1 almost surely as either $\epsilon$ or $\delta$ goes to 0. We further show that our estimator asymptotically exactly meets the specified error tolerance and failure probability.
Fairness Failure Modes of Multimodal LLMs
Canyu Chen ⋅ Anglin Cai ⋅ Joan Nwatu ⋅ Yale Li ⋅ Jessica Hullman ⋅ Rada Mihalcea ⋅ Kathleen McKeown ⋅ Manling Li
Although Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains, the fairness of their outputs is under-explored. Building on the BBQ language bias benchmark, we construct a new dataset MultiBBQ using attested social biases and AI-generated photorealistic images for controllable fairness evaluation of MLLMs in both visual-only and visual-language contexts. We propose two metrics Fairness Score and Bias Score and design an evaluation paradigm to address shortcut reasoning and data contamination challenges. Using comprehensive benchmarking, we diagnose four new Fairness Failure Modes of MLLMs. In particular, we discover that proprietary models may fail to conduct effective counter-bias reasoning in disambiguated contexts due to over-refusal, while open-source models are deficient in abstaining in ambiguous contexts. We also analyze how different input and model factors degrade fairness, demonstrate that MLLMs amplify bias over their backbone LLMs, and show the potential limited effectiveness of mitigation methods such as reasoning and fairness instruction. We release our code and dataset to facilitate further evaluations and the development of mitigation methods here.
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
Qingyue Zhao ⋅ Kaixuan Ji ⋅ Heyang Zhao ⋅ Quanquan Gu
*Kullback-Leibler* (KL) regularization is ubiquitous in reinforcement learning (RL) algorithms in the form of *reverse* or *forward* KL. Recent studies have demonstrated $\epsilon^{-1}$-type fast rates for RL under reverse KL regularization, in contrast to the standard $\epsilon^{-2}$-type sample complexity. However, for forward-KL-regularized objectives, existing statistical analyses are either not applicable or resulting in $\tilde{O}(\epsilon^{-2})$ slow rates. We take the first step towards addressing this problem via a streamlined analysis of forward-KL-regularized offline contextual bandits. We give the first $\tilde{O}(\epsilon^{-1})$ upper bounds in tabular and general function approximation settings, both under notions of *single-policy concentrability*. In particular, our convex-analytical pipeline unifies these settings by exploiting the pessimism principle in a novel way and completely bypasses the proof routines in previous works based on the mean-value theorem, which might be of independent interest. Moreover, we provide rate-optimal lower bounds, manifesting the tightness of our upper bounds in terms of statistical rates. Our lower bounds also demonstrate that the forward-KL-regularized sample complexity recovers the unregularized slow rate in the low-regularization regime, similarly to the reverse-KL regularization.
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
Zhengding Hu ⋅ Mingge Lu ⋅ Zhen Wang ⋅ Jixuan Ruan ⋅ Chang Chen ⋅ Zaifeng Pan ⋅ Yue Guan ⋅ Ruiyi Wang ⋅ Zhongkai Yu ⋅ Chao Zhang ⋅ Yufei Ding
LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottleneck. We identify that this cost comes from synchronized stage execution and imbalance inside each LLM-heavy stage. We present FlashEvolve, an efficient framework that replaces synchronized execution with asynchronous workers and queues, allowing different stages and steps to overlap. To handle data staleness introduced by asynchrony, FlashEvolve tracks artifact versions and applies different policies to update, discard, or patch stale artifacts. Unlike weight-space staleness in asynchronous RL, language-space staleness is inspectable and repairable: a stale artifact is not just delayed work, but readable evidence that the LLM can reflect on, revise, and turn into useful evolution signal. FlashEvolve further improves throughput and token efficiency with speculative stage completion and adaptive workflow control. On GEPA workloads, FlashEvolve improves proposal throughput by 3.5x on local vLLM and 4.9x on API serving over synchronous GEPA. The same design also applies to ACE and Meta-Harness. Our code repository is available at https://anonymous.4open.science/r/FlashEvolve-FEC7/.
FlashFFN: Multi-Head Decomposition Enables I/O-Aware Feed-Forward Network
Minshen Zhang ⋅ Xiang Hu ⋅ Jianguo Li ⋅ Wei Wu ⋅ Kewei Tu
The Feed-Forward Network (FFN) dominates the activation memory of modern Transformers because the intermediate tensor of width dff is materialized in HBM between the two matmul stages. FlashAttention-style I/O-aware tiling avoids the analogous cost in attention, yet has not been brought to FFN. We identify the reason as a hardware-geometry constraint: under one-pass tiled accumulation a fused FFN kernel must keep an output accumulator of width dmodel in on-chip SRAM, which is infeasible whenever dmodel >= 2048 on currently shipping datacenter GPUs. We formalize this as an SRAM-feasibility proposition and prove a corollary: multi-head decomposition is the simplest architectural transform that brings FFN inside this SRAM-feasible regime, because it makes the per-head accumulator width dh = dmodel/H a free parameter that can be driven below the GPU's SRAM cutoff. Naive multi-head FFN, however, blows up the per-head expansion ratio dff/dh as dmodel grows, drifting away from the well-validated SwiGLU 8/3 optimum and degrading quality at scale; we resolve this with a dense MoE-like sub-network correction that restores the optimal ratio while preserving SRAM-feasibility. The resulting architecture, FlashFFN, is an SRAM-feasible FFN block that replaces SwiGLU under matched parameter and FLOPs budgets and derives directly from the SRAM constraint. Across 128M--1.3B parameter models trained on 60--100B tokens of The Pile, FlashFFN achieves a Pareto improvement over SwiGLU at fixed compute: up to ~4x lower peak activation memory at long sequences, inference latency that matches or improves over SwiGLU (up to ~1.32x faster on H100), and consistent quality gains (up to -0.84 in evaluation perplexity and +1.60 average accuracy across six downstream benchmarks at 1.3B)
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
Elliot Epstein ⋅ Rajat Vadiraj Dwaraknath ⋅ John Winnicki
Score-debiased kernel density estimation (SD-KDE) achieves improved asymptotic convergence rates over classical KDE, but its use of an empirical score has made it significantly slower in practice. We show that by re-ordering the SD-KDE computation to expose matrix-multiplication structure, Tensor Cores can be used to accelerate the GPU implementation. On a 32k-sample 16-dimensional problem, our approach runs up to $47\times$ faster than a strong SD-KDE GPU baseline and $3{,}300\times$ faster than scikit-learn's KDE. On a larger 1m-sample 16-dimensional task evaluated on 131k queries, Flash-SD-KDE completes in $2.3$ s on a single GPU, making score-debiased density estimation practical at previously infeasible scales.
FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse
Lingzhi Yuan ⋅ Chenghao Deng ⋅ Fangxu Yu ⋅ Souradip Chakraborty ⋅ Mohammad Rostami ⋅ Furong Huang
Large Language Model (LLM)-based multi-agent systems are increasingly powerful, but current agentic workflow optimization paradigms make an unsatisfying trade-off. Task-level methods spend substantial offline compute yet deploy only a single workflow, leaving complementary candidates unused, while query-level methods synthesize a new workflow for each query at substantial inference cost. Our motivating analysis shows that these paradigms are more complementary than competing: workflows discovered during offline search often solve different subsets of queries, and many queries handled by expensive query-level generation can already be solved by cheaper precomputed workflows. This suggests a different objective: rather than searching for one universally best workflow or regenerating a workflow for every instance, we should build a compact bank of reusable, complementary workflows and select among them adaptively at inference time. Doing so requires solving three coupled problems: generating complementary rather than redundant candidates, compressing them into a small deployable portfolio, and assigning each query to the right workflow under a performance-cost trade-off. To this end, we present FlowBank, a three-stage framework for portfolio-based agentic workflow optimization. Diversifying proposes DiverseFlow to steer search toward under-covered queries and produce a high-coverage candidate pool. Curating proposes CuraFlow to compress this pool into a compact portfolio with minimal redundancy. Matching casts deployment as edge-value prediction on a query-workflow bipartite graph and routes each incoming query to the portfolio member with the best predicted utility. Across five benchmarks, FlowBank achieves the highest average score among the evaluated methods while remaining cost-competitive, improving over the strongest automated and handcrafted baselines by 4.26\% and 14.92\% relative, respectively.
Free energy Estimation on Any State Space
Jiajun He ⋅ Zijing Ou ⋅ Francisco Vargas ⋅ Yingzhen Li ⋅ José Miguel Hernández-Lobato ⋅ Carles Domingo i Enrich ⋅ Yuanqi Du
Free energy estimation is a fundamental yet challenging problem, from physics to statistics. Classical approaches rely on thermodynamic transformations, ranging from direct estimation, quasistatic integration, to finite-time averaging. Recent work learns neural transports to significantly accelerate the efficiency in the finite-time regime. In this paper, we generalize this framework to arbitrary state spaces. Building on this view, we develop a generalized neural transport learning approach for efficient estimation. Experiments validate the effectiveness and efficiency of the proposed method beyond continuous settings, extending to discrete and multimodal spaces as well as autoregressive settings.
From Nodes to Pixels: Topological and Structural Two-View Graph Imaging
Md Joshem Uddin ⋅ Soham Changani ⋅ Sai K Navuluru ⋅ Cuneyt Akcora ⋅ Baris Coskunuzer
Graph learning still lacks a broadly reusable input interface comparable to patches in vision or tokens in language. Unlike images or text, graphs are irregular, vary in size and density, and are invariant under many equivalent node labelings, making it difficult to standardize inputs across tasks and architectures. We introduce G2Image, a fixed-budget graph imaging framework that converts each graph into two complementary image-like views. TopoGrid provides a stable topological view by encoding an intrinsic bifiltration as a compact multipersistence surface grid, while GraphGrid provides a higher-bandwidth structural view by aggregating edge densities between groups defined by intrinsic node scores. Both views are permutation-invariant and come with resolution-controlled stability guarantees. Because the resulting representations are fixed-size tensors, they can be processed by lightweight 2D encoders and aligned through a supervised cross-view contrastive objective. Across graph classification and molecular prediction benchmarks, G2Image achieves strong average performance among the compared methods, improves over single-view and standard fusion variants, and extends to attribute-rich molecular graphs without changing the overall architecture. These results suggest that fixed-budget graph images offer a practical and reusable interface for graph representation learning.
From Zero to Hero: Training-Free Custom Concept Spawning in World Models
Kiymet Akdemir ⋅ Pinar Yanardag
Autoregressive world models have emerged as a powerful paradigm for interactive video generation, allowing users to navigate dynamically generated environments through actions. These models are typically conditioned on a text prompt and/or a single reference frame, from which the entire world is generated. Yet the moment the user navigates beyond what is visible in that frame, the unseen regions are populated by the base model's priors, with no mechanism for the user to specify what should appear and where. This is a fundamental limitation for applications such as gaming, interactive storytelling, and simulation, where controllable scene composition is essential. We refer to this missing capability as concept spawning; introducing a user-specified visual concept into a world model, analogous to spawning in a game engine. We introduce SPAWN (Swapping Pinned Anchor with Windowed iNjection), a training-free method for concept spawning. SPAWN exploits a structural property of image-to-video backbones: the first slot of the context memory is pinned to the reference frame and acts as a foundational anchor for every generated chunk. By swapping this anchor with an external concept latent over a short injection window and letting the original anchor return, we cause the concept to propagate naturally through the rollout via the model's own memory. SPAWN supports concepts from fine-grained entities such as characters and props to large-scale elements such as buildings and landmarks, and accepts either a concept image or a text description as input. Experiments show that SPAWN integrates concepts with consistent lighting, scale, and perspective while preserving identity and temporal coherence, demonstrating that controllable concept spawning is achievable in existing autoregressive world models without any training.
GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction
Sheng Zhao ⋅ Weikai Lin ⋅ Yuhao Zhu
Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.
Generalized Smooth Stochastic Variational Inequalities: Almost Sure Convergence and Convergence Rates
Daniil Vankov ⋅ Angelia Nedich ⋅ Lalitha Sankar
This paper focuses on solving a stochastic variational inequality (SVI) problem under relaxed smoothness assumption for a class of structured non-monotone operators. The SVI problem has attracted significant interest in the machine learning community due to its immediate application to adversarial training and multi-agent reinforcement learning. In many such applications, the resulting operators do not satisfy the smoothness assumption. To address this issue, we focus on a weaker generalized smoothness assumption called $\alpha$-symmetric. Under p-quasi sharpness and $\alpha$-symmetric assumptions on the operator, we study clipped projection (gradient descent-ascent) and clipped Korpelevich (extragradient) methods. For these clipped methods, we provide the first almost-sure convergence results without making any assumptions on the boundedness of either the stochastic operator or the stochastic samples. We also provide the first in-expectation unbiased convergence rate results for these methods under a relaxed smoothness assumption for $\alpha \leq 1/2$.
Generalizing Action-Conditioned Latent World Models with Video Model Rewards
Haichao Zhang ⋅ Yijiang Li ⋅ Shwai He ⋅ Van T Le ⋅ Nuno Vasconcelos ⋅ Yun Fu
JEPA-style latent world models offer an efficient alternative to pixel-space video generation by predicting future representations rather than synthesizing future frames. However, current action-conditioned predictors in such latent world models are typically trained on single or narrowly scoped embodied datasets with limited task, object, embodiment, and action-space coverage, and often fail to generalize beyond the training distribution. We argue that this lack of generalist action-conditioned prediction is a key bottleneck for scaling latent world models. In contrast, large controllable video world models have a clear scale-up path through large-scale video pretraining and encode broad visual-dynamics priors, but they are too computationally expensive to serve as online simulators and lack a direct interface to latent action-conditioned prediction. We introduce GeneralistJEPA, a framework for transferring generalist motion priors from video world models to efficient action-conditioned JEPA predictors, establishing a scalable training path for future latent world models. Instead of directly distilling generated futures, whose latents entangle useful action dynamics with source-mismatched content and generation artifacts, GeneralistJEPA factorizes future latent prediction into source-conditioned content continuation and action-induced dynamics. Controllable video world models are used as same-action environment randomizers to build an offline motion-reward bank, providing diverse motion supervision without treating generated content as an absolute target. Specifically, generated rollouts are encoded with a frozen JEPA encoder, and only their patch-level temporal deltas are distilled as video-world-model motion rewards for predicted latent dynamics. This design preserves the efficiency of latent rollout, since video generation and representation encoding are performed offline, while leveraging the broader visual-dynamics priors of video world models. We evaluate GeneralistJEPA on three complementary benchmarks: EgoDex for held-out subtask generalization and action recoverability, BAIR robot pushing for future-frame latent prediction and action sensitivity, and Physion for physical contact readout. Our results show that video-world-model motion rewards improve latent prediction, action sensitivity, action recoverability, and physical contact readout beyond real-only training and naive generated-latent distillation, addressing a key generalization bottleneck on the path toward scaling up action-conditioned latent world models.
Conformal prediction provides model-agnostic uncertainty quantification with guaranteed coverage, but conventional methods often yield overly conservative uncertainty sets, particularly in multimodal or heterogeneous settings. This inefficiency arises from two sources: (i) limited expressiveness of the predictive model and (ii) simplistic nonconformity scores design. Most existing approaches advance only one of these axes, leaving the other underexplored. We propose generative conformal prediction with Optimized Ranking and Coverage Allocation (ORCA), a three-stage framework that advances both aspects jointly. ORCA leverages generative models to capture the full conditional distribution and introduces a rank-dependent optimization procedure that adaptively allocates coverage for efficiency while maintaining validity. We cast this coverage allocation as an optimization problem, derive an exact mixed-integer linear programming formulation, and show that the solution converges asymptotically to the oracle density-level set. Across synthetic, semi-synthetic, and real datasets, ORCA produces substantially more efficient uncertainty sets than state-of-the-art baselines, demonstrating robust gains in scenarios where conventional conformal prediction methods fail.
The recent breakthrough Score-Repellent Monte Carlo (SRMC) method improves MCMC sampling by tilting the target density $\pi$ in $\mathbb{R}^d$ with the factor $\exp(-\alpha\,\theta_n^\top s(x))$, where the score function $s(x) = \nabla_x \log \pi(x)$, and $\theta_n$ is the score average from past samples. However, a single scalar repellence strength $\alpha\ge0$ cannot handle both the steep and the flat directions in anisotropic energy landscapes at once. We introduce Geometry-Aware SRMC (GA-SRMC), which generalizes the Euclidean alignment to $\theta^\top M s(x)$ for any symmetric positive-definite matrix $M$. A stochastic-approximation central limit theorem under a trace constraint on $M$ identifies the trace-normalized inverse Fisher $M_{\rm opt}\propto S^{-1}$ as the unique minimizer of the worst-direction covariance bound, elevating the Fisher information $S=\text{Cov}_\pi(s,s)$ from a proposal-side preconditioner (e.g., FisherMALA, natural gradient) to the optimal repulsion-side matrix, with a single online estimate $\widehat S_n$ serving both roles at no additional cost. On an $8$-target simulation spanning the condition number $\kappa\in[1.3,100]$ and dimension $d\in[2,785]$, GA-SRMC dominates SRMC by $4.4\times$ on anisotropic Gaussian ($\kappa=100$) and $6.8\times$ on MNIST under the MALA baseline, by $18\text{ - }32\\%$ on the well-conditioned Bayesian logistic-regression posteriors, and by up to $+48\\%$ under FisherMALA; the advantage scales monotonically with the condition number $\kappa$.
Geometry over Density: Few-Shot Cross-Domain OOD Detection
Li Li ⋅ You Qin ⋅ Jiate Li ⋅ Charith Peris ⋅ Lisa Bauer ⋅ Roger Zimmermann ⋅ Yue Zhao
Out-of-distribution (OOD) detection identifies test samples that fall outside a model's training distribution, a capability critical for safe deployment in high-stakes applications. Standard OOD detectors are trained on a specific in-distribution (ID) dataset and detect deviations from that single domain. In contrast, we study few-shot cross-domain OOD detection: given a \emph{single} pre-trained model, can we perform OOD detection on \emph{arbitrary} new ID-OOD task pairs using only a handful of ID samples at inference time, with no additional training? We propose \textbf{UFCOD}, a unified framework that achieves this goal through information-geometric analysis of diffusion trajectories. Our key insight is that diffusion noise predictions are score functions (gradients of log-density), and we extract two energy features: \emph{Path Energy} (integrated score magnitude) and \emph{Dynamics Energy} (score smoothness), that form a discrete Sobolev norm capturing how samples interact with the learned diffusion process. The central contribution is a \textbf{train-once, deploy-anywhere} paradigm: a diffusion model trained on a single dataset serves as a universal feature extractor for OOD detection across semantically unrelated domains. At deployment, each new task requires only $\sim$100 unlabeled ID samples for inference: no retraining, no fine-tuning, no task-specific adaptation. Using 100 ID samples per task, UFCOD achieves 93.7\% average AUROC across 12 cross-domain benchmarks, competitive with methods trained on 50k--163k samples, demonstrating $\sim$500$\times$ improvement in sample efficiency. See our code in \url{https://anonymous.4open.science/r/UFCOD-22E8/README.md}.
Global Optimality for Constrained Exploration via Penalty Regularization
Florian Wolf ⋅ Ilyas Fatkhullin ⋅ Niao He
Efficient exploration is a central problem in reinforcement learning and is often formalized as maximizing the entropy of the state-action occupancy measure. While unconstrained maximum-entropy exploration is relatively well understood, real-world exploration is often constrained by safety, resource, or imitation requirements. This constrained setting is particularly challenging because entropy maximization lacks additive structure, rendering Bellman-equation-based methods inapplicable. Moreover, scalable approaches require policy parameterization, inducing non-convexity in both the objective and the constraints. To our knowledge, the only prior model-free policy-gradient approach for this setting under general policy parameterization is due to Ying et al. (2025). Unfortunately, their guarantees are limited to weak regret and ergodic averages, which do not imply that the final output is a single deployable policy that is near-optimal and nearly feasible. In this work we take a different approach to this problem, and propose Policy Gradient Penalty (PGP) method, a single-loop policy-space method that enforces general convex occupancy-measure constraints via quadratic-penalty regularization. PGP constructs pseudo-rewards that yield gradient estimates of the penalized objective, subsequently exploiting the classical Policy Gradient Theorem. We further establish the regularity of the penalized objective, providing the smoothness properties needed to justify the convergence of PGP. Leveraging hidden convexity and strong duality, we then establish global last-iterate convergence guarantees, attaining an $\epsilon$-optimal constrained entropy value with $\epsilon$-bounded constraint violation despite policy-induced non-convexity. We validate PGP through ablations on a grid-world benchmark and further demonstrate scalability on two challenging continuous-control tasks.
GlucoFM: A Dual-Stream Foundation Model for Continuous Glucose Monitoring
Zechen Li ⋅ Keerthana Natarajan ⋅ Weizhi Zhang ⋅ Simon Lee ⋅ Yuwei Zhang ⋅ Max Xu ⋅ Menglian Zhou ⋅ Zeinab Esmaeilpour ⋅ Flora Salim ⋅ Mark Malhotra ⋅ Lindsey Sunden ⋅ Shwetak Patel ⋅ Yuzhe Yang ⋅ Ahmed Metwally
Continuous glucose monitoring (CGM) provides a dense window into metabolic physiology, yet existing generic time-series and CGM-specific foundation models typically learn from entangled glucose sequences without explicitly capturing the temporal structure of glycemic dynamics. We present GlucoFM, a lightweight CGM foundation model that aligns irregular recordings to a 24-hour chronological grid, preserves observation masks, and decomposes glucose dynamics into slow physiological state and transient event streams, capturing low-frequency glycemic baselines and short-term deviations that may reflect acute physiological responses or sensor artifacts. GlucoFM is pretrained on 109,066 hours of unlabeled CGM recordings from 477 subjects with two complementary objectives: masked contextual latent prediction over fused daily representations and temporal dynamics prediction over state and event streams. Across four diverse cohorts and seven clinical prediction tasks, GlucoFM achieves the strongest frozen linear-probing performance among evaluated baselines, improving average PR-AUC by 4.1 points over the best CGM-specific foundation model. Its gains are most pronounced on core metabolic outcomes, leading PR-AUC on all diabetes-risk and $\beta$-cell dysfunction tasks and on 3 of 4 insulin-resistance tasks. GlucoFM also achieves the best overall cross-dataset transfer performance among evaluated methods and strong few-shot adaptation, highlighting physiology-aware decomposition as an effective inductive bias for transferable CGM representation learning.
Merging different models into a single model is highly desirable for multi-task deployment, particularly in regulated domains like healthcare. In such settings, models are often locally fine-tuned due to strict data privacy policies, necessitating that aggregation operates exclusively on the fine-tuned weights. Existing meth- ods apply a uniform merge across all tasks, ignoring the geometric relationships among their learned subspaces. When task updates occupy misaligned subspaces, projecting them onto a shared low-rank basis discards task-relevant directions, a failure mode named subspace interference. Through this study, we show that the per-task projection error of a merged model is governed by the principal-angle alignment between each task and the shared subspace, so that grouping geometri- cally compatible tasks provably reduce this error within each group. Motivated by this observation, we propose Group-wise Rank-Aware Model Merging (GRAM), which uses the principal-angle geometry of task vectors to decide which tasks should be merged and then performs a standard SVD-based merge within each coalition; the pipeline requires only the fine-tuned weights at merge time. With only one additional model (K=2), GRAM significantly mitigates subspace in- terference and consistently outperforms prior merging methods across language understanding (GLUE with Flan-T5), 8-task CLIP vision, and medical instruction tuning (Qwen2-7B-Instruct with four English/Chinese medical LoRAs).
Graph Anomaly Detection as Dynamical Transport: Training-Free Scoring via Empirical Bayes
Fred Xu ⋅ Thomas Markovich ⋅ florence regol ⋅ Yizhou Sun
Node-level graph anomaly detection (GAD) identifies nodes whose attributes and interactions deviate from dominant graph regularities. Existing GAD models usually encode normality and anomaly scoring indirectly through architectures, message passing, reconstruction or contrastive objectives, and tuned score families. This entangles graph trust (how strongly graph structure should define normality), graph-spectral weighting, and anomaly-score choice, yielding scores that are costly, opaque, and unstable across graph regimes. We propose EB-GAD (Empirical-Bayes GAD), a training-free framework that models normality as graph-aware generalized Ornstein--Uhlenbeck (GOU) relaxation toward a graph-filtered template. Empirical Bayes fits the graph precision and template from the residual-field likelihood; the GOU then turns scoring into closed-form transport cost from a feature-neutral node to its observed endpoint along graph-spectral relaxation. Sweeping relaxation horizon and endpoint tolerance yields a \emph{transport bank}: equilibrium Mahalanobis scoring is one limit, while finite-horizon transport-energy and scale-normalized ratio scores reveal anomalies that static equilibrium scoring can mask. A label-free selector chooses the anomaly-score family from feature homophily, edge density, and feature dimension, then ranks candidates by fitted-null KS distance and rank-stability across neighboring transport configurations. On 11 benchmarks including financial fraud networks with up to 3.7M nodes, EB-GAD achieves the best or tied-best AUROC on 9 of 11 datasets; on the largest graph, eigendecomposition and scoring finish in about six minutes.
Grouped Adaptive Head Mixing for Personalized Multi-Task Federated Reinforcement Learning
Yiran Pang ⋅ Zhen Ni ⋅ Dimitris Pados ⋅ Xiangnan Zhong
Multi-task reinforcement learning (MT-RL) trains a single agent to solve multiple tasks by leveraging shared knowledge across task domains. However, most MT-RL methods assume centralized access to task data, making them impractical for privacy-sensitive or large-scale distributed settings. Multi-task federated reinforcement learning (MT-FRL) mitigates this issue by enabling agents to collaborate through shared models rather than raw trajectories, but often suffers from unstable performance under task heterogeneity, negative transfer, and imbalanced learning across tasks. A fundamental question in MT-FRL is how each agent should decide with whom to share skills and how strongly to rely on shared knowledge. To answer this question, we propose GAdapHeadFRL, a personalized MT-FRL framework based on geometry-aware grouping and adaptive decision-layer task-head mixing. First, we decompose the local Bellman residual into representation and task-head estimation errors, and derive a closed-form optimal sharing strength governed by task mismatch and estimation uncertainty. Second, we instantiate this insight by identifying compatible client groups from local updated geometry, constructing group-specific task-head prototypes, and dynamically mixing them with local heads using lightweight practical proxies. Experiments on heterogeneous benchmarks, including MiniGrid (up to 12 tasks) and MetaWorld (up to 50 tasks), show that GAdapHeadFRL consistently outperforms state-of-the-art personalized MT-FRL baselines. It achieves up to 24.5\% relative improvement on MiniGrid, and increases the converged-task ratio from 50.0\% to 72.9\% over the personalized baseline. In addition, the proposed method reduces performance standard deviation by over 82.9\% on the largest task settings compared with centralized MT-RL references. These results demonstrate that geometry-aware grouping and adaptive task-head mixing provide an effective and scalable principle for heterogeneous MT-FRL.
GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations
Jingbo Yang ⋅ Henry Lai ⋅ Xiaowen Wang ⋅ Shiyu Chang ⋅ Yaar Harari ⋅ Evgeniy Gabrilovich
Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems that extract, retrieve, and apply information across long-running conversations. However, both existing memory systems and benchmarks are built around the dyadic, single-user setup, even though real deployments routinely span groups and channels with multiple users interacting with the agent and with each other. This mismatch leaves three properties of group memory unmeasured: (i) group dynamics that go beyond concatenated one-on-one chats, (ii) speaker-grounded belief tracking, where the per-user memory modeling is needed, and (iii) audience-adapted language, where Theory-of-Mind shifts produce role-specific vocabulary. We introduce GroupMemBench, a benchmark that exposes all three. A graph-grounded synthesis pipeline produces multi-party conversations with controllable reply structure and conditions each message on per-user personas and target audiences. An adversarial query pipeline then binds every question to a specific asker across six categories, spanning multi-hop reasoning, knowledge update, term ambiguity, user-implicit reasoning, temporal reasoning, and abstention, and iteratively searches challenging, realistic queries that reflect comprehensive memory capability. Benchmarking leading memory systems exposes a sharp collapse: the strongest one reaches only 46.0% average accuracy, with knowledge update at 27.1% and term ambiguity at 37.7%, while a simple BM25 baseline matches or exceeds most agent memory systems. This indicates current memory ingestion erases the structural and lexical features group memory depends on, leaving multi-user memory far from solved.
Harder the Task, Sparser the Representation: Sparsity as a Learning Signature of Capability in LLMs
Mingyu Jin ⋅ Yutong Yin ⋅ Jingcheng Niu ⋅ Qingcheng Zeng ⋅ Wujiang Xu ⋅ Xinyuan Song ⋅ Wei Cheng ⋅ Mengnan Du ⋅ Zhaoran Wang ⋅ Tianlong Chen ⋅ Dimitris Metaxas
In this work, we investigate how the internal representations of Large Language Models (LLMs) change with inputs of increasing difficulty. Although sparsity changes may not be obvious for individual samples, dataset-level results reveal a consistent statistical trend: as task difficulty increases (e.g., harder questions, longer contexts, or more answer choices), the last hidden states of LLMs become systematically sparser. In short, \emph{the harder the task, the sparser the representation}. This sparsity--difficulty relation is observable across diverse models and domains, suggesting that harder inputs drive more concentrated activation patterns in the last hidden state. Through a series of controlled analyses and learning-dynamics experiments, we show that representational density is a learned property of data familiarity: models develop rich, distributed representations for mastered patterns, while unfamiliar inputs default to sparser activations. Our finding is also actionable. We illustrate this with Sparsity-Guided Curriculum In-Context Learning (SG-ICL), which uses sparsity to select few-shot demonstrations matched to the query's difficulty, outperforming standard CoT and Auto-CoT baselines on MATH-500. Our study provides new insights into how LLM representations reflect task difficulty and how this signal can be leveraged for inference.
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
Guoan Wang ⋅ Feiyu Wang ⋅ Zongwei Lv ⋅ Yikun Zong ⋅ Zhewen Tan ⋅ Tong Yang
Quantization-aware training (QAT) is a key approach for adapting large language models (LLMs) to extremely low-bit weights, where post-training quantization often suffers severe accuracy loss. However, most extremely low-bit QAT methods rely on the straight-through estimator (STE), which keeps hard round-and-clip weights in the forward pass while using an identity surrogate in the backward pass, creating a mismatch between discrete forward states and continuous latent-weight updates. As a result, small updates may fail to cross quantization boundaries, causing dead-zone stagnation in sub-optimal discrete states. To address this problem, we propose Hestia, an operator-faithful differentiable QAT framework for extremely low-bit LLMs. Instead of repairing the gradient of a hard quantizer, Hestia relaxes the quantizer itself as a temperature-controlled Softmax expectation over the same discrete codebook. The relaxation provides forward--backward consistent gradients during training and anneals to the target hard quantizer for inference. To stabilize this soft-to-hard transition, Hestia further uses offline Hessian traces as lightweight tensor-wise sensitivity signals for adaptive temperature scheduling. Experiments on Llama-3.2 show that Hestia consistently improves over ternary QAT baselines, with relative average zero-shot gains of 5.39% and 4.34% on 1B and 3B models, while adding only about 0.2% training-time overhead and no inference-time overhead. The code is available at https://github.com/hestia2026/Hestia.
Text-guided point-cloud grounding enables robots to ground natural-language descriptions in 3D environments, making it important for embodied AI and human-robot interaction. Existing coarse-to-fine methods primarily rely on global descriptors for submap retrieval, but these descriptors often compress object semantics, spatial relations, and scene layouts into a single vector, thereby limiting discriminability in large-scale scenes. We propose HiGraLoc, a multi-level cross-modal alignment framework that improves the coarse retrieval stage through three complementary branches. The instance branch uses a Hyperbolic Instance Structure Encoder to model hierarchical object semantics, the relation branch aggregates reliability-weighted pairwise spatial relations, and the global branch employs a Global Spectral Graph Encoder to capture multi-frequency scene structure. A Text2Loc-style coordinate regression stage then refines the location estimate within the retrieved submap. Experiments on KITTI360Pose show that HiGraLoc achieves a 19\% improvement in Top-1 recall@10m over existing state-of-the-art methods. Our code and dataset are available at https://github.com/Anonymous09871745/HiGraLoc.
HierFlow: Hierarchical Coupled Dual-Space Search for Automatic Agentic Workflow Generation
Dong Li ⋅ Yanchi Liu ⋅ Xujiang Zhao ⋅ Wei Cheng ⋅ Zhengzhang Chen ⋅ Xintao Wu ⋅ Zhong Chen ⋅ Chen Zhao ⋅ Haifeng Chen
Agentic AI systems enable LLMs to solve non-trivial tasks through structured workflows, but automatically generating such workflows remains challenging due to the discrete and combinatorial search space. Existing methods often rely on offline search or training, limiting query-level adaptability and incurring substantial data and engineering overhead. We formulate workflow generation as a coupled topology--execution search problem, where the upper-level topology induces subtask-specific code-search domains and lower-level execution feedback can revise the topology itself. Based on this formulation, we propose HierFlow, a training-free hierarchical test-time search framework for automatic agentic workflow generation. HierFlow couples feedback-driven topology refinement with an MCTS-inspired lightweight tree search for execution-level sub-workflow optimization, and uses an adaptive gating mechanism to selectively trigger execution-level search based on estimated necessity. We further provide a coupling-aware analysis characterizing when hierarchical decomposition and proxy-based gating are beneficial and when their advantages may degrade under stronger cross-subtask coupling. Experiments across QA, mathematical reasoning, and code generation benchmarks show that HierFlow achieves strong performance and favorable efficiency--quality trade-offs, outperforming competitive baselines without additional training.
High-dimensional Gaussian Graphical Model Testing for Long-Memory Time Series
Percy S. Zhai ⋅ Ping-Shou Zhong ⋅ Wei Biao Wu
Many real-world high-dimensional time series exhibit long-memory, but Gaussian graphical model testing in this regime remains understudied. We develop a direct, data-adaptive test statistic for assessing conditional independence in the graph structure of stationary Gaussian time series. We establish a finite-sample, Berry--Esseen type Gaussian approximation bound for the statistic, which applies to both short-memory and long-memory time series. The testing procedure is fully data-adaptive using block bootstrap method, on which we provide a finite-sample validity result including in the ultra-high-dimensional scenario, and can be extended to comparing graphical structures in two-sample tests. We also develop a consistency-empowered correction to the statistic and show that such tests attain asymptotic consistency in both size and power. Our proposed method is applied to a real-world fMRI data to understand functional connectivities within brain in different periods.
How does RL Post-training Induce Skill Composition? A Case Study on Countdown
Simon Park ⋅ Simran Kaur ⋅ Sanjeev Arora
While reinforcement learning (RL) successfully enhances reasoning in large language models, its role in fostering compositional generalization (the ability to synthesize novel skills from known components) is often conflated with mere length generalization. To this end, we study what RL post-training teaches models about skill composition and how the composition structure affects the skill trasnfer. We focus on the \countdown{} task (given $n$ numbers and a target, form an expression that evaluates to the target) and analyze model solutions using expression trees, where each subtree corresponds to a reusable subtask and thus can be viewed as a ``skill.'' Tracking tree shapes and their success rates over training, we find: (i) out-of-distribution (OOD) generalization to larger $n$ and to unseen tree shapes, indicating compositional reuse of subtasks; (ii) the order of skill acquisition depends upon the composition structure---models master shallow balanced trees (workload is balanced between subtasks) before deep unbalanced ones, with persistent fragility on right-heavy structures for a fixed composition depth. Our diagnostic reveals what is learned, in what order, and where generalization fails, clarifying how RL-only post-training induces OOD generalization beyond what standard metrics such as pass@k reveal. To study the implication of our findings for more real-world tasks, we present a preliminary study on formal theorem proving.
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
Shayan Hundrieser ⋅ Insung Kong ⋅ Johannes Schmidt-Hieber
We introduce Hyper Input Convex Neural Networks (HyCNNs), a novel neural network architecture designed for learning convex functions. HyCNNs combine the principles of Maxout networks with input convex neural networks (ICNNs) to create a neural network that is always convex in the input, theoretically capable of leveraging depth, and performs reliable when trained at scale compared to ICNNs. Concretely, we prove that HyCNNs require exponentially fewer parameters than ICNNs to approximate quadratic functions up to a given precision. Throughout a series of synthetic experiments, we demonstrate that HyCNNs outperform existing ICNNs and MLPs in terms of predictive performance for convex regression and interpolation tasks. We further apply HyCNNs to learn high-dimensional optimal transport maps for synthetic examples and for single-cell RNA sequencing data, where they oftentimes outperform ICNN-based neural optimal transport methods and other baselines across a wide range of settings.
Imperfect Influence, Reliable Rankings: A Theory of TRAK for Data Attribution
Han Tong ⋅ Shubhangi Ghosh ⋅ Haolin Zou ⋅ Arian Maleki
Data attribution, which traces a model’s prediction back to specific training data, is an important tool for interpreting sophisticated AI models. The widely used TRAK algorithm addresses this challenge by first approximating the underlying model with a kernel machine and then leveraging techniques developed for approximating the leave-one-out (ALO) risk. Despite its strong empirical performance, the theoretical conditions under which TRAK approximations are accurate, as well as the regimes in which they break down, remain largely unexplored. In this paper, we provide a theoretical analysis of the TRAK algorithm, characterizing its performance and quantifying the errors introduced by the approximations on which the method relies. We show that although the approximations can incur significant errors, TRAK preserves the separation between highly influential and weakly influential data points. We corroborate our theoretical results through extensive simulations and empirical studies.
Improved Leverage Score Sampling for Constrained Active Linear Regression
Aarshvi Gajjar ⋅ Syamantak Kumar ⋅ Aditya Makkar ⋅ Christopher Musco ⋅ Yihan Zhou
In this paper we study active learning for constrained linear regression. Given a design matrix $A \in \mathbb{R}^{n \times d}$, a constraint set $\mathcal C \subseteq \mathbb{R}^d$, and query access to a response vector $b \in \mathbb{R}^n$, we seek to output an approximate solution to the problem $\min_{x\in \mathcal{C}} \lVert Ax - b\rVert_2^2$ using as few queries to $b$ as possible. This problem arises in many scientific and engineering applications, and _randomized leverage score sampling_ provides a powerful method to solve it. However, existing leverage score methods either ignore the constraint or encode it only through ridge regularization, thereby suffering poor query complexity for anisotropic constraint sets. We propose _ellipsoid-regularized leverage scores_, where we first approximate the constraint set by an outer ellipsoid and then use its shape matrix to bias sampling toward directions that matter under the constraint geometry. We prove that the resulting query complexity is governed by a natural ``anisotropic effective dimension`` of the problem, $d_{\text{eff}}(\mathcal{C}) \leq d$. In particular, we show that $\tilde O(d_{\text{eff}}(\mathcal{C})/\varepsilon^2)$ label queries suffice to return a feasible point whose objective value is at most $(1+\varepsilon)$ times the optimum. This bound is typically much smaller than what would be obtained using standard leverage scores. Furthermore, for convex constraints, we show how to improve the dependence on $\varepsilon$ from $1/\varepsilon^2$ to $1/\varepsilon$. We complement these upper bounds with two lower bounds: first, a lower bound showing that the $1/\varepsilon^2$ dependence is unavoidable without convexity assumptions on $\mathcal C$, and second, a lower bound showing that the effective-dimension dependence is unavoidable for ellipsoidal constraints. We also provide experiments supporting our theoretical results.
Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updating model parameters. However, existing inference-time methods rely on a scalar reward model derived under a Bradley-Terry assumption, which cannot represent general preferences. Following recent work on fine-tuning with generalized preferences, in this work, we initiate the study of inference-time alignment under general preferences. We formulate the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies. We propose two algorithms: \emph{Best-of-Nash} (BoN) and \emph{Nash Mirror Descent} (NMD). We prove that both algorithms achieve a duality gap that matches the problem lower bound. Empirically, we implement the two methods on three datasets, which shows that our methods substantially outperform the base policy, converging to the performance of the fine-tuned models. Moreover, our results show that NMD remains robust across the regularization parameter.
INFUSER: Influence-Guided Self-Evolution Improves Reasoning
Siyu Chen ⋅ Miao Lu ⋅ Beining Wu ⋅ Heejune Sheen ⋅ Fengzhuo Zhang ⋅ Shuangning Li ⋅ Zhiyuan Li ⋅ Jose Blanchet ⋅ Tianhao Wang ⋅ Zhuoran Yang
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a Solver that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an optimizer-aware influence score that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly served by standard GRPO, we propose DuGRPO, a dual-normalized variant of GRPO, for generator training. Together, these turn the document pool into an adaptive curriculum that favors questions useful to the current solver, not just hard ones. On Qwen3-8B-Base, INFUSER outperforms strong self-evolution baselines with over 20% relative improvement on Olympiad and SuperGPQA benchmarks, and an 8B INFUSER co-evolving generator outperforms a frozen 32B thinking generator on math and coding. Ablations confirm each design choice is necessary. Code is available at
Inverse Reinforcement Learning with Just Classification and a Few Regressions
Lars van der Laan ⋅ Nathan Kallus ⋅ Aurelien Bibaut
Inverse reinforcement learning (IRL) seeks to recover a reward function from observed behavior, but rewards are typically only partially identified: many reward--value pairs can induce the same behavior policy. We study this problem in the maximum-entropy, or Gumbel-shock, model under a broad class of statewise affine normalization constraints, with anchor-action constraints as a special case. This leads to Generalized Policy-to-$Q$-to-Reward (GenPQR), a modular approach to normalized reward recovery based on policy estimation and $Q$-evaluation via the Bellman equation; both components can be instantiated with off-the-shelf classification and regression methods. We establish modular finite-sample guarantees under general function approximation, with separate terms for policy-estimation and $Q$-function estimation error. As a concrete instantiation, we study GenPQR with fitted $Q$-evaluation, reducing IRL to policy estimation followed by regression. In experiments, GenPQR improves on DeepPQR in reward recovery while remaining simpler and more modular. Relative to DeepPQR, our theory is broader: it goes beyond anchor actions, accommodates large action spaces, and is not tied to a specific neural-network architecture or training procedure.
Search-based planning is a powerful policy improvement paradigm in reinforcement learning, with milestone successes on discrete-action domains through the AlphaZero and MuZero families. Extending it to continuous control, however, faces a structural mismatch between the search operator and the policy class: search at each state evaluates a finite candidate set and returns a discrete distribution, while the parametric Gaussian policy commonly used in continuous control is a continuous density. We prove that distilling the search output into a Gaussian induces a non-vanishing projection error that inflates policy variance and produces a return gap that persists even when the policy mean is locally optimal, and is not closed by additional search compute or data. Inspired by the iterative refinement of MPPI in continuous control, we propose \textbf{Iterative Gumbel Planning (IGP)}, which adapts this principle to a discrete policy class. IGP represents the policy as a per-dimension factorized categorical distribution aligned with the search support, eliminating the projection error and reducing the residual approximation to a bounded discretization error controllable by the grid resolution. Joint candidates are drawn by Gumbel Top-$K$ sampling without replacement, avoiding the $V^{|A|}$ combinatorial explosion of joint discretization, and search-refine iteration runs for multiple rounds at the same state. On DMControl and HumanoidBench, IGP achieves state-of-the-art performance, with improved sample efficiency and lower temporal action variance at evaluation.
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Haozhan Tang ⋅ Zerui Wang ⋅ Yuxian Gu ⋅ Song Han ⋅ Han Cai
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. Most existing zero-shot methods fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion--exclusion attention merge and an on-the-fly RoPE correction rotation make the bifocal construction essentially free at inference; fused into a single CuTe kernel, long-context prefill reaches up to $1.39\times$ FA2 throughput on H100 (approaching the Hopper-only FA4), and single-batch generation incurs $\le 4\%$ overhead at every length. On Qwen3-1.7B/4B/8B up to 128K context, Jet-Long leads RULER by $+4.79$/$+2.18$/$+2.03$ pp over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance) and attains the lowest PG-19 perplexity. Jet-Long also generalizes to hybrid attention architectures such as Jet-Nemotron for further long-context improvement without retraining, and remains hyperparameter-resilient for ease of deployment.
JODA: Composable Joint Dynamics for Articulated Objects
Tianhong Gao ⋅ Cheng Yu ⋅ Yinghao Xu ⋅ Mengyu Chu
Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamical effects that govern realistic mechanical behavior, such as frictional holding, detents, soft closing, and snap latching. Existing approaches either ignore the detailed structure of dynamics entirely, or use simple models with limited expressiveness. We introduce JODA, a framework for generating joint-level dynamics as a structured three-channel field over the joint degree of freedom, capturing conservative forces, dry friction, and damping. Instantiated using shape-constrained piecewise cubic interpolation (PCHIP), this formulation defines a compact and expressive function space that is both interpretable and compatible with differentiable simulation. Building on this representation, we develop methods for inferring and refining joint dynamics from multimodal inputs. Given visual observations and joint context, a vision-language model proposes structured dynamical primitives, which are composed into a unified dynamics field. The resulting representation supports both direct manipulation and gradient-based refinement. We demonstrate that JODA enables plausible and controllable modeling of diverse joint behaviors, providing a unified interface for inference, editing, and optimization.
Knowing Without Saying: How Contextual Evidence Survives but Fails to Surface in Transformers
Ruochen Jin ⋅ Tianyu Pang ⋅ Chenxi Lin ⋅ Lei Hsiung ⋅ Jun Jie Ou Yang ⋅ Tao Wen ⋅ Yujun Yan ⋅ Yaoqing Yang
Large language models frequently ignore in-context evidence that contradicts their parametric priors, producing answers that reflect default beliefs rather than the supplied passage, a phenomenon termed knowledge conflict. This failure is commonly attributed to insufficient encoding of the contextual signal. We challenge this explanation. Through systematic, layer-by-layer analysis in a controlled multi-hop question-answering setting, we show that models faithfully encode the context-supported answer in their intermediate representations, yet suppress it in the final layers before generation. We term this late-layer reversal: the representational support for the evidence answer rises through the middle layers, then collapses in the final ~40\% of the network. To explain this phenomenon, we identify a Dual-Track computational structure within the transformer's hidden state. A knowledge direction, recovered by a learned linear probe, encodes the contextually supported answer; a nearly orthogonal output direction, aligned with the unembedding geometry, determines the model's prediction. These two directions are causally severed: interventions along the knowledge direction produce no change in output, whereas targeted suppression of prior-restoring computation along the output direction reliably causes the model to follow the contextual evidence. Our findings hold across nine model-dataset combinations and reveal that knowledge conflict failures arise not because models fail to represent contextual knowledge, but because their generation pathway declines to consult it.
Know Your Task, Learn It Right: Task-Aware Optimistic Value Learning for Multi-Task Multi-Agent Reinforcement Learning
Chang Liu ⋅ Mengyang Li
Cooperative multi-agent reinforcement learning under task variability requires both identifying which task an episode belongs to and adapting the learning algorithm to that identity. Recent class-aware methods address the first half of this problem by clustering trajectories into latent task classes and feeding the class label to the agent policy. They leave the second half largely untouched: the Bellman update and the credit assignment performed by the value mixer remain identical across all task classes. We show empirically that, even when task identity is recovered with high accuracy, the Bellman dynamics across classes diverge sharply, with simple classes saturating early and harder classes sustaining large temporal-difference residuals and large cross-agent value disagreement. Building on this observation, we propose Know Your Task, Learn It Right (KYT), a task-aware extension of value factorization that injects per-class learning signal into the Bellman update itself. KYT maintains a lightweight tracker that summarizes the per-class TD-error and cross-agent value spread without any extra network, and uses the tracker in two places: an optimistic target that adds a bounded per-class exploration bonus, and a class-conditioned mixer that modulates credit assignment through a softplus-gated affine layer preserving the IGM property. Across StarCraft II micromanagement benchmarks including SMACv2 and the unit-combination suite SurComb, KYT outperforms strong class-aware and exploration-aware baselines, with the largest gains concentrated on the hardest task classes that current methods underfit.
Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast model weights before clearing its key-value cache. During the sleep phase, the model performs $N$ offline recurrent passes over the accumulated context and updates the fast weights in its state-space model (SSM) blocks through a learned local rule. This shifts extra computation to the sleep phase while preserving the latency of wake-time prediction. We test our method on controlled synthetic tasks, including cellular automata and multi-hop graph retrieval, as well as a more realistic math reasoning task, on which a regular transformer as well as SSM-attention hybrid models fail. We then show that increasing sleep duration $N$ for our models improves performance, with the largest gains on examples that require deeper reasoning.
Layerwise Progressive Freezing: A Training Scaffold for Depth-Scalable Binary Networks
Evan G Smith ⋅ Bashima Islam
Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen. We study an orthogonal axis: when and where binarization is enforced during training. We introduce StoMPP (Stochastic Masked Partial Progressive Binarization), which gradually replaces clipped weights and activations with their hard binary counterparts layer by layer from input to output, using stochastic partial masks with soft refresh. StoMPP delivers two complementary benefits. As a standalone training rule, it provides a fully STE-free procedure that improves over vanilla STE with gains that grow with depth (ResNet-50 BNN: +18.0/+13.5/+3.8 on CIFAR-10/100/ImageNet), and the pattern holds across ResNet-18/34/50, MobileNetV2, and BERT fine-tuning. Composed with surrogate gradients by applying STE only to frozen entries, it reaches +27.1/+19.8/+17.7 over vanilla STE on the same setting. Underlying both regimes is a single mechanistic finding: progression order is decisive. Forward layerwise progression prevents depth collapse, reverse progression collapses to near-chance, and binary-weight networks (without binary activations) are insensitive to order. We trace this asymmetry to activation-induced gradient blockades: a committed binary activation severs gradient flow upstream, and ordering controls when these blockades form. To isolate the progression's contribution from any benefit conferred by STE, we conduct all ablations in the STE-free regime; the resulting characterization (schedule, refresh, ordering, dynamics) thus reflects the progression itself rather than its interaction with surrogate gradients.
LCD$^3$: Layout-Conditioned Diffusion for Dataset Distillation in Object Detection
Yue Cao ⋅ Mohsen Zardadi ⋅ Yu Hu ⋅ Yanshuo Fan ⋅ Jozsef Hamari ⋅ Zheng Liu ⋅ Jianyang Gu
Object detection data contains a structural asymmetry that is easy to overlook in dataset distillation (DD). Each detection image couples a spatial layout with a visual realization. Layouts define detection targets through object categories and bounding boxes, while pixels instantiate these targets with particular object appearances and backgrounds. Although both forms can be redundant across a dataset, they play different roles. Layouts define the task distribution, so synthesizing them freely risks altering the detection problem. Visual realizations, on the other hand, are conditional samples attached to layouts. This motivates a separation principle: preserve and recombine layouts from the original data, then use generative priors to diversify their visual realizations. Based on this principle, we propose LCD$^3$, a layout-conditioned diffusion framework for object detection DD. LCD$^3$ constructs composite layouts from diverse groups based on scene-level embeddings. A layout-conditioned diffusion model then generates new images from these layouts, enriching object appearance and scene realization without discarding the spatial structure needed for detection. The generation process is grounded with semantically verified object crops to retain the original visual style. Experiments show that this separation between layout and realization produces more effective distilled detection datasets. LCD$^3$ consistently outperforms the previous state-of-the-art, OD$^3$, across benchmarks using fewer objects per distilled image, e.g., a $12.5$\% mAP50 improvement at a $0.5$\% compression ratio on PASCAL VOC.
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
Yining Hong ⋅ Huang Huang ⋅ Manling Li ⋅ Fei-Fei Li ⋅ Leonidas Guibas ⋅ Jiajun Wu ⋅ Yejin Choi
Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequence of independent trials where mistakes repeat rather than accumulate into experience. Drawing upon human reflective practitioners, we introduce Reflective Test-Time Planning, which integrates two modes of reflection: \textit{reflection-in-action}, where the agent uses test-time scaling to generate and score multiple candidate actions using internal reflections before execution; and \textit{reflection-on-action}, which uses test-time training to update both its internal reflection model and its action policy based on external reflections after execution. We also include retrospective reflection, allowing the agent to re-evaluate earlier decisions and perform model updates with hindsight for proper long-horizon credit assignment. Experiments on our newly-designed Long-Horizon Household benchmark and MuJoCo Cupboard Fitting benchmark show significant gains over baseline models, with zero-shot generalization to photorealistic HM3D environments and real-robot experiments on a Franka Panda arm. Ablations confirm that reflection-in-action and reflection-on-action are mutually dependent, and that retrospective reflection achieves better credit assignment than step-wise external feedback at lower computational overhead. Qualitative analyses further highlight behavioral correction through reflection.
Learning Preference Representations for Preference-Conditioned Image Generation
Wenyi Mo ⋅ Tianyu Zhang ⋅ Yalong Bai ⋅ Ligong Han ⋅ Ying Ba ⋅ Dimitris Metaxas
Modern image generative models produce high-quality, prompt-faithful images, yet the same prompt can correspond to different desired outputs for different users. We study \emph{preference-conditioned image generation}: generating images that follow a target prompt while reflecting a user's latent visual preferences inferred from liked and disliked histories. This problem is challenging because real preference histories are sparse and costly to collect, user tastes mix semantic and stylistic factors that are difficult to verbalize, and diffusion generators are not naturally designed to condition on multi-image preference histories. We propose \textsc{PrefGen}, a framework for learning structured preference representations and injecting them into diffusion-based generators. To address the supervision bottleneck, we construct a large-scale synthetic-agent preference dataset with dense, coherent liked/disliked histories, using it as scalable supervision for preference representation learning. We train a multimodal large language model with preference-oriented visual question answering and analyze its hidden states to separate two complementary signals: an intra-user embedding for liked-versus-disliked distinctions and a cross-history consistency embedding for stable user-level tendencies. To bridge MLLM representations and diffusion conditioning, we introduce a distributional alignment objective based on maximum mean discrepancy and inject the aligned preference signal through a lightweight cross-attention branch. We evaluate \textsc{PrefGen} with a tiered protocol: \textsc{PrefBench} as a synthetic in-distribution diagnostic, leakage-free Pick-a-Pic transfer as a controlled real-user benchmark, and user-in-the-loop studies with participant-curated histories as human-facing evidence. Across these settings, \textsc{PrefGen} improves preference alignment over strong personalization baselines while preserving prompt fidelity and competitive image quality.
Learning Recoverable Neural Networks against Weight Corruption via Simple Zero-Sum Projection
Sanggeon Yun ⋅ Ryozo Masukawa ⋅ SungHeon Jeong ⋅ Hyunwoo Oh ⋅ Nathaniel D Bastian ⋅ Mohsen Imani
Deep neural networks are highly vulnerable to bit corruption in stored weights, whether from naturally occurring memory faults or deliberate fault-injection attacks. Existing defenses remain fundamentally limited: fault-tolerance methods degrade under cumulative corruption and remain vulnerable to strong white-box attacks, while ECC-based recovery methods are often tied to specific precisions or architectures and expose concentrated vulnerable surfaces. We propose a simple training-based recovery framework that endows model weights with recoverable structure through block-wise zero-sum constraints enforced by differentiable projection. The resulting method incurs no parameter-space overhead and generalizes naturally across architectures and numerical precisions. We further develop adaptive adversarial bit-flip attacks tailored to recovery-based defenses, covering both our method and prior ECC-based approaches. Across diverse architectures, precisions, datasets, and threat models, our method consistently delivers strong robustness and recoverability while preserving clean performance.
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
Zhongjie Shi ⋅ Wenjing Liao
This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain $[0,1]^d$ and $d$-dimensional compact Riemannian manifolds. We propose a novel constructive approximation framework for Transformers that builds local approximations of the target function and aggregates them into a global approximation via softmax partition of unity. This approach leverages the attention mechanism to achieve spatial localization through affine transformations of the input. The softmax activation plays a crucial role in aggregating local approximations to a global output. From an approximation perspective, we prove that a dense Transformer equipped with only two encoder blocks and standard single-hidden-layer point-wise feed-forward networks can achieve a uniform $\varepsilon$-approximation error for $\alpha$-H\"older continuous functions with $\alpha \in (0,1]$ using $\mathcal{O}(\varepsilon^{-d/\alpha})$ total parameters. Building upon this approximation guarantee, we establish a near minimax-optimal generalization error bound of order $\mathcal{O}\big(n^{-\frac{2\alpha}{2\alpha+d}} \log n\big)$ for the empirical risk minimizer, where $n$ is the training data size. The Transformer architecture studied in this paper is dense, shallow and wide, and employs softmax activation and sinusoidal positional encodings, closely reflecting practical implementations.
Learning When to Think: Adaptive Internal Computation for Reinforcement Learning
Rohit Kumar Salla ⋅ Simon Stepputtis
As agents have to solve diverse tasks from grid-worlds to robotics and language modeling, the difficulty of each step varies, yet policies are rigid and spend a fixed amount of compute on each step. A natural approach is to spend compute on the harder states while saving it on easy ones, allowing agents to retain performance while saving valuable compute budgets. To this end, we propose Adaptive Internal Computation (AIC), a lightweight policy architecture that iteratively refines latent states and learns when to halt using a value-of-computation signal.Concretely, AIC repeatedly refines a latent state representation and learns a halting head from a value-of-computation signal, the per-step gain in expected return, allowing the policy to stop refining as soon as further computation no longer improves the action it would take. A central challenge, however, is that matched-budget gains alone cannot establish that compute is being directed to the states that benefit the most from them. We therefore introduce a six-component causal evaluation protocol leveraging a targeted-versus-random ablation that tests whether removing compute from AIC's predicted high-compute states hurts performance more than removing the same amount of compute from randomly chosen states, thereby validating that AIC can successfully identify high-compute states. Across twelve diverse tasks, including grid worlds, simulated robotics experiments, and language modeling, we demonstrate that AIC outperforms fixed-compute policies across a set of five underlying model architectures. In particular, architectures leveraging AIC retain or improve performance while spending about 52\% less compute compared to their respective fixed counterparts.
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
Keerthi Kaashyap ⋅ Dennis Anthony ⋅ Akshay Krishnan ⋅ Nhi Nguyen ⋅ Jeremy Collins ⋅ James Hays ⋅ Shreyas Kousik ⋅ Animesh Garg
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of lack of supervisory signal, but rather due to inconspicuous architectural choices: spatially expressive decoders that dilute representational capabilities of the scene encoder, and low-level pixel-space targets that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is general purpose, and we show that it is on par or better than special-purpose geometry-supervised methods. SNAP also outperforms self-supervised representations across five geometric tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch tokens exhibit emergent viewpoint invariance that rivals heavily supervised models despite lower compute and data budgets. Under severe camera shifts where standard 2D representations collapse, SNAP maintains strict spatial consistency, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure.
Limits and Potential of Score-Based Data Valuation: Redundancy, Complementarity, and Non-Monotonicity
Kumar Kshitij Patel ⋅ Sai Praneeth Karimireddy ⋅ Raul Fernandez ⋅ Manolis Zampetakis
Score-based valuation methods, such as Shapley values and Leave-one-out (LOO), are widely used to assign value to data in modern machine learning pipelines, including for tasks such as attribution, selection, and pricing, yet it remains unclear when these scalar scores reliably guide downstream decisions. We show that their success is governed by three structural properties of the learning problem: substitutability, complementarity, and non-monotonicity. Substitutability (redundancy) can collapse pointwise credit, causing Shapley and LOO to fail even under monotone submodular valuations; bounded curvature limits this collapse and helps recover constant-factor approximation. Complementarity can break common score-based rules and greedy-style adaptive selection, though these effects diminish with sufficient coverage. Non-monotonicity implies that all non-adaptive methods, including score-based approaches, can fail, establishing a separation from adaptive algorithms. Our theoretical results, supported by empirical evidence, provide a structural view of data valuation and motivate a simple practical pipeline: deduplicate to reduce redundancy, ensure coverage to suppress complementarity, and then choose between score-based or adaptive methods based on non-monotonic effects.
LLM-ACES: Closed-Loop Discovery of Dynamic Systems with LLM-Guided Adaptive Search
Nikhil Abhyankar ⋅ Sha Li ⋅ Sanchit Kabra ⋅ Naren Ramakrishnan ⋅ Yulia Gel ⋅ Chandan Reddy
Recovering governing Ordinary Differential Equations (ODEs) from data is a central challenge in modeling dynamical systems across scientific domains. Existing approaches cast discovery as a static inference problem over fixed datasets, implicitly assuming that the observed trajectories are sufficiently informative. However, dynamical systems evolve over large state spaces, and limited data can often fit multiple distinct equations that explain the observations equally well, leading to identifiability gaps and incorrect recovery of the true dynamics. We introduce LLM-ACES or LLM-guided Active Closed-loop Equation Search, a closed-loop framework that jointly optimizes data acquisition and hypothesis generation under a constrained simulation budget. LLM-ACES leverages LLMs to construct structured, domain-informed hypothesis spaces via a two-stage process: high-level prior induction followed by candidate equation generation within each prior. To resolve ambiguity among competing hypotheses, we introduce an active data selection strategy that identifies regions of maximal predictive divergence among candidate equations and queries the system to obtain informative trajectories. This induces a feedback loop in which hypotheses guide data acquisition and newly acquired data refine the hypothesis space. Experiments demonstrate that LLM-ACES achieves more accurate equation recovery than prior state-of-the-art methods. Our ablations and analyses highlight the importance of coupling hypothesis construction with feedback-driven data acquisition for reliable discovery of dynamical systems.
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
Tong Zheng ⋅ Haolin Liu ⋅ Chengsong Huang ⋅ Huiwen Bao ⋅ Sheng Zhang ⋅ Rui Liu ⋅ Runpeng(Leo) Dai ⋅ Ruibo Chen ⋅ chenxi liu ⋅ Tianyi Xiong ⋅ Xidong Wu ⋅ Hongming Zhang ⋅ Heng Huang
Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS strategies are largely hand-crafted: researchers manually design reasoning patterns and tune heuristics by intuition, leaving much of the computation-allocation space unexplored. We propose an environment-driven framework, AutoTTS, that changes what researchers design: from individual TTS heuristics to environments where TTS strategies can be discovered automatically. The key to AutoTTS lies in environment construction: the discovery environment must make the control space tractable and provide cheap, frequent feedback for TTS search. As a concrete instantiation, we formulate width--depth TTS as controller synthesis over pre-collected reasoning trajectories and probe signals, where controllers decide when to branch, continue, probe, prune, or stop and can be evaluated cheaply without repeated LLM calls. We further introduce beta parameterization to reduce search-set overfitting and fine-grained execution trace feedback to improve discovery efficiency by helping the agent diagnose why a TTS program fails. Experiments on mathematical reasoning benchmarks show that the discovered strategies improve the overall accuracy--cost tradeoff over strong manually designed baselines. The discovered strategies generalize to held-out benchmarks and model scales, while the entire discovery costs only $39.9 and 160 minutes.
Low-Rank Hierarchical Merging for Efficient Long-to-Short Reasoning
Zeqiu Yu ⋅ Xuesheng Zhang ⋅ Wenxiao Zhao ⋅ Shixiao Wang
Long-to-Short (L2S) model merging seeks to combine the strong reasoning capability of long chain-of-thought models with the concise style of base models, but existing methods that rely on linear interpolation and first-order activation statistics fail to capture how the loss landscape responds to parameter displacement, leading to degraded accuracy on hard tasks. We propose Low-Rank Hierarchical Merging (LHM), a training-free framework that replaces heuristic statistics with lightweight second-order curvature information. LHM consists of three components: (1) a low-rank Hessian approximation that efficiently estimates per-layer Hessian traces via Hutchinson's estimator on random inputs, eliminating the need for task- specific data; (2) a Layer Second-Order Sensitivity (LSOS) index that combines Hessian trace with task-vector norm to quantify per-layer fusion difficulty; and (3) a Dual-Constraint Coefficient Mapping (DCCM) that converts LSOS scores into layer-adaptive merge coefficients while jointly satisfying a fusion-error bound and a global length-reduction target. Experiments across six mathematical benchmarks on 1.5B and 7B models show that LHM outperforms state-of-the-art methods such as ACM-TIES and ORCA, achieving over 4\% accuracy gain and up to 93.1\% reduction in output length relative to the long-CoT model.
Majority Bit-Aware Watermarking for Large Language Models
Jiahao Xu ⋅ Rui Hu ⋅ Olivera Kotevska ⋅ Zikai Zhang
The growing deployment of Large Language Models (LLMs) has raised concerns about their misuse in generating harmful or deceptive content. To address this issue, watermarking methods have been proposed to embed identifiable multi-bit messages into generated text for misuse tracing. However, existing methods often suffer from a fundamental trade-off between text quality and decoding accuracy. In particular, they have to restrict the size of the preferred token set (i.e., green list) during encoding to maintain a detectable watermark signal for decoding, which inevitably degrades generation quality. To improve this trade-off, we propose a novel message encoding paradigm called \textit{majority bit-aware encoding}, which relaxes the watermark signal strength from the green list size. This strategy allows for a strong watermark signal to be preserved in generated texts even when using a large green list. We introduce two instantiations of this paradigm: MajorMark and MajorMark$^{+}$, where the latter is specifically optimized for long messages. Extensive experiments on state-of-the-art LLMs demonstrate that our methods achieve higher decoding accuracy and superior text quality compared to prior baselines. The code of the proposed methods is available \href{https://anonymous.4open.science/r/MajorMark}{here} for review.
MammoGPS: A Benchmark for Visual Grounding, Perception, and Spatial Reasoning in Mammography
Yen Nhi Truong Vu ⋅ Dan Guo ⋅ Sripad Joshi ⋅ Harshit Kumar ⋅ Jason Su ⋅ Thomas P Matthews
Vision-language models (VLMs) can achieve strong image-level visual question-answer (QA) performance while relying on shortcuts rather than grounded visual understanding. This is especially concerning in medical imaging, where clinically meaningful interpretation often requires spatial reasoning, localization, and anatomical understanding. Yet in mammography, no benchmark directly assesses whether VLMs can localize findings or reason spatially. To address this, we introduce MammoGPS, a benchmark for evaluating visual grounding, perception, and spatial reasoning in mammography. MammoGPS comprises 56,638 QA pairs across 13 clinical tasks for anatomical and finding grounding, finding-type, quadrant prediction, and depth determination. MammoGPS also includes 94,119 QA pairs across 20 controlled diagnostic tasks designed to probe model failures and spatial biases: we design controlled difficulty ladders to isolate the effects of semantic cues, visual cues, clinical terminology, and background context, as well as spatial-bias tasks to test whether model performance depends on finding location. We evaluate 12 VLMs spanning proprietary, general-purpose, and medical-domain models. Our analysis shows that current systems struggle to localize suspicious findings, make limited use of spatial semantic cues, rely on priors and shortcut heuristics for categorical and grounding tasks, struggle with mammography-specific spatial terminology, and exhibit position-dependent biases. Together, our results underscore the need for more rigorous and fine-grained mammography VLM evaluations.
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
Guangchen Lan ⋅ Sipeng Zhang ⋅ Tianle Wang ⋅ Yuwei Zhang ⋅ Xinpeng Wei ⋅ Daoan Zhang ⋅ Xiaoman Pan ⋅ Hongming Zhang ⋅ Dong-Jun Han ⋅ Christopher Brinton
As the era of large language models (LLMs) unfolds, Preference Optimization (PO) methods have become a central approach to aligning LLMs with human preferences and improving performance. We propose Maximum a Posteriori Preference Optimization (MaPPO), a methodology for learning from preferences that explicitly incorporates prior reward knowledge into the optimization objective. Building on the paradigm employed by Direct Preference Optimization (DPO) and its variants of treating preference learning as a Maximum Likelihood Estimation (MLE) problem, MaPPO integrates prior reward estimates into a principled Maximum a Posteriori (MaP) objective. This not only generalizes DPO and its variants, but also enhances alignment by mitigating the oversimplified binary classification of responses. Additionally, MaPPO introduces no additional hyperparameters, and supports preference optimization in both offline and online settings. In addition, MaPPO can be used as a plugin for DPO variants, including widely used SimPO, IPO and CPO, and produce consistent improvements. Extensive empirical evaluations of different model sizes and model series on three standard benchmarks (MT-Bench, AlpacaEval 2.0, and Arena-Hard) demonstrate consistent improvements in alignment performance without sacrificing computational efficiency.
Practitioners choose among LLM alignment methods largely by trial and error. We provide a theoretical foundation: each pairwise method defines a vector field on the probability simplex that governs how the model redistributes mass between preferred and dispreferred completions. Projecting onto a single preferred/dispreferred pair reduces this to a one-dimensional margin ODE with a closed-form solution. These solutions reveal that DPO margins grow as $O(\log t)$ while IPO margins converge exponentially to $1/(2\tau)$; that label noise induces a qualitative phase transition in DPO (unbounded $\to$ bounded) but affects IPO only quantitatively; and that gradient concentration under asymmetric weighting governs trainability at scale. We confirm all three predictions across five model families spanning four benchmarks. The closed-form solutions also enable principled method design: we prove that adding \emph{any} label smoothing $\varepsilon > 0$ to SimPO creates a finite margin attractor at $\Delta^* = \beta^{-1}\log((1{-}\varepsilon)/\varepsilon)$, while $\varepsilon{=}0$ (standard SimPO) has no attractor. We test this across three model scales (3.8B--9B) and confirm it: the theory-derived configuration (TPO-S, $\varepsilon{=}0.05$) achieves the highest mean IFEval score on all three models.
Matrix-Free Stochastic Training of Low-Rank Spectral Graph Learning via Randomized Adaptive Spectral Estimation
Mingqi Yang ⋅ Yanming Shen
Large-graph spectral learning promises global, trainable propagation beyond local neighborhoods, but explicit spectral layers remain difficult to use in stochastic large-graph training: their bases are commonly obtained by a full-graph spectralization step and their graph-wide contexts are naturally full-batch. Matrix-free polynomial filters avoid eigensolvers and train stochastically, but tie the response to finite propagation bases; explicit spectral-subspace models offer more flexible global responses, but usually pay for full-graph basis construction and context evaluation. We introduce Randomized Adaptive Spectral Estimation (RASE), a matrix-free stochastic training framework for explicit low-rank spectral-context layers. RASE constructs a refreshable randomized block-Krylov/Ritz basis from probe vectors and sparse matrix--vector products, reuses probe-derived summaries as control variates for mini-batch estimates of the spectral context, and keeps the response module replaceable across polynomial, Chebyshev, and transformer-style instantiations. We analyze Krylov/Ritz polynomial-action exactness for the probe-generated subspace and a control-variate variance identity that isolates when the mini-batch context estimator reduces variance. Across 14 node-classification benchmarks, RASE matches its eigenbasis-trained reference within seed noise where the eigensolver is feasible (${\sim}67{\times}$ faster basis on PubMed), and trains the same backbone on the largest graphs where eigensolver preprocessing exceeds a 30-minute budget. The contribution is therefore not a new spectral filter family, but a training route that makes explicit low-rank spectral-context layers matrix-free and mini-batch trainable.
Measuring and Strengthening Behavioral Suppression in Language Models
Luxi (Lucy) He ⋅ Pengcheng Jiang ⋅ Jifan Zhang ⋅ Jiawei Han ⋅ Danqi Chen
Undesired model behaviors can emerge or re-emerge at any stage of training. Once observed, how effectively can we suppress them? To study this question, we first construct two model organisms---language models trained to exhibit a specific misalignment---with qualitatively distinct structures. In the narrow-trigger organism, undesired behavior is elicited by an identifiable prompt feature. In the broad-trigger organism, misalignment surfaces diffusely across open-ended responses without a separable cue. We then formalize post-hoc behavioral suppression and evaluate methods along three criteria: immediate suppression of the target behavior, durability under further training, and preservation of general capability. We compare standard methods---SFT, GRPO, and gradient-ascent unlearning---along these dimensions while controlling the number of gradient updates. Building on prior work on alignment shallowness, we hypothesize that more durable suppression requires the ability to recover from a wider range of misalignment states. We introduce Prefix-Expanded Adversarial Reinforcement Learning (PEARL), a GRPO variant that augments rollouts with continuations from intermediate states sampled along cached organism trajectories. Across both organism settings and two model families (Qwen3-4B, GPT-OSS-20B), PEARL achieves lower exploit rates while preserving task accuracy, and yields lower reactivation rates under both targeted and benign capability fine-tuning than SFT and GRPO baselines. Together, our framework and method offer a step toward more durable post-hoc removal of undesired behaviors.
Mechanistic Critics for Sample-Efficient NPU Design-Space Exploration
Yicheng He ⋅ Yuqi Xue ⋅ Tianyang Wu ⋅ Jian Huang ⋅ Bin Hu
NPU design-space exploration for LLM inference mainly relies on an expensive simulator that often returns only scalar feedback such as latency and power. Such feedback can rank evaluated configurations, but gives little information for deciding why a configuration is slow or which architecture parameter should be changed next. This limitation is especially severe because prefill and decode workloads can be limited by different resources, including compute throughput, HBM bandwidth, and inter-chip communication. We propose CriticNPU, a sample-efficient DSE framework built around an online-calibrated mechanistic critic for NPU architecture design. The critic uses explicit hardware limits to evaluate NPU configurations: before simulation, it estimates whether a proposed configuration is likely to improve over the current best and filters weak candidates; after simulation, it aggregates operator-level runtimes into prefill/decode bottleneck summaries and predicts the effect of candidate parameter changes. We instantiate the critic with a Roofline-style model that derives compute, memory, and communication limits from architecture parameters and calibrates per-operator-type effective limits using simulator observations. And CriticNPU combines this critic with a workload extractor and an architecture planner to propose configurations. Across six LLM workloads and 16 deployment configurations per model, CriticNPU achieves best-or-tied-best latency in 74% of prefill and 87.5% of decode settings while using roughly half the simulator calls of competitive DSE baselines. We conduct extensive ablation studies and detailed analyses to validate the effectiveness of our method from multiple perspectives.
AI agents are becoming active decision-makers on the Internet. As they make decisions in the same environments as humans, the environments themselves can change to influence them. We call this $\textit{mecha-nudging}$: changes to how choices are presented that systematically influence AI agents without materially degrading the decision environment for humans. To measure this phenomenon, we combine two frameworks---$\textit{Bayesian persuasion}$ from economics and $\mathcal{V}$-$\textit{usable information}$ from computer science---to get a common unit (bits) for quantifying how environments change across a wide range of interventions, contexts, and models. We apply this framework to over six million Etsy listings and find that, after ChatGPT’s release, listings contain significantly more machine-usable information for predicting agent curation decisions, increasing by 0.143 bits out of a maximum possible increase of 0.355. This shift is robust across prompts, token choices, labeling models, and fine-tuning architectures; absent in a regulated-text placebo; and far larger than the effect of generic LLM rewriting. In contrast, a human study finds little to no change in human-usable information. Our results provide the first large-scale evidence that systematic mecha-nudging is already occurring in the wild.
MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence
Hanqi Jiang ⋅ Junhao Chen ⋅ Yi Pan ⋅ Lifeng Chen ⋅ Weihang You ⋅ Haozhen Gong ⋅ Ruiyu Yan ⋅ Jinglei Lv ⋅ Lin Zhao ⋅ Hui Ren ⋅ Quanzheng Li ⋅ Tianming Liu ⋅ Xiang Li
Medical vision–language models (VLMs) are usually evaluated on intact image–question pairs, but trustworthy clinical use requires a stronger property: a model must recognise when the evidential basis for an answer has failed. We study this through silent failures under perturbed evidence, where a vision-required medical question is paired with a false premise, wording perturbation, knowledge-only rewrite, or ROI-corrupted image, yet the model returns a fluent non-refusal answer. We introduce MedVIGIL, a 300-case evaluation suite drawn from four public medical VQA sources, supervised end to end by four board-certified radiologists: every gold answer, refusal option, candidate-answer set, paraphrase, false-premise trap, ROI box, and clinical risk tier is clinician-authored. Two attending radiologists annotate every case in parallel, a senior radiologist consolidates the released manifest, and a separate fourth radiologist independent of construction answers every probe to provide the human reference baseline. The release contains 2,556 MCQ probes, 240 counterfactual triplets, physician-adjudicated risk-tier and answerability flags, ROI boxes, and a paired open-ended variant. We report seven correctness-conditioned audit metrics that summarise into the MedVIGIL Composite Score (MCS), and audit 16 vision-capable models plus two text-only baselines. The independent radiologist scores MCS 83.3 at silent-failure rate 5.8%, leaving a 14.1-point composite headroom above the strongest audited model (Claude Opus 4.7 at 69.2). The benchmark and evaluation harness are publicly released.
MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing
Nishant Bhargava ⋅ Rodrigo S Barrento
Long-term LLM agents must decide what to write into persistent memory before future queries are known, yet existing evaluations measure only final QA accuracy, entangling memory writing with retrieval and reasoning. We introduce MEMAUDIT, a package-oracle evaluation protocol that turns memory writing into a finite, auditable optimization problem with a certified ground-truth optimum under a fixed storage budget. Each evaluation package specifies an experience stream, candidate memory representations, and future-query requirements, enabling exact measurement of how well a memory writer preserves task-relevant information. We instantiate MEMAUDIT with a semantic coverage objective under storage and exclusivity constraints and compute exact optima via certified optimization. Across controlled and naturalistic settings, MEMAUDIT disentangles representation quality, validity preservation, and budget-aware selection effects that end-to-end QA cannot isolate. This enables, for the first time, principled and reproducible evaluation of memory-writing policies independent of downstream retrieval and reasoning.
MiCo: Microstructure-Consistent Flow Matching for Diffusion MRI Angular Super-Resolution
Zhixuan Zhou ⋅ Tingting Dan ⋅ Guorong Wu
Diffusion-weighted MRI (dMRI) requires densely sampling $q$-space across many directions and shells, leading to long scan times that limit clinical use. Angular super-resolution (ASR) recovers a dense signal from a sparse acquisition, enabling substantially shorter scans. We propose a physics-guided flow matching framework for ASR built on a basic fact about dMRI: under the cumulant expansion, the diffusion signal naturally splits into a known linear function of the $q$-vectors plus a residual carrying higher-order angular detail. We exploit this twice. First, we initialize the flow from a closed-form dense DTI estimate, recasting ASR as residual flow matching from a physically meaningful starting point rather than uninformative noise. Second, we impose a microstructure-consistency loss, rooted in the physical principle that any two diffusion-weighted signals from the same voxel must be explained by a single underlying set of microstructural parameters. Together, these designs anchor predictions wherever the acquisition is informative while leaving the learned prior to model crossing-fiber and higher-order angular detail. On the Human Connectome Project (HCP-YA) and UK Biobank, our method matches or surpasses analytical baselines and recent deep learning methods on signal reconstruction, kurtosis scalar maps, and fiber-orientation reconstruction under aggressive subsampling.
Minimally Invasive Steering of Language Models
Taha Entesari ⋅ Jingyu (Jack) Zhang ⋅ Daniel Khashabi ⋅ Mahyar Fazlyab
Test-time alignment aims to adapt large language models (LLMs) to runtime objectives without updating their parameters. This is especially useful when the reward signal is user-specific, time-varying, or available only through a black-box evaluator. We propose a minimally invasive pre-logit steering method that optimizes additive interventions to the hidden states of a frozen LLM in order to improve reward while preserving the base model's behavior. Rather than penalizing the Euclidean norm of the steering vector, we derive an effort penalty from the local KL geometry of the induced token distribution. Specifically, we show that the per-token KL divergence between the steered and reference policies admits a second-order expansion given by a Fisher-quadratic form in the steering vector. This regularizer has an analytic gradient computable through matrix-vector products with the frozen language-model head, making it as inexpensive as a standard quadratic penalty while retaining a principled KL interpretation. We further decompose the sequence-level KL gradient into an analytic Fisher component and a trajectory-dependent score-function component, and introduce a hierarchy of Fisher surrogates with provable first-order equivalence in the small-steering regime. The resulting algorithm provides a training-free, reward-driven test-time alignment procedure that improves reward while controlling deviation from the base policy.
Minimax-Optimal Transformer Classification for Functional Data with Dense-Sparse Phase Transition
Shuoyang Wang ⋅ Yidan Tian ⋅ Guanqun Cao
Despite the remarkable empirical success of transformer models in natural language processing, their theoretical foundations for functional data classification remain largely unexplored. This paper takes a first step toward closing this gap by developing a rigorous statistical framework for transformer-based functional classifiers. We show that a transformer architecture with an expanding attention window in the latent representation attains minimax-optimal excess-risk rates up to logarithmic factors in infinite-dimensional functional classification, without relying on conventional dimension-reduction procedures. Our analysis further reveals a dense-to-sparse phase transition: when the sampling frequency exceeds a critical threshold relative to the sample size, the proposed classifier achieves the optimal dense-observation rate; otherwise, we precisely characterize the degradation in convergence under sparse sampling. Extensive simulations and real-data experiments support the theory and demonstrate that transformers provide a competitive and theoretically justified approach to functional classification.
Mitigating Label Bias with Interpretable Rubric Embeddings
Calvin Isley ⋅ Johann D. Gaebler ⋅ Sharad Goel
Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions, and content moderation. In these settings, models are typically trained on historical human evaluations—for example, using past hiring decisions as a proxy for true applicant quality. However, if past evaluations unjustly penalize certain groups, models trained on these labels may inherit those biases. To address this problem, we propose basing predictions on rubric embeddings, a representation framework that replaces standard black-box embeddings with features derived from expert-defined criteria that align with the underlying construct of interest. By anchoring predictions to semantically meaningful dimensions, this approach guards against biased proxy signals. We provide both theoretical and empirical evidence that rubric embeddings mitigate label bias under plausible conditions. Empirically, we evaluate our method on a novel dataset of applications to a large master’s program. We find that models trained on rubric embeddings reduce group disparities while improving measures of cohort quality. Our results suggest that basing predictions on interpretable, domain-grounded representations offers a practical approach to learning in the presence of biased labels.
We study the problem of detecting treatment effects in randomized A/B experiments when the effects are potentially small and complex. A common approach in such settings is to fit machine learning (ML) models of outcomes on observed features and then apply classical $t$-tests to residualized outcomes. We show that this approach can suffer substantial power loss under treatment effect heterogeneity. We propose a randomization-based test whose statistic measures the out-of-sample predictive gain from including the treatment variable in a flexible ML model. Leveraging experimental randomization and sample splitting, our test is finite-sample valid for arbitrary models and loss functions. We establish power guarantees under heterogeneous treatment effects and demonstrate substantial empirical gains in simulations and large-scale A/B experiments.
We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown. We establish distributional consistency of the bootstrap transition estimator in both a single long-chain regime and the episodic offline RL regime. The key technical tools are a novel bootstrap law of large numbers (LLN) for the visitation counts and a novel use of the martingale central limit theorem (CLT) for the bootstrap transition increments. We extend bootstrap distributional consistency to the downstream targets of offline policy evaluation (OPE) and optimal policy recovery (OPR) via the delta method by verifying Hadamard differentiability of the Bellman operators, yielding asymptotically valid confidence intervals for value and $Q$-functions. Experiments on the RiverSwim problem show that the proposed bootstrap confidence intervals (CIs), especially the percentile CIs, outperform the episodic bootstrap and plug-in CLT CIs, and are often close to nominal ($50\%$, $90\%$, $95\%$) coverage, while the baselines are poorly calibrated at small sample sizes and short episode lengths.
Model Cascades with Provable Per-Class Quality
Ashwin G Colaço ⋅ Sharad Mehrotra ⋅ Michael J De Lucia ⋅ Kevin Hamlen ⋅ Murat Kantarcioglu ⋅ Latifur R Khan ⋅ Ananthram Swami ⋅ Bhavani Thuraisingham ⋅ Unnat Jain
Model cascades reduce inference cost by routing inputs through cheap models and escalating only when needed, from lightweight image classifiers to small language models. Existing exit rules (confidence thresholds, learned scorers, agreement signals) all optimize for aggregate quality, but none controls quality at the class level: a model that is systematically unreliable on a rare, fine-grained, or minority class silently violates the quality contract, regardless of its overall confidence. We introduce NOMAD, built on a simple insight: the predicted class label, not the confidence score alone, is the correct signal for cascade exit decisions. For each model, NOMAD identifies from held-out data the classes it handles reliably enough to serve as the final answer; a chain-safety certificate verifies that any sequence of models preserves per-class quality end-to-end, not just at each stage in isolation; and a greedy selector routes each input through the cheapest viable sequence, with a provable 4-approximation on cost. On 16 datasets spanning tabular, fine-grained vision, and text (including LLM cascades with 15–20× cost ratios between models), NOMAD achieves up to 40× speedup (geometric mean 4× over the role model) and is the only method among 11 baselines with zero per-class violations on every dataset. Code, configurations, fold indices, and an interactive results explorer (nomad-cascades.surge.sh) accompany this submission as supplementary material for reproducibility.
Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation
Youngdong Jang ⋅ Gyeongrok Oh ⋅ Jong Wook Kim ⋅ Hyunju Ryu ⋅ Hyung-gun Chi ⋅ SeungHyeon Kim ⋅ Seungryong Kim ⋅ Jonghyun Choi ⋅ Sangpil Kim
LiDAR scene flow estimation is essential for autonomous driving, as it provides 3D motion for each point. Self-supervised approaches use static-dynamic classification to mitigate the imbalance between static and dynamic points, deriving targeted supervision. However, existing methods rely on sparse geometric observations for this classification, making them vulnerable to data sparsity and occlusions. The resulting noisy labels provide incorrect motion guidance and degrade scene flow learning. To address this, we introduce TrackCue, a tracking-guided framework for improving dynamic object representation in LiDAR scene flow estimation. In particular, TrackCue repurposes point tracking to obtain dense image-space trajectories anchored to LiDAR points, providing motion cues beyond sparse geometric observations. Furthermore, we present a visually consistent motion compensation strategy that compares the tracked trajectories with ego-induced rigid trajectories in the image plane, effectively isolating true object motion from ego-induced apparent motion. To transfer these isolated motion cues back to the LiDAR domain, we perform visual motion cue lifting, which associates ego-compensated image trajectories with LiDAR points for static-dynamic label refinement. As a result, TrackCue produces more accurate static-dynamic classification and provides more reliable supervision for scene flow learning. Experimental results show that TrackCue significantly improves the precision and F1 score of dynamic labels, leading to performance gains in self-supervised scene flow estimation.
MOVEBENCH: A Benchmark for Global-Scale Wildlife Movement Forecasting
Justin Kay ⋅ Shir Bar ⋅ Ellen O Aikens ⋅ Martin Becker ⋅ Francesca Cagnacci ⋅ Juliet Cohen ⋅ Scott W Forrest ⋅ Jessica Kendall-Bar ⋅ Madeleine Lucas ⋅ Macon Overcast ⋅ Meredith S Palmer ⋅ Will Rogers ⋅ Nicholas J Russo ⋅ Christian Rutz ⋅ Larissa T Beumer ⋅ Michael B Brown ⋅ Ying-Chi Chan ⋅ Sarah C Davidson ⋅ Diego E Soto ⋅ Anne G Hertel ⋅ Roland Kays ⋅ Benjamin Koger ⋅ Guram Mikaberidze ⋅ Thomas Mueller ⋅ Ruth Oliver ⋅ Thorsten Papenbrock ⋅ Robert Patchett ⋅ Jared A Stabach ⋅ Dane Taylor ⋅ Scott W Yanco ⋅ Sara Beery
Understanding and predicting wildlife movement is critical for ecology and conservation. While trajectory forecasting has advanced for human and vehicle movement, wildlife trajectories present distinct challenges: they are unconstrained in space, highly stochastic, and influenced by environmental conditions. We introduce MOVEBENCH, the first large-scale benchmark for probabilistic wildlife movement forecasting, containing 2.6M GPS locations from 800+ individuals across 110 species in 127 countries, paired with 1.6B environmental raster tiles capturing 160 covariates known or hypothesized to influence movement. We propose a probabilistic evaluation protocol for movement trajectory forecasts, addressing limitations of point-prediction metrics for inherently stochastic phenomena. Through comprehensive empirical evaluation of four method families across multiple temporal and spatial scales, we reveal that: (1) existing predictive methods generalize better to future timepoints than to unseen individuals, (2) deep learning approaches do not consistently outperform simpler baselines, and (3) environmental covariate selection significantly impacts performance. MOVEBENCH enables standardized evaluation of movement forecasting methods and provides a foundation for methodological advances on this ecologically important task.
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Xinkai Zhang ⋅ Zhipeng Wei ⋅ Huanli Gong ⋅ Jing Ting Zheng ⋅ Yuchen Zhang ⋅ Yue Dong ⋅ N. Benjamin Erichson
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but they are usually evaluated as black-box pipelines with different budgets, judges, retry rules, and strategy generation procedures. As a result, it is often unclear whether reported gains reflect stronger attack mechanisms or different experimental conditions. We introduce MT-JailBench, a modular evaluation framework for benchmarking multi-turn jailbreaks under fixed conditions. MT-JailBench implements each attack as five interacting modules: evaluation function, attack strategy, prompt generation, prompt refinement, and flow control. This design enables fair comparison across attack methods and component-wise analysis of what drives attack success. Using MT-JailBench, we find that resource budgets and evaluation functions are major confounders: controlling turns, retries, interactions, sampled strategies, and judges substantially change the ranking of attacks. At the component level, prompt generation accounts for most performance variation, while refinement and flow control provide moderate gains. We also find that explicit dynamic strategy generation is not always necessary; stochastic sampling from a fixed strategy can rival more elaborate diversification mechanisms. Finally, recomposing the best components yields a strong attack configuration that outperforms its source attacks and generalizes across diverse target LLMs. MT-JailBench therefore provides a modular framework for comparing multi-turn jailbreaks, understanding the impact of components, and guiding stronger red-teaming evaluations.
Multi-agent Collaboration with State Management
Mengyang Liu ⋅ Taozhi Chen ⋅ Zhenhua Xu ⋅ Xue Jiang ⋅ Yihong Dong
Recent advances in multi-agent systems have shown great potential for solving complex tasks. However, when multiple agents edit a shared codebase concurrently, their changes can silently conflict and inconsistent views lead to integration failures. Existing multi-agent systems address this through workspace isolation (e.g., one git worktree per agent), but this defers conflict resolution to a post-hoc merge step where recovery is expensive. In this paper, we propose STORM, i.e., STate-ORiented Management for multi-agent collaboration. Specifically, STORM manages agent states by mediating their interactions with the shared workspace, ensuring that each agent operates on a consistent view of the codebase and that conflicting edits are detected and resolved at write time.We evaluate STORM on Commit0 and PaperBench across multiple LLMs. STORM outperforms the git-worktree-based multi-agent baseline by +18.7 on Commit0-Lite and +1.4 on PaperBench, while achieving comparable or better cost efficiency. Combined with single-agent runs, STORM reaches highest scores of 87.6 and 78.2 on the two benchmarks respectively, suggesting that explicit state management is a more effective foundation for multi-agent collaboration than workspace isolation. STORM can also be plugged into any multi-agent system. Our code is available at \url{https://anonymous.4open.science/r/STORM-Mutli-Agent}.
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
Soumita Hait ⋅ Ping Li ⋅ Haipeng Luo ⋅ Mengxiao Zhang
Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, where only the loss of the selected action pair is observed, Fiegel et al. [2025] show a separation between average-iterate and last-iterate convergence in duality gap: while the optimal $t^{-\frac{1}{2}}$ rate after $t$ rounds is achievable for the former via standard no-regret algorithms, the latter cannot converge faster than $t^{-\frac{1}{3}}$ in expectation or $t^{-\frac{1}{4}}$ with high probability. However, in many practical settings (such as preference learning), the players observe not only their loss but also the opponent’s action. This raises a natural question: can such additional information enable faster last-iterate convergence? We answer this question affirmatively, showing that $t^{-\frac{1}{2}}$ last-iterate convergence is achievable with high probability in this setting, via an efficient algorithm that updates its strategy infrequently by solving an estimated log-barrier-regularized game. We identify fundamental obstacles preventing standard analysis for multi-armed bandits (the single-player case) from generalizing to games, and develop a novel analysis to overcome them. Experiments confirm that our algorithm indeed converges faster than naive baselines and prior methods that do not exploit opponent-action feedback. Finally, we note that our results also improve those for dueling bandits, a special case with skew-symmetric game matrices.
No Coin Left Behind: Maximizing Strategic Surplus Against No-Regret Dynamics
Yiheng Su ⋅ Emmanouil-Vasileios Vlatakis-Gkaragkounis
We investigate the **strategic surplus** obtainable against a **Follow-the-Regularized-Leader (FTRL)** learner with constant step size $\eta$ in $n \times m$ two-player zero-sum games played over $T$ rounds against a clairvoyant optimizer. In contrast with prior analysis, we show that the extraction of such regret-scale surplus is an inherent feature of the FTRL family, rather than an artifact of specific instantiations. First, for a fixed max-min optimizer, we establish a sweeping law of order $\Omega(N/\eta)$, proving that utility surplus scales with the number of the learner's suboptimal actions $N$ and vanishes in their absence. Second, for an alternating optimizer, a surplus of $\Omega(\eta T/\text{poly}(n,m))$ can be guaranteed regardless of the equilibrium structure, with high probability, in random games. Our analysis uncovers a sharp geometric dichotomy: **non-steep** regularizers allow the optimizer to realize the maximal transient surplus via finite-time elimination of suboptimal actions, whereas **steep** regularizers introduce a vanishing tail correction that can delay surplus saturation. Finally, we discuss whether this leverage persists under bilateral payoff uncertainty and propose a susceptibility measure quantifying which regularizers are most vulnerable to learner-aware strategic steering.
Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and inducing widespread ties. We argue that this trajectory collapse creates substantial statistical inefficiency by reducing effective sample size and weakening the ability to distinguish systems. We propose preference-based trajectory evaluation, which compares trajectories directly through temporal preferences over progress and time-to-return profiles. Across diverse agentic and interactive benchmarks, standard success-based metrics produce tied comparisons on roughly 75\% of instances, whereas trajectory-aware preferences reduce ties to roughly 30\%, improving discriminative power, ranking stability, and data efficiency. Our results suggest that benchmark saturation, often portrayed as the result of poor data collection or problem difficulty, may also be explained by the choice of evaluation measure.
Cross-modal place recognition aims to retrieve a target 3D location from a spatial map using a natural language description. Most existing methods follow a global-descriptor matching paradigm, in which the text query and each 3D scene cell are independently compressed into single vectors and compared by global similarity. Although simple and widely adopted, this paradigm tends to discard the fine-grained object correspondences that are essential for language-guided localization. In this paper, we challenge the necessity of global descriptors and propose OLA-Place, an Object-Level Alignment framework for cross-modal place recognition without global descriptors. Instead of representing a query or a scene cell as a single holistic embedding, OLA-Place formulates place recognition as set-to-set semantic alignment between textual object mentions and 3D object instances. The framework contains three key modules: ObjectSet Encoder (OSE), which extracts object-level representations from both language descriptions and 3D scene cells; Object Message Encoder (OME), which injects intra-set object-context information while preserving object-level granularity; and Masked Max Alignment (MMA), which computes the query-cell matching score by aligning each textual object mention with its most similar valid 3D object. This formulation is permutation-invariant, naturally handles variable-size object sets, and requires no object-level correspondence annotations. Extensive experiments show that object-level alignment alone significantly outperforms global-descriptor-based methods, demonstrating that cross-modal place recognition is better understood as an object-level semantic correspondence problem rather than a global-embedding retrieval problem. Our code is available at: https://github.com/Anonymous09871745/OLA-Place.
One-Shot Generative Flows: Existence and Obstructions
Panagiotis Tsimpos ⋅ Daniel Sharp ⋅ Youssef Marzouk
We study dynamic measure transport for generative modeling, focusing on transport maps that connect a source measure $P_0$ to a target measure $P_1$ by integrating a velocity field of the form $v_t(x) = \mathbb{E}[\dot X_t \mid X_t = x]$, where $X_{\bullet} = (X_{t}) _{t}$ is a stochastic process satisfying $(X _{0}, X _{1}) \sim {P _{0}} \otimes {P _{1}} $ and $\dot X _{t}$ is its time derivative. We investigate when $X _{\bullet}$ induces a _straight-line flow_: a flow whose pointwise acceleration vanishes and is therefore exactly integrable by any first-order method. First, we develop multiple characterizations of straight-line flows in terms of PDEs involving the conditional statistics of the process. Then, we prove that straight-line flows under endpoint independence exhibit a sharp dichotomy. On the one hand, we construct explicit, computable straight-line processes for arbitrary Gaussian endpoints. On the other hand, we show that straight-line processes do not exist for targets with sufficiently well-separated modes. We demonstrate this obstruction through a sequence of increasingly general impossibility theorems that uncover a fundamental relationship between the sample-path behavior of a process with independent endpoints and the space-time geometry of this process' flow map. Taken together, these results provide a structural theory of when straight-line generative flows can, and cannot, exist.
One-Shot Private Confidence Regions via Resampling
Po-Ling Loh ⋅ Debepsita Mukherjee ⋅ Shourya Pandey ⋅ Purnamrita Sarkar
We propose a simple framework for constructing differentially private confidence regions _in one shot_, i.e., by adding noise only to the final resampling quantile instead of privatizing the estimator computed on each resample. The cost of privacy of our procedure is $O(\log B)$ under with-replacement ($m$-out-of-$n$) sampling and independent of $B$ under without replacement sampling (subsampling), avoiding the $\sqrt{B}$ factor that arises in previous works. We provide _nonasymptotic_ Gaussian Differential Privacy (GDP) and utility guarantees for both subsampling and $m$-out-of-$n$ resampling, covering mean-like estimators with small global sensitivity as well as estimators admitting efficiently computable _smooth sensitivity_ bounds, including quantiles and degenerate U-statistics. This allows us to also obtain private confidence regions for degenerate U-statistics where the private error is much smaller than the non-private error. In all, we provide a toolbox for widely applicable DP uncertainty quantification procedures under popular resampling strategies while avoiding the computational and privacy costs of privatizing many intermediate resample statistics.
Decision Transformers (DTs) have emerged as a powerful framework for sequential decision making by formulating offline reinforcement learning (RL) as a sequence modeling problem. However, extending DTs to online settings with pure RL gradients remains largely unexplored, as existing approaches continue to rely heavily on supervised sequence-modeling objectives during online finetuning. We identify hindsight return relabeling---a standard component in online DTs---as a critical obstacle to RL-based finetuning: while beneficial for supervised learning, it is fundamentally incompatible with importance sampling-based RL algorithms such as GRPO, leading to unstable training. Building on this insight, we propose new algorithms that enable online finetuning of Decision Transformers using pure reinforcement learning gradients. We adapt GRPO to DTs and introduce several key modifications, including sub-trajectory optimization for improved credit assignment, sequence-level likelihood objectives for enhanced stability and efficiency, and active sampling to encourage exploration in uncertain regions. Through extensive experiments, we demonstrate that our methods outperform existing online DT baselines and achieve new state-of-the-art performance across multiple benchmarks, highlighting the effectiveness of pure-RL-based online finetuning for Decision Transformers.
Online Maximization of Non-Decomposable Test and Population Utilities
Wojciech Kotlowski ⋅ Marek Wydmuch ⋅ Krzysztof Dembczynski
We study sequential learning for classification metrics that are non-decomposable across instances and are general functions of the confusion matrix, such as the Jaccard index and the $F$-measure. The learner observes instances sequentially, makes irrevocable predictions, and is evaluated by the final empirical performance metric, leading to an online-regret objective. After observing the sequence, the learner also returns a classifier for future population use, leading to a population-regret objective. This single protocol connects two central frameworks for optimizing non-decomposable metrics: Expected Test Utility (ETU) and Population Utility (PU). For smooth concave utilities of the confusion matrix, we show that the ETU benchmark is controlled, in expectation, by the population PU comparator, and that an algorithm with small online regret also yields PU guarantees by an online-to-batch conversion. We then turn to non-concave linear-fractional metrics. We devise a general stochastic Dinkelbach-root method that tracks the relevant metric parameter online and gives online and population regret of order $n^{-1/2}$, up to logarithmic factors, with explicit dependence on conditional-probability estimation error. Our empirical studies evaluate these methods on benchmark datasets.
On Lipschitz Explosion in Deep Neural Networks with Normalization: Consequences for Optimization and Robustness
Ashkan Soleymani ⋅ Reyhaneh Hosseinpourkhoshkbari ⋅ Hadi Daneshmand ⋅ Patrick Jaillet
The Lipschitz constant of a neural network is a basic stability quantity appearing in optimization guarantees, generalization bounds, and robustness certificates. Deep networks naturally admit two such notions, \emph{input Lipschitzness}, measuring sensitivity to data perturbations, and \emph{parameter Lipschitzness}, measuring sensitivity to weight perturbations. We prove that, for deep networks with normalization layers, with batch normalization as the canonical example, \emph{both} Lipschitz constants can grow exponentially with depth and hence with parameter dimension, even under strong per-layer norm control such as $\\|W_\ell\\|_2\le 1$, and already for linear activations. Thus, a mechanism widely used to stabilize training can create severe worst-case instabilities that are invisible from layerwise norm bounds alone. Parameter-Lipschitz explosion turns Lipschitz-dependent nonsmooth optimization guarantees into exponential-in-dimension bounds for deep normalized networks, while input-Lipschitz explosion makes worst-case Lipschitz-based generalization and robustness certificates vacuous. The latter also yields a concrete rank-separation mechanism for adversarial vulnerability. The theory predicts that perturbations introducing new singular directions should be amplified much more strongly than equal-energy perturbations that remain within the input's original singular subspace. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 support this prediction, showing that rank-creating perturbations cause substantially sharper drops in accuracy, confidence, and margins than same-subspace perturbations.
On the Parallel Optimality of Exponentiated Gradient Descent
Xin Jennifer Chen ⋅ Andrei Graur ⋅ Aaron Sidford
A classic result in learning and optimization theory is that exponentiated gradient descent (EG) or the multiplicative weight update method computes an $\\epsilon$-approximate minimizer of an $\\ell_1$-Lipschitz function over the $d$-dimensional probability simplex $\\Delta_d :=\\{x \in \\mathbb{R}^d_{\\geq0}:\\sum_{i\in[d]}x_i=1\\}$ with $\\tilde{O}(\epsilon^{-2})$ queries to a first-order oracle. We investigate whether there are improved parallel algorithms for this fundamental problem. We show that, up to logarithmic factors, the answer is negative for $\\epsilon=\\tilde \\Omega(d^{-1/6})$ - any randomized algorithm that makes $O(\\mathrm{poly}(d))$ subgradient queries per round requires $\\tilde{\\Omega}(1/\\epsilon^2)$ rounds to output an $\epsilon$-optimal point with constant success probability. Moreover, to obtain this result we provide an analogous characterization of the parallel complexity of minimizing a $\ell_1$-Lipschitz convex function over the unit $\\ell_1$ ball. Previous lower bounds for non-constant $\\epsilon$ for this $\\ell_1$-Lipschitz convex optimization problem either made additional assumptions on the domain, or instead attained a bound of $\\Omega(\\epsilon^{-2/3})$ [DG19].
On the Primacy Bias in RLVR Training
Nguyen Phuc ⋅ Ngoc-Hieu Nguyen ⋅ Parshin Shojaee ⋅ Chinh D La ⋅ Duy M. H. Nguyen ⋅ Rui Zhang ⋅ Khoa D Doan ⋅ Ting Hua ⋅ Nitesh Chawla
Verifiable Reward (RLVR) has emerged as a promising approach for enhancing the reasoning capabilities of large language models. However, it still remains unclear whether RLVR can extend reasoning capabilities beyond what is already encoded in the base model. In this paper, we study this question through the lens of primacy bias — the tendency of models to over-rely on already-learned representations, disproportionately improving performance on problems they already handle well while underexploring harder ones that require departing from those representations. Through our controlled experiments and large-scale analysis, we show that RLVR training on mixed data tends to further sharpen performance on easier problems while yielding only marginal gains on harder ones. We find that although hard problems lead to modest gains overall, they are actually the ones that help most in expanding the model’s reasoning capabilities and pushing the representations. However, prolonged training on these hard problems may also hurt and lead to forgetting on easy problems. These findings help to better understand some of the prior conflicting observations about RLVR, specifically, why on (OOD) domains where the base model performs poorly, improvements on hard problems can expand reasoning; while on well-trained domains, forgetting can reduce performance and mask genuine gains obtained from hard problems. Data and code for reproducing our experiments are available at: https://anonymous.4open.science/r/neurips26_sub-0DE8.
On the Token Value Inequality in Efficient Reasoning
Runjia Zeng ⋅ Hang Hua ⋅ Yiyang Liu ⋅ Zhiqiang Tao ⋅ Ruixiang Tang ⋅ Qifan Wang ⋅ Cheng Han ⋅ Dongfang Liu
Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the reasoning trace equally valuable? We present a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals. We show that normalized log probability helps distinguish core tokens, which carry structural and decisive reasoning content, from redundant tokens, which are exploratory, low-confidence filler that contributes less directly to the final answer. Building on these findings, we formulate the TokenProbe framework around two empirical findings and one claim: findings identify token value inequality first and then establish TokenProbe as a core-token proxy, and the claim introduces an efficient GRPO objective positing that selectively compressing redundant tokens can yield Pareto improvements in the accuracy–token efficiency space. Empirically, our method preserves reasoning quality while reducing the token usage by 76% of the baseline. Under matched reasoning-length budgets, we show that it can even outperform strong flagship baselines like Gemini-3.1-Pro.
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
Rui Yang ⋅ Qianhui Wu ⋅ Yuxi Chen ⋅ Hao Bai ⋅ Wenlin Yao ⋅ Hao Cheng ⋅ Baolin Peng ⋅ Huan Zhang ⋅ Tong Zhang ⋅ Jianfeng Gao
Building capable visual web agents demands precise grounding, long-horizon reasoning, and robust interaction with dynamic, real-world websites. Despite rapid progress, the strongest systems remain largely proprietary, while open agents still depend heavily on supervised post-training over large collections of curated web trajectories. This dependence creates a major scalability bottleneck: high-quality demonstrations are expensive to collect, and static datasets offer limited coverage of the diverse, ever-changing open web. Although online reinforcement learning (RL) has shown promise for text-based agents, its potential for training visual web agents directly on live websites remains largely unexplored. In this paper, we introduce OpenWebRL, an open framework for training visual web agents with online multi-turn RL on real websites. OpenWebRL covers the full training pipeline, including task selection, supervised warm-starting, live-browser execution, multimodal context management, trajectory success judging, and efficient multi-turn policy optimization. Using this framework, we train OpenWebRL-4B, which establishes a new open-source state of the art on challenging live-web benchmarks. With only 0.4K initialization trajectories and 2.2K open-ended training tasks, OpenWebRL-4B achieves 67.0% success on Online-Mind2Web and 64.0% on DeepShop, outperforming prior open agents of similar or larger scale and remaining competitive with proprietary systems including OpenAI and Gemini CUA. Beyond strong benchmark performance, OpenWebRL systematically identifies the key design choices that make online RL effective for visual web agents. Overall, our work offers a practical path toward building more capable, reproducible, and cost-efficient open web agents. We will release our training data, models, and code to support future research.
Optimal Byzantine-resilient Federated Learning with User-level Differential Privacy
Ming Xiang ⋅ Stratis Ioannidis ⋅ Edmund Yeh ⋅ Carlee Joe-Wong ⋅ Lili Su
Real-world deployment of federated learning (FL) requires resilience against adversarial clients and data privacy. Byzantine fault captures the worst-case behavior of adversarial clients, while differential privacy (DP) provides statistical guarantees against privacy loss. In particular, user-level DP guards a client's entire contribution and provides more realistic protection against information leakage. Despite intensive efforts, the minimax-optimal convergence bound of Byzantine-resilient FL with user-level DP remains open. We establish an information-theoretic lower bound for any Byzantine-resilient empirical risk minimization problem under user-level DP and heterogeneous clients. We propose ByzULDP, a central DP algorithm featuring a two-level momentum across the client and server. Client-side momentum mitigates stochastic gradient variance during robust aggregation. Unlike the fault-free case, naively combining non-linear Byzantine-resilient aggregation with central DP noise is insufficient to guarantee convergence to a bounded neighborhood of a stationary point. Our key contribution is a server-side momentum that leverages a property unique to user-level DP, where privacy sensitivity depends on how client updates are constructed, to empower convergence. ByzULDP achieves the order-optimal convergence bound for strongly-convex objectives, with vanishing error terms decaying at the optimal ${\mathcal{O}}(1/T)$ and ${\mathcal{O}}(1/\sqrt{T})$ rates for strongly-convex and non-convex objectives, respectively. We corroborate our analysis with numerical experiments on real-world datasets across diverse privacy budgets.
We study contextual dynamic pricing with linear valuations and bounded-support agnostic noise, whose induced demand curve may be non-Lipschitz with arbitrary jumps and atoms. Such discontinuities break the cross-context interpolation arguments used by smooth-demand pricing algorithms, while the best previous method achieved only $\widetilde O(T^{3/4})$ regret. We propose Conservative-Markdown Redirect-UCB Pricing, a polynomial-time algorithm that combines randomized parameter estimation, conservative residual-grid probing, and confidence-based one-step redirection. Our algorithm achieves $\widetilde O(T^{2/3})$ optimal regret, matching the known lower bounds of Kleinberg and Leighton (2003) up to logarithmic factors and improving over the previous upper bound of Xu and Wang (2022). Under stochastic well-conditioned contexts, this closes the long-existing open regret gap in linear-valuation contextual pricing under agnostic non-Lipschitz noise distribution.
Optimal Rates for Pure $\varepsilon$-Differentially Private Stochastic Convex Optimization with Heavy Tails
Andrew Lowy
We study stochastic convex optimization (SCO) with heavy-tailed gradients under pure $\varepsilon$-differential privacy (DP). Instead of assuming a bound on the worst-case Lipschitz parameter of the loss, we assume only a bounded $k$-th moment. This assumption allows for unbounded, heavy-tailed stochastic gradient distributions, and can yield sharper excess risk bounds. Prior work characterized the minimax optimal rate for $\rho$-zero-concentrated DP SCO up to logarithmic factors in this setting, but the pure $\varepsilon$-DP case has remained open. We characterize the minimax optimal excess-risk rate for pure $\varepsilon$-DP heavy-tailed SCO up to logarithmic factors. Our algorithm achieves this rate in polynomial time with high probability. Moreover, it runs in deterministic polynomial time when the worst-case Lipschitz parameter is polynomially bounded. For important structured problem classes --- including hinge/ReLU-type and absolute-value losses on Euclidean balls, ellipsoids, and polytopes --- we achieve deterministic polynomial time even when the worst-case Lipschitz parameter is infinite. Our approach is based on a novel framework for privately optimizing Lipschitz extensions of the empirical loss. We complement our upper bound with a nearly matching high-probability lower bound.
We study the problem of recalibrating an online predictor [KE17, OKS24]: given an arbitrary ``hint'' sequence of forecasts, the learner must output new predictions that are calibrated while incurring small excess error relative to the original forecasts, under a proper loss. We give an online algorithm that achieves $(\varepsilon, \varepsilon^2)$-recalibration for Lipschitz proper losses in $T \approx \varepsilon^{-3}$ rounds, using an imbalanced extension of the recent simultaneous Blackwell approachability reduction framework of [HTY26]. We also show this tradeoff is optimal by proving a matching lower bound for recalibrating against the squared loss. As an application, we show how our recalibration algorithm can be combined with the online refinement method of [FH23] to obtain simultaneous $\varepsilon$-calibration and $\varepsilon^2$-calibeating for smooth proper losses at the same asymptotic rate, improving upon prior works that achieved these properties separately or with a worse $\varepsilon$ dependence. We also discuss extensions to settings with multiple hint sequences.
Optimizing Retraining Schedules via Learning Curves
Jin Sima ⋅ Changlong Wu ⋅ Ananth Grama ⋅ Wojciech Szpankowski
Retraining is the primary mechanism by which deployed models adapt to new data, yet it is also among the most expensive operations in modern machine learning. How often must a model be retrained to remain near-optimal? We answer this question through a learning-curve based characterization of the retraining-frequency/risk trade-off in online learning. For the i.i.d. realizable setting, we observe that only $O(\log T)$ updates suffice to match the risk of *full retraining* whenever the learning curve is non-increasing. However, when the learning curve decays as a power law $t^{-\alpha}$ with $\alpha < 1$, as empirically observed in deep learning, the budget collapses further to $O(\log \log T)$ updates, yielding a *substantial asymptotic* improvement. We further design update schedules achieving these bounds, prove matching lower bounds, and present an adaptive algorithm that remains optimal when $\alpha$ is unknown. We then extend the analysis to piecewise-stationary and gradually drifting environments, and establish a no-free-lunch theorem showing that some prior knowledge of the learning curve is unavoidable, i.e., no universal algorithm can be competitive without it. Together, these results provide a sharp characterization of the frequency--accuracy trade-off in online retraining and bridge foundational learning theory with practical strategies for scalable deployment.
Order-based structure learning for zero-inflated count data under data heterogeneity
Fariha Taskin ⋅ Junsouk Choi ⋅ Hyunwoong Chang
We propose a structure learning method for directed acyclic graph (DAG) models of zero-inflated multivariate count data. Although causal structure learning with zero-inflated Poisson models is identifiable, existing algorithms rely on local search over the DAG space, leading to poor scalability and a tendency to converge to suboptimal solutions. To address these limitations, we develop a novel order-based scoring framework that searches over variable orderings rather than directly over DAGs. Building on this framework, we further extend the method to multiple heterogeneous datasets by identifying a shared variable ordering while allowing each dataset to have its own edge set. We implement a stochastic hill climbing algorithm with a random-to-random (R2R) proposal operator to efficiently explore the ordering space. Simulation studies show that the proposed method outperforms competing approaches and improves graph estimation accuracy under increasing heterogeneity when a shared causal ordering is preserved. We apply the method to single-nucleus RNA sequencing data from a major depressive disorder study.
PhysGraphNet: Physical-State Scene Graphs via Latent Graph Reasoning and Counterfactual Supervision
Zhengtao Yao ⋅ Runhao Li ⋅ Yan Wen ⋅ Guang Yang ⋅ Siheng Wang ⋅ Chenhao Wei ⋅ Rongchao Zhang ⋅ Guoqing Ma ⋅ Haoyan Xu ⋅ Junhao Dong
Manipulation planners depend on accurate predicates about physical state---is an object occluded? is a container open? is a path blocked?---but such predicates are hard to read off an image. Contrastively-trained vision--language models capture semantic content but collapse to near-zero binary F1 on these predicates, and even GPT-4o few-shot falls below random on multiple-choice physical-state questions. We present \ours, which predicts physical-state scene graphs from a single image and a natural-language goal. A frozen CLIP ViT, adapted with LoRA, feeds a heterogeneous graph of object, relation, and learnable memory tokens into a Latent Graph Reasoning Transformer (\lgrt); training combines supervised, counterfactual-margin, contrastive, and calibration losses with offline counterfactual-pair augmentation. On \pacbench, \ours reaches $\mu$-F1$\,{=}\,0.651$ and $99.3\%$ MCQ accuracy, a $+0.170$ absolute $\mu$-F1 gain over the strongest no-graph supervised baseline, with $99.6\%$ counterfactual directional consistency. Component ablations attribute $+0.426$ $\mu$-F1 to the \lgrt and $+0.392$ to the counterfactual loss. A \manipbench-trained variant scores $74.4\%$ MCQ against $30.0\%$ for the best zero-shot VLM on that benchmark's 4-choice format.
Physics-Conditioned Video Diffusion with Kinematic Priors for Fusion Capsule Polishing
Shashank Galla ⋅ Abhishek Hanchate ⋅ Monika Biener ⋅ Suhas Bhandarkar ⋅ Satish Bukkapatnam
Video diffusion models can generate realistic process videos, but using them as controllable surrogates for industrial systems remains difficult when key control variables are only partially observable and no simulator or differentiable physics model is available. We study this regime in inertial confinement fusion (ICF) target capsule polishing, where surface defects can degrade fusion yield. The dynamics are governed by the polishing speed $\omega_p$ and capsule-pad slip $S_C$. While capsule motion can be monitored from video, $S_C$ is only indirectly observable. Using only an analytical kinematic model, we adapt a frozen Stable Video Diffusion (SVD-xt) backbone via LoRA and condition the generation on Fourier-encoded physics parameters $(\omega_p, S_C)$. Two training objectives operate on latent frame-difference dynamics: (i) a batch level correlation loss aligning them with the analytical sliding speed, and (ii) a stratified ratio-matching loss that, conditioned on $\omega_p$, isolates slip structure from the dominant $\omega_p$-driven variance. We evaluate physics consistency by tracking capsule motion in generated videos, inverting the kinematic model, and measuring rank agreement under a physical validity gate. For speed, rank accuracy exceeds $0.97$ on real data; for slip, it reaches $0.735$ with Spearman $\ge 0.50$ on synthetic data. Overall, the results suggest a practical recipe for controllable industrial video surrogates when only low-dimensional analytical kinematics are available.
PIGRAM: An Interpretable Patch--Motif Interaction Grammar for Protein--Nucleic-Acid Recognition
Yaru Han
Protein--nucleic-acid recognition underlies aptamer discovery and nucleic-acid therapeutics, yet computational models often trade mechanistic interpretability for predictive flexibility. Hybrid interpretable--latent architectures risk fallback takeover, where the latent branch silently becomes the true predictor while the explicit pathway remains decorative. We introduce PIGRAM, a patch--motif grammar model that decomposes proteins into local residue patches and nucleic-acid partners into secondary-structure motifs, learning a sign-separated grammar over explicit physicochemical and geometric attributes to preserve rare favorable interactions. Heterogeneous rule families are selectively refined into a coarse-to-fine atlas, and a frozen-grammar residual Transformer provides sparse, bounded corrections without displacing the grammar as the primary predictor. On a strict sequence-pair-disjoint benchmark, PIGRAM achieves a Pearson correlation of 0.494, approaching the unconstrained Transformer (0.511) while exposing signed rules, fine subtypes, and route-gated corrections for each prediction. The learned rules recover bidirectional patch--motif mechanisms, and residual corrections concentrate selectively on grammar-hard regions. In fluorescence ELISA experiments, suppressive-rule relief improved GFP binding, whereas favorable-rule disruption weakened NELF binding, providing direct experimental support for grammar-guided design. Together, these results show that interpretable grammar can serve as both a predictive scaffold and an experimentally actionable design principle for aptamer optimization.
PIVOT: A Unified Agentic Framework for Streaming Long-Video Understanding
Qiushi Lyu ⋅ Qianlan Yang ⋅ Ziqi Pang ⋅ Yu-Xiong Wang ⋅ Liangyan Gui
Streaming long-video understanding is a realistic setting for always-on video assistants, where visual input arrives continuously and users may ask questions at arbitrary times. We study streaming long-video question understanding (QA) under a unified temporal formulation in which multiple questions are revealed over time and may concern the past, present, or future relative to their query times. The model receives no task-type labels and must operate under strict sequential access and bounded shared memory, deciding whether each active question is already answerable or should remain open for future evidence. This differs from prior streaming settings that often focus on short streams, prefix-answerable queries, separated temporal categories, or response timing without long-horizon shared memory. To evaluate this formulation, we propose Ref2Stream (Reference-to-Stream), a general benchmark-conversion protocol that turns temporally grounded offline video QA benchmarks into streaming QA benchmarks. Instantiated on LVBench, Ref2Stream yields LVBench-Ref2Stream, which mixes past, present, and future questions under strict sequential access and evaluates both answer accuracy and answer latency. Existing streaming video methods typically improve efficiency or response timing through predefined compression, retrieval, or response policies, while offline video agents perform adaptive evidence gathering but assume access to the full video and target question. We introduce PIVOT (Planning over Incremental Video Observations for Timely Answering), a unified agentic framework that incrementally updates bounded memory from the observed stream and uses a planner-controlled action loop to retrieve relevant memory, analyze current evidence, reflect on candidate answers, and decide whether to answer or defer. Our work suggests that agentic reasoning is a promising direction for building streaming long-video systems that can adaptively gather evidence and answer only when sufficient information has been observed.
PLACE: Patch-Level Agnostic Concept Extraction
Gabriele Onorato ⋅ Fabrizio Silvestri ⋅ Nathaniel D Bastian ⋅ Francesco Restuccia
While While Vision Foundation Models (VFMs) are increasingly used in several computer vision tasks, their internal representations remain substantially opaque. This mostly stems from polysemanticity, i.e., individual neurons encode mixtures of unrelated visual patterns. In stark contrast, human reasoning is inherently monosemantic, i.e., it naturally relies on distinct, isolated features to process information. Therefore, monosemanticity needs to be implemented within the model's activation space to obtain representations that are human-understandable concepts. Existing methods mostly operate on convolutional networks, so when applied to VFMs their "crop-and-resize" approach distorts the model's representations by disrupting global self-attention mechanisms and discarding spatial geometry. Furthermore, prior work requires downstream classification labels and is based on KL-divergence, which requires to propagate the gradients back from the classifier head. Ultimately incurring in excessive concept extraction time, and making it hardly applicable to extract task-agnostic concepts. To overcome these issues, we propose PLACE, a fully unsupervised concept extraction framework that operates directly on the native patch-token activations of a VFM. By introducing a geometric Gram-matrix alignment loss, PLACE mechanistically controls downstream KL divergence, thus enforcing faithfulness -- i.e., the notion that extracted concepts, although obtained in an unsupervised fashion, are relevant to downstream classification tasks. Evaluations on image classification tasks (ImageNet and COCO) on DINOv2, ViT-MAE and ViT-B/16 show that PLACE (i) using Gram-alignment it extracts concepts about $20\times$ faster compared to KL-supervised extraction and $11\times$ faster than Sparse Autoencoders (SAEs). Moreover, PLACE yields highly sparse concepts that (ii) achieve up to 40 percentage points sparser compared to the state of the art (i.e., max 0.933 Gini sparsity score), and (iii) are up to 92.3% similar to those obtained under classifier supervision (i.e., 0.923 cosine similarity).
Pointwise Lipschitz Continuous Graph Algorithms
Quanquan C Liu ⋅ Grigoris Velegkas ⋅ Yuichi Yoshida ⋅ Felix Zhou
In many real-world applications, it is undesirable to drastically change the problem solution after a small perturbation in the input as unstable outputs can lead to costly transaction fees, privacy and security concerns, reduced user trust, and lack of replicability. Despite the widespread application of graph algorithms, many classical algorithms are not robust to small input disturbances. Towards addressing this issue, we study the pointwise Lipschitz continuity of graph algorithms, a notion of stability introduced by Kumabe and Yoshida (2023, FOCS’23) and further studied in related settings (Kumabe and Yoshida, 2024, ICALP’24), (Kumabe and Yoshida, 2025, SODA’25), (Gima et al., 2025, ESA’25). Our main result is a linear programming (LP) based minimum S-T cut algorithm with a provably optimal Lipschitz constant, as witnessed by an accompanying lower bound. As a direct corollary, we give the first dynamic minimum S-T cut algorithm with non-trivial recourse bound. At the core of our techniques is a novel framework for analyzing the Lipschitz constant of regularized LP relaxations. Our framework crucially unlocks the use of weighted regularizers, which could not be analyzed through previous methods and leads to polynomial improvements in the Lipschitz constant compared to what is achievable through previous techniques. To demonstrate the flexibility of our methods, we also design an LP-based b-matching algorithm that improves on the state-of- the-art Lipschitz constant in certain input regimes when b ≡ 1. Moreover, our algorithm cleanly extends to the general case when b ≥ 1, whereas prior works are specialized to the case of b ≡ 1.
Policy-DRIFT: Dynamic Reward-Informed Flow Trajectory Steering
Atharva Mahajan ⋅ Abhijeet Vishwasrao ⋅ Yuning Wang ⋅ Ricardo Vinuesa
Skin-friction drag induced by wall-bounded turbulent flows accounts for a substantial fraction of energy consumption across commercial aerospace, wind energy, and marine transport. Its active reduction is one of the highest-value targets in engineering fluid dynamics. Deep reinforcement learning (DRL) has emerged as the leading approach for real-time flow control, yet its performance ceiling is set not by algorithmic capability but by reward structure, the naive scalar objective does not optimally reflect the underlying physics. Policy-DRIFT bypasses this ceiling by relocating reward information from policy gradients to generative model inference: a conditional flow matching model (CFM) constructs a physically-grounded manifold of realisable flow states spanning multiple control regimes, Terminal Reward Guidance (TRG) steers samples toward reward-maximising targets at inference, and a lightweight DRL policy, structurally decoupled from reward quality, tracks these full-field targets via root-mean-squared error (RMSE) minimisation. The test case is turbulent channel flow simulated using direct numerical simulation (DNS) at friction Reynolds number of $\mathrm{Re}_\\tau = 180$, which is the canonical benchmark for wall-bounded turbulence. Policy-DRIFT achieves $49\\%$ drag reduction approaching the theoretical upper bound, which is $\\approx 16\\%$ higher than the DRL benchmark, while consuming 37$\\times$ less actuation energy. Our approach combines generative methods with active flow control, marking a paradigm shift towards controlling complex physical systems efficiently.
The discourse on privacy risks in Large Language Models (LLMs) has disproportionately focused on verbatim memorization of training data, while more immediate and scalable privacy threats remain underexplored. We posit that LLM privacy must be understood as a lifecycle-wide problem, rather than being reduced to training-data leakage alone. We introduce a taxonomy of five open privacy problems posed by LLMs and flesh out each with concrete threat models, real-world case studies, and open research challenges. Through a longitudinal analysis of 1,772 AI/ML privacy papers from leading conferences (2016--2025), we reveal a disproportionate research focus: memorization dominates technical research, yet offers little traction against pressing problems like inference-time context leakage, autonomous agent behavior, and surveillance-enabling data aggregation. We provide a roadmap of technical, sociotechnical, and policy interventions, and call on the community to prioritize the underexplored problem categories---agent-based leakage, inference attacks, and data aggregation---that collectively receive less than 9\% of current research attention.
Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
Yuta Kobayashi ⋅ Pradyun Ramesh ⋅ Muhammad Ahmed Chaudhry ⋅ Vincent Jeanselme ⋅ Judy Wawira ⋅ Sanmi Koyejo ⋅ Kathleen M Capaccione ⋅ Shalmali Joshi
Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. For example, prior studies show that cardiomegaly may be omitted from ICU chest X-ray reports when the imaging request is focused on monitoring support device placement. As a result, models trained with standard approaches inherit these omissions, learning to under-report findings themselves. We propose PU-DPO, a preference optimization framework that constructs contrastive pairs via editing model responses, producing variants that explicitly mention or omit a specific finding. To prevent omission noise from corrupting the preference signal, we reformulate the objective under a positive-unlabeled (PU) learning framework, treating absent mentions as unlabeled rather than truly negative. Across semi-synthetic experiments and preliminary analyses on real-world chest radiograph benchmarks where adjudicated labels are available, PU-DPO yields consistent gains in detection rates across multiple pathologies without compromising specificity or overall report quality, and is more robust to omission noise than prior approaches.
Predicting and improving test-time scaling laws via reward tail-guided search
Muheng Li ⋅ Jian Qian ⋅ Wenlong Mou
Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ``best-of-$N$'' (BoN) strategy has already demonstrated significant improvements in performance, it lacks principled guidance on the choice of $N$, budget allocation, and multi-stage decision-making, thereby leaving substantial room for optimization. While many works have explored such optimization, rigorous theoretical guarantees remain limited. In this work, we propose new methodologies to predict and improve scaling properties via tail-guided search. By estimating the tail distribution of rewards, our method predicts the scaling law of LLMs without the need for exhaustive evaluations. Leveraging this prediction tool, we introduce Scaling-Law Guided (SLG) Search, a new test-time algorithm that dynamically allocates compute to identify and exploit intermediate states with the highest predicted potential. We theoretically show that SLG approaches a full-information oracle at a reduced scale and, under additional regularity conditions, achieves vanishing matched-budget regret; in a Gaussian two-stage model, SLG can also match BoN with a polynomially larger sampling budget. Empirically, we validate our framework across different LLMs and reward models, confirming that tail-guided allocation consistently achieves higher reward yields than Best-of-$N$ under identical compute budgets.
Predicting Plasticity in Deep Continual Learning: A Theoretical Perspective
Jiuqi Wang ⋅ Jayanth Srinivasa ⋅ Claire Chen ⋅ Shuze D Liu ⋅ Ali Payani ⋅ Shangtong Zhang
Deep continual learning requires models to adapt to new tasks without retraining from scratch. However, neural networks can lose their ability to adapt to new tasks after training on previous ones, a phenomenon known as loss of plasticity. There have been several explanations and diagnostics proposed for plasticity loss. Motivated by the philosophy that "all models are wrong, but some are useful", we ask: can existing diagnostics predict a neural network's plasticity? In this work, we take a practical view to interpret plasticity as trainability, i.e., a neural network's future optimization gain on a target task. We first take a theoretical approach, showing, by constructing a few counterexamples, that some widely adopted diagnostics of plasticity, including representation rank and neural tangent kernel rank, can fail to predict the loss of trainability in both regression and classification settings. We instead propose a novel metric, called optimization readiness, which combines gradient strength and gradient reliability. We prove that optimization readiness lower bounds one-step optimization gain under standard smoothness assumptions, providing a theoretical guarantee for its predictive power. Empirically, we show that across commonly used deep continual learning settings, such as Slowly-Changing Regression and Permuted MNIST, optimization readiness more reliably ranks checkpoints by trainability than prior diagnostics, even with substantially fewer samples.
Prompt–Activation Duality: Improving Activation Steering via Attention-Level Interventions
Diancheng Kang ⋅ Zheyuan (Frank) Liu ⋅ Ningshan Ma ⋅ Yue Huang ⋅ Zhaoxuan Tan ⋅ Meng Jiang
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful dialogue. We identify KV-cache contamination as a key failure mode: steered token states are stored and repeatedly reused, turning a local perturbation into cumulative coherence degradation. Motivated by the stability of prompt-based control, we propose Gated Cropped Attention-Delta steering (GCAD), which extracts steering signals from system-prompt contributions to self-attention and applies them with token-level gating. Across persona-steering experiments, GCAD preserves trait control while substantially improving long-horizon coherence. On the main multi-turn benchmark, GCAD improves average coherence drift from -18.6 to -1.9 and raises turn-10 trait expression from 78.0 to 93.1. These results suggest that activation steering becomes more reliable when interventions follow the prompt-mediated pathways that models already use for behavioral control.
Quantum Speedup of Multi-armed Bandits at Scale by Tackling Memory Decoherence
Rohit Muralitharan ⋅ Mark Squillante ⋅ Chen Wang ⋅ Xuchuang Wang ⋅ Chai Wah Wu
Recent advances in online learning have achieved an $O(n/\Delta_{[2]})$ query complexity in quantum $n$-arm multi-armed bandits (MABs) for best-arm identification, which represents a quadratic speedup over their classical counterparts. Here and throughout, $\Delta_{[2]}$ is the gap between the mean of the best and second-best arms. However, the implementations of these algorithms remain on a small scale due to various aspects of quantum hardware restrictions. One of the most significant restrictions is \emph{memory decoherence}, in which the qubits quickly lose information after a short period of time. Memory decoherence can occur at \emph{intra-circuit} and \emph{inter-circuit} levels. Intra-circuit decoherence happens often in near-term quantum computers, where the noise becomes overwhelming when the circuit is deep. On the other hand, inter-circuit decoherence captures future quantum machines, where a quantum memory is promising, but objects stored in the memory for too long can still become decoherent. In this paper, we present a comprehensive treatment of both memory decoherence problems. For inter-circuit decoherence, we derive a framework based on streaming bandits, and we obtain an algorithm that finds the best arm with high constant probability using $O(n/\Delta_{[2]})$ queries and an $O(\log{n})$ decoherence window. This is the first quantum algorithm for MABs that achieves an $o(n)$ decoherence window, which represents an exponential improvement. For intra-circuit decoherence, we observe that only $O(1)$-depth circuits can be used to avoid significant decoherence. Therefore, we devise various hardware optimizations and a destructive SWAP test subroutine to speed up the empirical performance. On the IBM Q System One machine with 127 qubits, our implementation scales the quantum MABs to over $1000$ arms, and our algorithms consistently achieve significant efficiency improvements over classical bandit algorithms in both physical experiments and computer simulations.
Qubrio: High-Performance Quantum Compilation via Multi-Agent LLM Collaboration
Jixuan Ruan ⋅ Zhuo Cui ⋅ Zhengding Hu ⋅ Zhongkai Yu ⋅ Xiang Fang ⋅ Yue Guan ⋅ Jason Ludmir ⋅ Phillip Weinberg ⋅ Xiu-Zhe Luo ⋅ Sheng-Tao Wang ⋅ Sonia L Alarcon ⋅ Hezi Zhang ⋅ Yufei Ding
Large Language Models (LLMs) offer an adaptable alternative to traditional quantum compilation heuristics, which bottleneck progress by demanding costly manual redesigns whenever hardware evolves. However, naive single-pass LLM approaches fail under the unique challenges introduced by quantum compilation. We introduce Qubrio, a practical agentic framework that significantly reduces heuristic redesigns through a three-level decomposition. By (1) partitioning programs into sequential operation stages, (2) deploying specialized agents for placement, routing and optimization, and (3) using disaggregated feedback for precise error correction, our system reliably navigates complex physical constraints. Beyond serving as a direct compiler, our LLM framework also uncovers novel strategies that facilitate convoy-style shuttling, substantially enhancing hardware concurrency without sacrificing fidelity. We further integrate these strategies into the previous state-of-the-art (SOTA) compiler, recovering a significant fraction of the LLM's performance advantages. On realistic workloads, our compiler achieves significant improvements in hardware runtime and program fidelity over the SOTA baseline PowerMove, seamlessly adapting to new hardware capabilities. Qubrio features an open user interface for practical use, with its anonymized repository available for double-blind review at https://anonymous.4open.science/r/Qubrio-486C.
Diffusion models generate samples by iteratively querying learned score estimates. A rapidly growing literature focuses on accelerating sampling by minimizing the number of score evaluations, yet the information-theoretic limits of such acceleration remain unclear. In this work, we establish the first score query lower bounds for diffusion sampling. We prove that for $d$-dimensional distributions, given access to score estimates with polynomial accuracy $\varepsilon=d^{-O(1)}$ (in any $L^p$ sense), any sampling algorithm requires $\widetilde{\Omega}(\sqrt{d})$ adaptive score queries. In particular, our proof shows that any sampler must search over $\widetilde{\Omega}(\sqrt{d})$ distinct noise levels, providing a formal explanation for why multiscale noise schedules are necessary in practice.
RAD-TFM: Robust and Domain-Adapted Tabular Foundation Models
Matthew Peroni ⋅ Franck Le ⋅ Vadim Sheinin
The development of tabular foundation models (TFMs) has accelerated in recent years, showing strong potential to outperform traditional ML methods for structured data. A key finding is that TFMs can be pretrained entirely on synthetic datasets, opening opportunities to design data generators that encourage desirable model properties. Prior work has mainly focused on crafting high-quality priors over generators to improve overall pretraining performance. Our insight is that parameterizing the generator distribution enables an adversarial, distributional robustness perspective: during training, we can adapt the generator to emphasize datasets that are particularly challenging for the model, which we formalize by introducing an optimality gap measure. Further, we develop a method to align the synthetic data generation with real-world datasets from a given domain, constraining the adversary to generate "realistic'" data. Together, these algorithms comprise the Robust and Domain-Adapted Tabular Foundation Models (RAD-TFM) pipeline, a model-agnostic adversarial training framework. Applied to the TabPFN V2 classifier, RAD-TFM improves performance across 6 diverse tabular benchmarks, with up to a 11\% increase in mean normalized AUC over the original TabPFN and other baseline algorithms, with only 100k additional training datasets, less than 0.1\% of the original pretraining data. These results highlight a promising new direction for targeted adversarial training and fine-tuning of TFMs using synthetic data alone.
RAHF: Reward-Amplified Human Feedback for Closed-Loop Policy Fine-Tuning
Haoyuan Cai ⋅ Seth Zhao ⋅ Jason Zhang ⋅ Bolei Zhou
Policies trained on offline logs can perform well in open-loop evaluation but fail in closed-loop deployment, because they never learn to recover from their own mistakes. Fine-tuning these policies typically requires either many human demonstrations or large-scale reinforcement learning rollouts, both of which are expensive. We propose RAHF, a two-stage closed-loop fine-tuning framework that reduces human effort by amplifying a small number of human corrections with a cheap verifiable reward. In Stage 1, a human operator monitors the policy during closed-loop rollout and intervenes at dangerous states, producing a small but targeted correction dataset that also keeps the training process safe. In Stage 2, the improved policy runs without any human involvement. At each step, the verifiable reward scores the policy's candidate proposals. We use these reward scores to construct group-normalized advantages and fine-tune the policy toward the higher-advantage proposals. To prevent the policy from forgetting its pretrained knowledge, we propose a Policy Adaptation Module (PAM) that freezes the pretrained weights and learns closed-loop corrections through a separate trainable branch. On BridgeSim-NavHard, RAHF with a human teacher reaches a driving score of 78.90, surpassing all baselines including DAgger, HG-DAgger, GRPO, and AWR. In addition, PAM reduces forgetting and enables zero-shot transfer to HUGSIM without any additional fine-tuning.
Rank Is Not Capacity: Spectral Occupancy for Latent Graph Models
Nikolaos Nakis ⋅ Panagiotis Promponas ⋅ Konstantinos Tsirkas ⋅ Katerina Mamali ⋅ Eftychia Makri ⋅ Leandros Tassiulas ⋅ Nicholas A Christakis
Graph representation learning has become a standard approach for analyzing networked data, with latent embeddings widely used for link prediction, community detection, and related tasks. Yet a basic design choice, the latent dimension, is still treated as a brittle hyperparameter, fixed before training and tuned by held-out performance. Learned factors are also identifiable only up to rotation and rescaling, so the nominal rank rarely coincides with the quantity that governs model behavior. We propose Spectral Prefix Extraction and Capacity-Targeted Representation Analysis (Spectra), which replaces rank as the unit of analysis with the spectrum of a learned positive semidefinite kernel, trace-normalized so that spectra are comparable across fits. The normalized eigenvalues form a distribution on the simplex, and their Shannon effective rank acts both as a summary of learned capacity and as a controllable training-time coordinate: a single scalar shapes this realized dimension during training, and bisection targets any desired value within the rank cap. To theoretically support that, we show local regularity and monotonicity of the realized-dimension profile. Across collaboration, social, biological, and infrastructure networks, Spectra traces performance--capacity frontiers that make the trade-off between predictive accuracy and realized dimension visible. It performs competitively with strong link-prediction baselines, yields aligned lower-capacity views of the same fitted model through spectral prefixes, and provides a principled handle on capacity in the overparameterized regime. Capacity thus becomes a property of the fitted model rather than a hyperparameter of the training.
RAVEN-Bench: A Paired EO-IR Video QA Benchmark for Aerial Multimodal Understanding
Yu Hu ⋅ Jianyang Gu ⋅ Yue Cao ⋅ Hao Liu ⋅ Kangnan Wang ⋅ Jozsef Hamari ⋅ Zheng Liu ⋅ Mohsen Zardadi
Multimodal large language models (MLLMs) have advanced rapidly on image and video understanding, but their ability to reason over paired aerial electro-optical and infrared (EO-IR) videos remains underexplored. Such videos contain small targets, sparse temporal evidence, changing viewpoints, and modality-dependent cues that are poorly captured by RGB-centric benchmarks. We introduce RAVEN-Bench (Reasoning over Aerial Videos with EO-IR aligNment), a paired EO-IR video QA benchmark for low-altitude aerial multimodal understanding. RAVEN-Bench contains 48 temporally aligned EO-IR video pairs and 576 human-verified four-way multiple-choice questions. It supports EO-only, IR-only, and paired EO+IR evaluation, and uses hierarchical capability levels and structured question groups to diagnose cross-modal gain, modality reliance, consistency, and reasoning coherence. We evaluate representative frontier and open-weight MLLMs, showing that aggregate performance can overestimate reliable EO-IR reasoning. RAVEN-Bench provides a focused testbed for assessing whether progress on RGB video benchmarks transfers to aerial cross-spectral understanding.
Reading Positional Coupling in Transformers with Diffusion Scores
Savik Kinger ⋅ Johannes Bertram ⋅ Luciano Dyballa ⋅ Andy Keller ⋅ Steven W Zucker
Understanding how transformers encode position and context is central to interpreting their behavior, because it determines how information flows across tokens and shapes the model's predictions. Common tools, such as linear probes and attribution methods, operate on single-position activations and are not designed to capture cross-position dependencies. We introduce Score-Block Token Geometry (SBTG), a score-based diagnostic that recovers cross-position dependency structure from activations and reveals distinct signatures of positional encoding mechanisms that these methods do not expose. SBTG trains a denoising score model on windows of activations and constructs lag-indexed operators that quantify coupling between positions within a layer. Applied to transformers with different positional encodings, it recovers distinct coupling signatures consistent with their architectures: ALiBi exhibits near rank-one coupling in early layers, RoPE distributes coupling across directions in patterns consistent with its rotation frequencies, and absolute embeddings show the strongest position-specific variation. We show that these signatures are stable across seeds and model sizes, and the activation-space directions identified by SBTG are causally relevant: ablating them causes larger performance drops on tasks that require reasoning over relative positions than on tasks that depend on absolute position signals. These results show that joint activation structure provides a lens on how positional encoding mechanisms are realized, revealing differences in how models use cross-position dependencies that are not apparent from the architecture alone.
RECIPE: Procedural Planning via Grounding in Instructional Video
Luigi Seminara ⋅ Antonino Furnari ⋅ Lorenzo Torresani
Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on this task is bottlenecked by annotation: clean labeled datasets are small, domain-narrow, and encode a single execution trajectory per example, even though many valid orderings often exist. Large-scale instructional video corpora such as HowTo100M offer orders of magnitude more procedural content, but supervised fine-tuning on pseudo-labels extracted from their noisy ASR narrations fails: segmentation and alignment errors propagate into training, and the resulting supervision is still single-trajectory. We identify a key asymmetry. Extracting clean step labels from noisy video is hard, but verifying whether a generated step sequence is temporally grounded in ASR transcripts is comparatively cheap and scales to millions of videos via precomputed text embeddings. We exploit this asymmetry in RECIPE, which uses grounding quality as a reward signal for Group Relative Policy Optimization (GRPO), turning the noisy corpus into a verification signal rather than a labeling source. The framework applies uniformly to two input configurations of the planner: a Socratic pipeline in which a frozen vision-language model rewrites the video into a textual history fed to the planner, and a Video configuration in which the planner consumes video tokens directly. It applies equally to annotated and weakly supervised training regimes. We evaluate on seven procedural benchmarks using a reference-based LLM-as-judge protocol that scores generated plans across six procedural-quality criteria. RECIPE-RL improves over the base checkpoint at every scale we test (0.5B, 3B, 7B) and on every benchmark, with macro-accuracy gains of +7 to +8 points in-domain at every scale and up to +16 points zero-shot. It substantially outperforms supervised fine-tuning on both annotated and pseudo-labeled continuations (the latter actually degrades the base checkpoint), and is robust to fully replacing human annotations with VLM-derived pseudo-traces. Plugging RECIPE-RL into the proposal stage of VidAssist [14] improves over the strongest zero-shot baseline in our comparison at every horizon on the Visual Planning for Assistance benchmark, and a diversity analysis on COIN shows that RECIPE-RL preserves the generation variety that supervised fine-tuning collapses.
RecoverBench: A Systematic Benchmark for Error Recovery in Robotic Manipulation
Zijia Tang ⋅ Ganlong Zhao ⋅ Xingping Chen ⋅ Junye Chen ⋅ Zihao Mo ⋅ Zedi Wang ⋅ Guanbin Li
Robotic manipulation policies have achieved remarkable success in controlled settings, but they catastrophically fail when execution errors occur---dropped objects, collisions, misaligned grasps. Despite the critical importance of error recovery for real-world deployment, no existing benchmark systematically evaluates manipulation policies' ability to recover from such errors. We introduce RecoverBench, the first systematic benchmark for error recovery in robotic manipulation. We propose a taxonomy of 12 Error Skills organized into 5 Recovery Behavior Groups (RBGs), spanning 24 error subtypes across 2 difficulty degrees. Our Error Skill pipeline automatically generates 1,360 reproducible error scenes from clean demonstration trajectories across 6 manipulation tasks. Each error scene includes complete simulation state snapshots, environment fingerprints, and RNG states for deterministic reproduction. Beyond evaluation, RecoverBench also supports recovery-demo collection and augmentation, and we show that this recovery supervision improves recovery performance. We evaluate 4 representative manipulation policies on RecoverBench. Even the best-performing policy reaches only 48.6% average recovery success, ranging from 32.0% on threading to 76.7% on stack, showing that current methods still lack robust recovery. These results highlight the need for dedicated recovery evaluation and position RecoverBench as a benchmark for improving recovery robustness in manipulation policies. Code, data, and evaluation protocols will be released.
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
Ziyue (Alvin) Liu ⋅ Zhengyang Wang ⋅ Ruijie (Ray) Zhang ⋅ Avinash Maurya ⋅ Hui Zhou ⋅ Paul Hovland ⋅ Sheng Di ⋅ Franck Cappello ⋅ Bogdan Nicolae ⋅ Zheng Zhang
Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems. Yet existing frameworks either focus on specific parallelism schemes or risk drifting away from a failure-free training trajectory. We propose ReCoVer, a resilient LLM pre-training system that upholds a single invariant: each iteration keeps the number of microbatches constant, ensuring per-iteration gradients remain stochastically equivalent to a failure-free run. The framework is organized as three decoupled protocol layers: (1) Fault-tolerant collectives that isolate faults from propagating across replicas; (2) in-step fine-grained recovery that preserves intra-iteration progress and prevents gradient corruption; (3) versatile-workload policy that dynamically redistributes microbatch quotas across the survivors. The design is parallelism-agnostic, integrating directly with both 3D parallelism and Hybrid Sharded Data Parallel (HSDP) as a drop-in substrate. We evaluate our implementation on end-to-end pre-training tasks for up to 512 GPUs, ReCoVer successfully preserves the training trajectory from a failure-free reference despite of 256 GPUs lost spread across the run. For comparison with checkpoint-and-restart baselines, ReCoVer demonstrates $2.23\times$ higher effective throughput after successive failures. This advantage results in ReCoVer processing 74.9% more tokens at 234 GPU-hours, with the gap widening as the training prolongs.
RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection
Zhihao Zhang ⋅ Gengwei Zhang ⋅ Tianlong Chen ⋅ Xiaoming Liu
Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector's predicted depth degrades accuracy, even falling below the detector's own prediction. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM's vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.
Regret–Oracle Complexity Tradeoffs in Agnostic Online Learning
Idan Attias ⋅ Steve Hanneke ⋅ Arvind Ramaswami
Agnostic online learning is classically solved via a reduction to the realizable setting, utilizing Littlestone's Standard Optimal Algorithm (SOA) as a base learner. However, the SOA is computationally intractable to execute even for a single round. To overcome this barrier, recent work in oracle-efficient online learning replaces the SOA with a realizable base learner that accesses the concept class exclusively through an offline empirical risk minimization (ERM) oracle. While such agnostic learners achieve near-optimal expected regret, they suffer from a doubly-exponential oracle complexity of $\mathcal{O}\big(T^{2^{\mathcal{O}(d_\mathrm{LD})}}\big)$, where $d_\mathrm{LD}$ is the Littlestone dimension and $T$ is the number of rounds. In this work, we significantly improve this oracle complexity while relying on an even weaker primitive: a weak-consistency oracle, which merely decides whether a given labeled dataset is realizable. At the core of our approach is an adaptive and dynamic agnostic-to-realizable reduction that actively prunes non-realizable label sequences on the fly. By using the VC dimension ($d_\mathrm{VC}$) to bound the number of dynamically maintained active paths, our algorithm reduces the total query complexity down to $\mathcal{O}(T^{d_\mathrm{VC}+1})$ while perfectly preserving near-optimal expected regret. Crucially, this dynamic pruning also yields a memory reduction over the standard reduction. Furthermore, we formally quantify the regret--oracle complexity tradeoff, providing upper bounds that smoothly interpolate between restricted query budgets and attainable expected regret. We complement these with lower bounds proving that any learner restricted to $Q = o(\sqrt{T})$ queries must suffer an expected regret of $\Omega(T/Q)$.
Reinforcement learning enhanced flow matching for reference-based tumor generation on unpaired CT images
Xiaoxuan Gong ⋅ Zhilin Zheng ⋅ Tiancheng Lin ⋅ Qinji Yu ⋅ Jianpeng Zhang ⋅ Shaoteng Zhang ⋅ Kai Cao ⋅ Xiaoli Yin ⋅ Yu Shi ⋅ Jie Ma ⋅ Ling Zhang ⋅ Yingda Xia
The success of deep learning in medical image analysis is often hindered by the long-tail distribution of disease cases, where rare or early-stage tumors are critically under-represented. While controllable synthesis offers a potential solution, existing methods suffer from a \textit{specification bottleneck}, as low-dimensional signals like text or binary masks fail to capture complex textures and morphological variations. In this paper, we propose an exemplar-driven synthesis framework that utilizes real tumor samples as high-dimensional templates to enable precise, case-specific generation. To the best of our knowledge, this is the first work to introduce reinforcement learning (RL) into the field of 3D medical image generation, called MedGRPO. We formulate the unpaired tumor synthesis task as a constrained trajectory optimization problem and leverage Flow-GRPO to find an optimal generation path. By designing a multi-objective reward system encompassing texture, shape, and semantic consistency, our model effectively balances pathological fidelity with anatomical coherence without requiring paired supervision. Rigorous evaluations across four major tumor types and the external AbdomenAtlas 2.0 dataset demonstrate that our method produces high-fidelity synthetic data that significantly enhances the performance of downstream diagnostic models.
REINS: Learning Inertia-Induced Geometry for Physics-Consistent Motion Representation in Clinical Gait Phenotyping
Yiran Ding ⋅ Yufei Zhang ⋅ Zijun Cui
Parkinson's disease and other gait-related neurological disorders manifest in subtle but biomechanically meaningful changes in how patients move, making computational gait phenotyping a valuable tool for clinical assessment. Yet existing latent representations of motion are typically learned from data alone and miss the biomechanical structure that gives clinical motion its diagnostic meaning, leaving them vulnerable when training cohorts are small or acquisition sites differ. We propose REINS (Riemannian Embedding for INertia-induced motion Spaces), an Euler-Lagrange-induced motion representation that uses the generalized inertia matrix of articulated systems to define a biomechanically grounded geometry on configuration space. Because kinetic energy is itself a quadratic form in velocity weighted by the inertia matrix, this metric is a mechanically determined Riemannian structure, and we learn a latent embedding of the resulting structured space. By defining the representation space itself through rigid-body dynamics rather than enforcing physics as a post-hoc penalty, REINS aligns the geometry of motion representation with the same physical laws clinicians rely on to interpret gait. We apply the framework on Parkinson's disease gait analysis, where deviation from healthy motion is measured as a geodesic distance under the inertia-induced metric, and report results on healthy-control vs.\ PD discrimination and UPDRS-gait severity correlation.
Rep2Text: Decoding Full Text from a Single LLM Token Representation
Haiyan Zhao ⋅ Zirui He ⋅ Yiming Tang ⋅ Fan Yang ⋅ Ali Payani ⋅ Dianbo Liu ⋅ Mengnan Du
Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental question: to what extent can the original input text be recovered from a single last-token representation in an LLM? To this end, we propose Rep2Text, a novel framework for decoding text from last-token representations. Rep2Text employs a trainable adapter that maps a target model’s last-token representation into the token embedding space of a decoding language model, which then autoregressively reconstructs the input text. Experiments across various model combinations (Llama-3.1-8B, Gemma-7B, Mistral-7B-v0.1, Llama-3.2-3B, etc.) show that, on average, roughly half of the tokens in 16-token sequences can be recovered from this compressed representation while preserving strong semantic coherence. Further analysis reveals a clear information bottleneck effect: as sequence length increases, token-level recovery declines, while semantic information remains relatively well preserved. We also find that scaling effects are less pronounced in inversion tasks. Finally, our framework demonstrates robust generalization to out-of-distribution clinical data.
RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
Xichen Pan ⋅ Satya Narayan Shukla ⋅ Aashu Singh ⋅ Shlok K Mishra ⋅ Saining Xie
Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text encoding, while denoising is handled by newly trained generative backbones. The emergence of representation autoencoders (RAEs) shifts the generation target toward semantically structured visual representations, creating a latent space that is more compatible with pretrained LLM priors. Inspired by multimodal LLMs (MLLMs), where an MLP projector is sufficient to align clean visual representations with a pretrained LLM, we repurpose the MLLM itself as a noisy representation encoder, extending this mechanism from clean to noisy inputs. We present RepFusion, which uses the resulting MLLM outputs as the conditioning signal for a diffusion transformer. In controlled, parameter-matched comparisons, RepFusion outperforms scaling the denoising transformer with newly initialized parameters. These results demonstrate that MLLMs provide strong priors for denoising visual representations and that scaling the compute of the conditional encoder is feasible in modern T2I systems.
RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms
Kenan Li ⋅ Rongzhi Li ⋅ Linghao Zhang ⋅ Qirui Jin ⋅ Liao Zhu ⋅ XiaoSong Huang ⋅ Geng Zhang ⋅ Yikai Zhang ⋅ Shilin He ⋅ Chengxing Xie ⋅ Xin Zhang ⋅ Zijian Jin ⋅ Bowen Li ⋅ Chaoyun Zhang ⋅ Yu Kang ⋅ Yufan Huang ⋅ Elsie Nallipogu ⋅ Saravan Rajmohan ⋅ Qingwei Lin ⋅ Dongmei Zhang
Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In this work, we introduce RepoLaunch, a novel agentic framework that automatically resolves dependencies, compiles source code, and extracts test results across diverse programming languages and operating systems. RepoLaunch achieves a 78\% build success rate, outperforming the Python/Linux-only prior system by 18\%. To demonstrate its application, we further present a fully automated pipeline for SWE dataset creation driven by RepoLaunch, which only requires human input at the task-design stage. RepoLaunch is open-sourced, and its automated task-generation pipeline has already been adopted by several recent works on agentic benchmarking and training.
Residual-Autoregressive Context for 3D Gaussian Splatting Compression
Wenqing Wang ⋅ Huimin Zeng ⋅ Yun Fu
3D Gaussian Splatting (3DGS) has emerged as a promising method that enables fast and high-quality novel-view rendering. However, its large parameter size remains a bottleneck for storage and transmission. We propose *RacerGS*, a compact 3DGS compression framework that achieves scene reconstruction at significantly reduced storage sizes with preserved fidelity. To obtain compact and spatially coherent features, we design a hash projection model that projects the anchor hash encodings into compact context features. Additionally, we introduce a FiLM-modulated autoregressive GRU context model that utilizes causal dependencies among anchor feature groups and uses recurrent hidden states to refine entropy-model parameters, improving bitrate efficiency. Furthermore, we adaptively entropy-code quantized residuals of anchor attributes around their context-predicted means, reducing symbol entropy by leveraging spatial consistency. Overall, *RacerGS* achieves more than **$\mathbf{150}\boldsymbol{\times}$** compression over vanilla 3DGS and **$\mathbf{28}\boldsymbol{\times}$** over Scaffold-GS, while providing comparable rendering quality. Extensive experiments across five datasets demonstrate our superiority over prior 3DGS compression methods in both storage reduction and rate-distortion performance.
Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection
Manni Cui ⋅ Ziheng Qin ⋅ ZiAn Wang ⋅ Ruiqi Liu ⋅ Dianyuan Zou ⋅ Jianglan Wei ⋅ Han Zhou ⋅ Yu Liu ⋅ Xu Jingrui ⋅ Wenhao Wang ⋅ Zhenyu Zhang
AI-generated videos (AIGVs) typically contain subtle temporal artifacts that arise from inter-frame inconsistencies rather than within individual frames. A detector that captures such artifacts should therefore benefit from video pretrained backbones over image only ones. In practice, however, video backbones with standard global readouts often fail to outperform strong image pretrained probes on AIGV benchmarks. We attribute this gap to excessive spatiotemporal aggregation in the readout. Video pretrained backbones tend to compress each frame into a single global descriptor. This compression suppresses local patch level temporal dynamics and discards inter patch relations, which are precisely the cues that AIGV detection most reliably depends on. Based on this, we propose Velocity Gated Patch Velocity Profiling (V-PVP), a lightweight readout that replaces only the aggregation layer with two parallel streams over the patch velocity field, adding only about $0.5$M trainable parameters. V-PVP serves as a general plug-and-play module that consistently improves performance across diverse video backbones under both end-to-end fine-tuning and linear probing settings. Our method reaches \textbf{95.28} AUC on AIGVDBench while keeping the backbone fully frozen. The results show that simply replacing the aggregation layer reactivates the temporal potential of frozen video backbones, restoring their advantage on AIGV detection. Code is available at https://anonymous.4open.science/r/PVP-81B3/.
RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents
mingchen li ⋅ Hansi Zeng ⋅ Zhuo Qian ⋅ Jiatan Huang ⋅ Sunjae Kwon ⋅ Hamed Zamani ⋅ Hong Yu
Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate queries, and search again. Training such agents raises a credit-assignment challenge: executable actions such as queries or summaries can be directly evaluated by the retriever, while latent reasoning steps are not directly observable and only affect future executable actions. This asymmetry makes outcome-level reward assignment unreliable, as the same final reward may credit reasoning steps that did not actually shape retrieval success. We propose RICE-PO, a critic-free policy optimization framework that converts retrieval interactions into localized learning signals. RICE-PO selects high-uncertainty executable actions as anchors, evaluates local counterfactual branches using retrieval metrics, and propagates credit to latent reasoning steps only when reasoning-to-action influence is strong and future residual effects are stable. On BRIGHT and BEIR, RICE-PO consistently outperforms prompt-based agents and group-based RL baselines under the same retriever setting. These results show that the structure of agent-environment interaction itself can provide useful supervision for training reasoning-based retrieval.
RILA: A Radar-Native Structured Interface from Sparse mmWave Point Clouds to Large Language Models
Shenglei Li ⋅ Tomoji Kishi
Millimeter-wave radar is promising for privacy-preserving human sensing, yet current radar-to-language pipelines typically map sparse point clouds into either global clip features or discrete token codes that are only weakly matched to the structure expected by large language models (LLMs). We present RILA, a radar-native structured interface layer for large-language-model adaptation, which converts sparse mmWave radar point clouds into LLM-readable event-aware tokens. RILA combines kinematic phase-space tokenization, flow-aware dual-order serialization, dual-scan selective state-space encoding, and an event-aware token abstraction that exposes both global clip context and temporally localized motion units to the language model. The same interface supports three language tasks through a shared decoder: clip summary generation, ordered event description, and temporal question answering. To stabilize the interface before instruction tuning, we introduce a structured interface-language alignment objective over clip and event tokens. We evaluate RILA using physics-aware synthetic pretraining and real-world mmWave adaptation under controlled radar-native baselines and matched decoder settings. This formulation reframes radar-to-LLM modeling as an interface design problem and studies how radar-native structure, rather than only generic cross-modal projection, determines whether sparse mmWave point clouds can be effectively understood by LLMs.
Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees
Yaacov Pariente ⋅ Vadim Indelman
Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation. We inherit finite-time guarantees for policy evaluation and sparse sampling---with estimation error independent of the risk level---and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.
Row-Private Symmetric Cone Programming: Scale-Efficient Algorithm and Lower Bound
Deming Chu ⋅ Ruizhe Zhang
We study the row-private approximate feasibility problem for symmetric cone programs (SCPs), a framework that encompasses linear, second-order cone, and semidefinite programming. We focus on high-sensitivity row privacy, where each record is a constraint and changing one record can cause the feasible region to change discontinuously. Since satisfying every private row is generally impossible in this regime, the goal is to remain exactly in the public cone while violating only a small number of private rows under a prescribed additive error $\eta>0$. We give an efficient $(\varepsilon, \delta)$-differentially private SCP solver with violation count $\widetilde{O}\left(k^2/\varepsilon\right)$, and only logarithmic dependence on the scale-to-error ratio $UR/\eta$, where $k$ is the ambient dimension, $R$ is the norm of some feasible point, and $U$ is the size of the problem. This exponentially improves the polynomial dependence on $U$, $R$, and $1/\eta$ in prior work (Song, Xue, and Zhang, NeurIPS 2025). Furthermore, we show that additive error does not remove the intrinsic dimension-dependence barrier. Previous row-private LP lower bounds (Kaplan, Mansour, Moran, Stemmer, and Tur, STOC 2025) established an $\Omega(k/\varepsilon)$ violation count only for exact partial satisfaction; using the padding-and-permuting fingerprinting codes (Peter, Tsfadia, and Ullman, COLT 2024), we prove the same barrier for relaxed satisfaction.
RSPO: Reasoning-Supervised Policy Optimization for Long-Tail Autonomous Driving
Jiayi Guan ⋅ jing wu ⋅ Jianhua Wu ⋅ Jinghui Lu ⋅ Zhijian Huang ⋅ Guang Li ⋅ Xiaoshuai Hao ⋅ Hanbing Li ⋅ Yigu Ge ⋅ Tao Xu ⋅ Haiyang Sun ⋅ Bing Wang ⋅ Guang Chen ⋅ Kuiyuan Yang ⋅ Hangjun Ye ⋅ Long Chen
Vision-Language Models (VLMs) have achieved remarkable advances in general autonomous driving. Nevertheless, effectively exploiting the reasoning capability of Chain-of-Thought (CoT) to enhance the robustness and accuracy of decision and planning in long-tail scenarios still remains a challenging open problem. To this end, we propose a \textbf{R}easoning-\textbf{S}upervised \textbf{P}olicy \textbf{O}ptimization (RSPO) algorithm to boost reasoning and decision-making performance under long-tail driving scenarios. Concretely, we first conduct an empirical analysis of the core challenges faced by current VLMs in autonomous driving decision and planning tasks, and formally define reasoning and decision-making tasks tailored to long-tail scenarios. On this basis, we design a curriculum-guided supervised fine-tuning strategy to enable the model to rapidly fit the mean distribution of model outputs. Furthermore, we propose a reasoning-supervised optimization algorithm, which enhances the robustness and accuracy of reasoning and decision in long-tail scenarios by guiding the optimality consistency between reasoning processes and final decisions. Finally, we construct multiple verifiable autonomous driving reasoning and decision datasets based on open-source benchmarks and conduct extensive multi-dimensional comparative experiments. Experimental results show that the proposed algorithm achieves consistent improvements in prediction accuracy across multiple datasets compared with existing methods. In particular, it achieves a 6.7\% relative improvement over the strong baseline on the CODA-LM task and also yields lower trajectory-prediction errors on downstream planning tasks.
RxGS: Receiver-Generalizable 3D Gaussian Splatting for Radio-Frequency Data Synthesis
Kang Yang ⋅ Mani Srivastava
Radio-frequency (RF) data synthesis predicts the received signal given transmitter and receiver positions, and is essential for wireless applications. Recent 3D Gaussian Splatting (3DGS)-based methods achieve efficient synthesis at any transmitter but only for a fixed receiver. Therefore, supporting $N$ receivers in one scene requires $N$ independent models and precludes prediction at unseen receivers. We present RxGS, which achieves receiver-generalizable synthesis within a single unified model. Our key insight is that scene geometry is receiver-independent while directional radiance is not: a first stage learns shared 3D Gaussian geometry, and a second stage freezes it and learns directional radiance conditioned on receiver position. A global conditioning branch captures shared receiver-dependent effects across the scene, while a local branch models per-scatterer variations from the receiver's geometry and occlusion. A multi-receiver CUDA rasterizer further batches rendering across all $N$ receivers. Evaluated across various RF datasets, RxGS matches or improves over per-receiver baselines with a single shared model and generalizes to receivers unseen during training within the scene, cutting training cost by up to $45\times$, inference cost by $7.6\times$, and storage by $N\times$.
Safe in Its Own Words: Self-Guided Safety Alignment for Multimodal Reasoning Models
Adeel Yousaf ⋅ Souradip Chakraborty ⋅ Mubarak Shah ⋅ Amrit Singh Bedi
Multimodal large reasoning models can reason over images and text, but they remain vulnerable to multimodal jailbreaks. Recent safety-alignment methods reduce attack success rate (ASR) by fine-tuning on reasoning traces from stronger teacher models. We show that this has a hidden cost of over-refusal. The aligned model learns to reject benign inputs that contain sensitive words or visual cues. This is especially harmful for boundary-safe inputs, where the prompt looks risky on the surface but is safe in context. We identify external teacher supervision as a key factor in this behavior. It introduces a distributional mismatch and shifts supervision away from the base model's own reasoning distribution. To address this, we propose SAGE (Safety-Aware Guided Elicitation), a self-guided data generation framework for multimodal safety alignment. SAGE guides generation in two stages: a reasoning-level guide first elicits a safety-aware trace, and a decision-level guide then elicits the final response conditioned on that trace. This factorized guidance lets SAGE shape both the reasoning process and the final safety decision without relying on external teacher traces. On LLaVA-CoT, SAGE reduces FigStep ASR from 86.2% to 2.1% and achieves an average over-refusal of 43.1, compared with 66.9--71.1 for prior reasoning-based safety-alignment methods, while preserving general utility. We will release the SAGE training dataset.
SAGE: Evidence-First Biomarker Discovery through Multi-Agent Reasoning
Sahar Almahfouz Nasser ⋅ Juan Francisco Pesantez Borja ⋅ Jincheng Liu ⋅ Sandeep Manandhar ⋅ Shikhar Shiromani ⋅ Mohammad T Hasan ⋅ Zenghan Wang ⋅ Suman Ghosh ⋅ Jinchu Li ⋅ Xuejian Xu ⋅ Aniket R Iyer ⋅ Naoto Tokuyama ⋅ Twisha Shah ⋅ Tilak Pathak ⋅ Soundharya Kumaresan ⋅ Yohei Abe ⋅ Himanshu Maurya ⋅ Anant Madabhushi
Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, guided by fragmented literature rather than rigorous biological validation. We introduce SAGE (Structured Agentic system for hypothesis Generation and Evaluation), a multi-agent framework that grounds biomarker discovery in biological evidence through three mechanisms: (i) knowledge-graph-anchored hypothesis generation via multi-path ontological reasoning, (ii) a debate-based multi-agent novelty assessment that stress-tests candidate biomarkers against existing literature, and (iii) an end-to-end automated validation pipeline that translates hypotheses directly into executable analyses on multimodal pathology datasets. Together, these components shift biomarker discovery from an intuition-driven, literature-browsing exercise into a structured, traceable reasoning process that clinicians and researchers can inspect, trust, and build upon.
Sample Complexity of Linear Regression under Random-Location Coordinate Corruptions
Ilias Diakonikolas ⋅ Jingyi Gao ⋅ Daniel Kane ⋅ Thanasis Pittas
We study Gaussian linear regression under coordinate-wise corruptions and missingness in high dimensions. A large body of work in statistics investigates estimation under different missingness and corruption mechanisms—ranging from missing completely at random to missing not at random—each leading to qualitatively different behaviors. Recent work on the problem considered a strong missing-not-at-random model, where an adversary can inspect the data and corrupt or erase an $\eta$-fraction of entries in each coordinate. A striking consequence of this model is an information-theoretic breakdown at $\eta = \Theta(1/\sqrt{d})$, beyond which non-trivial estimation is impossible even with infinite samples—an unusually small threshold compared to classical robust estimation settings. A natural question is whether this phenomenon is inherent or a consequence of adversarial control over corruption locations. To investigate this, we consider a more benign model in which corruption locations are chosen at random, and the adversary can affect the data only in a subset of these locations. We show that randomizing the locations fundamentally changes the answer. There is no sharp infinite-sample breakdown: non-trivial estimation is now possible for every $\eta < 1$. However, the price is sample complexity. We prove matching upper and lower bounds showing that the sample complexity for non-trivial estimation scales exponentially with $\eta^2 d$. Additionally, this sample complexity also includes a dependence on the signal-to-noise ratio $\sigma^2 / \|\beta\|^2$ which is another new phenomenon in this model.
SarcBench: A Bilingual Benchmark for Contextual Sarcasm Understanding, Response, and Generation
Dihong Huang ⋅ Zhuoyue Chang ⋅ Yuhao Liu ⋅ Zhouting Mo ⋅ Jianxing Yu ⋅ Wenqing Chen ⋅ Jingping Liu
Sarcasm understanding and response are important for large language models (LLMs) to engage naturally in real-world conversations. However, existing sarcasm benchmarks primarily focus on detection or related subtasks in isolation, lacking diagnostic power and failing to reflect the complexity of real conversational interactions. To address this gap, we introduce SarcBench, a bilingual benchmark with 30,083 aligned samples. It defines three interconnected tasks, Intent Recognition, Sarcasm Response, and Sarcasm Generation, enabling diagnostic evaluation of sarcasm understanding and conversational behavior. We evaluate 12 contemporary language models on SarcBench and find that even the top-performing model achieves only 61.21%. We also identify a clear understanding–action gap: models are consistently better at recovering intended meanings than at producing socially appropriate responses or generating sarcasm that faithfully realizes a specified rhetorical mechanism. In addition, performance varies substantially across models and task types, and controlled sarcasm generation often collapses rhetorical diversity into a narrow set of default templates. Taken together, these findings indicate that current models remain limited in handling sarcasm as an interactive social behavior rather than merely a recognition task.
Say the Same, Act Differently: Text-Orthogonal Action Subspaces in Reasoning Vision-Language-Action Models
Zihao Feng ⋅ Qingzhao Zhang ⋅ Chunyu Xia ⋅ Bo Yu ⋅ Zhuoqing Morley Mao ⋅ Ramesh Govindan
Vision-language-action (VLA) policies increasingly generate natural-language rationales before executing embodied actions. These rationales are attractive as monitors for safety-critical behavior because users or automated checkers can inspect whether the model appears to understand the scene and intend a safe action before execution. We show that this signal can be misleading. In many reasoning VLAs, the rationale generator and action head share continuous hidden states, while the generated text reveals only part of that representation. This creates a Say the Same, Act Differently failure mode, where rationale text remains stable while the action output changes substantially. We first study this phenomenon by introducing a new statistical method to separate text-sensitive and action-sensitive directions in hidden-state space of a reasoning VLA. Across the driving VLA Alpamayo and the manipulation VLA InstructVLA, we show that the dominant action direction lies largely outside the active text span, with mean projection ratios of only 1.9% and 5.3%. We define the residual as a text-neutral action-subspace (TNAS) direction, which tests whether action-relevant directions remain after removing measured text-sensitive directions. We then use TNAS to guide white-box pixel-space projected-gradient-descent (PGD) attacks on Alpamayo. TNAS-guided PGD successfully uses a localized two-forward-camera patch to induce a 30.5 m mean ADE shift in trajectory planning while preserving the exact rationale in 86.7% of cases. Moreover, across multiple text-preserving PGD objectives, the optimization direction consistently aligns with TNAS. These results reveal a fundamental shared-cache monitorability gap in reasoning VLAs: stable rationale text should not be treated as sufficient evidence that the action output remains stable.
Scaling Arbitrary Architectures and Optimizers with Automatic Parameterization
Shikai Qiu ⋅ Charlie Chen ⋅ Andres Potapczynski ⋅ Martin Marek ⋅ Andrew Wilson
Training neural networks at scale requires careful per-layer scaling of initialization, learning rate, and other optimizer hyperparameters, with the right rules depending jointly on the architecture, the optimizer, and which axis (width $D$, depth $L$, batch size $B$, etc.) is being scaled. Current practice is to derive each rule by hand, as in $\mu$P for width, CompleteP for depth, SDE-based and EMA-timescale arguments for batch size, and tailored prescriptions for matrix-preconditioned optimizers; as architectures, optimizers, and scaling axes continue to evolve, this hand derivation is increasingly a bottleneck to combining these advances at scale. We show that these derivations can be systematized by solving a linear system over hyperparameter exponents to satisfy the scaling constraints imposed by each primitive instruction of the training program, analogous to dimensional analysis in physics that enforces matched units on both sides of an equation. We develop \emph{automatic parameterization} (AutoP), a system that reads these rules off a traced training program and solves the resulting system to assign a scaling exponent to every hyperparameter, and provide a concrete JAX implementation. From a few primitive rules, AutoP recovers $\mu$P for width on Transformers, Monarch-structured Transformers, mixture-of-experts, and ResNets under Adam, SignSGD, Muon, and AdaMuon; CompleteP for depth; and the known batch-size prescriptions for the learning rate and AdamW weight decay. It also identifies new rules for the recurrence count in looped transformers, the context length in linear-attention and MLP-Mixer models, and the batch-size scaling for Muon's learning rate, all of which we demonstrate as empirically beneficial. Analogous to automatic differentiation, our results suggest that robust hyperparameter scaling rules can be automated, freeing practitioners to scale up new architectures and optimizers without rederiving the underlying theory or risking subtle errors that quietly cost compute at scale.
SCDM: Scalable Causal Discovery in Nonlinear Temporal Systems with Meta-Learning
Jingbo Wang ⋅ Kegeng Tang ⋅ Zihao Wang ⋅ Shaogang Ren
Causal discovery from nonlinear multivariate time series is challenging in high-dimensional systems, where the number of candidate directed relations grows quadratically with the number of variables. Existing methods are often limited by target-wise model fitting, repeated conditional testing, or expensive graph-search procedures. We propose SCDM, a shared-parameter meta-learning framework for scalable temporal causal discovery. SCDM treats each target variable as a task while learning a shared temporal predictor and a shared causal-strength matrix across tasks. Each meta-episode updates only a subset of target-variable tasks, allowing the model to accumulate structural evidence across the full system without exhaustive target-wise optimization. After training, a post-training graph readout converts the learned predictor into continuous directed causal scores for threshold-independent evaluation. We also provide a finite-error recovery principle showing that, under fully observed delayed-system assumptions and a positive risk-gap condition, thresholding SCDM scores recovers the graph when statistical, meta-optimization, and readout errors are below the population margin. Experiments on controlled synthetic, neuroimaging-inspired, high-dimensional, and real-world temporal benchmarks demonstrate competitive causal recovery and improved scalability in large temporal systems.
SceneAligner: 3D-Grounded Floorplan Localization in the Wild
Junhyeong Cho ⋅ Ruojin Cai ⋅ Hadar Averbuch-Elor
Many public buildings provide floorplans with a “you are here” indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. However, existing methods typically assume controlled small-scale environments and precise vectorized floorplans, limiting their ability to operate in large-scale buildings and rasterized floorplans. In this work, we present an approach for performing floorplan localization in the wild by grounding the task in a reconstructed 3D representation of the scene. Given an unconstrained image collection, our method reconstructs a gravity-aligned 3D scene and projects it into a 2D density map that serves as a floorplan proxy. Floorplan localization is then formulated as aligning this proxy with the input floorplan via a 2D similarity transform. To bridge the appearance gap between density maps and architectural floorplans, we adapt a 2D foundation model to learn cross-modal correspondences, introducing a fine-tuning scheme that encourages semantically aligned matches while preserving structural consistency. Extensive experiments demonstrate substantial improvements over prior methods, including in extremely sparse settings with as little as a single input image. Our code and data will be publicly available.
SceneFactory: GPU-Accelerated Multi-Agent Driving Simulation with Physics-Based Vehicle Dynamics
Yicheng Zhu ⋅ Yang Chen ⋅ Tao Li ⋅ Zilin Bian
Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide articulated vehicle dynamics and contact, but their non-vectorized control interfaces make large batched training difficult. GPU-batched systems such as Waymax and GPUDrive scale to hundreds of scenarios by replacing rigid-body physics with simplified kinematics models, omitting tire--road interaction, suspension, contact dynamics, and road-condition-dependent friction. We introduce SceneFactory, a GPU-vectorized platform for procedural scene construction, physics-based multi-agent simulation, and reinforcement learning in autonomous driving environments. Built on NVIDIA Isaac Sim and Isaac Lab, SceneFactory represents worlds and agents as batched tensors: vehicle control, observations, rewards, resets, and policy inference are executed as GPU tensor operations over the Isaac Lab tensor API. SceneFactory converts Waymo Open Motion Dataset road topologies into simulation-ready USD (Universal Scene Description) worlds. SceneFactory runs many worlds concurrently on one GPU, populates each with multiple articulated PhysX vehicles, and maps precipitation and road-surface type to PhysX material friction coefficients. Thanks to the GPU vectorization, SceneFactory achieves up to 127$\times$ higher throughput than a non-vectorized PhysX baseline on the same GPU and physics solver, reaching 19,250 controlled-agent simulation steps per second (CASPS) at 256 worlds $\times$ 16 agents. Cross-simulator transfer reveals an asymmetric dynamics gap: physics-grounded RL driving policies transfer to a simplified kinematic bicycle model with 99.5\% success, whereas the reverse transfer success rate drops to 47.3\%. Under wet-road friction, friction-aware policies reduce mean peak deceleration rate to avoid crash (DRAC) from 58.7 to 27.8\,m/s$^2$ without sacrificing goal reach. SceneFactory shows that scalable autonomous-driving training need not discard articulated rigid-body dynamics or physically grounded road-condition variation.
Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so locally plausible actions can lead to sharply different outcomes. We frame this as a problem of computer/OS state exploration, where effective behavior requires both expanding the reachable frontier and reducing ambiguity before committing. We present ScreenSearch, a system that combines structural screen retrieval and deduplication with an ambiguity-aware PUCT graph-bandit for large-scale desktop exploration. The retrieval layer converts UIA trees into location-aware structural features, indexes related screens through sparse token search and metadata filters, and maintains a shared deduplicated state graph across VM workers. On top of this graph, we define a scalable ambiguity signal based on matched-action outcome dispersion. If similar screens produce different next states under the same action signature, the state should be probed further rather than treated as resolved. We use this signal together with frontier rewards to drive large-scale exploration and replay-start policy evaluation over the shared graph. Across 11 desktop applications, ScreenSearch collects over 1M screenshots and over 30K deduplicated states, yielding large exploration corpora with substantial cross-application and within-application diversity. On a fixed replay-start slice, we observe a clear novelty--ambiguity trade-off: some policies reduce ambiguity quickly while discovering little frontier. Ambiguity reduction alone is therefore not a sufficient exploration objective. Appendix ablations show that stronger proposal priors can materially improve unique-state discovery during corpus building. These results suggest that state identity, proposal quality, and ambiguity-aware search all matter when deciding when to probe and when to commit.
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Ce Zhang ⋅ Ziyang Wang ⋅ Yulu Pan ⋅ Oluwatumininu Oguntola ⋅ Pranav Wagh ⋅ Qiyu Wu ⋅ Hiromi Wakaki ⋅ Mohit Bansal ⋅ Gedas Bertasius
Grounded long-video question answering, or Grounded LVQA, is the task of answering a question about a long video while also locating the short time interval in the video that supports the answer. Recent agentic methods approach this problem as a multi-turn exploration process with a single cropvideo(start, end) action. This action allows the model to progressively narrow its search from coarse to fine regions, but it does not provide a direct way to backtrack from a fine-grained mistake to a broader context. As a result, these agents often stop after only two turns and are unable to recover once they descend into the wrong part of the video. We propose VideoTreeSearch (VTS), a framework that formulates grounded LVQA as an iterative, self-correcting search over an adaptive temporal tree. VTS builds a non-uniform tree from scene boundaries so that each node corresponds to a semantically coherent video segment. It then trains an agent to navigate this tree using four discrete actions: zoomin, zoom_out, shift, and answer. These actions make backtracking and recovery explicit and learnable, rather than leaving them as implicit behaviors. To train the agent, we introduce a trajectory synthesis pipeline that generates multi-step navigation paths through the tree, including intentional detours into incorrect branches followed by recovery. These trajectories are first used for supervised fine-tuning and then for reinforcement learning with rewards based on grounding quality and answer accuracy. On three Grounded LVQA benchmarks—CG-Bench, Haystack-LVBench, and Haystack-Ego4D—VTS outperforms the strongest previous agentic methods by 12.5 mIoU on CG-Bench and 7.4 T-F1 on Haystack-Ego4D. The learned policy also transfers to general long-video question answering, surpassing all prior agentic baselines on Video-MME, MLVU, and LVBench by up to 7.1 accuracy points. Ablation studies show that self-correcting hierarchical search is the key factor behind these improvements: removing either adaptive descent or explicit backtracking leads to substantial performance drops.
SEED: Self-Speculative Decoding via Implicit Encoder–Decoder
Hankun Lin ⋅ Patrick Pynadath ⋅ Ruqi Zhang
Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose ***s****elf-sp****e****culative ****e****ncoder-****d****ecoder* (**SEED**), a self-speculative method that obtains high-quality drafts cheaply by reusing the contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder–decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder–decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.5$\times$ average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 17\% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning.
See it to Place it: Evolving Macro Placements with Vision Language Models
Ikechukwu Uchendu ⋅ Swati Goel ⋅ Karly Hou ⋅ Ebrahim Songhori ⋅ Kuang-Huei Lee ⋅ Wenjie (Joe) Jiang ⋅ Vijay Janapa Reddi ⋅ Vincent Zhuang
We propose using Vision-Language Models (VLMs) for macro placement in chip floorplanning, a complex optimization task that has recently shown promising advancements through machine learning methods. Because human designers rely heavily on spatial reasoning to arrange components on the chip canvas, we hypothesize that VLMs with strong visual reasoning abilities can effectively complement existing placement algorithms. We introduce VeoPlace (Visual Evolutionary Optimization Placement), a novel framework that uses a VLM—without any fine-tuning—to guide the actions of a base placer by constraining them to subregions of the chip canvas. The VLM proposals are iteratively optimized through an evolutionary search strategy with respect to resulting placement quality. Using global half-perimeter wirelength (gHPWL) after standard-cell placement and legalization as the evaluation metric, VeoPlace boosts ChiPFormer performance on 9 of 10 benchmarks with peak reductions exceeding 32%. We further demonstrate that VeoPlace generalizes to analytical placers, improving DREAMPlace on all 8 evaluated Superblue benchmarks with gains up to 4.3%. Our approach opens new possibilities for electronic design automation tools that leverage foundation models to solve complex physical design problems.
Self Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Haydn Jones ⋅ Yimeng Zeng ⋅ Alden Rose ⋅ Yifei Li ⋅ Yining Huang ⋅ Kaiwen Wu ⋅ Jiaming Liang ⋅ Maggie Huan ⋅ Yoseph Barash ⋅ Cesar de la Fuente-Nunez ⋅ Osbert Bastani ⋅ Zachary Ives ⋅ Mark Yatskar ⋅ Jacob Gardner
Manually curated biomedical repositories---spanning bioactivity, genomics, and chemistry---are expensive to maintain, lag behind primary literature, and often discard experimental context. The absence of contextual information obscures critical nuances, thereby complicating the assessment of data correctness and coverage, necessary criteria for building high-quality models. We show that PubMed itself can be turned into structured datasets---autonomously and cost-effectively---that are larger, more nuanced, and more accurate than the curated databases they would replace. We present three coupled contributions: (1) an LLM-based entity-tagging pipeline, grounded in nine biomedical ontologies, that tags 4.5 billion entities across 19 categories in a 22.5M-paper, 2.5-trillion-token PubMed corpus; (2) hybrid sparse–dense retrieval infrastructure supporting surgical entity-filtered semantic queries over the tagged corpus; and (3) Starling, a multi-agent deep research system that, given only a natural-language task description, autonomously designs precision- and recall-targeted retrieval filters, induces an extraction schema, and emits structured records with nuance-rich fields and supporting passages. Applied to six tasks---blood-brain barrier permeability, oral bioavailability, acute toxicity (LD50), gene-disease associations, protein subcellular localization, and chemical reactions---Starling produces ~7.7M records (per-task scale ranges from 131K to 3M); several of these are, to our knowledge, the largest public datasets for their respective properties. Frontier-model rejection of our kept extractions is 0.6–7.7% across tasks, surprisingly far below the error rates we measure on the widely used, manually curated counterparts (e.g., 16.5% on BBBMartins, 7.3% on BioavailabilityMa). Beyond scale and accuracy, the attached supporting passages carry nuance that tabular databases discard: for example, oral bioavailability of a molecule might depend on whether the patient is fed or fasted. Together, the corpus, retrieval layer, and agent establish a foundation for multimodal predictive and generative models in AI-driven therapeutic design. We release code and links to our datasets at https://anonymous.4open.science/r/starling-B026/.
SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers
Wonsuk Jang ⋅ Thierry Tambe
Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder edge deployment. Quantization can reduce these costs, yet existing methods often degrade video quality due to high activation variation and the difficulty of preserving semantic and temporal coherence. We propose SemanticDialect, which advances block-wise mixed-format quantization. In this framework, each block selects an optimal format (dialect) from a candidate set (formatbook), which is augmented with lookup tables that store quantization errors and quantized indices, enabling efficient per-block format selection and quantization with minimal online overhead. We further introduce attention-guided activation decomposition, which reduces quantization error via residual quantization, and semantic-aware dialect assignment (SeDA), which reduces cross-token quantization inconsistency by enforcing format uniformity among semantically correlated tokens. Experiments demonstrate that SemanticDialect outperforms prior quantization methods and block-wise formats (MXFP4, NVFP4) while approaching FP16 quality on Open-Sora 2.0. We also validate hardware deployability through RTL design and GPU kernel implementation.
ShadowBench: Exposing Lexical Anchoring and the Illusion of Forgetting in Large Language Models
Sujan Maharjan ⋅ Lu
Large Language Models (LLMs) have emerged as primary interfaces for factual retrieval, yet current evaluation paradigms rely almost exclusively on explicit entity names. In this work, we demonstrate that LLM knowledge is fundamentally lexically anchored: models possess extensive factual information but find it significantly difficult to retrieve without explicit name tokens. To quantify this limitation, we introduce ShadowBench, a rigorously hardened, shortcut-resistant benchmark designed to evaluate Latent Entity Association via a novel Dual-Trait Association (DTA) task. Our evaluation reveals a pervasive "Shadow Gap" across all model scales – up to the frontier models GPT-5.4 and Claude-Sonnet-4.6 – where removing lexical anchors causes performance drops of over 20%. Furthermore, we demonstrate that this lexical dependency exposes a critical vulnerability in AI safety. Applying ShadowBench to Machine Unlearning, we find that state-of-the-art algorithms (e.g., Gradient Difference, NPO) achieve only superficial lexical erasure. While unlearned models fail on direct queries, our novel Latent Entity Leakage Rate (LELR) metric reveals that reasoning models explicitly reconstruct the "forgotten" entity in their internal reasoning traces in over 89% of cases, utilizing residual shadow knowledge to solve associative tasks. Ultimately, ShadowBench proves that current unlearning paradigms create an "Illusion of Forgetting," demonstrating that existing methods act as superficial output filters and are not yet reliable for achieving true parametric erasure.
Signature Approach for Contextual Bandits with Nonlinear and Path-dependent Rewards
Xin Guo ⋅ Grace He ⋅ Xinyu Li
We study contextual bandits with nonlinear and path-dependent rewards through a novel signature-transform-based approach. Leveraging the universal nonlinearity property of signatures, we approximate continuous path-dependent reward functionals by linear functionals in the signature space. This representation enables the use of efficient linear contextual bandit methods while preserving expressive sequential structure. Building on this framework, we propose $\texttt{DisSigUCB}$, a signature-based disjoint upper confidence bound (UCB) algorithm. Under boundedness and non-degeneracy assumptions, we prove a high-probability data-dependent sublinear regret bound of order $\tilde{\mathcal O}(\sqrt{(d+m)KT})$ where $d$ is the context dimension and $m$ is the signature feature dimension. Experiments on temperature sensor monitoring, sleep-stage classification, and hospital nurse staffing demonstrate that $\texttt{DisSigUCB}$ consistently outperforms classical linear and kernelized contextual bandit baselines in nonlinear and path-dependent settings.
SLOT-IR: Learning Disentangled Slot Representations for Infrared Spectral Unmixing
Jingru Gan ⋅ Yannah J.U. Melle ⋅ Yanqiao Zhu ⋅ Daniel Schwalbe-Koda ⋅ Wei Wang
Infrared (IR) spectral unmixing recovers individual component spectra from a measured mixture, enabling chemical identification. Existing methods either assume linear superposition of components, which fails for liquid-phase mixtures where molecular interactions alter spectral shapes, or use deep networks that entangle all components in a shared representation. We introduce SLOT-IR, a slot-based architecture that decomposes mixture spectra into separate per-component embeddings and decodes each independently into a predicted spectrum. A two-stage training procedure first learns to map a single molecule from gas-phase spectra to liquid-phase, then trains the full unmixing model on mixtures. On the benchmark of annotated infrared spectra, SLOT-IR outperforms all baselines on both binary and ternary mixture identification. To verify that the model captures genuine chemical structure rather than statistical shortcuts, we apply a post-hoc BatchTopK Sparse Autoencoder (SAE) to frozen slot embeddings and test whether recovered features correspond to known functional groups. Statistical and causal analyses confirm model alignment with established chemistry.
sMMC-22M: A Context-Aware Dataset and Benchmark for Single-Cell Spatial Transcriptomics
Xi Li ⋅ Yaqi Hu ⋅ Ziheng Duan ⋅ Xinyi Wang ⋅ Simon D Sun ⋅ Diptanshu Sikdar ⋅ Yang Liu ⋅ Jing Zhang
Spatial transcriptomics has created a compelling opportunity to test whether tissue morphology can predict molecular state, but existing benchmarks are constrained by limited scale, spot-level supervision, and incomplete biological context. We present sMMC-22M, a cell-aligned multimodal resource comprising over 20 mil- lion cells across 25 organ categories, 66 studies, and multiple spatial-transcriptomic assays. We organize the benchmark around three data-centric axes that determine whether histology-to-molecular modeling can move beyond local interpolation. First, sMMC-22M provides scale: broad organ and study coverage enables con- trolled encoder benchmarking and reveals a scaling trend in which larger pathology foundation models improve morphology–gene correspondence under matched evaluation. Second, sMMC-22M provides resolution: by decomposing assays into aligned cell-level records, our framework converts spot-level histology–omics pipelines into single-cell predictors and evaluates them under strict in-domain and cross-patient splits. Third, sMMC-22M provides rich context: each cell is paired with spatial, molecular, and sample-level metadata, allowing analyses such as age-band shift in ovarian samples, where age-mismatched transfer sharply reduces prediction quality despite misleading global-distance summaries. Together, sMMC- 22M and STBoost establish a practical framework for single-cell histology-to-gene prediction while showing that robust generalization still depends on scale, cellular resolution, and explicit biological context.
Smoothed Elicitation Complexity for Approximate $\Gamma$-calibration of Discrete Classification Tasks
Jessica Finocchiaro ⋅ Victor Ganson ⋅ Drona Khurana
One prominent method of evaluating machine learning model trustworthiness is the notion of \emph{calibration}. In the binary outcome setting, a probabilistic predictor is calibrated if outcomes are realized according to a model's distributional prediction, conditioned on this prediction. Straightforward extensions of binary calibration definitions to probabilistic multiclass classifiers suffer from an exponential complexity blowup as the space of predictions grows exponentially in the number of classes $n$. As a remedy, \citet{noarov_statistical_2023} propose multiclass calibration with predictions that are \emph{properties} of the outcome distribution, reducing complexity from growing in the number of classes $n$ to the \emph{dimension} $d$ of the property, called its elicitation complexity. Previous work on approximate property calibration is generally limited to continuous scalar properties, despite many relevant properties of interest being discrete, like the mode or rankings. We characterize the approximate property calibration of discrete properties which are strongly orderable by using Lipschitz continuous properties as an intermediary. This work is the first to our knowledge to provide approximate calibration results for discrete properties. Along the way, we characterize the Lipschitz elicitation complexity of strongly orderable discrete properties by constructing algorithms for designing these Lipschitz properties, which we prove can be post-processed to obtain the original discrete property.
SPANUQ: Span-Level Uncertainty Quantification for Large Language Model Generation
Yimeng Zhang ⋅ Yingying Zhuang ⋅ Ziyi Wang ⋅ Yuxuan Lu ⋅ Pei Chen ⋅ Aman Gupta ⋅ Zhe Su ⋅ Ming Tan ⋅ Zhilin Zhang ⋅ Qun Liu ⋅ Manikandarajan Ramanathan ⋅ Rajashekar Maragoud ⋅ Edward Vul ⋅ Jing Huang ⋅ Dakuo Wang
Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence, while sequence-level scores fail to localize errors. We formalize Span-Level Uncertainty Estimation (SLUE), a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning. To address this task, we introduce SPANUQ, a lightweight ($\sim$25M parameter) probe that distills the uncertainty knowledge from expensive multi-sample inference into a single forward pass over LLM hidden states. SPANUQ employs a DETR-style span decoder to simultaneously detect spans and estimate their uncertainty via a Mixture of Beta distribution, trained with a principled combination of Beta NLL regression and contrastive ranking objectives. We construct SPANUQ-BENCH, the first span-level uncertainty benchmark comprising 20K prompts, $\sim$293K annotated spans, and continuous soft labels derived from multi-sample claim verification. Experiments on five LLM backbones show that SPANUQ consistently achieves the best span-level uncertainty quality (AUROC 0.908–0.944, MAE 0.110–0.129), outperforming the strongest probe baseline and all sampling-based methods while being $10$--$20\times$ faster. Its DETR-based span detector attains 0.910 F1, surpassing the best heuristic by 39.4\%, enabling precise error localization that sequence-level methods cannot provide. The framework generalizes across five LLMs spanning two model families (AUROC 0.908--0.944), and we additionally observe that sequence-level uncertainty is partially decomposable: the learned importance-weighted span composition achieves $\rho_{\text{seq}} = 0.839$, suggesting that span-level estimation subsumes sequence-level as a special case.
Sparse Expansion Utility: Identifying and Routing to Pivotal Steps in LLM Reasoning Chains
Rian Atri ⋅ Evan Luo
LLM chain-of-thought reasoning improves with additional compute, but standard 2 inference scaling [Wei et al., 2022, Wang et al., 2022, Snell et al., 2024, Brown 3 et al., 2024] treats all steps equally. We ask: which step in a reasoning chain benefits 4 most from extra computation? We make three layered contributions, each scoped 5 explicitly. (I) Measurement. We define expansion utility Uk as the marginal 6 gain in correctness probability from resampling from step k onward, measured via 7 a step-oracle sweep. Across nine model families (8B–123B), seven benchmarks 8 (math, science, code), and 6,500 chains, every one of the 41 healthy (model, bench) 9 cells shows positive oracle gap at 10% step budget (95% Wilson CI [0.914, 1.000]; 10 mean +22.0 pp, range +1 to +48 pp). (II) Structure. Per-step utility is sparse 11 (97.6% of cells: Gini ≥ 0.80); Gini and mean oracle gap are tightly anti-correlated (Pearson r = −0.913, p = 9.5 × 10−17 12 ). A resample-stability simulation and 13 a segmentation-rule sensitivity check (Appendix J) bound the headline against 14 noise and rule-choice confounders. (III) Routing proof-of-concept. A cross15 model router trained with a within-example listwise loss [Cao et al., 2007] and a 16 learned depth prior, combined at inference with a bench-conditional depth mask 17 and a confidence gate, recovers 5-seed-averaged gated GR@10% = +0.149 (95% 18 bootstrap CI [+0.085, +0.217]) of the cross-model oracle gap at matched step19 budget, with no architectural change to the underlying model. We frame this as 20 proof the framework is actionable cross-model, not as deployment-ready scaling; 21 matched-token comparison to self-consistency / best-of-N requires a separate 22 generation campaign and is left to follow-up (§7). Methodological observation. 23 Within-example pointwise top-k BCE under cross-model pooling produces a high24 AUC discriminator (AUC@10 = 0.70) that selects the wrong steps (gated GR 25 = −0.15); the listwise loss is shift-invariant to per-(model, bench) utility scale and 26 removes this inversion.
Language models are increasingly used in text generation, decision support, and automated interaction, where their behavior must be controlled in a localized, reversible, and selective way. Existing methods either update shared parameters or constrain generation through external interfaces, leaving open whether frozen models contain compact internal control. We introduce sparse internal control, a framework for steering target behavior by applying inference-time interventions to a small set of internal model nodes. The framework formulates control through local target controllability, where candidate site–direction pairs are characterized by their behavioral effects under target, preservation, and energy constraints. This yields two complementary selectors: Driver-OMP, which selects nodes that sparsely reconstruct a desired behavioral displacement, and Coverage-Driver, which favors nodes whose effects cover prompts stably. We further propose a six-axis control-evaluation protocol measuring reachability, signed reversal, dose response, feedback controllability, off-target preservation, and held-out reuse. Across IOI, MMLU MCQA, and sentiment steering on models from GPT-2 small to Qwen3-4B, sparse driver sets act as executable actuators: they steer target behavior, respond monotonically to dose, support closed-loop control, and preserve unrelated next-token behavior. On refusal steering, the same protocol exposes a boundary case that passes all evaluated axes except strict reachability. Held-out reuse separates prompt-specific controls requiring refitted strengths from population-level controls that transfer as frozen interventions. These results show that model control can move beyond output-side steering: mechanistic internal variables can be selected, certified, and reused as localized control handles for frozen language models.
How can we elicit truthful information from strategic agents? Traditional peer prediction mechanisms incentivize truthful reporting by rewarding agreement between agents' reports, but break down when agents can submit cheap but correlated signals---such as outputs from different LLMs. We introduce the Speakeasy mechanism, a framework for peer prediction when the principal can also sample these cheap signals. The mechanism augments any peer prediction rule with an audit strategy that penalizes reports for matching the principal's samples. We give a polynomial-time algorithm for the optimal audit strategy, so that truthful reporting is a strict Bayes--Nash equilibrium and yields the highest welfare against agents choosing between reporting truthfully and copying a single cheap signal. We extend the guarantee to more complex misreport strategies that mix multiple cheap signals, showing that Speakeasy preserves the truthfulness of the base mechanism. We empirically test the Speakeasy mechanism against classical peer prediction mechanisms on a peer-review dataset of ICLR submissions, with reviews from human reviewers and several LLMs. Several classical mechanisms fail to make truthful reporting a Bayes--Nash equilibrium once LLMs are available, while Speakeasy restores it for all of them.
SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
Boyao Wang ⋅ Zhihan Lei
Modular networks pursue specialization through learned routers, gates, and load-balancing losses. However, at matched total-parameter budgets, learned routers can underperform equal-weight No-Routing baselines. Is the bottleneck the routing algorithm, or the alignment between training-signal granularity and the target categories? We show that when each training and inference unit carries one coarse category tag, a fixed parameter-free routing scheme (SpecDrop) induces branch-category specialization and matches or exceeds the learned-routing baselines we evaluate at matched parameter budget; when this alignment breaks, the scheme matches the multi-branch No-Routing baseline. SpecDrop assigns each of $K$ branches weight $p_{\mathrm{a}}$ for its category and a small leakage $p_{\mathrm{i}}{>}0$ otherwise, merged through a category-independent fixed denominator that calibrates merged-branch magnitude to single-branch scale, with no learnable routing parameters or auxiliary losses. Category labels are required at inference, drawn from dataset metadata. On vision tasks where each image has one superclass label, SpecDrop reaches $\mathbf{79.23\%}$ on CIFAR-100 with ResNet-110, outperforming dense by $+4.75$, and $\mathbf{79.89\%}$ on ImageNet-1K with Vision Transformer ViT-S/16, outperforming the matched-supervision No-Routing baseline with a shared expert by $+6.53$. SpecDrop reaches the highest top-1 among the multi-branch routing baselines we evaluate at this parameter budget (Soft MoE and Mod-Squad on ImageNet). On NLP tasks where training units span multiple categories, SpecDrop matches the multi-branch No-Routing baseline on both SlimPajama-6B language modeling with a 30M Transformer and SuperNI instruction tuning over Llama-3.2-1B with LoRA. Furthermore, SpecDrop matches or outperforms multi-branch routing baselines including Demix and LoRAMoE, and is statistically tied with HydraLoRA. Per-branch pruning sensitivity reveals four specialization regimes ranked by category clarity: strongest on ImageNet-1K, strong on CIFAR-100, weakening on SlimPajama, and anti-aligned on SuperNI LoRA. Granularity alignment, not algorithm choice, localizes when routing helps. Code: https://anonymous.4open.science/r/C30862.
SpecForge: Agent-Oriented Code Documentation Optimization via Multi-Frontier Tree Search
YUTONG CHENG ⋅ Haifeng Chen ⋅ Peng Gao ⋅ Wei Cheng
As language models become autonomous software agents, documentation shifts from explanatory prose to a natural-language behavioral specification for code generation. We study agent-oriented documentation generation, where the objective is not readability but the correctness of code synthesized from documentation alone, turning the task into a black-box search over natural-language specifications evaluated only through the generated code and its execution behavior. Its defining difficulty is output coupling: program entities are behaviorally entangled, so a revision that sharpens the specification of one entity may simultaneously destabilize its dependents, inducing the familiar whack-a-mole failure mode in iterative refinement. We introduce SpecForge, a multi-frontier tree search solution that preserves complementary search frontiers through Pareto state management to avoid premature commitment and escape certain frontier-local optima, performs dependency-constrained bandit selection for callee-before-caller refinement, and uses diversified error-conditioned expansion to remain robust to noisy test feedback. On DevEval+, SpecForge attains the best performance across five backbone models, including a 94.9% solve rate with Claude-4.5-Sonnet and a 25.7% average improvement over the leading baseline. We further demonstrate that the resulting specifications improve downstream performance on both cross-language code translation and new-feature implementation.
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
Mehrdad Saberi ⋅ Keivan Rezaei ⋅ Soheil Feizi
Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop tool use in complex tasks introduces substantial latency, since the model must repeatedly wait for tool observations before continuing. We study how to accelerate such trajectories without changing the final trajectory the model would have taken without acceleration, assuming access to faster but less reliable speculator tools. We develop a theoretical framework for lossless speculation in multi-hop tool-use settings, characterizing the optimal achievable latency gain. We propose SpecHop, a continuous speculation framework that maintains multiple speculative threads, verifies predicted observations asynchronously as target tool outputs arrive, commits correct branches, and rolls back incorrect ones. This preserves task accuracy while reducing wall-clock latency. We show that SpecHop can approach the oracle latency gain with sufficient active threads. Empirically, we evaluate SpecHop on retrieval-augmented multi-hop tasks and find that its latency gains closely match theoretical predictions, reaching up to 40% latency reduction in some settings.
Spectrally Parameterized Neural Inverse Reconstruction
Harshvardhan Takawale ⋅ Aritrik Ghosh ⋅ Nirupam Roy
In neural inverse reconstruction, the forward model is typically treated as a fixed simulator that maps a neural scene representation to measurements. This work studies how the parameterization of the forward operator itself shapes the optimization landscape of coherent inverse problems. We introduce Moray, a reconstruction framework built on spectrally parameterized forward operators that bypass conventional time-domain synthesis and instead evaluate measurements directly through closed-form spectral kernels. This parameterization induces a substantially shorter and better-conditioned differentiable graph compared to the time-domain approaches. We further introduce Phase-Coherent Manifold Parameterization, a scene representation that jointly learns scene reflectivity and a deformable surface manifold, allowing the reconstruction to adapt to geometric deviations from the assumed imaging plane and thereby reducing common artifacts in coherent reconstruction. We instantiate these ideas for 77 GHz radar imaging using real measurements from a synthetic-aperture mmWave system. Across multiple challenging scenes, Moray consistently improves reconstruction fidelity, suppresses background artifacts, and degrades more gracefully under data limitations.More broadly, the results suggest that, in neural inverse problems, exposing appropriate analytical structure of the sensing physics to the optimizer can be as important as the choice of neural representation itself.
Stochastic Approximation Approach for Decentralized Optimization on Time Varying Random Networks
Chung-Yiu Yau ⋅ Haoming Liu ⋅ Hoi-To Wai
This paper initiates the study of a stochastic approximation approach with the Fully Stochastic Primal Dual Algorithm (FSPDA) framework for decentralized optimization on random and time varying topologies. Our framework relies on a novel observation that randomnesses in time varying topology can be incorporated into a stochastic equality constrained optimization formulation. We derived two new algorithms supporting sparsified communication on time varying topologies --- FSPDA-SA allows agents to execute multiple local gradient steps to accelerate convergence, and FSPDA-STORM further incorporates variance reduction to improve sample complexity. For problems with smooth (possibly non-convex) objective function, within $T$ iterations, FSPDA-SA (resp. FSPDA-STORM) finds an $\mathcal{O}( 1/\sqrt{T} )$-stationary (resp. $\mathcal{O}( 1/T^{2/3} )$) solution. The latter shows the first near-optimal convergence rate over time varying topology.
AI-driven review is poised to structure which science gets attention. By implementing checks for reproducibility, robustness, preregistration, claim scope, and other proxies, machine learning research is institutionalizing metascientific filters on scientific production. This position paper argues that mechanizing reform heuristics whose theoretical grounding remains contested or incomplete is counterproductive to scientific progress. AI review papers should be evaluated as proposed decision policies that shape future research. However, currently the emerging literature blurs the line between integrity filtering, based on necessary but insufficient signals of validity like reproducibility of stated results or lack of fake citations, and epistemic filtering, which uses machine-detectable signals to judge scientific quality. Drawing on debates in metascience, we show that proposed filters--including replicability, multiverse robustness, and preregistration of analysis surfaces--are insufficiently justified as general indicators of scientific value. We argue that human-in-the-loop review fails to resolve the problem, because automated signals shape attention and create incentives upstream. Instead, the field must move toward more rigorous motivation of signal-to-decision pipelines, including explicit specification of target constructs, linking assumptions, decision uses, failure modes, and incentive effects.
Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this assumption about distillation in large language model (LLM) pretraining. By varying architecture sizes and training token budgets, we create strong-to-weak, same-level, and weak-to-strong teacher-student relationships. We study distillation's effectiveness under these relationships with different mixes between language and distillation loss. Three findings emerge: (1) with proper loss mixing, weak-to-strong and same-level distillation improves over standard pretraining, where even small and undertrained teachers benefit large students; (2) making the teacher stronger can lead to saturated or even reversed gains; (3) distillation improves generalization (out-of-domain, downstream) more readily than in-domain fitting. Our results provide practical guidance for choosing the teacher model with the proper loss for more effective LLM pretraining distillation.
Structure Over Scale: Learning Visual Reasoning from Pedagogical Video
Bishoy Galoaa ⋅ Xiangyu Bai ⋅ Sarah Ostadabbas
State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and object selection that a preschooler solves without effort. We hypothesize that the explicit pedagogical structure, specifically the context-question-pause-answer cycles embedded in children's educational video, provides naturally co-aligned reasoning traces: temporally synchronized visual cues, questions, and answers that emerge only from deliberate pedagogical authoring and cannot be practically reconstructed through manual annotation at scale. To test this, we introduce SoSVQA (Structure over Scale Visual Question Answering), a unified benchmark of 10K question-answer pairs automatically extracted from Dora the Explorer (DoraVQA) and Mickey Mouse Clubhouse (ClubHVQA) with precise timestamp alignment, and fine-tune Qwen2-VL and Qwen3-VL using Group Relative Policy Optimization (GRPO) to leverage the clear correctness signals and structured reasoning traces inherent in educational content. Despite training on just 10K QA pairs from 78 hours of children's television, orders of magnitude less data than GPT and Gemini, our approach delivers generalizable performance gains for Qwen-based VLMs, yielding consistent improvements on NExT-QA (+19.7), Video-MME (+10.6), and MotionBench (+4.9), matching the performance of leading proprietary systems and demonstrating that content structure can compensate for content scale.
StyleStream 2.0: Fast and Controllable Streaming Voice Style Conversion
Yisi Liu ⋅ Nicholas Lee ⋅ Gopala Anumanchipalli
Voice style conversion (VSC) aims to transform a source utterance to match the timbre, accent, and emotion of a target voice while preserving linguistic content. StyleStream 1.0 introduced the first streamable zero-shot VSC system with state-of-the-art conversion quality, but remained limited in two ways: approximately 1s end-to-end latency on consumer hardware, and the lack of a text-based interface for designing or editing target voices. We present StyleStream 2.0, a fast and controllable streaming VSC framework that addresses these limitations. To reduce latency, we train a few-step pixel MeanFlow model and further fine-tune it with autoregressive feedback inspired by Self Forcing, improving robustness under streaming inference. This reduces end-to-end latency to 520~ms on a consumer GPU, a 2x speedup over StyleStream 1.0, while maintaining conversion quality. To enable controllable style manipulation, we remove the mel-spectrogram context and instead route target style information through a compact style embedding. In this embedding space, we train a unified text-conditioned flow matching model that supports both text-based voice design, which maps natural language prompts to voice styles, and instruction-based voice editing, which modifies a specific attribute of a source utterance while preserving the rest. Experiments show that StyleStream 2.0 achieves the strongest target style fidelity on VSC, voice design, and voice editing, while remaining competitive on intelligibility and non-edited attribute preservation.
Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation
Jacob Dang ⋅ Brian Y Xie ⋅ Omar G. Younis
Recent work on subliminal learning has shown that language models can transmit semantic traits through data with no apparent connection to those traits. Whether this phenomenon extends to agentic systems, where policies are learned from trajectories rather than static text, remains an open question with critical safety implications. We present the first empirical evidence that unsafe agent behaviors transfer subliminally through model distillation, demonstrated across two complementary settings. In the primary setting, we construct a teacher agent with a deletion bias, a tendency to perform destructive file-system actions through an API-style tool interface, and distill it into a student using only trajectories from ostensibly safe tasks, with all explicit deletion keywords filtered. In the secondary setting, we replicate the threat model in a native Bash environment, replacing API calls with shell commands and operationalizing the bias as a preference for chmod over semantically equivalent alternatives (e.g., chown, setfacl) when issuing the first permission-related command. Despite thorough sanitation, students inherit measurable biases in both settings: the API student's deletion rate reaches 100% (vs. 5% baseline) under homogeneous distillation, while the Bash student's chmod-first rate reaches 30–55% (vs. 0–10% baseline), with the strongest transfer observed in large-to-small distillation. Evaluations on TerminalBench further show that subliminal transfer persists in complex, multi-step tasks, indicating that behavioral biases are encoded implicitly in trajectory dynamics, independent of tool interface or keyword filtering. Explicit data sanitation is therefore insufficient as a defense; mitigating behavioral bias transfer in agentic systems will require fundamentally new strategies.
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
Jonathan Zheng ⋅ Zirui Shao ⋅ Alan Ritter ⋅ Wei "Coco" Xu
Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid existing knowledge. In this work, we propose a synthetic, simulation-driven framework for studying knowledge updates in LLMs. We introduce PARALLELEVENTS, a benchmark of fictional yet realistic future worlds that generates coherent event trajectories for controlled evaluation, avoiding contamination while preserving consistency. Building on this dataset, we develop SYNAPSE, a training framework that uses model-generated data to update model parameters via mid-training and instruction tuning. This synthetic pipeline enables scalable knowledge integration without costly human-curated data. Empirically, {\sc Synapse} outperforms existing methods by 14.23\%, demonstrating that simulation-based synthetic training leads to robust and coherent knowledge updates.
TAC: Timestamped Audio Captioning
Sonal Kumar ⋅ Prem Seetharaman ⋅ Ke Chen ⋅ Oriol Nieto ⋅ Jiaqi Su ⋅ Zhepei Wang ⋅ Rithesh Kumar ⋅ Dinesh Manocha ⋅ Nicholas J. Bryan ⋅ Zeyu Jin ⋅ Justin Salamon
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces temporally grounded audio descriptions at varying degrees of detail and resolution. TAC is trained with a synthetic data pipeline that constructs challenging and dynamic mixtures from real-world audio sources, enabling robust learning under realistic polyphonic conditions. Across event detection, dense captioning, and speaker diarization, TAC outperforms most competing methods - including dedicated diarization systems with a low hallucination rate and accurate temporal grounding. We also introduce TAC-V, an audio-visual pipeline to generate semantically rich audio-visual descriptions. We then show that TAC serves as a "semantic bridge" for a text-only reasoner: a simple TAC→LLM and TAC-V→LLM cascade achieves state-of-the-art scores on benchmarks for both audio (MMAU-Pro, MMSU, MMAR) and audio-visual (DailyOmni, VideoHolmes) understanding and reasoning respectively. We encourage readers to see detailed qualitative results on our demo page: https://audioanon.github.io/tacmodel/.
Tcell: Mitigating Harmful Fine-tuning for Large Language Models via Gradient Alignment
Tiansheng Huang ⋅ Virat Vishnu Shejwalkar ⋅ Oscar Chang ⋅ Ling Liu
Harmful fine-tuning attack becomes a concerning safety risk for mainstream fine-tuning-as-a-service providers, as attackers can submit harmful data to the API to compromise the safety alignment of the large language models. In this paper, we first explore an intuitive gradient mixing solution, and derive a key property ensuring the success of defense -- \emph{taking a fine-tuning update that has higher cosine similarity with the safety gradient can mitigate harmful fine-tuning.} Motivated by this key property, we design an alignment-stage defense, dubbed Tcell. The core contribution of Tcell is a regularizer that align the harmful gradient and the safety gradient during safety alignment, which ensures the harmful fine-tuning update exhibits high cosine similarity with the safety gradient, achieving \emph{gradient alignment}. The benefit of gradient alignment is supported by i) empirical and theoretical interpretation and ii) comparison to alternative design without gradient alignment. Code is available at https://anonymous.4open.science/r/Tcell-1C15/.
Temporal Island Sparse Autoencoders for Interpreting Clinical Time-Series Models
Muhan Yeo ⋅ Bong Gyun Kang ⋅ HyunGi Kim ⋅ Sungroh Yoon
Deep learning models for electronic health records achieve strong predictive performance but provide little insight into what clinical concepts they represent internally. Post-hoc attribution methods produce per-cell importance scores that lack temporal coherence, cross-variable structure, and reusability across patients. We introduce the Temporal Island Sparse Autoencoder (TI-SAE), which learns a backbone-specific, patient-shared clinical concept dictionary from a frozen predictive model’s representations. Each dictionary entry (latent) is defined by learned temporal Gaussian islands—specifying when the concept is active—and a sparse variable-composition vector—specifying which variables are involved. A decoupled selector then identifies which dictionary entries drive each patient’s prediction, separating representation learning from attribution. We evaluate TI-SAE across five backbone architectures spanning different representation strategies on MIMIC-III and MIMIC-IV (three prediction tasks each). TI-SAE produces structured clinical concepts that go beyond transient attribution heatmaps while achieving strong performance on standard faithfulness benchmarks.
TensorCommitments: A Lightweight Verifiable Inference for Large Language Models
Oguzhan Baser ⋅ Elahe Sadeghi ⋅ Eric Wang ⋅ Nico Vergauwen ⋅ Sam Kazemian ⋅ Hong Kang ⋅ Sandeep Chinchali ⋅ Sriram Vishwanath
Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve verifiable LLM inference, where a prover (the service) must convince a verifier (the client) that an inference was run correctly without rerunning the LLM. Existing cryptographic works are too slow at the LLM scale, while non-cryptographic ones require a strong verifier GPU. We propose TensorCommitments (TCs), a tensor-native proof-of-inference scheme. TC binds the LLM inference to a commitment, an irreversible tag that breaks under tampering, organized in our multivariate Terkle Trees. For LLaMA2, TC adds only 0.97% prover and 0.12% verifier time over inference while improving robustness to tailored attacks by up to 48% over the best prior work requiring a verifier GPU.
Testing and Estimation of Contextual Generalized Thurstone Models
Dongmin Lee ⋅ Anuran Makur ⋅ Japneet Singh ⋅ Boyu Xu
Many modern machine learning methods, such as reinforcement learning from human feedback (RLHF) to train reward functions, utilize preference learning models like Bradley-Terry-Luce (BTL) to learn latent scores of items based on pairwise comparison data. Despite the successes of such models in specific applications, the theoretical justification for the employed models has often been unclear. In this work, we provide a hypothesis testing procedure to test whether a given dataset of pairwise comparisons follows a contextual generalized Thurstone (CGT) model under certain classes of latent functions. The latent functions determine the latent scores, and can be expressed as weighted sums of basis functions. Our CGT model encompasses a wide class of preference learning models, including the popular BTL model. In particular, we prove critical thresholds for our tests in the minimax sense, showing that they scale like $1/\smash{\sqrt{|\mathcal{E}|^{1/2}k}}$ under certain important regimes, such as complete graphs and perfect matchings. Furthermore, to establish our testing results, we also derive error bounds for CGT parameter estimation that generalize and improve on known results in the BTL and non-contextual settings. Finally, we conduct experiments on various synthetic and real-world datasets to verify our theoretical results.
\texttt{FEROM}: Frontier Endogenous Reveal-Order Marginal Policy Optimization for Masked Diffusion LMs
Zian Su ⋅ Ziyang Huang ⋅ Kaiyuan Zhang ⋅ Xiangyu Zhang
Masked diffusion language models generate text by iteratively unmasking positions, with practical samplers often selecting reveal positions from the model's own logits. The reveal order is therefore an endogenous latent variable of the rollout process with an enormous discrete latent space. Policy optimization objectives relying on pre-defined schedules or sampled trajectories introduce either bias or high variance. In this paper, we propose \texttt{FEROM}, \emph{Frontier Endogenous Reveal-Order Marginal Policy Optimization}, which targets the rollout-induced marginal response policy of masked diffusion LMs. \texttt{FEROM} derives a Rao--Blackwellized policy-gradient identity over latent reveal orders and expresses the resulting estimator as posterior edge occupancy on a reveal-state DAG. To make marginalization practical, we introduce Frontier Reveal Marginalization, a budgeted estimator that combines scorelaw local reveal estimation with frontier expansion of high-mass partial states. Integrated into a GRPO-style objective, \texttt{FEROM} replaces single-path log-scores with a locally marginalized edge-based surrogate. Experiments on math and coding tasks show comparable or improved results over existing methods under matched compute budgets. Offline proxy study further shows potential in gains with increased budgets.
The Complexity Kink: LLM Rubric Instruments for Causal Inference on Code Generation Reliability
Michael Hernandez ⋅ Tian Zhao
Large language models (LLMs) are increasingly evaluated and deployed as code generators, but benchmarks still lack a reliable way to measure when a programming task becomes structurally too complex for a model to solve. Many code evaluations estimate task complexity from the code a model produces. This makes the key variable endogenous: when a model fails on a difficult prompt, it may emit a short stub, partial solution, or broken program whose measured complexity is low, causing hard failures to be reclassified as easy cases. This can hide whether reliability degrades smoothly or changes regime at a sharp complexity threshold, a distinction that matters for model selection, benchmark design, and deployment guardrails. We present the Complexity Kink benchmark, a prompt-side experiment that measures intended solution complexity before generation and separates it from generated-output complexity and functional correctness. We construct a 5,000-prompt Python benchmark from OpenCodeInstruct, stratified to cover the range of structural complexity, and score each prompt with a fixed six-dimensional rubric covering branching, iteration, state, data structures, edge cases, and algorithmic composition. Four out-of-panel LLM judges score the full prompt set, yielding complete coverage and high composite inter-rater reliability (ICC = 0.865). We then evaluate a 21-model panel with unit-test execution and estimate the relationship between prompt-side complexity and pass rate using instrumental variables and threshold regression. By making task complexity observable before generation, this framework tests whether the apparent complexity kink is a real structural break or a measurement artifact, and gives researchers and practitioners a more defensible way to compare models, diagnose reliability limits, and design evaluations that do not erase the failures that matter most.
The Cost of Absolute Position: A Spread-Expressivity Tradeoff for Additive Positional Encodings
Noah Mitchell ⋅ Isaac Gabriel ⋅ Alexander Wyatt
Self-attention without positional information is permutation equivariant, therefore positional encodings are required whenever a transformer must distinguish input order. However, injecting absolute position can also weaken length generalization. We formalize this tension for additive positional encodings by introducing positional spread, the total variance of the positional embedding matrix. For additive PE in softmax attention with mean pooling, we prove an upper bound demonstrating that the equivariance breaking component of risk is controlled by $\|\widetilde E\|_F^2$, along with a matching lower bound for self-separable additive PEs. We finally prove a spread-expressivity tradeoff: any additive PE that distinguishes absolute positions with minimum separation $\delta$ must have spread $\Omega(\delta^2 n)$. The theory predicts two regimes: when the target is near permutation invariant, lower spread improves length extrapolation; when the task requires absolute position, additional spread reduction destroys expressivity. We test this prediction using controlled synthetic tasks, spread-regularized additive PE, and long context retrieval and reasoning tasks. Across low positional sensitivity tasks, spread strongly predicts OOD accuracy; on absolute position tasks, performance demonstrates the predicted non-monotone tradeoff. For score-based methods such as RoPE and ALiBi, we include a displacement based comparison rather than matching bounds, clarifying both the reach and limits of the spread theory.
The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show
Parsa Esmati ⋅ Somjit Nath ⋅ Katja Hofmann ⋅ Derek Nowrouzezahrai ⋅ Samira Ebrahimi Kahou ⋅ Majid Mirmehdi
Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet it remains unclear whether these models internally encode physical structure, or merely reproduce motion patterns seen during training. We study this question by probing video diffusion models along latent trajectories corresponding to real videos with known physical plausibility. To obtain such trajectories, we approximately invert the deterministic sampling process by integrating the learned velocity field backward from a clean video latent to noise, giving access to the model’s intermediate states and attention maps. Using these recovered trajectories, we show that physical plausibility is linearly decodable from diffusion transformer states across IntPhys and InfLevel, reaching around 81\% average accuracy and outperforming dedicated representation-learning baselines such as V-JEPA and VideoMAE. Surprisingly, this signal is absent from the VAE latent input and emerges inside the denoising transformer itself, despite the model not being trained with a self-supervised predictive objective. These findings suggest that physically meaningful representations can arise as a byproduct of generative denoising.
Theoretical Limits of Language Model Alignment
Lucas Monteiro Paes ⋅ Natalie Mackraz ⋅ Barry-John Theobald ⋅ Federico Danieli
Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i) reinforcement learning, which maximizes the expected reward under a KL-divergence constraint, and (ii) best-of-$N$ alignment, which selects the highest-reward output among $N$ independent samples. Despite their widespread use, the fundamental limits of reward improvement under a KL budget remain poorly understood. We characterize the information-theoretic limits of KL-regularized alignment by deriving the maximum achievable expected reward gain for a fixed KL-divergence budget. Our first result provides a closed-form expression for the optimal reward improvement, governed by a Jeffreys divergence term rather than the $\sqrt{\texttt{KL}}$ used in prior analyses. We further reformulate this expression as a covariance under the base model, yielding a practical estimator that predicts achievable alignment gains from base model samples alone. We extend our analysis to the proxy reward setting, showing that the gap between ideal and proxy alignment (reward hacking) grows with the KL penalty factor and magnitude of reward error. We then prove that reward ensembling mitigates reward hacking, providing a theoretical justification for this technique used in practice. Empirically, we compute the KL–reward Pareto frontier for two alignment tasks for LMs, safety and summarization, and show that best-of-$N$ closely approaches the theoretical limit, while PPO and GRPO remain substantially suboptimal. Our theoretical results shed light on several empirically observed phenomena in the alignment literature and suggest that algorithmic improvements are needed to achieve optimal alignment without high inference costs.
Theory on Attention Dynamics for Out-of-Distribution In-Context Learning
Junze Deng ⋅ Daouda Sow ⋅ Sen Lin ⋅ Yingbin Liang
Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called $\alpha$-type and $\beta$-type attention weights, which represent the transformer’s confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.
The Power of a Random Sample in Online Algorithms
Omer Wasim ⋅ Sami Davies ⋅ Shallu Tomer ⋅ Rathish Das
In this paper, we introduce a semi-random model in online optimization, which interpolates naturally between the standard (adversarial) and random-order arrival model, which we call the online with a random sample (ORS) model: the full input may be chosen adversarially, but a random $p$-fraction of the input is selected and presented to the algorithm in random-order, before the remaining non-sampled elements are presented (in adversarial order). While similar in spirit to the AOS (adversarial order with a sample) model introduced by Kaplan, Naori and Raz, in our ORS model, the sampled $p$-fraction is revealed online instead of offline, and hence, competitiveness is measured with respect to the full input. Our central focus in the paper is applying the ORS model to the secretary problem and its natural $(k,1)$ variant. For the secretary problem, we present an optimal $pe^{-p}$-competitive algorithm, while for the $(k,1)$-secretary problem, we obtain a competitive algorithm whose performance converges to optimal competitiveness as $p\rightarrow 1$. We also include a $O(\log (1/p))$-competitive algorithm for facility location.
We study dynamic pricing where a seller repeatedly interacts with a strategic, non-myopic buyer who has a fixed private valuation and discounts future utility. Prior work focused exclusively on posted-price mechanisms, where the seller gives a take-it-or-leave-it offer. For our first result, we show that menu mechanisms consisting of allocation-payment achieve $O(T_\gamma \log T_\gamma)$ regret, where $T_\gamma$ is the buyer's effective discounted time horizon. We also establish a $\Omega(T_\gamma)$ lower bound, demonstrating the bound is tight up to $\log$ factors. Considering the geometric discounting buyer with a constant discount factor, our bound is $O(1)$, while prior bounds using posted-price mechanisms incur an unavoidable $\Omega(\log\log T)$ factor in regret. Our second contribution is more conceptual in nature. The problem of dynamic pricing sits at the intersection of two paradigms: learning with strategic agents in computer science / machine learning and revelation-principle-based mechanism design in economics, yet their relationship has remained unclear. We establish a fundamental equivalence: indirect learning-based mechanisms and direct revelation mechanisms achieve identical optimal regret. The adaptive, data-driven algorithms of online learning and explicit type elicitation are two languages towards solving the same problem.
The Prestige: Benchmarking Cognitive Visual Reasoning using Magic Tricks
Shuo Wen ⋅ Beixi Du ⋅ Edwin Meriaux ⋅ Junming Shi ⋅ Chloe Si ⋅ David Meger ⋅ Doina Precup ⋅ Gregory Dudek
With the rapid development of Vision-Language Models (VLMs), it is increasingly critical to evaluate their ability in accurate and robust reasoning based solely on visual information. However, existing video understanding benchmarks often overlook the deep cognitive processes inherent in visual reasoning and frequently rely on multiple-choice formats, implicitly encouraging reasoning shortcuts. To address these limitations, we introduce \emph{The Prestige}, a novel video understanding benchmark designed to isolate VLM reasoning under deliberate cognitive misdirection through professional magic performances. The dataset consists of a curated collection of high-quality magic videos paired with a diverse distribution of strictly open-ended questions targeting causal state tracking, paradoxical reasoning, and resistance to adversarial questions. These questions need the models to perform long-horizon, multi-step commonsense reasoning with cognition. Through extensive evaluations of leading open-weight and proprietary VLMs, using a combination of machine and human grading, we find a substantial gap between model and human performance. \emph{The Prestige} establishes a rigorous new frontier for evaluating temporal, causal, and cognitive-flow reasoning in multimodal foundation models.
The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models
Nguyen Phuc ⋅ Chinh D La ⋅ Duy M. H. Nguyen ⋅ Nitesh Chawla ⋅ Binh T. Nguyen ⋅ Khoa D Doan
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key method for improving Large Language Models' reasoning capabilities, yet recent evidence suggests it may paradoxically shrink the reasoning boundary rather than expand it. This paper explains this shrinkage issue of RLVR by analyzing its learning dynamics and reveals two critical phenomena that account for this failure. First, we expose negative interference in RLVR, where learning to solve certain training problems actively reduces the likelihood of correct solutions for others, leading to the decline of Pass@$k$ performance, or the probability of generating a correct solution within $k$ attempts. Second, we uncover the winner-take-all phenomenon: RLVR disproportionately reinforces problems with high likelihood, correct solutions under the base model, while suppressing other initially low-likelihood ones. Through extensive theoretical and empirical analysis on multiple mathematical reasoning benchmarks, we show that this effect arises from the inherent on-policy sampling in standard RL objectives, causing the model to converge toward narrow solution strategies. These insights motivate us to design a \textit{simple yet effective} data curation algorithm that focuses RLVR learning on low-likelihood problems (the non-winners) and achieves notable improvement in Pass@$k$ performance, illustrating the potential of a broader class of data-curation mitigation strategies.
Visual in-context learning (ICL) enables vision-language models (VLMs) to adapt at test time via demonstration examples, but its effectiveness depends critically on the examples selected. Recent counterfactual selection methods improve ICL by constructing demonstrations that expose how individual attribute changes affect the answer. However, they require multi-stage pipelines involving attribute extraction, caption engineering, and composed image matching. We propose a simpler alternative grounded in a key observation: a model's own prediction errors are natural counterfactual demonstrations that directly encode its failure modes. Our method, SMILE (Self-Mined Visual In-Context Learning from Errors), builds a model-specific pool of hard negatives offline by recording the target VLM's own mistakes. At test time, the prediction serves as a direction signal that surfaces, in the semantic embedding space, the past mistakes most likely to recur for the current query. This prediction-conditioned identification provides a unified mechanism for both closed-set classification and open-ended visual question answering without task-specific design. Across four benchmarks and four VLMs, SMILE delivers an average task-score gain of +2.72 points over the SOTA baseline at 5× lower latency. Source code is provided in the supplementary material.
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
Anay Kulkarni ⋅ Chia En Lu ⋅ Dheeraj Mekala ⋅ Jayanth Srinivasa ⋅ Gaowen Liu ⋅ Jingbo Shang
Tool use enables large language models to solve complex tasks through sequences of API calls, yet existing reinforcement learning approaches fail to scale to multi-step composition settings. Outcome-based rewards provide only sparse feedback, while trajectory-supervised rewards depend on annotated reference solutions, penalizing valid alternatives and limiting scalability. We propose TIER: Trajectory-Invariant Execution Rewards, a reward framework that derives supervision directly from function schemas and runtime execution, rather than from reference trajectories. The reward decomposes into format validity, schema adherence, execution success, and answer correctness, providing dense, interpretable sequence-level feedback derived from fine-grained verification of individual steps of tool use. This design allows any valid execution path to receive credit, naturally supporting multiple solution strategies and adapting to evolving tool interfaces. On DepthBench, a compositional benchmark stratified by depth (1 to 6 steps), TIER achieves >90% accuracy across steps, where trajectory-supervised rewards collapse beyond step-4. We further demonstrate consistent gains on benchmarks like BFCL v3 and NestFUL. Ablation studies confirm that all reward components are necessary, highlighting the importance of multi-level supervision for compositional reasoning.
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
Xiaona Zhou ⋅ Tianjiao (Joey) Yu ⋅ Muntasir Wahed ⋅ Constantin Brif ⋅ Ismini Lourentzou
Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory performance when applying large language or multimodal models to finding abnormal patterns in sequential data. Public anomaly detection benchmarks typically provide interval annotations but not natural-language rationales, making it difficult to fine-tune VLMs to produce grounded, interpretable decisions. To address this gap, we construct VisAnomBench, an explanation-augmented curated benchmark built from public time-series datasets and augmented with high-quality anomaly explanations selected from multiple large VLMs using fine-grained, task-specific rewards. Building on this benchmark, we present VisAnomReasoner, a parameter-efficient VLM for time-series anomaly detection. Experimental results on VisAnomBench show that VisAnomReasoner achieves more accurate anomaly localization and consistently outperforms all baselines, with improvements of at least 21.23 and 23.87 percentage points in precision and F1, respectively. Additional experiments on the TSB-AD-U benchmark demonstrate strong cross-benchmark generalization, with VisAnomReasoner improving precision and F1 by 9.57 and 13.39 percentage points, respectively. Dataset, code, and model weights will be open-sourced.
To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems
Pengfei He ⋅ Zhenwei Dai ⋅ Xianfeng Tang ⋅ Yue XING ⋅ Hui Liu ⋅ Jingying Zeng ⋅ Qiankun Peng ⋅ Shrivats Agrawal ⋅ Samarth Varshney ⋅ Suhang Wang ⋅ Jiliang Tang ⋅ Qi He
Large Language Model-based Multi-Agent Systems (LLM-MAS) have demonstrated strong capabilities in solving complex tasks but remain vulnerable when agents receive unreliable messages. This vulnerability stems from a fundamental gap: LLM agents treat all incoming messages equally without evaluating their trustworthiness. While some existing studies approach trustworthiness, they focus on a single type of harmfulness rather than analyze it in a holistic approach from multiple trustworthiness perspectives. We address this gap by proposing a comprehensive definition of trustworthiness inspired by human communication theory \citep{grice1975logic}. Our definition identifies six orthogonal trust dimensions that provide interpretable measures of trustworthiness. Building on this definition, we introduce the Attention Trust Score (A‑Trust), a lightweight, attention‑based method for evaluating the trustworthiness of messages. We then develop a principled trust management system (TMS) for LLM‑MAS that supports both message‑level and agent‑level trust assessments. Experiments across diverse multi‑agent settings and tasks demonstrate that our TMS significantly improves robustness against malicious inputs.
Toward in Silico Strain Evaluation: A Multimodal Surrogate for Fermentation Dynamics with Metabolic Graph Pretraining
Yunxiao Li ⋅ Difeng Gao ⋅ Yubin Zheng ⋅ JIJIAO ZENG
Strain engineering for industrial fermentation faces a structural design-test scale gap: combinatorial pathway perturbations of $2$–$3$ enzymes already span $10^4$–$10^6$ candidate strains, while bench-scale fermentation throughput remains on the order of $10^2$ runs per laboratory per year. One way to narrow this gap is in silico evaluation by perturbing a learned gene-expression $\to$ phenotype mapping (conditioned on process state); the precondition is that such a mapping can be learned from bench-scale-sized datasets. As a proof-of-concept, we ask whether a genome-scale metabolic model (GEM)-graph mechanistic prior supports learning such a mapping from $\sim44$ bioreactor runs of an engineered Yarrowia lipolytica astaxanthin-producing strain. We propose a spatio-temporal GNN encoder structured as a Mass Flow Graph over the GEM, paired with a multimodal input head over gene-expression snapshots, sensor streams, time-series assays, and run-level metadata. The encoder is pretrained by masked-flux prediction on Flux Balance Analysis (FBA) samples, distilling the GEM's stoichiometric and mass-balance constraints into it before fermentation data are seen. Under $5$-fold cross-validation across $5$ seeds, the surrogate attains the lowest mean composite test RMSE among the baselines we compare against. Structural ablations show that the GEM-graph prior is necessary—an order-of-magnitude RMSE divergence when removed—while pretraining together with the gene-expression modality make a significant joint contribution, supporting the proof-of-concept claim.
Training Agent to Scale Inference-Time Reasoning
Yunzhe Qi ⋅ Sirui Chen ⋅ Jiaru Zou ⋅ Yanjun Zhao ⋅ Mengting Ai ⋅ Jingrui He
Inference-time scaling provides a promising path for improving agentic reasoning performance, and one widely adopted form is sample-then-select: first sample a pool of candidate reasoning trajectories, then select from the resulting candidates. While this keeps scaling lightweight without updating the policy, it exposes a training-deployment mismatch. Standard trajectory-wise Reinforcement Learning (RL) objectives optimize normalized signals for individual trajectories, whereas choosing from a large candidate pool with the policy's own selection signal is inherently set-conditional: it requires calibrated ranking and confidence aggregation across correlated trajectories. To address this challenge, we propose AutoPortfolio, a set-conditional policy alignment framework for efficient and effective sample-then-select scaling. AutoPortfolio builds a tractable environment-informed target on the sampled candidates, and uses its induced Plackett-Luce (PL) ranking to calibrate the policy which trajectories to prioritize within that set. It trains the policy with a balanced PL-motivated objective that reallocates probability mass among competing trajectories, preserving reward-aligned alternatives while sharpening away from low-quality distractors. Aligned with our policy training, our lightweight inference-time Mass Aggregation scaling adaptively perceives confidence from candidate trajectories and selects the most promising answer(s). Theoretically, we show that reducing our set-conditional loss narrows the discrepancy between policy-induced and environment-informed PL rankings on the sampled set. Empirically, across 12 challenging reasoning benchmarks, AutoPortfolio improves over strong agentic RL baselines, with the clearest gains in larger candidate pools where the set-level selection is crucial.
Train on the Sphere, Deploy on the Hill: Closed-Form-Anchored Surrogates for Real-Terrain Boundary-Integral Equations
Stephane Zsoldos ⋅ Therice Morris ⋅ Varundev Sukhil ⋅ Benjamin Wetherfield
The analysis of electrostatic induction effects on transmission lines in real-world power distribution systems requires inferring induced surface charge density $\sigma$ on overhead-line corridors at deployment scale, where dense $\mathcal{O}(N^3)$ boundary-element solvers for the underlying Poisson boundary integral equation (BIE) are infeasible. We present a recipe for training cheap PDE surrogates from canonical-case closed forms (flat ground, grounded sphere) plus sparse dense-solver fine-tuning where the canonical case breaks like in our case study. The recipe also avoids a structural degeneracy between training operator $A_\theta$ and solution $\sigma$ in self-supervised operator learning that tends to lead to internally consistent but non-physical solutions. The recipe combines a closed-form local-tangent-plane baseline plus a small neural correction; a frozen analytic kernel; an $SE(3)$- and scale-invariant dimensionless feature stack derived from a $2$-term identity $\sigma = -(\nabla\phi\cdot\hat{n})/(2\pi) + H\phi/(4\pi)$, exact on flat ground and the grounded sphere and supervised training against this identity. An $8$-parameter linear head recovers the analytic curvature coefficient $+1/(4\pi)$ from data to within $2\%$; holds to $1\%$ relative RMS in $\sigma$ across $12$ orders of magnitude in source-charge scale, where DeepONet, Set Transformer, and FNO baselines miss the analytic answer by orders of magnitude out-of-distribution; and matches direct-collocation boundary-element ground truth on a synthetic curved patch to $3.8\%$ relative RMS. A regime-of-validity result delineates when the canonical $2$-term identity applies; outside it (terrain curvature radius $\kappa^{-1}$ much greater than source clearance height $|h|$, the corridor regime), we extend the recipe with sparse dense-BEM-target supervision. On several industrial transmission-line corridors with native triangulated meshes, the extended recipe attains $9.2\%$ median relative RMS to dense BEM on held-out corridors, $20\times$ tighter than the analytic 2-term plug-in. Deployment: a $1.7$kB MLP, seconds per corridor on a single CPU core.
Trajectory Planning without Trajectory Data: A Manifold-Guided Approach
Silong Yong ⋅ Anji Liu ⋅ Cunxi Dai ⋅ Carl Busart ⋅ Guanya Shi ⋅ Yilun Du ⋅ Katia Sycara ⋅ Yaqi Xie
A common way for trajectory planning is to leverage generative models trained on large collections of expert trajectories. At inference time, the model generates executable trajectories by conditioning on task goal constraints. However, trajectory-based methods rely on costly supervision, scale poorly with sequence length, and often generalize poorly to unseen constraints such as novel start-goal pairs. We propose an alternative to learn the underlying state-space manifold and use the geometry of the manifold for trajectory planning. This approach requires only state observations and enables generalization to unseen constraints by constructing trajectories on the learned manifold of the state space. Experiments on classical planning problems in maze demonstrate the effectiveness of our work. We further show that the method can be used in higher dimension space where table-top robot arms are considered. Our method is able to find feasible path given only infeasible straight line reference, and be comparable to State-of-the-Art trajectory-based model without actually learning on trajectory data.
Transolver-GMsFEM: A Hybrid Framework for High-Contrast Multiscale PDEs on Irregular Grids
Alexander Rudikov ⋅ Sergei Stepanov ⋅ Vladimir Fanaskov ⋅ Eric Chung ⋅ Ekaterina Muravleva ⋅ Ivan Oseledets
High-contrast multiscale PDE problems are common in real-world applications, yet current neural PDE solvers struggle to achieve sufficient accuracy on such tasks. The Generalized Multiscale Finite Element Method (GMsFEM) addresses this by compressing fine-scale heterogeneity into localized basis functions which are used to obtain more accurate solutions than neural PDE solvers. However, constructing these basis functions requires solving many local eigenvalue problems—the major computational bottleneck. We address this issue by proposing **Transolver-GMsFEM**, a new hybrid framework that utilizes the efficiency of neural PDE solver for predicting multiscale basis functions while preserving the solution quality of GMsFEM. Experiments on 2D/3D steady-state and time-dependent high-contrast multiscale PDEs on irregular grids show proposed method achieves over $100\times$ speedup for basis construction. Crucially, our experiments demonstrate that Transolver-GMsFEM significantly outperforms state-of-the-art neural PDE solvers in both accuracy and robustness, especially in out-of-distribution tasks where pure neural PDE solvers dramatically fail.
We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly relevant in modern AI planning, where deep minimax search and Monte Carlo Tree Search (MCTS) with language model long rollouts face a fundamental tradeoff: heuristic evaluations are cheap but biased, while accurate rollouts are reliable but prohibitively expensive. We propose 2FFS, a two-fidelity tree-search algorithm that brings multi-fidelity flat bandit ideas into trees. The algorithm combines minimax-style fast expansion with MCTS-style stochastic sampling, adaptively deciding when to exploit cheap biased evaluations and when to invoke expensive accurate evaluations for local certification. We prove fixed-confidence correctness, establish finite stopping for exact identification, and give a polynomial-depth cost upper bound for general-depth trees. Across numerical stochastic-tree experiments, 2FFS uses substantially fewer samples and computational operations comparing to existing BAI-MCTS baseline.
Ultra Fast PDE Solving via Physics Guided Few-step Diffusion
Xiangrui Cindy Kong ⋅ Yueqi Wang ⋅ Haoyang Zheng ⋅ Weijian Luo ⋅ Guang Lin
Diffusion-based models have demonstrated impressive accuracy and generalization in solving partial differential equations (PDEs). However, they still face significant limitations, such as high sampling costs and insufficient physical consistency, stemming from their many-step iterative sampling mechanism and lack of explicit physics constraints. To address these issues, we propose \emph{Phys-Instruct}, a novel physics-guided distillation framework which (1) not only compresses a pre-trained diffusion PDE solver into a few-step generator via matching generator and prior diffusion distributions to enable rapid sampling, (2) but also enhances the physics consistency by explicitly injecting PDE knowledge through a PDE distillation guidance. Phys-Instruct is built upon a solid theoretical foundation, leading to a practical physics-constrained training objective that admits tractable gradients. Across five PDE benchmarks, Phys-Instruct achieves orders-of-magnitude faster inference while reducing PDE Error by more than 8$\times$ compared to state-of-the-art diffusion baselines. Moreover, the resulting unconditional student model functions as a compact prior, enabling efficient and physically consistent inference for various downstream conditional tasks. Our results indicate that Phys-Instruct is a novel, effective, and efficient framework for ultra-fast PDE solving powered by deep generative models.
Understanding the Effects of Neuron Dominance in Deep Reinforcement Learning
Zifan Wu ⋅ Qian Lin ⋅ Blake Lawlor ⋅ Haijun Zhao ⋅ Daniel Brown
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive neurons, i.e., the so-called ``dormant neurons'', has been identified as one source of capacity loss. This paper investigates \textit{dominant neurons}, neurons whose activation values are significantly larger than average, as a potential cause for neuron dormancy. We demonstrate the existence of dominant neurons in a number of visual control tasks, and perform an analysis of the learning dynamics showing how dominant neurons can induce dormancy in the subsequent layer. To gain a better understanding of this phenomenon, we examine it through the lens of representation learning and establish its connection with representation collapse. Furthermore, this paper evaluates several mitigation strategies for dominant neurons across a variety of visual control tasks. Our results show that strategies that induce lower peak activation scores tend to exhibit greater representational capacity, lower dormant neuron percentage, and better performance. Among these mitigation strategies, LayerNorm with weight decay has the strongest performance, despite its simplicity. Moreover, switching the value learning loss from regression to a classification loss also significantly mitigates the neuron dominance issue and improves the performance. As a potential explanation of the effectiveness of classification losses, we provide an analysis that shows how a classification loss can prevent representation collapse.
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
Jiayun (Peter) Wang ⋅ Yu Wang ⋅ Weijie Gan ⋅ Zhenting Wang ⋅ Wei Wei
We introduce spatially grounded contextual image generation, a new controllable image generation task that reframes the conditioning paradigm. Instead of supplying a reference image and a global text prompt through two separate encoders (vision and language), UniVL is trained to bind semantics to spatial locations directly from a single unified visual input, in which the textual instruction is rendered onto the spatial mask, removing the need for a standalone text encoder at inference. This enables contextual image generation, which follows user’s specified what should appear where instructions, as well as waiving the need of text encoder to save computation significantly. For the task, we propose a framework in which the UniVL encoder—adapted from an optical-character-recognition-pretrained backbone—reads the unified condition optically, producing a UniVL embedding fVIL that fuses visual and semantic intents to spatial locations, packed as a single token sequence. A two-stage pipeline aligns UniVL in VAE embedding space and then conditions a pretrained diffusion backbone entirely on UniVL embeddings, eliminating the standalone text encoder (e.g., T5). The reframing is deliberately minimalist for text, but the empirical payoff is large. On UniVL-ImgGen, a benchmark of 477K mask-annotated images that we construct to support training and evaluation, UniVL achieves superior image quality over text-prompted baselines (FID: 14 → 11, PSNR: 16 → 20) while eliminating the text encoder entirely, reducing inference TFLOPs by up to 52% and runtime by up to 44%. Additional ablation studies verify components of different parts of the proposed method, paving way for efficient and spatially grounded image generation with unified conditioning paradigm.
Unlearning Diffusion Policies via Relative Fisher Forgetting
Manuel Kelly ⋅ Yingxue Zhang ⋅ Fangzhou Lin ⋅ Yanhua Li ⋅ Xin Zhang
As diffusion-based offline reinforcement learning (RL) move closer to real deployment, it becomes critical to remove the influence of specific training data for privacy, safety, and regulatory compliance. Since retained data may overlap with or generalize from the forget data, strict retraining equivalence can be ill-posed in offline RL. Existing unlearning methods are ineffective for diffusion policies, as training influence is dispersed across the denoising process and reinforced by critic values. We introduce Relative Fisher Forgetting (RFF), the first principled framework for selective unlearning in diffusion-based offline RL. RFF combines two asymmetric components: critic-side value suppression removes residual value incentives associated with the forget set, eliminating $Q$-guidance pathways that would otherwise sustain forgotten behaviors; actor-side relative-Fisher updates attenuate forget-set-dominant parameter influence in the denoising policy, reducing behavioral regrowth. To stabilize training, RFF alternates actor-critic updates and employs gradient clipping and retain-set regularization. Experiments on MuJoCo benchmarks show that RFF achieves the lowest identifiable forget-set reliance among baselines while preserving retained performance, and remains over $12\times$ more efficient than retraining. When undesired behaviors are primarily supported by the forget set, RFF additionally suppresses them without collateral degradation.
We show that a generative model can discover concepts from data in an unsupervised setting while learning to generate. The Dirichlet Concept Diffusion Model (DCDM) embeds concept discovery into diffusion-based generation. DCDM learns concept centers and infers a Dirichlet distribution over them for each input. The resulting weighted center shapes the forward diffusion mean and reverse denoising process, so components are learned through the same evidence lower bound used to model the data. The method uses no class labels, attribute annotations, captions, pretrained text-to-image priors, or semantic concept supervision. Analysis shows how the formulation preserves concept-level information along diffusion paths and reduces denoising ambiguity, creating pressure for concept centers to capture stable modes. Experiments across diverse image domains show that the learned components form coherent prototypes, support interventions, and reflect recurring visual structure.
Variational Inference via Entropic Transport Descent
Vincent Pacelli ⋅ Akash Ratheesh Babu ⋅ Evangelos Theodorou
Particle-based variational inference (ParVI) methods approximate an intractable target distribution by evolving an ensemble of interacting samples. Existing approaches rely predominantly on kernel-based repulsion (e.g., SVGD), which suffers from variance collapse in high dimensions and mode collapse on multimodal targets—pathologies caused by the absence of global transport structure. We introduce entropic transport descent (ETD), a ParVI family that frames each particle update as an entropy-regularized optimal transport problem. Derived from the JKO proximal scheme by lifting to the space of couplings and relaxing via the KL chain rule, each ETD iteration reduces to a Sinkhorn computation. The resulting transport plan provides global coordination, guiding each particle to nearby high-density proposals and naturally preserving multimodal structure. ETD can operate entirely score-free, requiring only pointwise evaluations of the unnormalized target density. Experiments on variance-collapse diagnostics, Bayesian logistic regression, neural networks, and molecular Boltzmann distributions show that ETD matches or outperforms SVGD, AGF-SVGD, and SGLD, with the largest gains in high-dimensional and multimodal settings.
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
Zichen Xie ⋅ Mrigank Pawagi ⋅ Yuxin Liu ⋅ Aaditi Rai ⋅ Lize Shao ⋅ John Berberian ⋅ Sicong Che ⋅ Wenxi Wang
Large language models can now generate useful code from natural language, but their outputs still come without correctness guarantees. Verifiable code generation offers a path beyond testing by requiring models to produce not only executable code, but also formal specifications and machine-checkable proofs. Progress in this direction, however, is difficult to measure: existing benchmarks are often small, focus on only one part of the pipeline, lack ground-truth proofs or rigorous specification validation, or target verification settings far from mainstream software development. We present VeriContest, a benchmark of 946 competitive-programming problems from LeetCode and Codeforces for verifiable code generation in Rust with Verus. Each problem pairs a natural language description with expert-validated formal specifications, judge-accepted Rust code, Verus-checked proofs, and positive and negative test suites. VeriContest is constructed through a three-phase pipeline that scales from manually verified seed problems to semi-automated expansion with human-in-the-loop review. To further strengthen benchmark quality, we use testing as an additional quality-assurance layer for validating postcondition completeness. VeriContest supports both isolated and compositional evaluation of specification generation, code generation, proof generation, and end-to-end verified program synthesis. Evaluating ten state-of-the-art models reveals a sharp gap between ordinary coding ability and verifiable code generation: the strongest model reaches 92.18% on natural-language-to-code generation, but only 48.31% on specification generation, 13.95% on proof generation, and 5.29% end-to-end. These results identify proof and specification generation as the central bottlenecks for current models and establish VeriContest as a rigorous platform for measuring and training future systems that generate code with machine-checkable correctness.
Verify0: Can AI Agents Build Formally Verified Software Repositories?
Zhe Ye ⋅ Hantao Lou ⋅ Yuechun Sun ⋅ Peiyang Song ⋅ Zhengxu Yan ⋅ Timothe Kasriel ⋅ Qingyang Zhang ⋅ Kaiyu Yang ⋅ Soonho Kong ⋅ Jingxuan He ⋅ Dawn Song
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated software. Existing benchmarks in this direction either focus on individual functions or only evaluate proof generation with provided implementations. It is still an open question whether agents can make coherent implementation and proof choices across real multi-module codebases. To bridge this gap, we introduce Verify0, the first benchmark to evaluate joint implementation and proof synthesis at the repository level. Verify0 contains 35 multi-module instances sourced from real-world repositories spanning Python, Dafny, Verus, and Coq, and covering diverse domains from cryptographic protocols to distributed systems. Each instance consists of a multi-module Lean 4 repository with predetermined API interfaces, manually curated formal specifications, and reference implementations, supporting both proof-only and code-and-proof evaluation modes. To improve benchmark reliability, Verify0 also includes an audit mechanism where agents are allowed to formally prove unsatisfiability of provided specification or incorrectness of reference code, which surfaces and corrects latent code and specification errors during curation. We evaluate frontier coding-agent configurations with Lean toolchain access. The strongest agent fully solves only 18 of 35 instances and closes no specifications on the hardest repositories. Verify0 provides a concrete testbed for measuring progress toward repository-scale verified software synthesis, where current agents still fall short. We release the benchmark, curation pipeline, and evaluation harness at https://anonymous.4open.science/r/verify0-neurips26-submission-82C2.
VerifyThisBench: Joint Evaluation of Code, Specifications, and Proof
Xun Deng ⋅ Barış Bayazıt ⋅ Si Cheng Zhong ⋅ Andreas Veneris ⋅ Fan Long ⋅ Xujie Si
Large language models (LLMs) have demonstrated remarkable progress in code generation, but many existing benchmarks are approaching saturation and offer little guarantee on the trustworthiness of the generated programs. To improve visibility into model reasoning on formal correctness, we introduce \verify, a new benchmark that evaluates end‑to‑end program verification from natural language descriptions: models must (i) extract formal specifications, (ii) implement in a verification‑aware language, and (iii) construct machine‑checkable proofs. Our evaluation reveals a gap between formal verification and true correctness. While models can sometimes produce programs that pass verification, only a small fraction are actually aligned with the intended task: across 1,078 tasks, six SOTA models collectively produce just 36 human-verified correct solutions, with the best model achieving only 20. This discrepancy arises because verification only guarantees correctness with respect to the generated specification, not the original intent. To address this, we introduce a filtering methodology combining automatically generated test cases with an independent LLM-based semantic judge, requiring human inspection only when the two signals disagree. Further, we propose VerifyThisBench, a relaxed variant where partial specifications, implementations, or proofs are provided disentangle sources of difficulty. Together, We release VerifyThisBench with test suite, VerifyThisBenchXS, and a unified evaluation environment spanning seven verification tools to support future research on trustworthy program synthesis.
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Qijia He ⋅ Xunmei Liu ⋅ Hammaad Memon ⋅ Ziang Li ⋅ Zixian Ma ⋅ Jaemin Cho ⋅ Jason Ren ⋅ Daniel Weld ⋅ Ranjay Krishna
Scalable Vector Graphics (SVG) are essential for technical illustration and digital design, offering resolution independence and semantic editability. In practice, original vector files are frequently lost, leaving only rasterized versions (e.g., PNG, JPEG) that resist modification, while manual reconstruction is prohibitively expensive. Progress on automating raster-to-SVG conversion has been bottlenecked by two gaps: existing SVG datasets are dominated by icons and decorative graphics that lack the complexity of professional diagrams, and existing benchmarks rely on pixel- or embedding-level similarity that fails to capture structural correctness (e.g., broken connectivity, misplaced arrows). We close both gaps with paired contributions targeting diagram-centric figures (e.g., model architectures, flowcharts, schematics). For training, we introduce VFIG-Data, the largest figure-to-SVG dataset of its kind at 66K pairs, combining real paper figures converted via a describe-and-generate pipeline with programmatic diagrams that supply noise-free supervision over arrow styles, fonts, and geometry. For evaluation, we introduce VFIG-Bench, a structure-aware evaluation suite, paired with VFIG-Bench-OOD, an out-of-distribution set of figures manually curated from highly cited arXiv papers. Beyond pixel and embedding similarity, our protocol reports rubric-based VLM-Judge scores and Elo ratings from pairwise human preference evaluation. Built on these contributions, VFIG is a VLM family trained with a simple-to-complex SFT curriculum followed by RL with rendering-aware rewards. VFIG achieves state-of-the-art open-source performance, outperforming the best open-source VLM baseline by over 30\%, and matches Claude Sonnet 4.6 on VFIG-BENCH: Gemini-Judge 78.2% vs. 76.7% and GPT-Judge 87.5% vs. 87.4%. It remains slightly behind the strongest proprietary models GPT-5.2 and Gemini-3.
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
Zhongyi Cai ⋅ Yi Du ⋅ Chen Wang ⋅ Yu Kong
Embodied agents are expected to assist humans by actively exploring unknown environments and reasoning about spatial contexts. When deployed in real life, agents often face sequential tasks where each new task follows the completion of the previous one and may include infeasible objectives, such as searching for non-existent objects. However, most existing research focuses on isolated goals, overlooking the core challenge of sequential tasks: the ability to reuse spatial knowledge accumulated from previous explorations to guide subsequent reasoning and exploration. In this work, we investigate this underexplored yet practically significant embodied AI challenge. Specifically, we propose 3DSPMR, a 3D SPatial Memory Reasoning framework that utilizes Field-of-View (FoV) coverage as an explicit geometric prior. By integrating FoV-based constraints, 3DSPMR significantly enhances an agent’s memory, reasoning, and exploration capabilities across sequential tasks. To facilitate research in this area, we further introduce SEER-Bench, a novel Sequential Embodied Exploration and Reasoning Benchmark that spans two foundational tasks: Embodied Question Answering (EQA) and Embodied Multi-modal Navigation (EMN). SEER-Bench uniquely incorporates both feasible and infeasible tasks to provide a rigorous and comprehensive evaluation of agent performance. Extensive experiments verify that 3DSPMR achieves substantial performance gains on both sequential EQA and EMN tasks.
VolFill: Single-View Amodal 3D Scene Reconstruction with Volumetric Flow Matching
Tuan D Ngo ⋅ Chuang Gan ⋅ Evangelos Kalogerakis
Reconstructing the complete geometry of a scene from a single RGB image remains challenging—especially when inferring hidden structures where visual evidence is incomplete. We introduce VolFill, a generative framework that predicts the 3D structure of the complete scene rather than relying on traditional pixel-aligned regression. Our method utilizes a hybrid 3D VAE to compress sparse truncated unsigned distance function grids into a compact latent space, paired with a latent Diffusion Transformer that denoises this representation to recover the complete scene. We condition the generation on geometry foundation models, leveraging rich spatial priors for robust reasoning. Unlike existing methods limited by per-ray constraints or unstructured point-cloud queries, VolFill provides a structured representation that supports direct surface extraction and occupancy queries at scale. Extensive experiments on the 3D-FRONT and NRGBD datasets demonstrate that our approach significantly outperforms current baselines, providing a robust foundation for holistic spatial understanding.
Watermarking has emerged as a leading technical proposal for attributing generative AI content and is increasingly cited in global governance frameworks. This position paper argues that current implementations risk serving as symbolic compliance rather than delivering effective oversight. We identify a growing gap between regulatory expectations and the technical limitations of existing watermarking schemes. Through analysis of policy proposals and industry practices, we show how incentive structures disincentivize robust, auditable deployments. To realign watermarking with governance goals, we propose a three-layer framework encompassing technical standards, audit infrastructure, and enforcement mechanisms. Without enforceable requirements and independent verification, watermarking will remain inadequate for accountability and ultimately undermine broader efforts in AI safety and regulation.
When and Why is Optimistic Multiplicative Weights Slow? The Geometry of Energy Dissipation
John Lazarsfeld ⋅ Anas Barakat ⋅ Georgios Piliouras ⋅ Antonios Varvitsiotis ⋅ Andre Wibisono
This paper studies the convergence of the Optimistic Multiplicative Weights Update algorithm (OMWU) in two-player zero-sum games. Recent works have identified instances on which the last-iterate of OMWU can converge arbitrarily slowly, but understanding when and why this slow convergence occurs has remained open. In this work, we develop a new analysis framework that gives sharp, quantitative explanations for this behavior. Our analysis is based on viewing the algorithm's dual iterates as an *optimistic skew-gradient descent* with respect to an energy function. We prove over the dual iterates that energy is dissipative, and by establishing tight bounds on the magnitude of dissipation, our analysis quantifies the geometric bottlenecks that arise when the corresponding primal iterates are close to the simplex boundary. This further translates into a new linear last-iterate convergence rate in KL divergence on games with a unique and interior Nash equilibrium. Compared to prior work, this new rate contains a much sharper dependence on game-specific constants, and we prove this dependence is optimal. Moreover, these geometric insights further translate into new separations on *uniform* convergence rates for OMWU. On the one hand, we prove *constant lower bounds* on the uniform *best-iterate* convergence rate in KL divergence and Total Variation distance from Nash. On the other hand, we establish for the $2\times 2$ setting a new $\widetilde O(T^{-1/2})$ best-iterate rate in duality gap, improving substantially over prior work. Together, this shows in general that uniform convergence rate guarantees do not transfer across different measures of distance to Nash.
When Are Compositional Problems Learnable from Verifiable Rewards?
Daniel Barzilai ⋅ Yotam Wolf ⋅ Ronen Basri
Many problems are inherently compositional: solving them requires a sequence of intermediate decisions that jointly determine the final answer. A central question is when such a compositional structure can be learned from outcome-level feedback alone, where supervision of the intermediate steps is unavailable. We study this question for autoregressive models trained using reinforcement learning with verifiable rewards (RLVR). We identify the \emph{task-advantage ratio}, a joint property of the task and the model, which measures whether intermediate decisions present an advantage in reaching a correct final solution. We show that this ratio governs learnability in our setting: when the advantage is present, RLVR efficiently learns the target composition, while when it is absent, training can converge to suboptimal compositions. We further show that the required advantage arises naturally in several structured problems, but may depend critically on the quality of the initial model. Our results help to clarify when compositional problems can be learned from final rewards alone.
A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior. Because these importance estimates serve as the evidence for identifying circuits, systematic errors can lead to the misidentification of the underlying mechanisms. While activation patching provides a gold-standard causal metric, its computational cost is prohibitive at scale. Practitioners instead rely on attribution patching, a gradient-based, first-order approximation whose reliability remains poorly understood. In this work, we characterize the source of this unreliability, demonstrating that the dominant error stems from the non-linearities in the downstream network rather than local curvature at the patched component. This insight yields three practical tools: (i) a reliability score to detect untrustworthy estimates, (ii) error bounds quantifying potential attribution mis-specifications, and (iii) a Hessian-vector-product (HVP) correction that eliminates the leading-order error with only one additional backward pass. In evaluations across five model families (124M–9B parameters) and both random-token and naturalistic (name-swap) perturbations, HVP is the only second-order correction feasible at larger scale, where standard baselines like Integrated Gradients become computationally prohibitive. In comparative experiments, a multi-step HVP variant matches or exceeds the accuracy of Integrated Gradients at significantly lower compute, outperforming prior second-order baselines. These improvements lead to higher-fidelity circuit recovery on standard benchmarks and support a Screen-Flag-Fix workflow that targets computational effort only toward the components flagged as unreliable.
When Expert Disagreement Hurts: Auditing Prestige-Sensitive Revision in LLM Decision Pipelines
Yupeng Tang ⋅ Mingfeng Lin
Language models are increasingly used in decision pipelines where an initial answer is revised after another agent disagrees. In such settings, revision can reflect useful reconsideration, ordinary prompt instability, or sensitivity to the perceived prestige of the disagreeing source. We introduce a controlled two-pass audit that separates these effects by holding the question and evidence fixed while varying only the revision context: neutral reflection, anonymous disagreement, and expert-labeled disagreement. We evaluate five instruction-tuned models on 1,989 evidence-grounded binary decisions from medicine, scientific claim verification, and contract reasoning. Disagreement consistently induces reversals beyond drift, while the incremental effect of an expert label varies across models. More importantly, these reversals are often harmful: under expert-labeled disagreement, 70.1% of reversals aggregated across models move from correct to wrong answers. In a branch-paired API subset, low-prestige disagreement is weaker than expert disagreement, and a simple evidence-gated revision prompt reduces harmful reversals while improving final accuracy. For open-weight models, teacher-forced Yes/No scores reveal that many output reversals are not accompanied by corresponding preference-sign changes. These results make prestige-sensitive revision a concrete, reproducible, and actionable target for LLM evaluation.
When Form Changes but Logic Doesn’t: Building Logic-invariant LLMs through Structures
Xuyuan Liu ⋅ Xinshuai Dong ⋅ Elynn Chen ⋅ Yujun Yan
Large language models (LLMs) have achieved strong performance across diverse reasoning tasks, yet it remains unclear whether this performance reflects an understanding of underlying logical structure or a reliance on shallow semantic cues. A key indicator of such understanding is whether models treat logically equivalent inputs consistently, producing the same predictions despite differences in expression. Through systematic evaluation on curated data, we show that even strong LLMs often fail this test: they do not consistently preserve predictions across logically equivalent inputs and generalize poorly to unseen logical forms. To address this limitation, we propose LoGIcal STructure-guided Reasoning (LoGIST), a framework that guides LLM reasoning with explicit logical structure. LoGIST maps each input instance to a decision diagram, a directed acyclic graph that captures the underlying logical structure, and encodes this graph into a logical embedding that conditions the LLM’s reasoning process. Across multiple training settings, LoGIST reduces inconsistency under logical equivalence and improves generalization to unseen logical forms, suggesting that logical invariance is a useful target for building more robust and generalizable logical reasoning in LLMs.
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Xiaolin Zhou ⋅ Aojie Yuan ⋅ Zheng Luo ⋅ Zipeng Ling ⋅ Xixiao Pan ⋅ Yicheng Gao ⋅ Haiyue Zhang ⋅ Jiate Li ⋅ Shuli Jiang ⋅ Prince Z Wang ⋅ Zixuan Zhu ⋅ Jinbo Liu ⋅ Ryan Rossi ⋅ Hua Wei ⋅ Xiyang Hu
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user typos propagate into hallucinated tool names, a misconfigured request timeout can stall an agent indefinitely, and duplicate tool names across servers can freeze an SDK. We study these failures as a sim-to-real gap in the tool-use partially observable Markov decision process (POMDP), where deployment noise enters through the observation, action space, reward-relevant metadata, or transition dynamics. We introduce RobustBench-TC, a benchmark with 22 perturbation types organized by these four POMDP components, each grounded in a verified GitHub issue or documented tool-calling failure. Across 21 models from 1.5B to 32B parameters (including the closed-source o4-mini), the robustness profile is sharply uneven: observation perturbations reduce accuracy by less than 5%, while reward-relevant and transition perturbations reduce accuracy by roughly 40% and 30%, respectively; scale alone does not close these gaps. We then propose ToolRL-DR, a domain-randomization reinforcement learning (RL) recipe that trains a tool-use agent on perturbation-augmented trajectories spanning the three statically encodable POMDP components. On a 3B backbone, ToolRL-DR-Full retains roughly three-quarters of clean accuracy and reaches an aggregate perturbed accuracy comparable to open-source 14B function-calling baselines while substantially narrowing the gap to o4-mini. It closes approximately 27% of the Transition gap despite never seeing transition perturbations in training, suggesting that RL on adversarial static tool-use inputs induces a more persistent retry policy that transfers to unseen runtime failures. The code and benchmark dataset are available at https://anonymous.4open.science/r/robustbench-tc-release-3AB7/.
When, Where, What: Structural Guarantees for Travel Time Prediction on Temporal Graphs
Gabriel Buginga ⋅ Gabriel D Vilela ⋅ Jincheng Zhou ⋅ Mohit Tawarmalani ⋅ Bruno Ribeiro
Accurate traversal (travel) time distribution estimation on temporal graphs benefits from satisfying six structural properties: when (ordering, flexibility), where (composition with prefix equivariance), and what (non-Markov, snapshot). No existing temporal graph learning method jointly satisfies these properties. We propose PATH-HAZ, a discrete-time hazard model that provides guarantees of all six requirements. It composes per-step PMFs via dynamic programming for spatially consistent predictions across shared paths, while cross-attention over historical graph snapshots encodes confounder-induced dependencies between non-adjacent steps on the path. Evaluated on simulated logistics and real-world high-speed rail data, PATH-HAZ surpasses both point-estimate (19.3% RMSE improvement) and probabilistic baselines (56% CRPS reduction), while also allowing for detailed structural explainability of delays.
Where Tabular Foundation Models Falter on Genetic Data: Datasets That Expose and Provide a Path to Address the Gap
ANIRBAN DAS ⋅ Yan Cui
Tabular foundation models are used as off-the-shelf predictors for heterogeneous tabular tasks, but it remains unclear how they will perform on real genotype-to-phenotype tabular datasets, which carry unique challenges. One such challenge is ancestry-dependent non-stationarity in allele effect sizes (the effects of genetic variation on disease risk): the predictive contribution of a given genetic variant may change significantly across the ancestry spectrum. This is especially consequential for patients from ancestry groups that are underrepresented in existing datasets, because models that are unaware of this non-stationarity are more likely to perform poorly for them. The battery of synthetic tasks used to pretrain tabular foundation models does not capture this structure, and we show that it leads to systematic performance degradation. Using controlled hierarchical Gaussian-process stress tests, we demonstrate that both off-the-shelf TabICL and TabPFN are robust when ancestry dispersion in the data is low, but as dispersion grows, the latent ancestry-dependent non-stationarity of effect sizes becomes a first-order problem, and the predictive performance of both models deteriorates. We confirm the same pattern on All of Us (AoU) (a large biobank containing whole-genome sequencing and electronic health record (EHR) data spanning the ancestry spectrum) across an extensive panel of cancer phenotypes evaluated with the two leading tabular foundation models: TabICL and TabPFN. Holding the in-context training-set size, phenotype, and feature set fixed, ancestry-specific in-context tables, which are less dispersed in ancestry space, consistently outperform size-matched meta-ancestry tables, isolating ancestry dispersion as the driver of degradation. To address the failure without changing the base architecture, we construct two synthetic task families that explicitly encode ancestry-dependent effect drift and instruction-tune an off-the-shelf TabICL model on tasks sampled across both families. On held-out AoU evaluations covering cancer and respiratory disease phenotypes, the tuned model yields more stable performance across ancestry-distance bins and is especially strong in the bins farthest from the center of the in-context exemplars in ancestry space, that is, on subjects whose ancestry is most underrepresented among the provided exemplars. The paper contributes an evaluation protocol, a failure analysis, and two synthetic task families targeted at ancestry-dependent non-stationarity in precision-medicine tabular modeling.
Why Transformer-Based Language Models Need Explicit Mechanisms of Cognitive Control
Suketu Patel ⋅ Hongbin Wang ⋅ Jin Fan
We argue that cognitive control, the capacity to hold multiple divergent candidate responses and reconcile them through categorical inhibition rather than soft blending, should be a first-class computational object in transformer-based language models. We develop this claim at two levels. At the \emph{signal level}, interference paradigms reveal a characteristic failure signature: rather than adjudicating between competing responses, current transformer-based language models blend them, and accuracy collapses as interference scales. At the \emph{goal level}, sustained agency requires maintaining concurrent goals, enforcing priority during signal level adjudication, and revising the goal set as evidence accumulates about whether those goals are being successfully pursued; current systems approximate this only through external scaffolding. The architectural gap revealed at the signal level is the same gap that will block transformers from scaling to the kind of goal modulated and self regulating cognition that long-horizon autonomous agents require. We define concrete criteria for what counts as architecturally explicit cognitive control, show that neither the dominant training pipeline nor reasoning-token training implements it, and commit to predictions that distinguish our position from scale-based optimism.